You know the feeling. You spend an hour with a coding agent establishing the architecture: why the rate limiter sits at the gateway, why cache keys are versioned, which subsystem is frozen. The session ends. Tomorrow the agent is a stranger again — and your AGENTS.md has quietly grown into a 400-line monolith that the model skims, drifts past, or ignores.

That is the exact wound OKF Agent Memory is built to close. The project (okf-memory/okf-agent-memory, MIT) launched on September 5, 2026, has gathered 700+ GitHub stars in under four weeks, and topped the weekly trending charts for AI agent repositories. Its proposition is deliberately unglamorous: your agent's durable knowledge lives in your repository as plain Markdown files with YAML frontmatter — a knowledge/ directory you can read, git diff, and review in a pull request like any other source file.

What makes it more than a folder of notes is the tooling wrapped around it: a single pure-Go binary with zero dependencies that adds BM25 search (the project reports sub-300-microsecond lookups with no vector database and no embedding API costs), bundle validation, and a stdio MCP server so Claude Code, Cursor, Codex, and other agents can query the memory with tools like okf_search and okf_show. No accounts, no API keys, no services to run.

This tutorial installs v0.4.4 (released September 27, 2026), bootstraps a memory bundle into a demo project, records two real architectural concepts, links them, validates the bundle, and then connects the whole thing to an agent over MCP — verifying the protocol handshake with real JSON-RPC calls. Every command below was executed and every output shown is real.

What you'll need #

  • Linux, macOS, or Windows — the tutorial uses the prebuilt Linux binary (7.6 MB); macOS and Windows builds ship in the same release, and macOS users can also brew install okf-memory/tap/okf.
  • A terminal and a scratch project directory. We'll scaffold a demo project, so you don't need to touch your own code.
  • No API key, no GPU, no database, no account. Retrieval is local BM25 in a compiled binary; the bundle is Markdown on disk.
  • Git (optional but recommended) — the bundle is designed to be committed, so memory gains history, blame, and code review for free.

Step 1 — Install okf #

Grab the v0.4.4 binary and its checksum file from the GitHub release, then verify the download before trusting it:

curl -sL -o okf https://github.com/okf-memory/okf-agent-memory/releases/download/v0.4.4/okf-linux-amd64
curl -sL -o checksums.txt https://github.com/okf-memory/okf-agent-memory/releases/download/v0.4.4/checksums.txt
grep okf-linux-amd64 checksums.txt
# 073c393644034080713fca2c9c814da5fcd568e85f0a49c99f187d3448424a9f  bin/okf-linux-amd64
sha256sum okf
# 073c393644034080713fca2c9c814da5fcd568e85f0a49c99f187d3448424a9f  okf

The hashes match, so the binary is the one the maintainers published. Make it executable, put it on your PATH, and confirm the version:

chmod +x okf && sudo mv okf /usr/local/bin/okf
okf version
# okf version v0.4.4 (OKF v0.2 specification)

A look at the full command surface shows the whole workflow in one screen:

okf --help
# OKF Agent Memory CLI (vv0.4.4)
#
# Commands:
#   validate               Validate an OKF bundle for conformance and graph health
#   search                 Search concepts using in-memory BM25 scoring or frontmatter filter
#   show                   Display full concept details, frontmatter, and links
#   create                 Create a new concept with automated bookkeeping
#   update                 Update an existing concept
#   relate                 Connect two concepts with a relative link and context
#   init                   Initialize a new OKF v0.2 bundle (index.md, log.md)
#   bootstrap              Scaffold complete memory stack (skill, AGENTS.md, knowledge, Makefile)
#   agents                 Manage AGENTS.md, lint AAG rules, and maintain SSoT tool symlinks
#   mcp                    Run as a Model Context Protocol (MCP) server over stdio
#   hub                    Zero-knowledge sync and vault management (push, pull, sync, serve)
#   version                Print version information

Step 2 — Bootstrap memory into a project #

okf bootstrap scaffolds the complete memory stack into any directory — new or existing — in one command. We'll use a disposable demo project (replace demo with your own repo path when you're ready):

mkdir -p demo
okf bootstrap demo --name "Demo API"
# Successfully bootstrapped OKF Agent Memory in 'demo'!
# Created:
#   - knowledge/ (index.md, log.md)
#   - .agents/skills/okf-memory/ (SKILL.md, discovery, update, etc.)
#   - AGENTS.md
#   - Makefile
#
# Run 'okf validate knowledge' or 'make validate' to verify.

Four things landed, and each has a job:

  • knowledge/ — the OKF v0.2 bundle itself: index.md (root progressive-disclosure index) and log.md (dated changelog, starting with today's initialization entry).
  • .agents/skills/okf-memory/ — an agent skill definition (SKILL.md plus discovery.md, remember.md, update.md, relationships.md, examples.md) that teaches any agent the search-before-write protocol.
  • AGENTS.md — project instructions for AI coding agents in the project's Dual-Memory Agent Architecture: a compact behavioral codex with hard invariants like "MUST execute okf_search(query=keywords, limit=3) before proposing architecture" and "NEVER scan knowledge/ via list_dir, grep_search, find, or raw file readers" — agents must use the search tools, never brute-force the files.
  • Makefile — conveniences: make validate runs strict validation, make search q="..." queries memory.

This is the project's central design bet, and it's worth stating plainly: instead of choosing between a bloated prompt monolith (AGENTS.md stuffed with everything) and a vector-database black box (agents never semantically search for operational rules mid-task), memory is split into a tiny push layer of invariants loaded at session start and a pull layer of knowledge retrieved on demand. The bundle costs zero tokens at baseline.

Step 3 — Record decisions as concepts #

A concept is one atomic unit of knowledge: a Markdown file with YAML frontmatter, addressed by an id like decisions/rate-limiting. Let's record two, the way you would after a design discussion:

cd demo
okf create decisions/rate-limiting knowledge --type Decision \
  --title "Rate limiting strategy" \
  --desc "Standardized on token-bucket rate limiting at the API gateway, 100 req/s per key."
# Created concept 'decisions/rate-limiting.md' (Rate limiting strategy) in 'knowledge'

okf create architecture/cache-layer knowledge --type Concept \
  --title "Redis cache layer" \
  --desc "Redis sits between the API and Postgres for hot read paths; cache keys are versioned by schema revision."
# Created concept 'architecture/cache-layer.md' (Redis cache layer) in 'knowledge'

Concept types are semantic (Decision, Architecture, Fact, Requirement, Bug; default Fact), and create also accepts --body for longer Markdown, --tags, and --status (draft, stable, deprecated). The file it writes is worth reading in full, because the frontmatter is the whole trust model:

cat knowledge/decisions/rate-limiting.md
# ---
# type: Decision
# title: Rate limiting strategy
# description: "Standardized on token-bucket rate limiting at the API gateway, 100 req/s per key."
# generated: { by: agent/cli, at: "2026-09-28T21:36:14Z" }
# status: stable
# ---

Note the generated: provenance stamp — who wrote it and when. This is the lower of the two trust tiers. The higher tier, verified: human:..., is reserved for human-confirmed decisions and (per the project's convention, enforced in the generated AGENTS.md) agents are forbidden from forging it: "NEVER forge human verification (verified: is human-only)". An agent can draft; only a human can stamp. That single rule, sitting in plain YAML where git blame can see it, is more governance than most agent-memory systems ship.

Step 4 — Search before you ask #

The protocol's first rule is search before write: always query memory before authoring, so concepts never duplicate and the agent never hallucinates a divergent copy of a decision. Search is in-memory BM25 — no embeddings, no network:

okf search "rate limiting" knowledge
# Found 1 matching concept(s) in 'knowledge':
#
#  1. [context]    [11.78] decisions/rate-limiting (Decision)
#     Standardized on token-bucket rate limiting at the API gateway, 100 req/s per key.
#     Matches: title, description, id

Each hit carries a BM25 score (11.78), the concept type, and which fields matched. The [context] tag is the governance tier for this concept — the bundle can also mark concepts as constraint (must be obeyed) or hold (subsystem frozen, stop and ask a human), and okf search --for-path pkg/auth/ finds the concepts governing a specific file before an agent edits it.

Search returns summaries. The full concept — frontmatter, body, and links — comes from show, and only on demand. That's progressive disclosure: the agent reads 300-token atoms instead of a 20,000-token monolith:

okf show decisions/rate-limiting knowledge
# ID:          decisions/rate-limiting
# Type:        Decision
# Title:       Rate limiting strategy
# Description: Standardized on token-bucket rate limiting at the API gateway, 100 req/s per key.
# Generated:   2026-09-28T21:36:14Z by agent/cli
#
# --- Body ---

Both commands also speak JSON (--json) for scripting — okf search "redis cache" knowledge --json returns structured records with concept_id, score, matched_on, and link lists, ready to pipe into other tools.

Updating is symmetric with creating:

okf update decisions/rate-limiting knowledge \
  --desc "Standardized on token-bucket rate limiting at the API gateway, 100 req/s per key; burst capacity 200."
# Updated concept 'decisions/rate-limiting.md' in 'knowledge'

Concepts become a graph with relate, which connects two concepts with a semantic description of the relationship. The flag is --desc (required):

okf relate decisions/rate-limiting architecture/cache-layer knowledge \
  --desc "Cache hit rate affects effective gateway throughput under rate limits."
# Linked 'decisions/rate-limiting' -> 'architecture/cache-layer' in 'knowledge'

Now the part no folder-of-notes can do: validate checks OKF v0.2 conformance, graph connectivity, and description drift. Run it strict, the way CI would:

okf validate knowledge --strict --drift
# OKF v0.2 check of "knowledge" (v0.2): 2 concept(s), 0 error(s), 0 gate finding(s),
# 0 warning(s); 0 broken link(s), 0 orphan(s), 0 stale [--strict]. Conformant.

Fully conformant. To see the gates working, try it before linking: with two unconnected concepts, the same command reports gate architecture/cache-layer.md: orphan (no concept links in or out) and concludes "Conformant, but the producer gate failed." Orphan detection, broken-link detection, and stale-review gating (--stale flags concepts past their stale_after date) turn memory hygiene into something a pre-commit hook can enforce.

Two pieces of bookkeeping happened automatically along the way. First, create and relate maintain per-directory index.md files — the progressive-disclosure layer:

cat knowledge/decisions/index.md
# # Decisions
# * [Rate limiting strategy](rate-limiting.md) - Standardized on token-bucket rate
#   limiting at the API gateway, 100 req/s per key; burst capacity 200.

An agent reads the index (titles plus one-line descriptions), then pulls full concepts only for what matters. Second, every mutation appends to knowledge/log.md — a dated changelog of creations, updates, and links. Between the index, the log, and git log, the memory's entire history is auditable without any special tooling.

Conceptual diagram: a knowledge/ folder tree with index.md, log.md, and concept subdirectories of markdown files with YAML frontmatter, linked into a knowledge graph with a magnifier highlighting one node
Image: AI-generated illustration for AI Frontier Post.

Step 6 — Hand it to your agent: the MCP server #

Everything so far works from the terminal. The reason this is trending with agent-tooling people is the next command: okf is also a native MCP server over stdio. No separate process to install, no port to manage — the binary speaks the Model Context Protocol on standard input/output:

okf mcp knowledge

Point Claude Desktop or Cursor at it with the standard MCP server config (from the project's README):

{
  "mcpServers": {
    "okf-memory": {
      "command": "/usr/local/bin/okf",
      "args": ["mcp", "/path/to/your/project/knowledge"]
    }
  }
}

To prove this isn't just a documented claim, I verified the protocol handshake directly — spawning the server and speaking JSON-RPC to it the way any MCP client would. The initialize handshake identifies the server, tools/list exposes the tool surface, and tools/call runs a real search:

# initialize  ->  {"name": "okf-agent-memory", "version": "v0.4.4"}
# tools/list  ->  okf_search, okf_show, okf_validate, okf_create, okf_update, okf_relate
# tools/call okf_search {"query": "rate limiting", "limit": 3}  ->
# [{
#   "concept_id": "decisions/rate-limiting",
#   "title": "Rate limiting strategy",
#   "type": "Decision",
#   "description": "Standardized on token-bucket rate limiting at the API gateway, 100 req/s per key; burst capacity 200.",
#   "governance": "context",
#   "score": 13.17,
#   "matched_on": ["title", "description", "id", "body"],
#   "outbound": ["architecture/cache-layer"]
# }]

Note the outbound field: the search result carries the link we created with relate, so the agent can walk the graph from a hit to its related concepts. Once configured, an agent with the bootstrapped skill follows the loop the AGENTS.md codex mandates: okf_search before proposing changes, okf_show only for relevant hits, okf_update preferred over okf_create when the concept already exists.

Conceptual illustration: a laptop code editor sends a question through a glowing terminal pipe to a small local binary that returns ranked memory cards, while a large unused database server fades in the background
Image: AI-generated illustration for AI Frontier Post.

OKF vs the alternatives #

ApproachStrengthWhere OKF wins
Vector-DB memory (Mem0, Letta)Semantic similarity across large corporaNo embedding API costs or latency, no service to operate, deterministic lexical ranking; agents retrieve operational rules they would never think to semantically search for
Prompt monolith (giant CLAUDE.md)Zero tooling, always loadedBundle costs zero tokens at baseline; progressive disclosure avoids attention drift from 20k-token dumps
Raw markdown notesSimple, git-nativeAdds search, validation gates, provenance tiers, and an agent protocol on top of the same plain files
OKF Agent MemoryGit-native, local, validated, agent-queryable—

The honest limitation: BM25 is lexical. If your agent phrases things with entirely different vocabulary than your concepts used, keyword search can miss where embeddings would connect. The project bets — reasonably, for operational knowledge like decisions, constraints, and runbooks — that shared project vocabulary dominates, and that determinism and zero cost are worth the trade. There is also a hub subcommand for zero-knowledge encrypted sync of the bundle across machines; I did not exercise it here, and it needs a password and secret key, so treat it as a documented-but-unverified extra.

The takeaway #

The reason OKF Agent Memory is climbing isn't a benchmark — it's that the artifact matches the workflow developers already trust. Memory is a folder of Markdown in the repo. It gets committed, diffed, blamed, and reviewed. The binary around it does the three things folders can't: search it in microseconds, validate it in CI, and serve it to agents over MCP. The trust model is one YAML key (generated vs verified), and the governance model is one flag (--for-path before you edit).

Run okf bootstrap on a real project tonight. Record the three decisions your agent keeps forgetting — the ones you re-explain every session. Then ask the agent about one of them in a fresh session and watch it okf_search first instead of guessing. If it answers from your memory instead of its priors, you'll understand the 700 stars.