Give your AI agent a real memory: hands-on with Hindsight, the local-first memory layer
Every agent demo remembers everything until you restart it. Hindsight — 42,000+ stars on GitHub, MIT licensed — is a persistent memory layer with three verbs: retain, recall, reflect. We ran the whole thing locally on Ollama with a 0.6B model on a 2-core CPU box, and measured every step.

Agents without memory are goldfish with tools. Every serious agent deployment eventually hits the same wall: the context window is not a memory. It forgets between sessions, it cannot distinguish a durable fact from a passing mention, and stuffing history into the prompt is how you burn money and still miss the thing the user told you three weeks ago.
Hindsight (vectorize-io/hindsight, MIT license, ~42,500 stars as of late September 2026) attacks this directly. It is a local-first memory service for agents with a deliberately small API surface — three verbs: retain (extract and store facts from text), recall (hybrid retrieval over those facts), and reflect (an agentic loop that reasons over memories to answer questions). Under the hood: PostgreSQL with pgvector-style ANN search, a cross-encoder reranker, entity resolution, and background consolidation that merges and deduplicates memories over time. It speaks MCP, so Claude Code, Cursor, and friends can plug straight in.
In this tutorial you will run Hindsight end to end on your own machine: install, point it at local Ollama models, boot the server, create a memory bank, retain facts, recall them with scored hybrid search, and run the reflect loop. Every command and output below was executed and verified on a 2-core CPU VM — including the two places where the setup fought back.
What you'll need#
- Python 3.10+ and Ollama installed locally.
- About 1 GB of disk for two small models:
qwen3:0.6b(~522 MB) for the LLM calls andnomic-embed-text(~274 MB) for embeddings. No GPU required; everything below ran on CPU. - About 30 minutes the first time (mostly model downloads). No API keys — the entire stack is local.
1. Install the server#
Hindsight ships as the hindsight-api-slim package; the [local-ml] extra pulls in the local reranker and ML dependencies (sentence-transformers and friends). On a CPU-only box, install a CPU torch first so pip does not reach for CUDA wheels:
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install "hindsight-api-slim[local-ml]"
ollama pull qwen3:0.6b
ollama pull nomic-embed-text
We verified version 0.10.1 of the package. The install is the heavy step — torch plus the ML extras took the bulk of our setup time. After that, everything is configuration.
2. Boot the server against local Ollama#
The server binary is hindsight-local-mcp. It reads its configuration from environment variables, spins up an embedded PostgreSQL (via pg0 — no separate database to install), loads the cross-encoder reranker (cross-encoder/ms-marco-MiniLM-L-6-v2, downloaded from HuggingFace on first boot), and verifies the Ollama connection:
export HINDSIGHT_API_LLM_PROVIDER=ollama
export HINDSIGHT_API_LLM_MODEL=qwen3:0.6b
export HINDSIGHT_API_LLM_BASE_URL=http://127.0.0.1:11434
export HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai
export HINDSIGHT_API_EMBEDDINGS_OPENAI_BASE_URL=http://127.0.0.1:11434/v1
export HINDSIGHT_API_EMBEDDINGS_OPENAI_MODEL=nomic-embed-text
export HINDSIGHT_API_EMBEDDINGS_OPENAI_API_KEY=x
export HINDSIGHT_API_EMBEDDINGS_OPENAI_DIMENSIONS=768
hindsight-local-mcp
Two notes from our run. First, the embeddings provider is set to openai with a base URL pointing at Ollama's OpenAI-compatible /v1 endpoint — that is the supported way to get local embeddings; the API key can be any placeholder. Second, embedded PostgreSQL refuses to run as root (initdb will not initialize a cluster for the superuser). If you are root in a container, create an unprivileged user and launch from there.
Once it is up, the version endpoint confirms the build and its feature flags:
curl -s http://localhost:8888/version
{"api_version":"0.10.1","features":{"observations":true,"mcp":true,
"worker":true,"bank_config_api":true,...}}
3. Create a memory bank#
Memories live in banks — isolated namespaces, one per agent or project. The Python client is hindsight_client:
from hindsight_client import Hindsight
hs = Hindsight(base_url="http://localhost:8888")
hs.create_bank("agent-memories")
4. Retain: teach it facts#
retain is the write path. You hand it raw text; the server extracts structured facts (what/when/where/who/why), resolves entities, embeds the facts, and stores them. It runs asynchronously through a worker queue, so the call returns immediately with an operation id:
r = hs.retain("agent-memories",
"Priya Nair prefers TypeScript over JavaScript for all new projects.",
retain_async=True)
print(r.operation_id) # d5aea680-...
We retained three memories. All three completed, and the stored facts show the extraction quality — including entity resolution turning a bare mention into a named entity:
Priya Nair prefers TypeScript over JavaScript for all new projects.
| When: 2026-09-29 | Involving: Priya Nair | for all new projects
Marcus recommended Postgres 16's JSONB features for the new analytics API.
| When: Tuesday, September 29, 2026 | Involving: Marcus
The office wifi password is on a sticky note under the router.
| When: 2026-09-29 | Involving: none
Behind the scenes each retain runs fact extraction (LLM, ~40s per memory on our 2-core box with the 0.6B model), entity resolution against Postgres trigram indexes, embedding via Ollama, and then consolidation — a background job that merges and deduplicates. Both consolidation jobs for our bank completed cleanly.

5. Recall: hybrid search with scored results#
recall is the read path: semantic vector search plus keyword search, fused and reranked by the local cross-encoder. Our query returned in 1.3 seconds with per-signal scores exposed on every hit:
r = hs.recall("agent-memories", "What does Priya prefer for new projects?")
for hit in r.results:
print(hit.scores["final"], hit.text[:60])
1.0997 Priya Nair prefers TypeScript over JavaScript for all new projects.
0.0000 Marcus recommended Postgres 16's JSONB features for the new ana...
... The office wifi password is on a sticky note under the router.
The correct memory ranked first with a reranker score of 0.9997 (semantic 0.77, keyword 0.70) while the unrelated memories scored near zero. That score breakdown is genuinely useful in production: when recall misbehaves you can see which signal misled it instead of guessing.
6. Reflect: let the agent reason over its memories#
reflect is the interesting verb — an agentic loop (up to 3 iterations by default) that retrieves memories and reasons over them to answer a question, rather than just returning hits:
r = hs.reflect("agent-memories",
"What do I know about the team's technology preferences?")
print(r.text)
Ours completed in 49.8 seconds. The honest caveat: with a 0.6B local model the reasoning is shallow — it retrieved the right context but hedged instead of synthesizing. The mechanism (retrieval → multi-step reasoning → answer) is verified working; the quality of reflection scales with the model you point it at. For a serious deployment you would run reflect against a 7–30B model and keep the 0.6B model for the cheap extraction path. That split — small model for retain, big model for reflect — is the cost-efficient shape this architecture invites.

Troubleshooting: the proxy gotcha#
The one real fight in our setup is worth documenting because it fails silently. In a sandboxed or corporate environment with HTTP_PROXY/ALL_PROXY set, the embeddings client (the OpenAI SDK pointed at Ollama) honors the proxy for 127.0.0.1 — and the proxy cannot reach your loopback interface. Symptom: retains hang forever after fact extraction succeeds, with the OpenAI client retrying /embeddings indefinitely and zero memories stored. The LLM calls work fine (that code path ignores proxy env), which makes it deeply confusing.
Fix: export a clean NO_PROXY that covers loopback before launching the server — and keep it simple. The stock sandbox value containing IPv6 [::1] breaks httpx's URL parsing, so overwrite it:
export NO_PROXY="localhost,127.0.0.1"
export no_proxy="localhost,127.0.0.1"
Keep the proxy itself enabled, though — the cross-encoder model downloads from HuggingFace at first boot, and direct TLS fails in the same sandboxes. Proxy for the outside world, bypass for loopback.
The takeaway#
Hindsight earns its star count on architecture, not hype: three verbs, local-first, Postgres-backed, MCP-native, with the score transparency and background consolidation that separate a toy from infrastructure. Our full loop — install, boot, 3 retains, hybrid recall at 1.3s, agentic reflect at ~50s — ran on a 2-core CPU box with sub-GB local models and no API keys. The two things to know going in: run it as a non-root user (embedded Postgres insists), and mind the proxy bypass for loopback. Start with the 0.6B model to learn the API for free, then scale the reflect model when the reasoning quality matters.