Your agent forgets everything between sessions — fix it with mem0, hands-on
mem0 is the open-source memory layer for AI agents — 66,000+ stars, Apache 2.0. We wired it into a chatbot with real API calls: store, search, and retrieve memories per user. Every command verified.

Every conversational AI you have ever used has the memory of a goldfish: close the tab and it forgets who you are, what you were working on, and what it promised you. That amnesia isn't a quirk of one model — it's architectural. LLM context windows are temporary, and when the conversation ends, everything evaporates. The fix is a memory layer: a piece of infrastructure that extracts facts from conversations, stores them, and retrieves the relevant ones the next time you talk. Mem0 (mem0ai/mem0) is the open-source project that made this idea boring in the best way — and on September 30 it was being actively committed, held 66,000+ GitHub stars, and shipped a new single-pass memory algorithm (LoCoMo score up from 71.4 to 92.5, LongMemEval from 67.8 to 94.4) the README publishes proudly.
In this tutorial you'll install the Mem0 Python library, teach a chatbot three facts about a user with one API call, retrieve them later by semantic search with another, wire both into a persistent chat loop, and isolate memories per user. Every command below is taken from the project's README and verified against the actual library. It takes about thirty minutes.
1. What mem0 actually is#
Mem0 isn't a database you dump strings into. It's a memory extraction and retrieval pipeline: you hand it a conversation, an LLM distills that conversation into discrete facts ("Alex prefers dark mode", "Alex is allergic to peanuts"), each fact is embedded with text-embedding-3-small, and stored with metadata like user_id. Later, a query is embedded the same way and the most relevant memories come back by hybrid semantic search — so "what can't Alex eat?" finds the peanut allergy even if you ask about restaurant orders.
Three numbers worth knowing from the README. The default LLM is OpenAI's gpt-5-mini — memory extraction needs a model to reason about the conversation, so this is not a $0 tool like a local embedding cache. The new (April 2026) algorithm retrieves with single-pass retrieval: one call, no agentic loops, at a top_200 retrieval budget, p50 latency around one second. And the fine print that matters: those headline benchmark numbers (92.5 LoCoMo, 94.4 LongMemEval) were measured on Mem0's managed platform, which includes proprietary optimizations not in the open-source SDK — open-source users should expect "directionally similar gains", the README says, not identical numbers.
2. What you'll need#
- Python 3.10+ and
pip. - An OpenAI API key set as
OPENAI_API_KEYin your environment. Mem0's defaults — the extraction LLM (gpt-5-mini) and the embedding model (text-embedding-3-small) — both live there. (The library supports other LLMs and embedders; the README links a full supported-components list, but the defaults are what this guide runs.) - About 30 minutes. Each memory operation calls an LLM, so this is a paid-API tutorial — budget a few cents, not dollars.
No Docker, no accounts, no dashboard for the library path.
3. Step 1 — Install#
The README's install is one line:
pip install mem0ai
If you want enhanced hybrid search with BM25 keyword matching and entity extraction, the README offers the NLP variant — it needs a spaCy model too:
pip install mem0ai[nlp]
python -m spacy download en_core_web_sm
The plain install is enough for this tutorial. Verify it imports:
python -c "from mem0 import Memory; print('mem0 ready')"
Honest note: the first Memory() call needs that OpenAI key in the environment — without it you'll get an authentication error, not a helpful one. export OPENAI_API_KEY=sk-… before anything else.
4. Step 2 — Store your first memories#
Instantiation is one line — the README uses no config for the default stack:
from mem0 import Memory
memory = Memory()
Now teach it facts. add takes a list of role/content messages and a user_id — the LLM extracts the durable facts itself, you don't hand it a database row:
messages = [
{"role": "user", "content": "I'm Alex. I prefer dark mode in every app."},
{"role": "assistant", "content": "Noted — dark mode everywhere."},
{"role": "user", "content": "Also, I'm allergic to peanuts. Please remember that."},
{"role": "assistant", "content": "Got it — peanut allergy noted."},
]
result = memory.add(messages, user_id="alex")
print(result)
Mem0 returns a structured result of what it stored. The point to internalize: you gave it conversation, and it stored facts — "Alex prefers dark mode" and "Alex is allergic to peanuts" as separate retrievable memories. That extraction step is the whole product. It also handles updates and dedup itself: tell it "I moved to Chicago", then "actually, I'm back in Toronto", and the memory gets reconciled rather than duplicated.

5. Step 3 — Retrieve them by meaning#
The other half of the loop. search takes a natural-language query, a filters dict scoping results to one user, and top_k:
relevant = memory.search(
query="What can't Alex eat?",
filters={"user_id": "alex"},
top_k=3,
)
for entry in relevant["results"]:
print("-", entry["memory"])
Ask about food and get back the peanut allergy — even though the stored fact says nothing about food. That's semantic retrieval earning its keep. The response is a dict with a results list; each entry carries the memory text plus metadata (creation time, categories). The README's production numbers put this whole path at roughly one second p50 — you're paying an LLM round-trip for the extraction, but retrieval itself is a single vector lookup pass.
6. Step 4 — Wire it into a chat loop#
This is the README's canonical pattern, and it's the design worth stealing: every user turn searches first (inject memories into the system prompt) and adds after (persist anything new from the exchange):
from openai import OpenAI
from mem0 import Memory
openai_client = OpenAI()
memory = Memory()
def chat_with_memories(message: str, user_id: str = "default_user") -> str:
# Retrieve relevant memories
relevant_memories = memory.search(query=message, filters={"user_id": user_id}, top_k=3)
memories_str = "
".join(f"- {entry['memory']}" for entry in relevant_memories["results"])
# Generate Assistant response
system_prompt = f"You are a helpful AI. Answer the question based on query and memories.
User Memories:
{memories_str}"
messages = [{"role": "system", "content": system_prompt}, {"role": "user", "content": message}]
response = openai_client.chat.completions.create(model="gpt-5-mini", messages=messages)
assistant_response = response.choices[0].message.content
# Create new memories from the conversation
messages.append({"role": "assistant", "content": assistant_response})
memory.add(messages, user_id=user_id)
return assistant_response
Drive it with a tiny REPL — the README's main() loop reads stdin until you type exit. Then the test that proves the whole point: in a fresh Python process tomorrow, ask "should I order the pad thai?" and the assistant remembers the peanut allergy — because the memory lives in Mem0's store, not in anyone's context window. No context-window stuffing, no hand-rolled vector DB plumbing, no "summarize our chat so far" hacks.

7. Step 5 — Per-user isolation, and where to run it in production#
Notice every call took user_id="alex". That filter is Mem0's isolation primitive: memories are namespaced per user (or per agent, per app — any string), and a search filtered to one user can never surface another's. For a multi-tenant chatbot this is the security boundary, so make the id your authenticated user id, not a nickname, and don't let clients set it themselves.
The README lays out three deployment paths, and which to pick is honest guidance worth repeating:
- Library (
pip install mem0ai) — testing and prototyping. This is what you just built; the vector store runs locally by default. - Self-hosted server —
cd server && docker compose up -din the repo for your own infrastructure, with a browser wizard atlocalhost:3000and auth on by default. - Mem0 Platform — the managed cloud (zero ops, dashboard, the headline benchmark numbers).
There's also a CLI for poking at memories from the terminal — mem0 init, then mem0 add "Prefers dark mode" --user-id alice and mem0 search "What does Alice prefer?" --user-id alice. Nice for debugging what the extraction actually stored, without writing a script.
8. Honest limitations#
- It needs an LLM, so it costs money and makes mistakes. Memory extraction is one model judging what matters in a conversation. It can misremember, over-extract trivia, or miss the thing you actually cared about. Treat retrieved memories as hints the model reasons about, not ground truth — never feed them directly into safety-critical decisions.
- The headline benchmarks are platform numbers. The README says explicitly: LoCoMo 92.5 / LongMemEval 94.4 reflect Mem0's managed platform with proprietary optimizations; the open-source SDK you just installed gets "directionally similar" results, not identical ones.
- Your conversations flow through your configured providers. Every
addandsearchround-trips an LLM and an embedding model — with the defaults that's OpenAI, i.e. your users' chat content leaves your network. Self-hosting the LLM/embedder stack changes the privacy story, at the cost of the defaults' quality. - Latency is real. Roughly a second p50 per the README's own numbers — fine for a chat turn, wrong for a hot inference loop. If you need memory consulted inside a per-token decision path, this isn't it.
The takeaway#
The stat that should end the goldfish era: a chatbot built on this pattern can answer "what can't Alex eat?" in a fresh session, because "allergic to peanuts" was extracted, embedded, and retrieved as a fact — not pasted into a prompt. That's the entire architecture: add distills conversation into facts, search brings back the relevant ones, and a user_id filter keeps everyone's memories separate. Two API calls, both from the README, and your agent stops having amnesia.
Start where this tutorial did: pip install mem0ai, three facts about one user, one semantic search proving they come back by meaning. Then try the moves that matter in production — per-user ids you control server-side, the self-hosted Docker stack for your own data, and a skeptical eye on what the extractor chooses to remember. That's the whole idea: context that persists.
Sources: the mem0ai/mem0 repository README (verified September 30, 2026: 66,389 stars, Apache 2.0, Quickstart Guide, Basic Usage chat_with_memories pattern, CLI reference, self-hosted and platform deployment notes), and the Mem0 paper arXiv:2504.19413. Every API call in this guide (Memory(), add, search with filters/top_k) comes from the README; benchmark figures are the README's own managed-platform numbers, labeled as such.