Stop feeding your agent the whole repo: a hands-on tutorial for code-review-graph
Your AI coding agent burns most of its context budget doing one thing: reading code. code-review-graph flips that cost — one local build of a structural map, then your agent asks questions instead of scanning files. Here is the complete loop, verified command by command.

The most expensive thing your AI coding agent does is read code. Not generate it — read it. Every review, every refactor, every “who calls this function?” starts with the agent scanning files into context, and on a real codebase that scan costs thousands of tokens before a single line is written. You pay for those tokens twice: once in money, once in accuracy, because the more irrelevant code in context, the worse the output gets.
That is the problem code-review-graph attacks, and it is why the project is blowing up right now: 26,588 stars on GitHub, with 6,423 added in a single week per this week’s GitHub trending digest — one of the fastest-moving developer tools in open source this month. The idea is disarmingly simple. Instead of letting your agent grep its way through the repo on every question, you build a persistent local knowledge graph of the codebase once — every function, class, import, call edge, and test linkage — and serve it to any AI coding tool over MCP. The agent asks structural questions (“who calls this?”, “what breaks if I change this?”) and gets back hundreds of tokens of answers instead of whole files.
In this tutorial you will install code-review-graph (version 2.3.9, the current release), build a graph for a real project, query it from the command line, wire it into Claude Code as an MCP server, and run a full change-impact review. Every command below was run and checked on a Linux sandbox before publication — including the places where the tool pushed back.
What you'll need#
- Python 3.10 or later — check with
python3 --version. pip, pipx, or uv; pick whichever you already use. - A git repo to experiment on. Small is fine to start; the tool is built for large ones, but the commands are identical.
- Any AI coding tool for step 4 (Claude Code, Cursor, Codex, Copilot, Windsurf, Zed, Continue, OpenCode, Antigravity, Gemini CLI, Qwen, Kiro — the installer auto-detects all of them). If you have none installed, steps 1–3 and 5 still work from the terminal.
- About 20 minutes. No GPU, no API keys, no accounts. Everything runs on your machine: the project is MIT-licensed, ships with no telemetry, and stores the graph in
.code-review-graph/graph.dbinside your repo.
1. Install it#
The package is a normal PyPI release, so there is nothing to compile and no model weights to download:
pip install code-review-graph==2.3.9 # or: pipx install code-review-graph
# or: uv tool install code-review-graph
Confirm the install and survey the command surface:
code-review-graph --help
You should see the full vocabulary — install, build, update, watch, status, visualize, wiki, detect-changes, query, impact, search, architecture, dead-code, communities, refactor, register, serve, daemon — all local, all offline. That one help screen is already the honest scope of the project: it builds a graph, queries it, and serves it.
2. Build your first graph#
Move into any project directory and run:
cd your-project
code-review-graph build
On a three-file Python demo I ran for this article (a tiny billing module with a test file), the output was:
INFO: FTS index rebuilt: 10 rows indexed
INFO: Loaded 7 unique nodes, 18 edges
Full build: 3 files, 10 nodes, 18 edges (postprocess=full)
Under four seconds. The README states the initial build takes about 10 seconds for a 500-file project — plausible given what I measured, since parsing is tree-sitter AST work, not embedding. After the first build, code-review-graph update re-parses only changed files (my first update after editing fell back to a full rebuild with a “no usable incremental base” notice — expect one rebuild, then increments), and code-review-graph watch keeps the graph live as you save.

3. Ask the graph questions#
This is where the tool earns its name. status gives you the headline numbers:
code-review-graph status
Nodes: 10
Edges: 18
Files: 3
Languages: python
Last updated: 2026-09-29T15:15:31
search finds entities by name or meaning, with a built-in full-text index — no embeddings required:
code-review-graph search invoice
{
"status": "ok",
"query": "invoice",
"search_mode": "fts",
"summary": "Found 2 node(s) matching 'invoice'",
"results": [
{
"name": "create_invoice",
"qualified_name": "api.py::create_invoice",
"kind": "Function",
"file_path": "api.py",
"line_start": 3,
...
Note search_mode: fts — out of the box, search is full-text. True semantic search needs the optional embeddings extra (pip install code-review-graph[embeddings], then code-review-graph embed); the local provider is a small MiniLM model, and cloud providers are opt-in per call. For step-by-step navigation you do not need embeddings at all.
The query command is the precision instrument. It takes one of eight patterns — callers_of, callees_of, imports_of, importers_of, children_of, tests_for, inheritors_of, file_summary — and a target in file.py::function_name form. I learned the target syntax the hard way: billing.Billing.invoice_total returned not_found, while the file-scoped form worked:
code-review-graph query callers_of api.py::create_invoice
{
"status": "ok",
"pattern": "callers_of",
"target": "api.py::create_invoice",
"description": "Find all functions that call a given function",
"summary": "Found 0 call site(s) across 0 node(s) for callers_of('api.py::create_invoice')",
"confidence": "python getattr dispatch and decorator-based registration are not
statically traced, so callers can be missing here"
}
Two things worth noticing in that output. First, the answer is structured JSON your agent can consume directly — this is the “instead of files” payoff. Second, the tool is honest about its limits: static analysis cannot trace dynamic dispatch, and it says so on every result rather than silently omitting callers. A tool that tells you when it might be wrong is a tool you can trust in a review loop.
Two more commands worth running now. architecture shows the project as communities of related code — my demo split into a billing community and an invoice community with cohesion scores, the kind of map that answers “what is this repo actually made of?” in one screen. And dead-code listed every function with no callers or test references — three hits on the demo, all correct.
4. Wire it into your AI coding tool#
The graph becomes powerful when your agent can query it mid-task. One command does the wiring:
code-review-graph install --platform claude-code
Drop the --platform flag and it auto-detects every supported tool on your machine. Running it inside my demo repo produced a full, verifiable setup — and this is the part of the tutorial most worth reading closely, because it shows what “MCP integration” concretely means here:
.mcp.json— a stdio MCP server entry pointing at the installed package:
{
"mcpServers": {
"code-review-graph": {
"command": "/path/to/venv/bin/python3",
"args": ["-m", "code_review_graph", "serve"],
"cwd": "/path/to/your-project",
"type": "stdio"
}
}
}
CLAUDE.mdinstructions — a graph-usage section injected into the project rules, telling the agent to query the graph before scanning files, and to verify conclusions against source (the injected text explicitly says: do not change code from graph output alone).- Four generated skills under
.claude/skills/—debug-issue,explore-codebase,refactor-safely,review-changes. - PostToolUse hooks in
.claude/settings.jsonplus a git pre-commit hook, so the graph updates as you work. .gitignoreentry for.code-review-graph/— the graph database stays local to your machine.
Restart your coding tool after installing so it picks up the new MCP server. From then on, the agent has roughly two dozen graph tools — get_minimal_context_tool (the docs say to call it first: an ultra-compact ~100-token summary), query_graph_tool, get_impact_radius_tool, detect_changes_tool, get_review_context_tool, semantic_search_nodes_tool, get_architecture_overview_tool, refactor_tool, generate_wiki_tool among them — and the injected instructions route it through the graph before it reaches for the file system.
5. Run a real change-impact review#
Here is the end-to-end workflow the project is designed for. Make an edit — I appended a refund method to the demo billing class — then ask for the impact before you commit:
code-review-graph detect-changes --brief
Analyzed 3 changed file(s):
- 7 changed function(s)/class(es)
- 1 affected flow(s)
- 3 test gap(s)
- Overall risk score: 0.57
- Untested: create_invoice, refund, invoice_total
┌─────────────────────── Token Savings ────────────────────────┐
│ Full context would be: 158 tokens │
│ Graph context used: 158 tokens │
│ Saved: 0 tokens (~0%) │
└──────────────────────────────────────────────────────────────┘
Read that output like a reviewer. It tells you what changed, what it touches, where the test gaps are, and an overall risk score — the skeleton of a code review, generated from structure rather than prose. The Token Savings panel is the project’s core economic argument rendered as a number: what the agent would have spent dumping files versus what the graph query cost.
One honest caveat, from my own run: on a three-file demo the panel showed zero savings, because there is nothing to save when the whole repo fits in a tweet. These are the project’s estimates, and the tool only wins once the codebase is large enough that dumping files costs more than a graph query — which is exactly the regime it was built for. The project’s own benchmark diagrams make the same point at scale: reading one corpus whole cost 143,594 tokens versus a 2,196-token graph answer. Treat both numbers as the vendor’s claim, not gospel; the mechanism is sound, and you can measure it yourself with eval, which runs the project’s own benchmark suite against your setup.
For targeted blast-radius analysis, impact takes changed files explicitly:
code-review-graph impact --files billing.py --depth 2
Pair it with the agent: “review my uncommitted changes” now routes through detect_changes_tool and get_review_context_tool, and the MCP responses carry the savings metadata so you can see the cost of every answer.

6. Going further#
Three commands I verified that are worth knowing beyond the core loop. visualize writes an interactive D3.js graph to .code-review-graph/graph.html — open it in a browser to explore the codebase as a literal map. wiki generates a Markdown wiki from the community structure (.code-review-graph/wiki, three pages for my demo). And register / repos enroll multiple repositories in one registry, with daemon running a multi-repo watch process — the answer to “what about my monorepo?”
When to use this vs alternatives#
vs semantic code search (zvec-grep, covered here): complementary, not competing. zvec-grep finds code by meaning — “where do we handle retries?” — while code-review-graph maps structure — “what calls the retry handler?” Run both: search with zvec, impact-analyze with the graph.
vs RTK (the Rust CLI output proxy, covered here): different ends of the token pipe. RTK cuts what the agent outputs and streams; code-review-graph cuts what it reads. They stack — use both if token spend is your metric.
vs built-in repo indexing (Copilot workspace, Cursor’s index): this is free, local, and tool-agnostic — no subscription, no cloud upload of your code, and it works with seventeen platforms from one install. The trade: you own the build/update step, and the index lives in your repo.
When not to use it: tiny repos (my demo proved the overhead exceeds the savings below a few dozen files), heavily dynamic code (the tool warns you that getattr dispatch and decorator registration are not traced — believe it), and codebases where you cannot commit to keeping the graph fresh.
The takeaway#
code-review-graph is not another agent framework or another prompt trick. It is infrastructure: a cheap, local, structural index of your code that every AI coding tool you own can share. The 6,400-stars-in-a-week surge makes sense once you see the economics — agents that ask the graph read less, cost less, and, because they see callers and tests instead of raw files, review better. Install it on your largest repo, run detect-changes --brief on your next PR, and look at the savings panel. If the number convinces you, wire in the MCP server and stop feeding your agent the whole repo.