Ask an AI agent to explain a math paper and it will confidently invent the proof structure. It will tell you Theorem 3 depends on Lemma 2 when Lemma 2 does not exist, cite equations by numbers it made up, and summarize a proof it never located in the source. The failure is not reasoning — it is evidence. The agent read a PDF as flat text and guessed at the scaffolding.

PaperGraph MCP (lotchuazzz-crypto/papergraph-mcp) is the open-source project attacking exactly that failure, and it is climbing fast: 275 GitHub stars as of September 30, 2026, up from about 100 in early September — a repository created on September 2 that has shipped 20 releases in under four weeks, with commits landing as recently as today. It is an MIT-licensed MCP server, written in Python, that turns arXiv papers, local LaTeX projects, and born-digital PDFs into a local, theorem-centered workspace: extracted results, proof evidence with source spans, explicit dependency edges, and deterministic Markdown reading reports. No model weights, no API key, no account, no cloud. In this tutorial you will install it, wire it into an MCP client, load a real arXiv paper through the full pipeline, trace a theorem's proof evidence, and export a reading report — every command verified by running it.

1. What you'll need#

  • uv — the install path the README documents. curl -LsSf https://astral.sh/uv/install.sh | sh gets it on macOS and Linux; Windows users take the PowerShell installer from the uv docs.
  • Python 3.10+ — uv manages the interpreter itself, so any system Python works as the launcher.
  • An MCP-capable client (optional) — Claude Code, Claude Desktop, or anything that accepts a JSON stdio server config. If you do not have one, every step below also works from the CLI, which we demonstrate throughout.
  • Internet access to arXiv — PaperGraph downloads paper sources only from arXiv's fixed e-print endpoint; it never scrapes arbitrary URLs or bypasses paywalls.
  • No API key, no account, no GPU, no cost. Extraction is deterministic LaTeX/PDF parsing into a local SQLite workspace — the intelligence comes from your agent, not from PaperGraph calling a model.

2. Step 1 — Install the pinned release#

Pin the release tag so your install is reproducible. The current stable release is v1.1.7 (a maintenance cleanup; v1.1.6 fixed source re-import data loss):

uvx --from git+https://github.com/lotchuazzz-crypto/[email protected] papergraph-mcp --version
papergraph-mcp 1.1.7

This installs without cloning the repo. Now check the environment diagnostics:

uvx --from git+https://github.com/lotchuazzz-crypto/[email protected] papergraph-mcp doctor
{
  "dependency_extraction_basis": "statement_explicit_latex_refs_only",
  "package_name": "papergraph-mcp",
  "recommended_source": "git+https://github.com/lotchuazzz-crypto/[email protected]",
  "release_tag": "v1.1.7",
  "version": "1.1.7"
}

That first field matters more than it looks: statement_explicit_latex_refs_only is PaperGraph's evidence contract. Dependency edges are built only from explicit LaTeX references (\ref, \eqref, \autoref, \cref, \Cref) inside theorem-like statements. An empty dependency result means "no resolvable label references found under this rule" — not "this theorem has no mathematical dependencies." The tool refuses to guess, and it says so in its own diagnostics.

3. Step 2 — Validate before you load#

PaperGraph separates validating a paper request from loading it. The validator accepts bare IDs, URLs, Markdown links, and prose descriptions, and it refuses to proceed on ambiguity:

papergraph-mcp validate-arxiv-request "math/0307200"
{
  "action": "safe_to_load",
  "selected_id": "math/0307200",
  "status": "single_input",
  "message": "The request identifies one paper. It is safe to load math/0307200."
}

If validation returns action: ask_user_to_choose, you must pick a candidate — detecting a conflict and plowing ahead anyway is, per the docs, a failure. This is the first place PaperGraph's evidence-first discipline shows up in the UX: disambiguation is the caller's job, not something the tool silently resolves.

4. Step 3 — Wire it into your MCP client#

Add PaperGraph to any MCP client that accepts JSON-style stdio configuration (Claude Code, Claude Desktop, and most others):

{
  "mcpServers": {
    "papergraph": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/lotchuazzz-crypto/[email protected]", "papergraph-mcp"]
    }
  }
}

Restart the client after changing its configuration. The server speaks stdio, so running the bare command in a terminal just waits quietly for a client connection — that is normal.

We verified the server end to end over a real MCP stdio session: tools/list returns 63 tools, spanning the full workflow surface — paper loading, paper maps, result and proof inspection, reading sessions and queues, reference resolution, and report export. The core verbs you will use most:

  • open_workspace — open or initialize a persistent multi-paper workspace (a plain local SQLite file you choose).
  • workspace_add_arxiv_paper — download and parse an arXiv LaTeX project into the workspace.
  • workspace_get_paper_map — the evidence-first overview: main-result candidates, structure, reading route, external risks.
  • workspace_list_results / workspace_get_result — the extracted theorem-like results with source spans.
  • workspace_get_result_proof / workspace_get_proof_dependencies — proof evidence and explicit dependency edges.
  • workspace_export_paper_reading_report — a deterministic Markdown report you can commit to Git or hand to another session.
Four-stage workflow diagram: arXiv paper document, LaTeX parsing stage, local database, and verified reading report
The pipeline this tutorial walks through: arXiv source → LaTeX parse → local SQLite workspace → deterministic reading report. Every arrow is a real command you will run.

5. Step 4 — Load a real paper#

Open a workspace — a plain SQLite file you choose the location of — and add a paper. We use a paper published yesterday (September 29, 2026): Chao Ma's Lazard's Realization Problem for N-Series: Bar Obstructions and Finite-Quotient Descent (arXiv 2609.38002), a group-theory paper written in standard LaTeX with numbered theorems, lemmas, and proof environments:

// MCP: open_workspace
{"path": "/path/to/papers.sqlite3"}
// MCP: workspace_add_arxiv_paper
{"arxiv_id": "2609.38002"}

The add call returns the stored paper record. Ours came back as paper_id: "arxiv:2609.38002" with the title extracted, 26 theorem-like results, 23 proofs, and 33 citations detected. The whole import — download, LaTeX parse, extraction — took seconds on a 2-core VM. Nothing was sent to any model; the workspace file is yours to keep, diff, and commit.

6. Step 5 — Read the paper map#

workspace_get_paper_map is the first thing to call on any loaded paper. It returns the evidence-first overview: main-result candidates, structure, a candidate reading route, external risks, and an evidence-quality verdict. For our paper:

// MCP: workspace_get_paper_map
{"paper_id": "arxiv:2609.38002"}
  • 5 main-result candidates, scored — the top three (score 3) are all corollaries; the two score-1 entries are a proposition and a lemma.
  • Recommended starting point: arxiv:2609.38002::corollary-1.2-canonical-minimality — with an explicit caution attached: "This is not a claim that the result is mathematically central." The map ranks by evidence structure, not by importance, and it says so.
  • A 3-item reading route: start at Corollary 1.2, then its proof (priority: required), then stop at the corollary's unresolved citation mentions — the map tells you exactly where the evidence runs out.
  • Evidence status: usable — with warnings listing the three theorems that have no proof evidence (2.1, 2.3, 2.5) rather than silently skipping them.

7. Step 6 — Trace proof evidence and dependencies#

This is the feature that justifies the whole project. Ask for the proof of Proposition 1.1:

// MCP: workspace_get_result_proof
{"result_id": "arxiv:2609.38002::proposition-1.1-equivalence-with-lazards-dimension-equality"}

You get the proof text with its provenance: association_basis: "immediately_follows_result" (confidence 0.8 — the proof was matched because the \begin{proof} environment directly follows the proposition), method: "latex_proof_environment", confidence 1.0, and source spans down to byte offsets in the .tex file (8264–8694). If the association were weaker, the basis and confidence would say so — the uncertainty is data, not hidden.

Now the dependency graph. Lemma 2.2's proof contains the line Theorem~\ref{theorem-2.1-loseys-abelian-injectivity-theorem} gives ... — an explicit LaTeX cross-reference. Query its dependencies:

// MCP: workspace_get_proof_dependencies
{"result_id": "arxiv:2609.38002::lemma-2.2-kernel-transgression"}
{
  "result_id": "arxiv:2609.38002::lemma-2.2-kernel-transgression",
  "known": {
    "resolved_local_results": [
      {
        "result_id": "arxiv:2609.38002::theorem-2.1-loseys-abelian-injectivity-theorem",
        "display_kind": "Theorem",
        "method": "latex_environment",
        "confidence": 1.0
      }
    ]
  }
}

Lemma 2.2 depends on Theorem 2.1 — resolved with confidence 1.0, because the LaTeX says so. No embedding similarity, no LLM inference, no vibes. The dependency graph below shows the real extraction: the eight results of the paper's first two sections, with the single verified \ref edge drawn solid and the recommended reading-route start dashed. Everything else is deliberately unconnected — PaperGraph found no explicit-reference evidence there, and drawing more would be fabrication.

Diagram of the real PaperGraph extraction for arXiv 2609.38002: eight theorem-like results with the single verified explicit-reference dependency edge from Lemma 2.2 to Theorem 2.1
The verified extraction for arXiv 2609.38002: real result labels, one explicit-\ref edge (Lemma 2.2 → Theorem 2.1, confidence 1.0), and the reading-route start. Unconnected nodes have no explicit-reference evidence — and the tool does not invent any.

8. Step 7 — Export the reading report#

The payoff for agent workflows is a deterministic Markdown report — same paper, same report, every time:

// MCP: workspace_export_paper_reading_report
{"paper_id": "arxiv:2609.38002"}

The export (about 15 KB for this paper) is structured for handoff: Paper Map (starting point, evidence status, counts), Evidence Triage (what is safe to use, what needs review, next actions), Main-Result Candidates, Candidate Reading Route, Supported Local Logic Chain, External Reading Risks, Evidence Quality, Evidence Boundaries, and Next Commands. The triage section is the honest core: 1 supported local chain, 41 external references flagged as blocked because PaperGraph will not auto-import external papers without user-supplied identifiers. An agent receiving this report knows exactly what it can trust and where it must stop.

9. Honest limits — what it refuses to do#

A tool is defined by its refusals, so we stress-tested one. Loading Baez and Lauda's classic Higher-Dimensional Algebra V: 2-Groups (math/0307200) extracted 7 theorems at confidence 1.0 — but the paper contains zero \begin{proof} environments (we checked the source: its proofs are inline prose). PaperGraph's response to workspace_get_result_proof was:

{"known": {}, "inferred": [], "unresolved": {"proof": "not_found"},
 "warnings": ["No proof evidence was found for this result."]}

Not a guessed proof. Not a "proof sketch assembled from nearby paragraphs." A machine-readable not_found. That is the contract working as designed — and it tells you something practical: PaperGraph shines on papers written in standard amsthm style and degrades gracefully, loudly, on everything else.

Three more limits worth knowing before you adopt it:

  • arXiv only, for downloads. Paper fetching goes through arXiv's fixed e-print endpoint. Local LaTeX projects and born-digital PDFs load from disk, but there is no generic web fetching and no paywall bypassing.
  • External references stay external. Citations to other papers are detected (33 in our test paper) but not auto-imported — resolving them is a separate, user-driven step. Cross-paper dependency chains require you to load each paper.
  • Parsing is LaTeX-literal. Author names came back with raw LaTeX commands still attached (Chao Ma\thanks{...}); exotic macro packages can confuse the extractor. The 52 unresolved risks on the 2-groups paper are the tool showing its work, not hiding it.

10. When to reach for it#

Reach for PaperGraph when an agent needs to work a math paper, not just summarize it: literature-review agents that must cite the exact theorem a claim rests on, proof-checking assistants that need the dependency chain before the statement, and reading-plan tools that should start a newcomer at the right corollary. It is the wrong tool when you want a quick natural-language summary of a paper — a chat model does that faster — and it is not a general PDF reader; its evidence model is built for LaTeX-structured mathematics.

The deeper point is architectural. The current generation of research agents fails on papers the way early code agents failed on repositories: by operating on flat text and hallucinating structure. PaperGraph is the equivalent of giving the agent a parsed AST with source spans instead of a raw file dump — 63 tools over a local SQLite workspace, deterministic reports, and a dependency rule (statement_explicit_latex_refs_only) that would rather say not_found than guess. At 275 stars in four weeks, the open-source community seems to have decided that evidence-first is the right layer to build on. We ran every command above against v1.1.7 and the outputs are exactly as shown — the workspace files are yours to reproduce.