PageIndex: vectorless, reasoning-based RAG you can run on your own models — hands-on
Tencent's WeKnora went from roughly 26,700 GitHub stars on September 18 to 30,743 by the time I checked on September 28 — about 1,100 new stars a day, enough to top GitHub's weekly trending chart for open source. The reason is simple: it is not another chatbot wrapper. WeKnora is a full knowledge platform — document ingestion with hybrid retrieval, a RAG Q&A layer, an agent runtime, auto-generated wikis, and an MCP server — in one MIT-licensed codebase. I built it from source, wired it to local models, ingested real documents, and got cited answers back. Here is the whole thing, hands-on.

PageIndex hit 37,000 GitHub stars this week by attacking the most boring part of RAG: it throws away the vector database. Instead of chunking your PDFs and embedding them, it builds a hierarchical tree of the document's actual structure — sections, subsections, pages — has an LLM write a summary for every node, and then sends an agent with tools to walk that tree when you ask a question. I ran the whole loop locally with my own model: indexed a five-page report, got a real section tree with summaries, and asked questions that came back with correct, section-cited answers. Here is the whole thing, hands-on.
Why this is trending now#
The numbers when I read the repository on September 29, 2026: 37,158 stars, 3,250 forks, 480 commits, 18 releases, MIT-licensed, from VectifyAI. It was sitting at #3 on GitHub's Python trending page the day I checked. The star count alone does not explain the excitement — the idea does.
Every production RAG system today is a vector pipeline: parse, chunk, embed, store in a vector DB, retrieve top-k chunks, stuff them into a prompt. It works, but it has well-known failure modes — chunk boundaries that split a thought in half, embeddings that match on vibe rather than meaning, and retrieval that cannot see document structure at all. PageIndex's bet is that the structure of the document is the index. A financial report already has sections, headings, and pages; an agent that can navigate that hierarchy like a human reader does not need cosine similarity to find the revenue table.
This is part of a broader 2026 swing away from "embed everything" retrieval. But PageIndex is the most complete open-source implementation of the alternative I have tested: deterministic layout-based tree extraction, LLM-written node summaries, and an agentic chat loop with real document tools. The August 2026 release added the piece that makes it tutorial-worthy: SDK local mode, which runs indexing, retrieval, and chat against your own LLM provider — including a model server on your own machine — with no PageIndex cloud account at all.
What PageIndex actually is#
PageIndex is a Python library (pip install pageindex), not a platform. It has two halves:
- Indexing —
client.submit_document("report.pdf")parses the PDF, extracts a structural tree, and (optionally) has an LLM summarize every node. The default engine is called flash: tree extraction from layout statistics is fully deterministic and uses no LLM at all; the LLM only writes node summaries and runs a merge/expand optimization pass. Astandardmode exists as the alternative engine. Only PDFs are accepted — PDF/A and PDF/X are explicitly not supported. - Chat —
client.chat("question", doc_id=doc_id)runs an agent loop. The agent gets tools (browse_documents,get_document_structure,get_page_content) and walks the tree to answer, instead of receiving a pile of retrieved chunks. You can run it over OpenAI's Responses API, Anthropic's Messages API, or plain Chat Completions — the protocol is declared explicitly, never inferred.

The model routing is the part I want to highlight, because I verified it from the source. A bare model name like "gpt-4o" is routed to openai/gpt-4o; a provider prefix like "anthropic/claude-..." selects the vendor. But the escape hatch is what matters for self-hosting: any OpenAI-compatible endpoint works. Name your model "openai/<model-id>" and pass {"base_url": "http://your-server:port/v1", "api_key": "x"} as the backend, and both indexing and chat run against your server. I read this in pageindex/local_api.py and pageindex/local_chat.py before trusting it — and then proved it with a real server.
Setup: the library and a local model#
This tutorial runs everything locally — no PageIndex account, no API keys to a hosted provider. You need Python 3.9+ and any OpenAI-compatible model server. I used llama-server from llama.cpp serving Qwen2.5-1.5B-Instruct (Q4_K_M) on http://127.0.0.1:8080. A 1.5B model is the floor, not the recommendation: it is slow (more on timings below) and you should use the biggest model your hardware allows. The point I am proving is that the architecture works with a small local model, not that a 1.5B model is ideal.
pip install pageindex
# verify
python -c "import pageindex; print(pageindex.__version__)" # 0.2.20
Start your model server with an OpenAI-compatible endpoint. With llama.cpp:
llama-server -m Qwen2.5-1.5B-Q4_K_M.gguf \
--host 127.0.0.1 --port 8080 -c 4096 --jinja -t 2
Check it answers before going further:
curl http://127.0.0.1:8080/health
# {"status":"ok"}
NO_PROXY, but the stock value in many sandboxes ([::1]) does not cover IPv4 localhost the way httpx expects. Export NO_PROXY="localhost,127.0.0.1" (and lowercase no_proxy) before running anything below. This cost me twenty minutes.Step 1: index a document, get a tree back#
Indexing needs two configurations: which model writes the summaries, and where that model lives. Here is the exact code I ran:
from pageindex import PageIndexClient
BACKEND = {"api_base": "http://127.0.0.1:8080/v1", "api_key": "x"}
client = PageIndexClient(
index={"model": "openai/qwen", "backend": BACKEND},
)
res = client.submit_document("meridian_solar_2025.pdf")
doc_id = res["doc_id"] # e.g. "pi-a6ef5619e01a41d0874b2add66d1de2f"
A note on the model string: "openai/qwen" does not mean OpenAI hosts it. The openai/ prefix selects the OpenAI-protocol driver, and api_base in the backend redirects that driver to your local server. The api_key can be any non-empty string for a local server that does not authenticate.
My test document was a five-page operations report I generated for this tutorial (text-based PDF with real headings and PDF bookmarks). Flash indexing took just under five minutes on a 2-vCPU machine with the 1.5B model — the layout analysis is instant; the time is all LLM summary calls, one per node. Then:
tree = client.get_tree(doc_id, node_summary=True)
def walk(nodes, depth=0):
for n in nodes:
print(" " * depth + f"{n['node_id']} | {n['title']}")
walk(n.get("nodes", []), depth + 1)
walk(tree["result"])
And the output — a genuine document tree, each node carrying a summary written by the local model:
0000 | Preface
0001 | Generation Performance
0002 | Maintenance and Reliability
0003 | Financial Summary
0004 | Risk Factors
For example, node 0001's summary read: "In 2025, the generation fleet achieved a total net generation of 842 gigawatt-hours, exceeding the 792 gigawatt-hour target by 6.2 percent. Plant-level results…" — accurate against the source PDF. Two details worth knowing: get_tree hides summaries unless you pass node_summary=True (I stared at summary-less output for a while before reading the signature), and the Preface node is PageIndex's own convention — when the detected hierarchy starts after page 1, the leading pages get a Preface node, exactly as documented in the flash API.
One more honest observation: flash consumed my PDF's embedded bookmarks as the section frame (its use_embedded_toc behavior), which is why the tree is this clean. On a PDF without bookmarks it falls back to pure layout detection, and on a document with no detectable hierarchy at all it degrades to one node per page. The tree quality follows the document's own structure — garbage in, flat tree out.
Step 2: ask questions — the agent walks the tree#
Chat is configured the same way, with one extra decision: the protocol. protocol="chat_completions" is the plain answer lane (it is also the default); "responses" and "messages" drive the agent natively over OpenAI's Responses API or Anthropic's Messages API, which matters if you want tool-call history to round-trip. For a local server, Chat Completions is the lane:
client = PageIndexClient(
chat={"model": "openai/qwen",
"backend": {"base_url": "http://127.0.0.1:8080/v1",
"api_key": "x"}},
)
answer = client.chat(
"Which plant produced the most electricity in 2025 and how much?",
doc_id=doc_id,
protocol="chat_completions",
)
Two things I learned here. First, protocol is a parameter of the chat() call, not of the client configuration — putting it in the chat={...} dict raises PageIndexAPIError: Unknown chat keys. Second, with protocol="chat_completions" the return value is the full Chat Completions envelope (choices/usage), not a bare string — pull the text out of resp["choices"][0]["message"]["content"]. Omit protocol entirely and you get the bare answer string; the docs describe the no-protocol lane as "thin sugar over the same engine."

The answers, from the 1.5B local model, after the agent walked the tree with its tools:
Q: Which plant produced the most electricity in 2025 and how much?
A: "The 'Generation Performance' section of the document shows that the Copper Flats facility in Arizona generated 391 gigawatt-hours, which surpassed its target by 9.1 percent due to an inverter retrofit completion in March…" — correct, and it cited the section it read.
Q: What happened at Del Sol in June and how many hours did it account for?
A: "…hail damage in June, which caused two strings to be offline for eleven days. Therefore, the total downtime accounted for 96 hours in June." — correct on both facts.
Both answers are grounded in the document text, both name their source section, and both survived my fact-check against the PDF. The agent did not receive any chunks — it navigated the tree with get_document_structure and get_page_content and composed from what it read.
Timings and the small-model caveat#
Full honesty on cost, measured on a 2-vCPU cloud VM with Qwen2.5-1.5B-Q4_K_M on CPU:
| Step | Time | Notes |
|---|---|---|
| Flash indexing (5 pages, 5 nodes) | ~5 min | Layout analysis is instant; all time is LLM summary calls |
| Question 1 (multi-section answer) | ~7 min | Several agent turns, each a full local inference |
| Question 2 (single-section answer) | ~2.5 min | Fewer turns needed |
A GPU, a larger model, or a hosted API would collapse these numbers — the architecture parallelizes summary calls (summary_concurrency is a real parameter) and I was running everything on two CPU cores. But the small-model caveat is real and you should hear it: during an earlier run, the same 1.5B model wrote a node summary that said "2024" where the document said "2025". The tree and the retrieved text are deterministic and exact; the summaries and answers are model-generated and inherit whatever model you plug in. Size your model to your accuracy bar. Nothing about PageIndex fixes a weak model — it just gives a good model better scaffolding to climb.
Limitations worth knowing before you commit#
- PDF-only. The local pipeline accepts PDF files and nothing else. Scanned/image PDFs need OCR upstream — flash will tell you the tree is unreadable.
- Structure-dependent. The tree is only as good as the document's own hierarchy. Well-structured reports shine; a 200-page dump of unstructured text degrades toward one-node-per-page.
- Agentic chat needs tool-calling. The chat loop is a real agent — your model must reliably emit tool calls. A 1.5B model managed it here, but barely; this is the strongest argument for using a capable model for chat even if you index with a small one (the library supports different models per slot).
- The cloud path exists too. If local operation is not your goal, the managed API (
PageIndexClient(api_key=...)) has a free tier with trial credits and is the documented quick start. I verified the local path because "runs on your own models" is the interesting claim.
The takeaway#
PageIndex earned its trending spot with a genuinely different answer to document QA: no embeddings, no vector store, no chunking heuristics. Parse the document into the tree it already is, summarize every node, and let an agent walk it. I ran that entire loop on my own hardware with a 1.5B local model — indexing produced a clean five-node tree with accurate summaries, and both test questions came back correct with section citations.
It is not magic and it is not model-independent: the summaries and answers are only as good as the LLM you attach, and the tree is only as good as your document's structure. But as scaffolding, it is the most convincing alternative to chunk-and-embed RAG I have run hands-on. If you have a folder of structured documents people keep asking questions about — reports, manuals, filings — this is worth an afternoon. Start local, prove the tree quality on your own PDFs, and only then decide whether you need the managed cloud.
Links: VectifyAI/PageIndex on GitHub · docs.pageindex.ai · I verified this tutorial against pageindex 0.2.20 (PyPI) and the repository state of September 29, 2026, running SDK local mode against a local llama.cpp server (Qwen2.5-1.5B-Q4_K_M).