AI Frontier Post
AI News

LlamaIndex launches OpenDocRouter: every document model under one API, priced at cost

Document parsing is a model-selection problem nobody has time for. On October 7, LlamaIndex launched OpenDocRouter, a hosted API that puts five frontier and five open-source document-parsing models behind a single endpoint — each benchmarked on the company's ParseBench for both quality and cost, so you can see the price-performance tradeoff before you parse a page.

LlamaIndex announced OpenDocRouter on October 7, 2026: a single API for turning PDFs and images into markdown, backed by ten models — five frontier, five open-source — each run under a versioned parsing recipe and benchmarked on ParseBench. The hook is transparency. Instead of picking an OCR vendor on vibes, you get a public scorecard: every model ranked on quality and on cost per thousand pages, served at the provider's own token price with no markup.

The problem it targets is real and getting worse. LlamaIndex says there are now 130+ models on its ParseBench leaderboard, a Hugging Face search for "ocr" surfaces thousands more, and frontier labs ship new document-capable models nearly every month. Evaluating each one means the same grunt work: prompts, rate limits, deployments, integrations, and re-benchmarking every new candidate. OpenDocRouter exists to absorb that work once, for everyone.

Ten models, one endpoint

At launch, the API offers five frontier models — Claude Opus 5.5, Gemini 3 Flash, Gemini 3.8 Flash, GPT-5.6 Terra, and GPT-6 Luna — alongside five open-source ones: Infinity-Parser2-Flash, MinerU2.5-Pro, TeleOCR, dots.mocr, and PaddleOCR-VL-1.6. The company says the launch set was chosen to cover every corner of its benchmarks, and that new models get run through ParseBench — prompts, costs, and settings calibrated — before they're added. Automated routing is on the roadmap.

A chart listing OpenDocRouter's ten launch models: five frontier models and five open-source models, with ParseBench quality and cost figures.
OpenDocRouter's launch lineup: five frontier and five open-source models, all benchmarked on ParseBench. Graphic: AI Frontier Post, data from LlamaIndex.

Benchmarked before you bill

ParseBench reports scores in five categories — Tables, Charts, Faithfulness, Formatting, and Grounding — with Overall as their mean, plus a price per 1,000 typical pages. The spread is the story: Claude Opus 5.5 leads at 84.20 overall (93.53 on tables, 91.03 on faithfulness) but costs $48.82 per 1,000 pages, while GPT-6 Luna scores 71.34 at $0.80, and the open-source MinerU2.5-Pro lands at 70.05 for $0.86. If your documents are mostly simple pages, the frontier model is roughly 60x the cost for a 13-point quality gain — the kind of tradeoff that's hard to evaluate without exactly this scorecard.

Comparison cards showing Claude Opus 5.5 at 84.20 ParseBench overall and $48.82 per 1,000 pages versus GPT-6 Luna at 71.34 and $0.80.
The quality-per-dollar spread: the top ParseBench scorer costs roughly 60x more per 1,000 pages than the cheapest option. Graphic: AI Frontier Post, data from LlamaIndex.

The mechanics

The API is a POST /v1/parse endpoint that accepts PDF, PNG, and JPEG files — or URLs pointing at them — up to 50 MB or 500 pages, with inline base64 data up to about 3 MB. Requests run synchronously up to 50 pages; larger documents go asynchronous. A pages field like "1-3,7" selects specific pages, and every page returns its own markdown, status, and charge. Failed pages can be retried on their own. Typed Python and TypeScript SDKs install via pip and npm, with source published in the run-llama/opendocrouter-py and run-llama/opendocrouter-ts repositories.

Retention is opt-in: nothing from a document is kept unless caching is enabled, cached results are stored encrypted for 24 hours, a delete call removes them sooner, and cache hits for the same pages, model, and layout settings are served free. The service retries each page once on timeouts, rate limits, provider errors, and capacity issues. Account limits start at 10 concurrent requests and 300 parse calls per minute.

Layout as a service

The most interesting piece is the grounding engine. Models differ wildly in what they guarantee about bounding boxes and layout — some output boxes natively, some need prompting, some can't do it at all. Setting layout: true gives any model the same treatment: markdown with grounded bounding boxes and layout elements in reading order, under a shared set of classes (title, section_header, text, list_item, table, picture, chart, formula, caption, footnote, and more), with coordinates as fractions of the page and confidence values attached. A page whose layout fails keeps its markdown without a layout charge — and layout adds $0.20 per million tokens only on pages where it succeeds.

At-cost pricing, and what it isn't

OpenDocRouter bills purely per token. Accounts top up prepaid credits starting at $25 with a 5% fee on each top-up; frontier models are charged at their providers' token prices — Claude Opus 5.5 at $4.00 per million input and $20.00 per million output tokens, GPT-6 Luna at $0.10 and $0.50 — with no markup. Failed, cached, and blank pages are free.

LlamaIndex is explicit about where this sits next to LlamaParse, its managed document platform. OpenDocRouter is the fast lane: the latest models, quickly hosted, switchable behind one API, pay only for what you use. LlamaParse keeps the enterprise tiering — hand-tuned parsing, enterprise controls, self-hosted deployments, plus schema extraction and indexing APIs. Whether the "at-cost" router pulls customers up-market or cannibalizes the platform's margins is the question to watch.

Sources: LlamaIndex announcement (Oct. 7, 2026); LlamaIndex blog — "Introducing OpenDocRouter: every document model under one API"; Unite.AI.