China Telecom's AI lab just released a model that is tiny by 2026 standards and claims to beat some of the biggest names in the business. The Xingchen AGI Lab — the research arm of China Telecom Artificial Intelligence Technology Co., Ltd. — has open-sourced TeleOCR, a document-parsing model with roughly 1.2 billion parameters that posts new state-of-the-art scores on three international document benchmarks and won the ICDAR 2026 Sci-ImageMiner Challenge. The weights, code and technical paper are all public.

All the performance claims come from China Telecom itself, announced via a press release dated September 30 and corroborated by the team's arXiv paper. They have not been independently reproduced yet — but the numbers are specific, and the repos are live, so anyone can go check.

The numbers#

TeleOCR's headline result is on OmniDocBench v1.6, described as the field's most comprehensive document-parsing benchmark: an overall score of 96.87 out of 100 across 10 document types, 11 layouts and five languages — the highest of all evaluated models, per the company. It ranked first in text recognition, table reconstruction and reading-order restoration.

The other results are in the same vein:

  • Wild-OmniDocBench v1.5 (camera-captured documents): 88.53 overall — about a point ahead of the runner-up and 4 to 10 points ahead of most end-to-end models.
  • PureDocBench: an average of 78.41 across three tracks, with a four-point lead over the second-place model on the most challenging "real degradation" track.
  • ICDAR 2026 Sci-ImageMiner Challenge: first place in the scientific-figure-to-table task, with a TEDS score more than two percentage points above the runner-up's.

Most pointedly, China Telecom says TeleOCR outperformed not just specialized document models like MinerU 2.5-Pro and PaddleOCR-VL-1.6, but general-purpose giants including Gemini 3 Pro and GPT-5.2 on document parsing tasks — at roughly 1.2 billion parameters versus models dozens of times its size. "TeleOCR proves that precision engineering and targeted training can trump sheer scale," a Xingchen AGI Lab spokesperson said in the release. "We built a 1.2-billion-parameter model that beats models dozens of times its size in document parsing, and we are sharing it with the global community."

Editorial illustration: a smartphone scanning a wrinkled paper contract, with glowing detection lines tracing text and a table on the photographed page
Illustration: AI Frontier Post.

How it works#

The core problem TeleOCR attacks is one that anyone who has worked with document AI will recognize: existing systems tend to excel at either clean digital documents or camera-captured ones, but not both. Pipeline-based approaches handle clean PDFs well but choke on the geometric distortion of a photographed page; end-to-end models are more robust to distortion but stumble on high-resolution structured content like tables and formulas.

The team describes three innovations fused into a single vision-language framework:

  • Deformation-aware learning embeds geometric perception directly into the model, eliminating the need for external dewarping modules — so a photo of a wrinkled contract or a tilted whiteboard can be parsed accurately in one pass, according to the release.
  • Content-structure decoupling separates structural reasoning from content generation, enabling precise reconstruction of tables, formulas and scientific charts.
  • Multi-model consensus voting generates high-quality training labels by aggregating predictions from heterogeneous models, avoiding the systematic bias of single-model labeling.

The model handles eight document-parsing tasks: digital layout detection, camera-captured layout segmentation, text recognition, formula recognition (LaTeX output), table recognition (OTSL/HTML output), code block recognition, scientific figure analysis and seal recognition.

Open and available#

Unlike many corporate model announcements, this one ships with the artifacts. The code is on GitHub, the weights are on Hugging Face, and the technical paper is on arXiv — the paper's abstract matches the press release's scores point for point (96.87, 88.53 and 78.41). A production API is also offered on China Telecom's Tianyi AI Open Platform.

Editorial illustration: a scientific bar chart and a table being reconstructed by geometric lines into clean digital data blocks
Illustration: AI Frontier Post.

Why this matters#

Document parsing is one of those unglamorous bottlenecks that quietly determines how far enterprise AI can go. Converting contracts, invoices, medical records and research papers into structured, machine-readable data is the precondition for almost every back-office automation pitch — and the "photographed document" case is the one that breaks most systems in production.

A 1.2B open-weights model that handles both cases in a single pass, and does it better than frontier generalists, is genuinely interesting to anyone building on this layer. China Telecom says the next phase targets financial document processing, medical record digitization, academic research workflows and government archives. The caveats stand: company-reported benchmark claims, independently unverified for now. But the repos are public and the paper is out — the verification can start today.