The Allen Institute for AI announced today that it is open-sourcing AstaBrief 8B — the report-generation model behind Fast mode in Asta, its agentic platform for scientific work. Alongside the weights, Ai2 is releasing the training data, an example workflow in its ai2-scholarqa-lib GitHub repository for generating reports from your own PDFs, and a long technical write-up of how the model was built.

AstaBrief takes a research question plus retrieved literature excerpts and writes a cited report. The experiment’s goal, in Ai2’s telling: find out whether a small open model trained specifically for scientific report generation could match the report quality of the proprietary models it had been using, while cutting generation time and serving cost.

Fast mode, live today#

AstaBrief is available from today in Asta’s “Generate a report” feature as Fast mode, sitting alongside the existing Claude-powered Thinking mode. The headline number: across the full Asta pipeline, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode — about 3.5× faster, and what Ai2 describes as “nearly an order-of-magnitude reduction in generation time” relative to the proprietary models it tracked.

The early usage data is worth a look: of 374 Asta users who tried Fast mode, 29.1% used it on two or more days, and 23% never switched back to Thinking mode for later report threads — another 18% alternated, using Fast mode for roughly 40% of their threads. Positive feedback ran at 84.2% versus 85.2% for Thinking mode, a similar rate, though Ai2 cautions the feedback is too sparse to support strong conclusions.

Built for citations, not just prose#

AstaBrief starts from Qwen3-8B and is licensed under Apache 2.0, with most of the effort poured into post-training data rather than a fancier optimizer. Ai2 considered reinforcement-learning-based training — its earlier DR Tulu work had shown RL helps long-form report generation — but judged RL too unstable and expensive, and bet it could get there with a cheaper recipe: supervised fine-tuning followed by direct preference optimization (DPO).

The data pipeline is the real story. Training began with 90,000 real research queries from scientists using the ScholarQA framework, after stripping beta-tester and bot traffic, queries too short to be meaningful, and an LLM-filtering pass that caught non-English queries, non-scientific requests, and prompts containing personal information. Full-report SFT targets were generated through the multi-step ScholarQA pipeline backed by a mix of proprietary systems — Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1 — and quality filtering left 47,000 usable examples. For DPO, two judge models (GPT-4.1 and DeepSeek-R1, aligned with human preferences 95% of the time) picked winners on paired reports, keeping only pairs where both judges agreed: about 6,000 final examples, fine-tuned in Ai2’s open-instruct framework on 8×H100 GPUs.

One design decision did most of the speed work: AstaBrief was trained to write the full report in one pass, bypassing the snippet-summarization and clustering stages the Claude-based pipeline uses — and skipping section-by-section writing. Ai2 says it found this possible without sacrificing performance.

Scientific papers flowing as light streams into a cited report document
Citation grounding is the design target: AI-generated illustration — AI Frontier Post

What the numbers say#

The main development target was SQABench-CS2, a set of 200 user-written computer science research questions, tracked on four metrics: rubric score (coverage of necessary content), answer precision (paragraph relevance), citation precision (does each citation support its claim), and citation recall (are claims fully supported). The model card reports AstaBrief-8B averaging 87 across the metrics, versus 83.7 for the SFT-only checkpoint and 77.3 for base Qwen3-8B — including 90.5 on citation precision. LLM-judged win rates against the Asta ScholarQA pipeline hit 72% on the test split, and a small human study found that while DR Tulu won overall preference, two of the three researchers preferred AstaBrief on citation accuracy.

The most instructive result is about data, not scale: testing four statistics-based filters on the synthetic training reports, Ai2 found the strongest gains came from simply filtering out reports with low citation density. More aggressive filtering, filter combinations, and learning-rate sweeps added nothing meaningful. Scientific specialization, in other words, looks less like “add more scientific text to pretraining” and more like a post-training data-composition problem — a useful signal for anyone training domain models on a budget.

Open weights, your firewall#

Because the weights are open, institutions can run AstaBrief on their own hardware — including behind their own firewall — which Ai2 describes as necessary when research questions touch on sensitive or unpublished work. The released example workflow in ai2-scholarqa-lib shows how to adapt it for reports from your own PDFs. The effort sits inside NSF OMAI, the U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery.

A glowing stopwatch in front of a stack of research papers symbolizing fast report generation
51 seconds vs 178: speed as a design constraint, not a side effect. AI-generated illustration — AI Frontier Post

One honest caveat from the announcement: most of the training and evaluation was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at that time. Ai2 hasn’t rerun the full evaluation against today’s frontier models — read the numbers as evidence about training and system design choices, not as a claim about the state of the art.

The weights and training data are on Hugging Face, and Fast mode is live in Asta today. Coverage: Unite.AI’s summary.