HypoArena: run a full AI co-scientist pipeline entirely offline
AI co-scientists that generate, debate, and evolve research hypotheses are moving from lab demos to real workflows. HypoArena — an open-source workbench that collected hundreds of GitHub stars within hours of its launch — decomposes that whole generate-debate-evolve machinery into testable stages you can run on a laptop with no API keys, no network, and no GPU. Here is the complete hands-on tour.
Every few months, a headline announces that an AI system has proposed a genuinely novel scientific hypothesis — a new drug combination, a materials recipe, a causal mechanism nobody had written down. The systems behind those headlines share a common skeleton: generate candidate claims, debate them with critics, and evolve the survivors into sharper ideas. It is a compelling loop, but until now it has been nearly impossible to study as an outsider. The code is proprietary, the runs cost real money in API calls, and the evaluations are bespoke.
HypoArena attacks exactly that gap. It is a fully offline workbench that decomposes the mechanism layer of generate-debate-evolve pipelines into individually testable components: a synthetic literature factory, span-level grounding verification, pluggable agent adapters, Bradley–Terry/Elo tournaments, paraphrase deduplication, hypothesis evolution operators, and Bayesian evidence accumulation. It went public today and picked up hundreds of stars within hours — and unusually for a viral launch, the thing actually runs. The whole pipeline executes in about two seconds on a laptop, with NumPy as its only runtime dependency.
The crucial design decision: nothing here touches a real model. Every experiment runs on synthetic corpora with planted ground truth, and every agent in the loop is scripted, replayed from a fixture, or pointed at a local mock server. That sounds like a limitation until you see what it buys: deterministic, reproducible experiments about the machinery of discovery pipelines — does the Elo tournament recover a known skill ordering, does grounding catch fabricated citations, does the debate loop actually improve claims. Those are properties you can verify with zero API spend, and they are exactly the properties that are hardest to test in production co-scientist systems.
In this tutorial you will install HypoArena, run the end-to-end demo, walk each of the nine pipeline stages, learn to read its reports, swap the agent adapters, and make runs reproducible with seeds and checkpoints.
What you will need#
- Python 3.11+ — the project declares
requires-python = ">=3.11"; I tested on 3.12. - pip and git — nothing else. No Docker, no GPU, no API keys, no accounts.
- About five minutes — cloning and installing takes a couple of minutes; every command after that finishes in seconds.
- $0 — the runtime dependency set is NumPy alone. PyTorch is an optional extra used only for a small demonstration ranker.
Step 1 — Install HypoArena#
Clone the repository, create a virtual environment, and install in editable mode:
git clone https://github.com/OpSafari/hypoarena.git
cd hypoarena
python3 -m venv .venv
.venv/bin/pip install -e .
Verify the install:
.venv/bin/hypoarena --version
# hypoarena 0.1.4
.venv/bin/hypoarena --help
usage: hypoarena [-h] [--version] COMMAND ...
Offline workbench for hypothesis-discovery pipelines: grounded claim graphs,
synthetic literature, debate loops and Elo tournaments. Every command runs on
synthetic data with planted ground truth and claims no real benchmark.
positional arguments:
COMMAND
corpus generate a synthetic corpus with planted ground truth
generate propose candidate claims from the corpus
verify grade every claim against the corpus
dedup find near-duplicate claims
debate run the propose-critique-revise loop
rank rank claims with an Elo tournament
evolve expand the graph with evolution operators
accumulate
update beliefs from graded evidence
report run the full pipeline and write reports
demo run a small offline end-to-end demonstration
Ten subcommands, nine of them pipeline stages plus the demo shortcut. Note the honesty baked into the help text: "Every command runs on synthetic data with planted ground truth and claims no real benchmark." Keep that sentence in mind — it is the project's entire epistemic contract, and we will return to it.
Step 2 — The 30-second demo#
Before touching any stage individually, run the whole thing once:
.venv/bin/hypoarena demo --chains 2 --chain-length 2 --out ./demo-run
hypoarena demo - offline synthetic corpus with planted ground truth
planted links: 2 recovered: 2 rate: 1.0
[recovered] gene G3 causes apoptosis rate
[recovered] gene G3 decreases apoptosis rate
note: a synthetic demonstration only; no claim about real discovery.
Here is what just happened. HypoArena's synthetic literature factory generated a small corpus of paper-like documents, and while generating them it planted two causal chains as ground truth (here, claims about gene G3 and apoptosis). The pipeline then ran end to end — proposing claims, checking them against the corpus, debating them, ranking them — and recovered both planted links. The recovery rate is the workbench's core scoreboard: on data where the truth is known by construction, did the machinery find it?
The demo also wrote report.md and report.html into ./demo-run/run/. Open the Markdown report in your editor now; we will learn to read it in Step 4.
Step 3 — The nine stages, one pipeline#
Each subcommand is one stage, and every stage reads the previous stage's artifacts and writes its own JSONL files. Run the full pipeline with explicit control:
.venv/bin/hypoarena report --chains 3 --chain-length 3 \
--seed 42 --out ./full-run --run-id tutorial
report: run tutorial
documents=22 claims=15 evidence=18
recovered 6/6 planted links (rate 1.0)
json: /home/user/full-run/tutorial/report.json
markdown: /home/user/full-run/tutorial/report.md
html: /home/user/full-run/tutorial/report.html
Twenty-two synthetic documents, fifteen candidate claims, eighteen evidence links — processed in under two seconds. Here is what each stage contributes:
| Stage | What it does | Key artifact |
|---|---|---|
| corpus | Generates paper-like documents with planted causal chains, competing hypotheses, and paraphrase clusters as ground truth | corpus.jsonl, truth.jsonl |
| generate | Proposes candidate claims from the corpus via a DiscoveryAgent (scripted by default) | candidates.jsonl |
| verify | Grades every claim against the corpus with span-level grounding checks | grounding.jsonl |
| dedup | Finds near-duplicate claims with Jaccard, TF-IDF cosine, and MinHash LSH | dedup.jsonl |
| debate | Runs the propose → critique → revise loop between scripted agents | debates.jsonl |
| rank | Ranks claims with an Elo (Bradley–Terry) tournament and a rubric judge | tournament.jsonl |
| evolve | Expands the claim graph with evolution operators behind a novelty gate | evolution.jsonl |
| accumulate | Updates Bayesian beliefs from graded evidence | beliefs.jsonl |
| report | Renders everything as JSON, Markdown, and self-contained HTML | report.json/md/html |
You can also run stages individually — hypoarena debate --out ./full-run --run-id tutorial reuses the existing artifacts — which is how you experiment with one mechanism at a time. Every stage accepts the same flags: --seed for the RNG, --chains and --chain-length for the planted truth, --run-id to name the run, and --resume to reuse completed stage checkpoints instead of recomputing them.
Peek at a raw claim to see the schema the whole pipeline shares:
{
"claim_id": "clm_53e6521e7236",
"statement": "batch increases cells through 7 in the assayed population, measured by dose response",
"subject": "batch", "relation": "increases", "object": "response",
"provenance": {"origin": "agent", "generation": 0, "seed": 42, "notes": "proposed"},
"scope": {"population": "synthetic corpus", "conditions": []}
}
Claims carry their provenance — which agent proposed them, in which generation, under which seed. That provenance trail is what makes every number in the final report auditable.
Step 4 — Read the report like a scientist#
Open ./full-run/tutorial/report.md. It is organized as a lab notebook: counts, then one section per mechanism. Here are the sections that matter, from my seed-42 run.
Grounding tells you whether claims are actually supported by corpus spans — the workbench's hallucination detector:
| metric | value |
|---|---|
| total | 12 |
| grounded | 9 |
| ungrounded | 3 |
| fabricated | 0 |
| grounded_rate | 0.75 |
Span-level grounding verification is one of the sharpest tools here: it checks whether a claim's citations point at real text spans, and flags fabricated references, drifting numbers, and negations flipped from the source. A grounded rate of 0.75 with zero fabricated citations is the kind of baseline you would want before trusting any downstream ranking.
Ranking shows the Elo tournament outcome — every claim starts at 1500 and plays pairwise matches judged by a rubric:
| # | claim | elo | played | wins | win_rate |
|---|---|---|---|---|---|
| 1 | clm_f50b76ce85fd | 1536.6 | 11 | 3 | 0.64 |
| 2 | clm_f7a23517fbaf | 1535.9 | 11 | 3 | 0.64 |
| 3 | clm_ba3983081f84 | 1535.4 | 11 | 3 | 0.64 |
Beliefs are the Bayesian posteriors after evidence accumulation — the workbench's answer to "how confident should we be in each claim now?"
| claim_id | posterior |
|---|---|
clm_ba3983081f84 | 0.872 |
clm_bc4031d4acb9 | 0.860 |
clm_f50b76ce85fd | 0.831 |
And Cost is my favorite section in the whole report:
| metric | value |
|---|---|
| calls | 24 |
| prompt_tokens | 1545 |
Tokens are counted, never billed. The package ships no price table and writes no dollar amounts into any artifact — a deliberate design choice the docs call out explicitly. In a year when everyone is anxious about agent API spend, a workbench that meters usage without ever needing a key is a small act of sanity.
Step 5 — The agents are swappable (and fake on purpose)#
The discovery loop never talks to a model directly. It talks to a DiscoveryAgent protocol — propose, critique, revise — with three bundled adapters, all of which run offline:
| Adapter | What it is | Use it for |
|---|---|---|
| ScriptedAgent | Responds by rule, with a quality tier: vague, focused, or mechanistic | Default for every example; tiers implant a known skill ordering the Elo tournament must recover |
| ReplayAgent | Reads canned responses from a JSONL fixture | Regression tests — "run once, replay forever" |
| HttpAgent | Speaks OpenAI-compatible chat-completions | Plugging in a real model later; refuses non-loopback endpoints by default |
The quality tiers are the clever bit. A vague agent ignores context and generalizes; a focused one uses context entities; a mechanistic one adds mechanisms and measurement methods. Because higher tiers are strictly more specific on the same request, whether the Elo tournament recovers the implanted skill ordering is a decidable property — a genuine test of the ranking mechanism that depends on no model's capabilities at all.
Watch a debate in debates.jsonl to see the loop breathe:
{
"proposal": "batch increases cells through 7 in the assayed population, measured by dose response",
"agents": ["scripted-0.9", "scripted-0.5", "scripted-0.2"],
"rounds_run": 2, "converged": true,
"turns": [{
"round_index": 0,
"critiques": [
{"agent": "scripted-0.5", "text": "the claim about batch does not name a testable assay"},
{"agent": "scripted-0.2", "text": "the claim needs more support"}
],
"revised": "batch increases cells through 7 in the assayed population, measured by dose response (unchanged)"
}]
}
Three scripted critics at different quality tiers, two rounds, convergence. The low-tier critic can only manage "needs more support"; the mid-tier one correctly demands a testable assay. That gradient is the whole point: the machinery is being exercised, not the models.
Step 6 — Make it yours: seeds, scale, and checkpoints#
Everything is seeded and checkpointed, so experiments are reproducible down to the byte. Re-run the same command twice and diff the artifacts — they will be identical:
.venv/bin/hypoarena report --chains 3 --chain-length 3 \
--seed 42 --out ./full-run --run-id tutorial
.venv/bin/hypoarena report --chains 3 --chain-length 3 \
--seed 42 --out ./full-run2 --run-id tutorial
diff -r ./full-run/tutorial ./full-run2/tutorial && echo "byte-identical"
Change the seed and you get a fresh synthetic literature with different planted chains — a new exam for the same machinery. Scale up the difficulty with --chains and --chain-length. And if you interrupt a run, --resume picks up from per-stage checkpoint files instead of starting over:
ls ./full-run/tutorial/checkpoints/
# corpus.done generate.done verify.done dedup.done debate.done
# rank.done evolve.done accumulate.done report.done
Three bundled examples in examples/ go deeper on individual mechanisms: an Elo-recovery tournament, a novelty cycle for the evolution operators, and a grounding-verification walkthrough on a planted corpus. Each is a small Python script you can read in one sitting — start with examples/tournament/elo_recovery.py if the ranking stage intrigued you.
The honesty contract#
HypoArena ships a document called docs/honesty.md — "what this repository measures, and what it does not" — and the authors ask you to read it before anything else. Its core claim: the workbench verifies mechanism correctness, not model capability. Graph invariants, grounding precision on synthetic data, Elo recovery of implanted orderings, dedup recall on planted paraphrase clusters, Bayesian monotonicity, checkpoint byte-identity. Every one of those is checkable with deterministic tests. None of them says anything about whether any real language model can do science.
That distinction is worth internalizing, because it is the exact line most AI evaluations blur. A recovery rate of 1.0 on planted synthetic chains tells you the pipeline's machinery works when the truth is known — it is a necessary condition for trusting the machinery on unknown truth, not a sufficient one. The authors never claim otherwise, and the CLI reminds you on every run. In a field drowning in benchmark hype, a tool that prints its own epistemic limits next to its scores deserves attention.
Caveats before you commit#
- It is hours old. The repository went public today; expect APIs to shift, docs to be reorganized, and sharp edges. I verified version 0.1.4; pin your clone to a commit if you build on it.
- The deep docs are in Chinese. The README and CLI are in English, but the detailed mechanism docs under
docs/(architecture, grounding, debate, tournament, belief) are written in Chinese with English technical terms preserved. Machine translation handles them fine, but it is a real speed bump. - Synthetic only, by design. Nothing here will discover anything about the real world. If you want the
HttpAgentto talk to a real model, you will point it at your own OpenAI-compatible endpoint — and at that point you inherit every evaluation problem the workbench was built to sidestep. - The optional torch ranker needs PyTorch. Everything in this tutorial used the NumPy-only default path; the
.[dev,torch]extra pulls CPU PyTorch for a small trainable-ranker demo I did not run.
The takeaway#
HypoArena is not a co-scientist. It is something arguably more useful right now: a wind tunnel for co-scientist machinery. By replacing models with scripted agents, literature with synthetic corpora, and ground truth with planted chains, it turns the generate-debate-evolve loop from an expensive black box into nine inspectable, testable, reproducible stages that run in two seconds on a laptop.
If you build agent pipelines, the ideas transfer directly: provenance on every artifact, span-level grounding before ranking, Elo tournaments with implanted orderings as a mechanism test, token counting without billing, and an honesty doc that states what your numbers do and do not mean. That is a better evaluation culture than most production systems have — and you can install it with pip.
HypoArena is MIT-licensed at github.com/OpSafari/hypoarena. This tutorial was verified against v0.1.4; every command above was executed and every output shown is real.