AI-generated editorial illustration: a robot scientist examining a glowing holographic network of hypothesis nodes connected by evidence links above a dark laboratory bench

Every few months, a headline announces that an AI system has proposed a genuinely novel scientific hypothesis — a new drug combination, a materials recipe, a causal mechanism nobody had written down. The systems behind those headlines share a common skeleton: generate candidate claims, debate them with critics, and evolve the survivors into sharper ideas. It is a compelling loop, but until now it has been nearly impossible to study as an outsider. The code is proprietary, the runs cost real money in API calls, and the evaluations are bespoke.

HypoArena attacks exactly that gap. It is a fully offline workbench that decomposes the mechanism layer of generate-debate-evolve pipelines into individually testable components: a synthetic literature factory, span-level grounding verification, pluggable agent adapters, Bradley–Terry/Elo tournaments, paraphrase deduplication, hypothesis evolution operators, and Bayesian evidence accumulation. It went public today and picked up hundreds of stars within hours — and unusually for a viral launch, the thing actually runs. The whole pipeline executes in about two seconds on a laptop, with NumPy as its only runtime dependency.

The crucial design decision: nothing here touches a real model. Every experiment runs on synthetic corpora with planted ground truth, and every agent in the loop is scripted, replayed from a fixture, or pointed at a local mock server. That sounds like a limitation until you see what it buys: deterministic, reproducible experiments about the machinery of discovery pipelines — does the Elo tournament recover a known skill ordering, does grounding catch fabricated citations, does the debate loop actually improve claims. Those are properties you can verify with zero API spend, and they are exactly the properties that are hardest to test in production co-scientist systems.

In this tutorial you will install HypoArena, run the end-to-end demo, walk each of the nine pipeline stages, learn to read its reports, swap the agent adapters, and make runs reproducible with seeds and checkpoints.

What you will need#

  • Python 3.11+ — the project declares requires-python = ">=3.11"; I tested on 3.12.
  • pip and git — nothing else. No Docker, no GPU, no API keys, no accounts.
  • About five minutes — cloning and installing takes a couple of minutes; every command after that finishes in seconds.
  • $0 — the runtime dependency set is NumPy alone. PyTorch is an optional extra used only for a small demonstration ranker.

Step 1 — Install HypoArena#

Clone the repository, create a virtual environment, and install in editable mode:

git clone https://github.com/OpSafari/hypoarena.git
cd hypoarena
python3 -m venv .venv
.venv/bin/pip install -e .

Verify the install:

.venv/bin/hypoarena --version
# hypoarena 0.1.4
.venv/bin/hypoarena --help
usage: hypoarena [-h] [--version] COMMAND ...

Offline workbench for hypothesis-discovery pipelines: grounded claim graphs,
synthetic literature, debate loops and Elo tournaments. Every command runs on
synthetic data with planted ground truth and claims no real benchmark.

positional arguments:
  COMMAND
    corpus    generate a synthetic corpus with planted ground truth
    generate  propose candidate claims from the corpus
    verify    grade every claim against the corpus
    dedup     find near-duplicate claims
    debate    run the propose-critique-revise loop
    rank      rank claims with an Elo tournament
    evolve    expand the graph with evolution operators
    accumulate
              update beliefs from graded evidence
    report    run the full pipeline and write reports
    demo      run a small offline end-to-end demonstration

Ten subcommands, nine of them pipeline stages plus the demo shortcut. Note the honesty baked into the help text: "Every command runs on synthetic data with planted ground truth and claims no real benchmark." Keep that sentence in mind — it is the project's entire epistemic contract, and we will return to it.

Step 2 — The 30-second demo#

Before touching any stage individually, run the whole thing once:

.venv/bin/hypoarena demo --chains 2 --chain-length 2 --out ./demo-run
hypoarena demo - offline synthetic corpus with planted ground truth
planted links: 2   recovered: 2   rate: 1.0
  [recovered] gene G3 causes apoptosis rate
  [recovered] gene G3 decreases apoptosis rate
note: a synthetic demonstration only; no claim about real discovery.

Here is what just happened. HypoArena's synthetic literature factory generated a small corpus of paper-like documents, and while generating them it planted two causal chains as ground truth (here, claims about gene G3 and apoptosis). The pipeline then ran end to end — proposing claims, checking them against the corpus, debating them, ranking them — and recovered both planted links. The recovery rate is the workbench's core scoreboard: on data where the truth is known by construction, did the machinery find it?

The demo also wrote report.md and report.html into ./demo-run/run/. Open the Markdown report in your editor now; we will learn to read it in Step 4.

Step 3 — The nine stages, one pipeline#

AI-generated editorial illustration: nine glowing pipeline stages labeled corpus, generate, verify, dedup, debate, rank, evolve, accumulate, report connected by arrows

Each subcommand is one stage, and every stage reads the previous stage's artifacts and writes its own JSONL files. Run the full pipeline with explicit control:

.venv/bin/hypoarena report --chains 3 --chain-length 3 \
  --seed 42 --out ./full-run --run-id tutorial
report: run tutorial
  documents=22 claims=15 evidence=18
  recovered 6/6 planted links (rate 1.0)
  json: /home/user/full-run/tutorial/report.json
  markdown: /home/user/full-run/tutorial/report.md
  html: /home/user/full-run/tutorial/report.html

Twenty-two synthetic documents, fifteen candidate claims, eighteen evidence links — processed in under two seconds. Here is what each stage contributes:

StageWhat it doesKey artifact
corpusGenerates paper-like documents with planted causal chains, competing hypotheses, and paraphrase clusters as ground truthcorpus.jsonl, truth.jsonl
generateProposes candidate claims from the corpus via a DiscoveryAgent (scripted by default)candidates.jsonl
verifyGrades every claim against the corpus with span-level grounding checksgrounding.jsonl
dedupFinds near-duplicate claims with Jaccard, TF-IDF cosine, and MinHash LSHdedup.jsonl
debateRuns the propose → critique → revise loop between scripted agentsdebates.jsonl
rankRanks claims with an Elo (Bradley–Terry) tournament and a rubric judgetournament.jsonl
evolveExpands the claim graph with evolution operators behind a novelty gateevolution.jsonl
accumulateUpdates Bayesian beliefs from graded evidencebeliefs.jsonl
reportRenders everything as JSON, Markdown, and self-contained HTMLreport.json/md/html

You can also run stages individually — hypoarena debate --out ./full-run --run-id tutorial reuses the existing artifacts — which is how you experiment with one mechanism at a time. Every stage accepts the same flags: --seed for the RNG, --chains and --chain-length for the planted truth, --run-id to name the run, and --resume to reuse completed stage checkpoints instead of recomputing them.

Peek at a raw claim to see the schema the whole pipeline shares:

{
  "claim_id": "clm_53e6521e7236",
  "statement": "batch increases cells through 7 in the assayed population, measured by dose response",
  "subject": "batch", "relation": "increases", "object": "response",
  "provenance": {"origin": "agent", "generation": 0, "seed": 42, "notes": "proposed"},
  "scope": {"population": "synthetic corpus", "conditions": []}
}

Claims carry their provenance — which agent proposed them, in which generation, under which seed. That provenance trail is what makes every number in the final report auditable.

Step 4 — Read the report like a scientist#

Open ./full-run/tutorial/report.md. It is organized as a lab notebook: counts, then one section per mechanism. Here are the sections that matter, from my seed-42 run.

Grounding tells you whether claims are actually supported by corpus spans — the workbench's hallucination detector:

metricvalue
total12
grounded9
ungrounded3
fabricated0
grounded_rate0.75

Span-level grounding verification is one of the sharpest tools here: it checks whether a claim's citations point at real text spans, and flags fabricated references, drifting numbers, and negations flipped from the source. A grounded rate of 0.75 with zero fabricated citations is the kind of baseline you would want before trusting any downstream ranking.

Ranking shows the Elo tournament outcome — every claim starts at 1500 and plays pairwise matches judged by a rubric:

#claimeloplayedwinswin_rate
1clm_f50b76ce85fd1536.61130.64
2clm_f7a23517fbaf1535.91130.64
3clm_ba3983081f841535.41130.64

Beliefs are the Bayesian posteriors after evidence accumulation — the workbench's answer to "how confident should we be in each claim now?"

claim_idposterior
clm_ba3983081f840.872
clm_bc4031d4acb90.860
clm_f50b76ce85fd0.831

And Cost is my favorite section in the whole report:

metricvalue
calls24
prompt_tokens1545

Tokens are counted, never billed. The package ships no price table and writes no dollar amounts into any artifact — a deliberate design choice the docs call out explicitly. In a year when everyone is anxious about agent API spend, a workbench that meters usage without ever needing a key is a small act of sanity.

Step 5 — The agents are swappable (and fake on purpose)#

AI-generated editorial illustration: two AI robot agents debating with speech bubbles, a glowing Elo rating ladder ascending between them

The discovery loop never talks to a model directly. It talks to a DiscoveryAgent protocol — propose, critique, revise — with three bundled adapters, all of which run offline:

AdapterWhat it isUse it for
ScriptedAgentResponds by rule, with a quality tier: vague, focused, or mechanisticDefault for every example; tiers implant a known skill ordering the Elo tournament must recover
ReplayAgentReads canned responses from a JSONL fixtureRegression tests — "run once, replay forever"
HttpAgentSpeaks OpenAI-compatible chat-completionsPlugging in a real model later; refuses non-loopback endpoints by default

The quality tiers are the clever bit. A vague agent ignores context and generalizes; a focused one uses context entities; a mechanistic one adds mechanisms and measurement methods. Because higher tiers are strictly more specific on the same request, whether the Elo tournament recovers the implanted skill ordering is a decidable property — a genuine test of the ranking mechanism that depends on no model's capabilities at all.

Watch a debate in debates.jsonl to see the loop breathe:

{
  "proposal": "batch increases cells through 7 in the assayed population, measured by dose response",
  "agents": ["scripted-0.9", "scripted-0.5", "scripted-0.2"],
  "rounds_run": 2, "converged": true,
  "turns": [{
    "round_index": 0,
    "critiques": [
      {"agent": "scripted-0.5", "text": "the claim about batch does not name a testable assay"},
      {"agent": "scripted-0.2", "text": "the claim needs more support"}
    ],
    "revised": "batch increases cells through 7 in the assayed population, measured by dose response (unchanged)"
  }]
}

Three scripted critics at different quality tiers, two rounds, convergence. The low-tier critic can only manage "needs more support"; the mid-tier one correctly demands a testable assay. That gradient is the whole point: the machinery is being exercised, not the models.

Step 6 — Make it yours: seeds, scale, and checkpoints#

Everything is seeded and checkpointed, so experiments are reproducible down to the byte. Re-run the same command twice and diff the artifacts — they will be identical:

.venv/bin/hypoarena report --chains 3 --chain-length 3 \
  --seed 42 --out ./full-run --run-id tutorial
.venv/bin/hypoarena report --chains 3 --chain-length 3 \
  --seed 42 --out ./full-run2 --run-id tutorial
diff -r ./full-run/tutorial ./full-run2/tutorial && echo "byte-identical"

Change the seed and you get a fresh synthetic literature with different planted chains — a new exam for the same machinery. Scale up the difficulty with --chains and --chain-length. And if you interrupt a run, --resume picks up from per-stage checkpoint files instead of starting over:

ls ./full-run/tutorial/checkpoints/
# corpus.done  generate.done  verify.done  dedup.done  debate.done
# rank.done  evolve.done  accumulate.done  report.done

Three bundled examples in examples/ go deeper on individual mechanisms: an Elo-recovery tournament, a novelty cycle for the evolution operators, and a grounding-verification walkthrough on a planted corpus. Each is a small Python script you can read in one sitting — start with examples/tournament/elo_recovery.py if the ranking stage intrigued you.

The honesty contract#

HypoArena ships a document called docs/honesty.md — "what this repository measures, and what it does not" — and the authors ask you to read it before anything else. Its core claim: the workbench verifies mechanism correctness, not model capability. Graph invariants, grounding precision on synthetic data, Elo recovery of implanted orderings, dedup recall on planted paraphrase clusters, Bayesian monotonicity, checkpoint byte-identity. Every one of those is checkable with deterministic tests. None of them says anything about whether any real language model can do science.

That distinction is worth internalizing, because it is the exact line most AI evaluations blur. A recovery rate of 1.0 on planted synthetic chains tells you the pipeline's machinery works when the truth is known — it is a necessary condition for trusting the machinery on unknown truth, not a sufficient one. The authors never claim otherwise, and the CLI reminds you on every run. In a field drowning in benchmark hype, a tool that prints its own epistemic limits next to its scores deserves attention.

Caveats before you commit#

  • It is hours old. The repository went public today; expect APIs to shift, docs to be reorganized, and sharp edges. I verified version 0.1.4; pin your clone to a commit if you build on it.
  • The deep docs are in Chinese. The README and CLI are in English, but the detailed mechanism docs under docs/ (architecture, grounding, debate, tournament, belief) are written in Chinese with English technical terms preserved. Machine translation handles them fine, but it is a real speed bump.
  • Synthetic only, by design. Nothing here will discover anything about the real world. If you want the HttpAgent to talk to a real model, you will point it at your own OpenAI-compatible endpoint — and at that point you inherit every evaluation problem the workbench was built to sidestep.
  • The optional torch ranker needs PyTorch. Everything in this tutorial used the NumPy-only default path; the .[dev,torch] extra pulls CPU PyTorch for a small trainable-ranker demo I did not run.

The takeaway#

HypoArena is not a co-scientist. It is something arguably more useful right now: a wind tunnel for co-scientist machinery. By replacing models with scripted agents, literature with synthetic corpora, and ground truth with planted chains, it turns the generate-debate-evolve loop from an expensive black box into nine inspectable, testable, reproducible stages that run in two seconds on a laptop.

If you build agent pipelines, the ideas transfer directly: provenance on every artifact, span-level grounding before ranking, Elo tournaments with implanted orderings as a mechanism test, token counting without billing, and an honesty doc that states what your numbers do and do not mean. That is a better evaluation culture than most production systems have — and you can install it with pip.

HypoArena is MIT-licensed at github.com/OpSafari/hypoarena. This tutorial was verified against v0.1.4; every command above was executed and every output shown is real.