C5R claims GPT-6 Astra runs its research lab end-to-end — Reddit's top replies push back
A five-month-old startup says a frontier model runs its research lab end to end. The r/singularity thread that surfaced the announcement on September 25 isn't buying the pitch — and C5R's own benchmark page quietly documents the physical-lab mistakes its models still make.

On September 24, the lab-automation startup C5R announced itself on X: "In 12 weeks, we built a research facility that is run entirely by AI," spanning biology, chemistry, and materials science, with a new benchmark called SciUniverse to prove it. The next morning, a video-and-link post of the announcement landed on r/singularity. By midday it had 123 points and 36 comments — and the most visible replies were not the victory lap the announcement was going for.
This is the social-signal pattern we cover here: not the press release, but the informed community verdict on it. And in this case the verdict is a split screen — genuine curiosity on one side, and on the other, working scientists saying the demo does not match what bench work actually feels like.
What C5R announced#
C5R is the project of founder Michael Akilian, who says the company formed five months ago to move AI research out of the purely virtual world and into physical labs. As profiled by Ryan Merket in RuntimeWire on September 24, Akilian began studying biology 18 months ago, ran worm experiments out of a small apartment lab, later worked in a lab at UCSF, and previously did hardware work at Apple and Misfit Wearables before co-founding and selling the AI company Clara Labs.
The facility, dubbed Facility-0, is described by the company as a model-driven lab where a model inspects equipment inventories and specifications, writes an experiment as code, and dispatches instructions to instruments and people — then collects measurements to analyze results and pick the next step. The equipment list runs from hydraulic presses to pipettes, and the company says it integrated more than 40 instruments with its own action, inventory, and scheduling systems.
To score this work, C5R published SciUniverse, a benchmark of scientific work that, in the company's words, "spans the entire research process across chemistry, biology, and materials science." The current release is Level 1: tasks that are straightforward for working scientists and take no more than a few hours each. Crucially, the company's own description says models direct work at Facility-0 "by controlling machines and giving instructions to human operators" — so the "entirely by AI" framing in the launch post coexists with humans still executing parts of the loop.

The numbers C5R published#
SciUniverse Level 1 reports pass@1 averaged across 17 task families, with per-task API inference costs. The headline scores:
| Model | Pass@1 | Cost per task attempt |
|---|---|---|
| Claude Fable 5.1 | 45.3% | $40.61 |
| GPT-6 Astra | 32.5% | $52.37 |
| Claude Opus 5 | 30.5% | $46.31 |
| Grok 4.6 | 26.2% | $13.41 |
| Gemini 3.8 Flash | 14.6% | $4.55 |
| GPT-5.6 Sol | 9.4% | $16.53 |
Two things worth noting before the Reddit takes. First, even the best model clears less than half of "straightforward for scientists, a few hours at most" tasks — this is an honest-looking baseline, not a solved problem. Second, the reported costs are API inference only: lab and labor costs are not reported, which matters because one commenter put the physical side at "$1M+ equipment and $100/hr labor" — at that scale, the model's token bill is a rounding error. (That cost argument, from u/Recoil42, was itself the answer to another commenter asking why C5R didn't just use a cheaper model like GPT-5.6 Sol.)
Why Reddit isn't buying it#
The r/singularity thread (posted 11:09 AM EDT on September 25 by u/Distinct-Question-16, carrying a 1:46 embedded video and the link to C5R's X post) is the fresh phenomenon here — and its most visible comments are skeptical, mocking, or doing the math:
- The safety angle. u/tskir pointed out the irony that the entire AI Box thought-experiment literature — decades of arguments about whether an AI could talk its way out of containment — was rendered "null and void" because, as they put it, "we kinda skipped the box part entirely and voluntarily." Their framing: the model gets software sandboxes while being wired into equipment that affects the physical world.
- The bench-worker's verdict. u/kellogg76, who says they own the Opentrons Flex liquid handler shown in the video, wrote that they "would not trust a thing it makes" and that the machine's "verified" protocols don't run in their experience.
- The VC-theater theory. u/jc2046 called the presentation staged and theatrical — a funding demo dressed as a breakthrough.
- The economics. Beyond the cost argument above, u/Edgezg was the optimistic counterpoint: this is the hopeful use case they'd been waiting for — AI doing deep research toward cancer or HIV cures. And u/cultureicon claimed, without evidence, that Eli Lilly is already doing this "1,000x faster."
A note on sourcing: r/singularity hides comment scores for 120 minutes after posting, so exact rankings weren't available. The comments above are the most visible ones in the thread's top sorts — what the community itself has surfaced so far. The OP made no replies.

The strongest evidence against the hype came from C5R itself#
The most damning material in the whole story is not a Reddit comment — it's the company's own failure catalog. In its launch materials, C5R admits its models "pipette frozen samples, contaminate DNA, fail to account for evaporating solvents, and vortex open containers." A self-driving lab that contaminates DNA is not a small detail; it's the entire difficulty of physical science, compressed into one sentence. Virtual benchmarks fail gracefully. Wet labs fail with real materials, real cross-contamination, and real costs.
This is also not a new idea, just a newly instrumented one: self-driving labs go back years — Carnegie Mellon's Coscientist system, demonstrated with Emerald Cloud Lab's robotic facilities, showed AI models designing and executing chemistry experiments end to end. What C5R is really selling is a measurement harness: an instrumented physical environment where model failures are legible and scoreable. That is genuinely useful — if the scoring survives contact with independent replication.
What to watch#
The Reddit thread is barely hours old; the verdict can still shift as more bench scientists show up. What would change the story: independent researchers reproducing the SciUniverse tasks on their own hardware; Level 2+ results on work that isn't straightforward-for-scientists; and anything resembling a real, repeatable physical result rather than a 1:46 highlight reel. And tskir's point deserves to outlive this thread — if models are going to be wired into physical labs, the safety conversation can't stay in the digital world. The box has been skipped. The question now is who writes the protocol for what's plugged in next.
Sources: the r/singularity thread posted September 25, 2026; C5R's SciUniverse page (c5r.net); Ryan Merket's RuntimeWire profile of C5R (September 24, 2026); C5R's launch post on X (@c5rcorp, September 24, 2026).