LiveNerf: the community-built watchdog tracking whether Opus 5.5 gets quietly nerfed
A Reddit user's open-source benchmark measures Claude Opus 5.5 against its own launch-day baseline every day for 30 days — so the internet's favorite argument, 'did they nerf it?', finally gets an answer.

Every few weeks the same argument erupts somewhere online: the model feels dumber than it did last week. Someone suspects a quiet swap, a squeezed thinking budget, quantized weights. The lab says nothing changed — and nobody can prove anything either way, because nobody measured the model on launch day.
Built by Reddit user u/TheOnlyVibemaster and announced on r/ClaudeAI on September 27, it measures Claude Opus 5.5 against its own launch-day baseline, once a day for 30 days, under a pre-registered, open-source protocol. The thread drew 269 upvotes and 58 comments in its first two hours; the top-voted comment captured the mood in five words: "downdetector but for AI nerfs."
#The argument it's trying to end
The suspicion has been building for months: that frontier models get quietly worse after release — a quantization here, a smaller model behind the same name there, a shaved reasoning budget. It could equally be nothing at all: humans are excellent at pattern-matching on noise, and a model that feels different is not a model that measures different. The problem is that nobody has ever had a clean day-zero baseline — so every dispute collapses into vibes versus vibes.
When asked in the thread whether there's objective evidence Anthropic has ever done this, the author answered plainly — "There isn't evidence yet." The point is to find out, not to assume.
#How the watchdog works
Opus 5.5 launched on September 22, 2026. The LiveNerf repository was created the same day, and the measurement series began on September 24. Once a day, the full panel runs through a Claude Max subscription using headless Claude Code — no API key involved.
The panel is the clever part. From 2,336 questions drawn from GPQA Diamond, MMLU-Pro, competition math, and the 2025–26 AIME sets, the author kept only the 78 that Opus 5.5 gets right sometimes — questions it always aces or always flubs carry no information about drift. Those 78 were calibrated, locked, and validated under a pre-registered protocol before the series began.
Since you can't make a frontier model's sampling deterministic, LiveNerf makes everything else deterministic instead: frozen prompts, a pinned Claude Code CLI version, exact-match graders — no LLM judge, which would itself drift — and raw logs kept forever. The README puts it well: a Claude Code update changes the harness, "and a changed harness looks exactly like a changed model," so the daily runner refuses to execute if the CLI version has drifted. An Opus 5 control arm runs the same questions daily, to separate harness changes from model changes.
The statistics follow Anthropic's own "Adding Error Bars to Evals" methodology: paired per-item score differences against the baseline, with clustered standard errors. A change only counts if it clears a 99% interval in both 10-day follow-up windows, reaches at least 3 points, and doesn't show up in the control arm. The instrument can detect an accuracy shift of roughly 7.5 points per 10-day window. There's also a subtle secondary signal: output token counts. If the model quietly starts thinking less, that shows up in tokens before it shows up in accuracy.
As of September 27, four of thirty days are in — all 90 samples a day on the same harness hash, none missed. The first results row lands after day 20; the first verdict around October 24.

#What it can't see
The author publishes the instrument's limits alongside its strengths — which is exactly why the project deserves to be taken seriously. In validation, secretly swapping Opus 5 in for Opus 5.5 was not distinguishable at 99% confidence: a same-family model swap of that size would sail straight through undetected.
The thread hashed out the other confounds — dynamic throttling during peak hours, a changed system prompt rather than a changed model. The author's defenses: a locked app version with auto-updates off, and the control arm to catch harness-side drift. One person, one harness, one subscription — a community instrument, not an independent lab audit, but the first of its kind pointed at a launch-day baseline.

#The bigger pattern: users are building their own instruments
LiveNerf isn't alone: the same thread surfaced at least two more community-built nerf trackers, and the reaction suggests an appetite for many more. The pattern is bigger than one model — users no longer trust release notes, so they're building open measurement infrastructure to check the labs' work themselves.
Benchmarks used to be how labs sold models. Increasingly, they're how users audit them. A public, pre-registered, honest null result is the best possible outcome for Anthropic here — and if a dip does appear in late October, it will arrive with receipts: frozen prompts, pinned versions, raw logs, and a control arm.
#What to watch
- The first results row, due after day 20 — roughly mid-October — and the first formal verdict around October 24.
- Replication. The design is documented so others can run their own series; watch whether the genre spreads to other models and other labs.
- The labs' response. Publishing their own launch-day baselines would be the cheapest way for a lab to end the argument before the community does it for them.
Either way, the nerf argument is about to become a data argument. That's new — and it's the community, not the labs, that made it happen.