Claude Opus 5.5 tops SimpleBench at 88.4% — first model to beat the human average
Two days after Anthropic shipped Claude Opus 5.5, the model has taken the top spot on SimpleBench. A leaderboard screenshot posted to r/singularity on Thursday shows Opus 5.5 scoring 88.4% — ahead of the 83.7% average posted by untrained humans on a benchmark designed to humble AI. The post drew 448 upvotes and 91 comments in about four hours, and the top-voted reaction was half disbelief, half celebration: a joke asking whether the "Highest Human Score" row was still available.

Two days after Anthropic shipped Claude Opus 5.5, the model has taken the top spot on SimpleBench. A leaderboard screenshot posted to r/singularity on Thursday shows Opus 5.5 scoring 88.4% — ahead of the 83.7% average posted by untrained humans on a benchmark designed to humble AI. The post drew 448 upvotes and 91 comments in about four hours, and the top-voted reaction was half disbelief, half celebration: a joke asking whether the "Highest Human Score" row was still available.
We covered the Opus 5.5 launch on Tuesday. This is the week-after story: the benchmark receipts are starting to arrive, and the first one is a statement result.
What SimpleBench actually tests#
SimpleBench is deliberately the awkward benchmark. It is a multiple-choice text test of just over 200 questions covering spatio-temporal reasoning, social intelligence, and what its creators call linguistic adversarial robustness — trick questions, in plain English. The premise, stated on the benchmark's own site, is that individuals with unspecialized high-school knowledge outperform state-of-the-art models on it. The site's front page lists a non-specialized human baseline of 83.7%, measured across nine participants, and names Claude Fable as the previous top model at 81.9%.
That framing is exactly why an 88.4% matters more than another saturated leaderboard. Most text benchmarks now fall to frontier models by default; SimpleBench was built as the holdout. Opus 5.5 is the first model to clear the human average on it.
The numbers#

The original poster linked the live leaderboard at simple-bench.com and noted the sting in the tail for competitors: Opus 5.5 took the record, they wrote, at less than half the price of Claude Fable 5.1, the previous record holder. Price per unit of capability has been the quiet theme of this release cycle, and the thread kept returning to it.
The community verdict#
The comments read like a milestone being acknowledged in real time. The most upvoted comment joked about the surviving "Highest Human Score" row while conceding "it looks like a crazy good model." Another highly-voted reply argued the 5.5 numbering is justified after all — "a real jump across benchmarks and real-world use" — a direct answer to skeptics who had called the release a point update. A third mocked the model-collapse doomers: where, they asked, is the promised wall, the rot from synthetic-data training?
One long sub-thread got into why this particular benchmark carries weight. SimpleBench, a commenter noted, has never seen the kind of saturating jump that flattened other benchmarks — it has staying power as a measure precisely because it resists gaming. A jump like this one, on this test, reads as signal rather than noise.
The caveats#
The thread was not pure celebration, and the corrections are worth keeping. Two commenters pointed out that the single highest recorded human score on the leaderboard still sits above 88.4% — "for now." Opus 5.5 beat the human average; the best individual human is still ahead of it.
Another noted a real-world wrinkle: Fable handles very long contexts better, with Opus making small errors well before the 500k-token mark — a reminder that a multiple-choice reasoning score does not capture everything a model is asked to do. And one commenter asked the question benchmark-watchers always should: does the lead survive on messy, real-world coding and support sets outside clean evaluations?
There is also a housekeeping oddity worth flagging: the benchmark's own front page still names Fable as the top model. The leaderboard is moving faster than the site's copy. Treat the shared screenshot and the live leaderboard — not the marketing text — as the current state.
What to watch#
- Official confirmation. Anthropic has not published the 88.4% figure itself; the score comes from a community-shared leaderboard. Watch for it in Anthropic's own eval reports.
- The agentic benchmarks. A multiple-choice reasoning test is one thing; the coding and agent evaluations the industry actually buys on are the next receipts to arrive.
- The response. A rumor also circulating today claims Opus 5.5 caught OpenAI off guard and compressed the GPT-6.1 Astra timeline — we tracked that rumor separately. If true, scores like this one are the reason.
- The front page. Whether SimpleBench updates its copy to name the new leader will tell you how fast third-party evals are keeping up with release week.
Benchmarks are snapshots, and snapshots get retaken. But the symbolism is hard to miss: the 5.5 that skeptics called a point release now owns the one leaderboard specifically designed to humble frontier models — at less than half the price of the model it dethroned.