Anthropic’s release of Claude Opus 5.5 has pushed the company’s most capable model line forward while sharpening a debate over how frontier AI performance should be measured. Introduced on September 22, the model is the first in Anthropic’s 5.5 family and is designed for difficult coding, research and knowledge-work tasks. Anthropic says it performs at roughly the level of Claude Fable 5.1 on many workloads while costing 40% less to run than the prior Opus generation. The combination of stronger results and lower operating cost makes the launch commercially important even before benchmark arguments are settled.
On Anthropic’s published tests, Opus 5.5 scored 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode, both ahead of Opus 5. The company also reported a 67.7% tools-assisted result on Humanity’s Last Exam, a difficult multidisciplinary benchmark built from expert-level questions. Anthropic cautioned that results depend on evaluation settings and that production safeguards can route some sensitive tasks to other models, a reminder that a single headline score does not fully describe how the deployed system behaves.
Independent and alternative benchmark views make the picture more complicated. BenchLM’s September 24 table for the Artificial Analysis version of Humanity’s Last Exam listed Opus 5.5 first at 61.4%, ahead of other tracked models but below 65%. That is not necessarily a contradiction: the two figures come from different evaluation configurations, including different rules for tool use. The official Humanity’s Last Exam site also released a new HLE-Diamond benchmark on September 22, adding another lane for measuring advanced systems. Comparisons remain meaningful only when the question set, harness, tools and scoring rules match.
That methodological gap was visible in Polymarket trading. A contract on whether any Claude model will record at least 65% on the official Humanity’s Last Exam leaderboard by year-end traded near a 27% implied probability, down about 18 percentage points over the latest day. The contract resolves from the benchmark’s designated official source, so Anthropic’s own tools-assisted result does not automatically settle it. The price is a trader estimate, not a judgment that the published 67.7% figure is false. It reflects uncertainty over whether a qualifying score will appear under the exact resolution standard before December 31.
For developers and enterprise buyers, the practical questions extend beyond one benchmark. They will test whether Opus 5.5’s gains hold across long coding sessions, scientific research, automation and routine production work, and whether the lower price changes which tasks justify using Anthropic’s top tier. Benchmark maintainers, meanwhile, will face pressure to explain tool policies and evaluation harnesses clearly as models become more agentic. The next decisive evidence will come from reproducible independent runs and updates to the official leaderboard. Until then, Opus 5.5 is plainly a stronger release, but the exact distance between a company-reported score and a benchmark used for resolution remains an open question.



