AI news, models and products, with sources中文
NewsMajor Models22 Sept 2026

Claude Opus 5.5

Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work while costing about 40% less to run than Opus 5, and scores best of any model on its alignment audit. The same day, OpenAI releases GPT-6 Sol and GPT-6 Luna.

Anthropic chief executive Dario Amodei being interviewed on stage
Anthropic co-founder and chief executive Dario Amodei (right) at TechCrunch Disrupt in 2023. File photo. Photo: TechCrunch / Wikimedia Commons, CC BY 2.0

What happened

On 22 September, the US AI company Anthropic released Claude Opus 5.5, the first model in its Claude 5.5 family. Smaller and cheaper versions, Sonnet 5.5 and Haiku 5.5, were promised for the following weeks; Sonnet 5.5 arrived on 28 September.

Anthropic summed up the release in one line: Opus 5.5 performs at the level of Fable 5.1 on most work and costs 40% less to run than Opus 5.

OpenAI released GPT-6 Sol and GPT-6 Luna the same day. The comparisons with OpenAI below come from Anthropic’s announcement, which compares against GPT-6 Astra and the older GPT-5.6 Sol. It does not include the two models OpenAI released that day.

How much better is it?

Anthropic compares models on a set of public tests. The one it leads with is Terminal-Bench 4.0, which asks an AI to complete a chain of professional tasks on its own in a computer’s command line: setting up servers, fixing programs, processing data. It measures less “can it answer a question” and more “can it finish a job by itself”.

Terminal-Bench 4.0: completing complex tasks in the command lineShare of tasks completed; higher is better
  • Claude Opus 5.566.4%
  • GPT-6 Astra57.9%
  • Claude Fable 5.155.8%
  • Claude Opus 552.3%
  • GPT-5.6 Sol37.3%

GPT scores are as reported by OpenAI and quoted by Anthropic. Opus 5.5's standard error is about ±2.6 points. Source: Anthropic, “Introducing Claude Opus 5.5”, 22 Sep 2026

Opus 5.5 also came first on several other tests:

Test What it measures Opus 5.5 Comparison
OSWorld 2.1 Using desktop software like a person 81.8% Fable 5.1: 80.7%
Humanity’s Last Exam (with tools) Very hard questions written by experts across subjects 67.7% Fable 5.1: 65.6%
GDPval-AA v2.1 Real work tasks from 44 occupations 1846 Fable 5.1: 1735
CursorBench 4.0 Vague, multi-file requests from real coding sessions 57.8% GPT-5.6 Sol: 41.7%

It did not win everything. On Zapier’s AutomationBench, which tests business workflows across apps, GPT-6 Astra scored 41.4% to Opus 5.5’s 40.0%. On Terminal-Bench-Science, a test of scientific research tasks, GPT-6 Astra scored 64.6% to Opus 5.5’s 58.7%.

GDPval-AA v2.1: real work from 44 occupationsElo rating; higher is better. Maintained by the independent firm Artificial Analysis
  • Claude Opus 5.51846
  • Claude Fable 5.11735
  • Claude Opus 51708
  • GPT-5.6 Sol1588
  • GPT-6 Astra1542

Source: Anthropic, “Introducing Claude Opus 5.5”, 22 Sep 2026

Cheaper and faster

The most tangible part of the upgrade is cost. Anthropic says Opus 5.5 needs less computing power to run than Opus 5, and has cut prices to match. It also uses fewer tokens to finish the same job. Together, Anthropic estimates, typical use costs about 40% less.

Typical cost vs Opus 5
−40%
Output speed vs Opus 5
30%+ faster
API price per million tokens (in / out)
$4 / $20
Cache-read price vs Opus 5
−60%

Source: Anthropic announcement. Opus 5 costs $5 per million input tokens and $25 per million output tokens.

For subscribers, Anthropic also raised the five-hour usage limits on its Pro, Max, Team and seat-based Enterprise plans, and gave users a one-off limit reset they can save and use when they choose.

What it can do: examples from Anthropic and early testers

Beyond the scores, Anthropic’s announcement gives concrete examples. They come from the company and its partners and have not been independently checked, but they help show what the model can handle:

  • Large code changes. One tester completed a 680,000-line code migration in less than a day, work Anthropic says would have taken an engineering team weeks.
  • Rewriting well-known software. Anthropic asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own tests; Opus 5.5 finished in 9.5 hours, against 12 for Fable 5.1, and cost 51% less.
  • Reports backed by sources. Models were asked to write a report on a company’s quarterly results using only information they could find on a copy of the web. An automated grader checked every figure and quote, and a single invented number meant failure. Opus 5.5 cleared the bar in 16 of 18 attempts; Fable 5.1 and Opus 5 never did.
  • Business analysis. Asked to assess a proposed merger between two fictional companies, building an Excel model and then a presentation for executives, Opus 5.5 took 63 minutes against 93 for Opus 5, at half the cost.

Anthropic also highlights the model’s writing. Users had complained that Opus 5 was hard to follow; Opus 5.5 puts the most important information first and uses less jargon.

Safety

A week before the launch, Anthropic’s chief executive Dario Amodei argued in an essay that AI progress should be paced so that safety work stays ahead of what models can do. Opus 5.5 is the company’s first release since then. (An earlier public call along the same lines was a July open letter signed by more than 1,100 people from AI companies; see Pacing the Frontier.)

The safety points in the announcement:

  • Outside evaluators, including METR and Frontier Design, tested the model before release.
  • On Anthropic’s main alignment test, an automated audit across nearly 2,000 simulated scenarios, it did better than recent Claude models on nearly every measure of misbehaviour, and is less likely to take hard-to-reverse actions or step outside the permissions it has been given.
  • On a new test of whether a model tries to get past containment boundaries, it tried about 85% less often than Opus 5 or Mythos 5.1. Every attempt was low severity, and the model reported each one itself.
  • Because its biology and cybersecurity abilities are close to Mythos 5.1, it ships with safeguards similar to Fable 5.1’s. When they intervene, cybersecurity tasks are handed to Opus 4.8 and biology tasks to Opus 5. Vetted organisations can apply for fewer restrictions.

Things to keep in mind

  • The scores come from the company. Every figure above is from Anthropic’s announcement. The GPT scores are numbers OpenAI published, quoted by Anthropic, and test conditions differ.
  • Anthropic itself says the margins overstate the gap. In its words, at this level benchmark margins “have become a less reliable guide to real-world differences”, and in its own use the gap between Opus 5.5 and Fable 5.1 is smaller than the scores suggest.
  • No comparison with the models OpenAI released the same day. GPT-6 Sol and Luna are not in the tables.
  • For an independent comparison, see the LMArena leaderboards on our Models page, which are based on blind votes by users.

How to try it

  • Website and apps. Paid users can choose Opus 5.5 in the model menu on claude.ai and in the Claude apps. For plan prices, see our Claude product page.
  • Coding. It is available in Claude Code, which also offers a “fast mode” up to 2.5 times faster, at a higher price ($8 per million input tokens, $40 per million output tokens).
  • Developers. Through the Claude API at $4 per million input tokens and $20 per million output tokens.

Model summary

Developer
Anthropic
API price
$4 / $20 per million input / output tokens; cache reads $0.20

Reported benchmark results

Terminal-Bench 4.066.4%GPT-6 Astra: 57.9%
Humanity's Last Exam (with tools)67.7%Fable 5.1: 65.6%
OSWorld 2.181.8%
GDPval-AA v2.11846Knowledge work; GPT-6 Astra: 1542

Data source: Anthropic announcement. Scores are as published by the developer at launch. Test conditions differ between companies, so compare with care.

Sources