AI news, models and products, with sources中文
NewsMajor Models28 Sept 2026

Claude Sonnet 5.5 nears Opus 5.5 at half the price, scoring 70.6% on a coding test where Sonnet 5 managed 10.3%

Anthropic's mid-tier model keeps Sonnet 5's $2/$10 price but scores 70.6% on Terminal-Bench 4.0 and is 2 points behind Opus 5.5 on office work, matching Sonnet 5's best scores at about a tenth of the cost per task. The scores are Anthropic's own; on LMArena it beats GPT-6 Sol but only ties GPT-6 Astra.

What happened

On 28 September Anthropic released Claude Sonnet 5.5, the second model in its Claude 5.5 family, six days after the flagship Opus 5.5. The small Haiku 5.5 followed on 7 October.

Anthropic’s division of labour: Opus 5.5 is for complex work that needs sustained judgment; Sonnet 5.5 is for well-defined everyday tasks such as fixing bugs and producing documents, slides and spreadsheets.

API price per million tokens (in / out)
$2 / $10
Output speed vs Sonnet 5
30%+ faster
Cost per task vs Sonnet 5
up to −30%
Cache-read price from 7 Oct
$0.10

Sources: Anthropic, “Introducing Claude Sonnet 5.5”; Claude API pricing page. For comparison, Opus 5.5 is $4 / $20 and Fable 5.1 is $10 / $50.

The price per token did not change; the saving comes from using fewer tokens. Anthropic says Sonnet 5.5 needs far fewer tokens for the same job, and in its testing costs up to 30% less per task than Sonnet 5. On 7 October Anthropic also halved its cache-read price from $0.20 to $0.10 per million tokens, which it says makes most agent tasks about 20% cheaper again.

How much better than Sonnet 5?

The headline result is Terminal-Bench 4.0, which asks a model to complete chains of tasks in a command line on its own: configuring servers, fixing programs, processing data. Sonnet 5 was weak here. Sonnet 5.5 tops every model Anthropic lists, including the more expensive Opus 5.5 and Fable 5.1.

Terminal-Bench 4.0: completing complex tasks in the command lineShare of tasks completed, higher is better; each model's best score
  • Claude Sonnet 5.570.6%
  • Claude Opus 5.566.4%
  • GPT-6 Astra57.9%
  • Claude Fable 5.155.8%
  • Claude Opus 552.3%
  • GPT-5.6 Sol37.3%
  • Claude Sonnet 510.3%

GPT scores are as reported by OpenAI and quoted by Anthropic. OpenAI has not published a GPT-6 Sol score on this test. Sources: Anthropic, “Introducing Claude Sonnet 5.5” (28 Sep 2026) and “Introducing Claude Opus 5.5” (22 Sep 2026)

On other tests it sits close to Opus 5.5:

Test What it measures Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
GDPval-AA v2.1 Real work from 44 occupations (Elo) 1844 1449 1846 1487
AA-Briefcase v1.1 Long-running knowledge work (Elo) 1811 1359 1822 1483
OSWorld 2.1 Using desktop software like a person 80.1% 57.0% 81.8% —
Humanity’s Last Exam (with tools) Expert-written questions 64.5% 54.9% 67.7% —
CursorBench 4.0 Vague requests from real coding sessions 55.5% 34.1% 57.8% —
FrontierCode 1.1 Whether code changes could be merged as is 52.1% 42.4% 54.4% 49.3%
Chartography Reading charts (no tools) 61.6% 15.6% 64.4% 53.6%

Figures from Anthropic’s announcement; GDPval-AA and AA-Briefcase are run by the independent firm Artificial Analysis. “—” means not reported.

For context: on GDPval-AA the more expensive Fable 5.1 scored 1735 and OpenAI’s GPT-6 Astra 1542 (figures from the Opus 5.5 announcement). On this office-work test, Sonnet 5.5 beats Fable 5.1 by about 100 points at a fifth of the per-token price.

What a task actually costs

Anthropic published each model’s score against the real cost per task at every effort level, using actual API prices. A representative selection:

Test Model (effort) Score Cost per task
FrontierCode 1.1 Sonnet 5 (High) 39.4% $6.10
Sonnet 5.5 (High, API default) 49.4% $0.42
GPT-6 Sol (Max, its best) 49.3% $2.07
Opus 5.5 (Medium) 54.6% $0.80
CursorBench 4.0 Sonnet 5 (Max, its best) 34.1% $7.17
Sonnet 5.5 (Low) 35.8% $0.50
Sonnet 5.5 (Max) 55.5% $9.67
Opus 5.5 (High) 56.0% $3.97
Terminal-Bench 4.0 Sonnet 5 (Max) 10.3% $11.62
Sonnet 5.5 (Medium, app default) 28.8% $0.83
Sonnet 5.5 (Max) 70.6% $12.54
Opus 5.5 (Xhigh) 66.4% $7.35

Data from the cost charts in Anthropic’s announcement, at the launch cache-read price of $0.20. After the 7 October cut, Anthropic recalculated Terminal-Bench in its Haiku 5.5 announcement: Sonnet 5.5’s cost fell by roughly 20–26% at each level, for example from $14.15 to $10.44 at Max (the two announcements’ baseline figures differ slightly).

Three things stand out:

  • The same score for an order of magnitude less. On CursorBench, Sonnet 5.5 at its lowest setting beats Sonnet 5’s best score for $0.50 against $7.17, about 1/14 of the cost. On FrontierCode at the same High setting, it scores 10 points more for about 1/15 of the cost.
  • Cheaper than OpenAI’s comparable model. On FrontierCode, Sonnet 5.5 at High matches GPT-6 Sol’s best score for about a fifth of the cost.
  • At maximum effort it is not cheaper than Opus. On CursorBench, Opus 5.5 reaches 56.0% at High for $3.97; Sonnet 5.5 needs Max to reach 55.5% and spends $9.67. Anthropic says Sonnet 5.5 is best value at low and medium settings; at high settings the two cost about the same.

“Up to 30% cheaper” and “about a tenth of the cost” are not contradictory: the first is the typical saving at the same settings; the second compares Sonnet 5.5 at low effort with Sonnet 5 at its highest.

Did it overtake GPT-6?

Some reports say Sonnet 5.5 overtook GPT-6 in the rankings. LMArena asks users to vote between answers from two anonymous models and turns the votes into Elo scores. In the data we collected on 6 October (boards updated 1–2 October), the answer depends on which GPT-6:

LMArena web development leaderboard (WebDev)Blind-vote Elo score, higher is better; effort setting in brackets
  • Claude Opus 5.5 (max)1815
  • GPT-6 Astra (max)1788
  • Claude Sonnet 5.5 (xhigh)1786
  • GPT-6.1 Sol (max)1758
  • Claude Fable 5.1 (max)1749
  • Claude Opus 5 (max)1695
  • GPT-6 Sol (max)1689

Sonnet 5.5's margin of error is ±18 points, so its 2-point gap to GPT-6 Astra is a tie. Source: LMArena WebDev leaderboard, updated 1 Oct 2026

  • Against the standard GPT-6 Sol: yes. About 100 points ahead on web development, and 1471 to 1457 on the overall text leaderboard.
  • Against GPT-6 Astra and GPT-6.1 Sol (released 29 September): no. A 2-point tie with Astra on web development; behind Astra (1477) and 6.1 Sol (1483) on overall text.
  • Not at the top. The overall text leader is Google’s Gemini 4 Argon (1525, preliminary); Opus 5.5 scores 1504.

What early testers report

These examples come from partner companies quoted in Anthropic’s announcement and have not been independently checked:

  • Building apps. Across 118 real app builds, Base44 found Sonnet 5.5’s apps scored level with Opus 5’s, taking 3.6 rounds of revision per app on average against Opus 5’s 7.7.
  • Finance. On 2,441 finance tasks, the hedge fund Balyasny found Sonnet 5.5 scored above Sonnet 5 using about 121,000 tokens per answer, against 497,000.
  • Customer support. Zendesk ran hundreds of real tickets and found them processed 20% faster, with fewer wrong decisions.
  • Slides. Anthropic gave it a listed company’s quarterly results materials, call transcripts and a slide template and asked for a 10-slide operating review. Two experts judged the first draft ready to send.

Things to keep in mind

  • The benchmark numbers are the vendor’s. All test and cost figures come from Anthropic; GPT scores are OpenAI’s own numbers as quoted by Anthropic, under different conditions. There is no like-for-like comparison with GPT-6.1 Sol, released the day after.
  • Anthropic says Opus 5.5 is still clearly stronger at complex, open-ended work that needs sustained judgment, in its own testing and external testers’. Close scores do not mean equal ability.
  • Maximum effort is not always best. On FrontierCode, Sonnet 5.5 scored 46.2% at Max but 52.1% at Xhigh. Anthropic says at Max it more often split code review across many sub-agents, leading to timeouts or out-of-scope edits.
  • Tighter safeguards. Because its cybersecurity ability is close to Opus 5’s, it is the first Sonnet model with cyber safeguards: higher-risk security requests are handed to Sonnet 5, while routine bug finding and fixing is unaffected.

How to use it

  • Everyday users. Choose it in the model menu on claude.ai and the desktop and mobile apps, where effort defaults to Medium. Claude’s pricing page lists Sonnet-tier models on every individual plan from Free to Max; Pro is $20 a month ($17 a month billed annually).
  • Coding. Available in Claude Code. For well-defined work, such as changing a button or fixing a bug you have already located, it is faster and lighter on usage limits than Opus 5.5. For designing the overall architecture, Anthropic and its testers suggest Opus 5.5.
  • Developers. The model ID is claude-sonnet-5-5, also available on Amazon’s, Google’s and Microsoft’s clouds. Code that turned thinking off with thinking.type: disabled must switch to the new between_tools setting, as the migration guide explains; changing the model name is not enough.

Model summary

Developer
Anthropic
API price
$2 / $10 per million input / output tokens

Data source: Anthropic announcement

Sources