AI news, models and products, with sources中文
News Models29 Sept 2025

Claude Sonnet 4.5

Anthropic says Sonnet 4.5 can stay focused on complex tasks for more than 30 hours, and leads coding and computer-use benchmarks at launch.

Model summary

Developer
Anthropic

Reported benchmark results

SWE-bench Verified77.2%
OSWorld61.4%Sonnet 4 four months earlier: 42.2%

Data source: Anthropic announcement. Scores are as published by the developer at launch. Test conditions differ between companies, so compare with care.

Sources