AI news, models and products, with sources中文

Overall (text)LMArena ↗

#ModelDeveloperScoreVotes
1gemini-4-argon-highPrelim.Google1525 ±94,932
2claude-opus-4-6-highAnthropic1505 ±377,636
3claude-fable-5-highAnthropic1504 ±438,387
4claude-opus-5.5-highAnthropic1504 ±94,552
5claude-opus-4-7-highAnthropic1501 ±464,946
6claude-fable-5.1-maxAnthropic1501 ±611,800
7claude-opus-4-6Anthropic1497 ±382,189
8gemini-3.8-flash-highPrelim.Google1495 ±526,298
9muse-spark-1.3-maxMeta1494 ±612,343
10claude-opus-4-7Anthropic1494 ±466,058

As of Oct 2, 2026 · 8,626,731 votes · 413 models

Open-weight models (text)LMArena ↗

#ModelDeveloperScoreVotes
1kimi-k3-maxOpenMoonshot1488 ±528,280
2mimo-v2.6-proOpenXiaomi1480 ±94,056
3glm-5.3-maxOpenZ.ai1478 ±617,857
4glm-5.2-maxOpenZ.ai1476 ±445,399
5deepseek-v4.1-flash-maxOpenDeepSeek1474 ±78,728
6glm-5.3-flashOpenZ.ai1473 ±522,971
7mimo-v2.5-proOpenXiaomi1468 ±470,871
8glm-5.1OpenZ.ai1465 ±458,358
9deepseek-v4-pro-high-20260813OpenDeepSeek1464 ±79,780
10kimi-k2.6OpenMoonshot1461 ±439,936

As of Oct 2, 2026 · 8,626,731 votes · 413 models

CodingLMArena ↗

#ModelDeveloperScoreVotes
1gemini-4-argon-highPrelim.Google1560 ±171,230
2claude-fable-5-highAnthropic1552 ±710,038
3claude-opus-4-6-highAnthropic1551 ±620,235
4claude-opus-4-7-highAnthropic1551 ±618,521
5claude-opus-4-6Anthropic1548 ±522,901
6claude-opus-4-7Anthropic1546 ±618,728
7gpt-6-astra-maxOpenAI1543 ±132,135
8gpt-6.1-sol-maxOpenAI1542 ±24609
9kimi-k3-maxOpenMoonshot1541 ±87,285
10mimo-v2.6-proOpenXiaomi1540 ±181,103

As of Oct 2, 2026 · 1,848,296 votes · 408 models

Building web appsLMArena ↗

#ModelDeveloperScoreVotes
1claude-opus-5.5-maxAnthropic1815 +16/-162,062
2gpt-6-astra-maxOpenAI1788 +10/-106,123
3claude-sonnet-5.5-xhighAnthropic1786 +18/-181,531
4gpt-6.1-sol-maxOpenAI1758 +17/-171,620
5claude-fable-5.1-maxAnthropic1749 +10/-106,318
6claude-sonnet-5.5-highAnthropic1715 +13/-132,527
7claude-opus-5-maxAnthropic1695 +6/-616,954
8gpt-6-sol-maxOpenAI1689 +11/-113,411
9gemini-4-argon-highPrelim.Google1680 +13/-132,422
10qwen3.8-maxPrelim.Alibaba1671 +12/-123,454

As of Oct 1, 2026 · 838,388 votes · 138 models

Understanding imagesLMArena ↗

#ModelDeveloperScoreVotes
1claude-fable-5-highAnthropic1309 ±713,709
2qwen3.8-maxAlibaba1301 ±710,040
3claude-opus-4-6-highAnthropic1299 ±622,328
4claude-opus-4-7Anthropic1298 ±723,169
5claude-opus-4-7-highAnthropic1298 ±722,755
6gemini-3.7-flash-highPrelim.Google1296 ±103,814
7claude-opus-4-6Anthropic1294 ±626,869
8muse-sparkMeta1294 ±95,868
9muse-spark-1.2 (xHigh)Meta1293 ±141,951
10gpt-6.1-sol-maxOpenAI1291 ±161,408

As of Oct 2, 2026 · 1,436,923 votes · 159 models

Image generationLMArena ↗

#ModelDeveloperScoreVotes
1gpt-image-2.5-sunburstPrelim.OpenAI1425 ±716,652
2gpt-image-2.5-flarePrelim.OpenAI1397 ±715,792
3gpt-image-2 (medium)OpenAI1384 ±492,549
4grok-imagine-image-2.0 (canvas)Prelim.SpaceXAI1336 ±123,070
5mai-image-2.6Microsoft AI1333 ±523,918
6reve-2.1Reve1302 ±88,250
7grok-imagine-image-2.0 (20260801)SpaceXAI1297 ±613,278
8muse-imageMeta1274 ±541,071
9reve-2.0Reve1269 ±615,774
10gemini-3.1-flash-image (nano-banana-2) [web-search]Google1261 ±459,301

As of Oct 5, 2026 · 6,583,920 votes · 81 models

Image editingLMArena ↗

#ModelDeveloperScoreVotes
1gpt-image-2.5-sunburstPrelim.OpenAI1524 ±558,172
2gpt-image-2.5-flarePrelim.OpenAI1481 ±555,745
3gpt-image-2 (medium)OpenAI1462 ±3304,260
4grok-imagine-image-2.0 (canvas)Prelim.SpaceXAI1437 ±910,821
5mai-image-2.6Microsoft AI1427 ±550,147
6grok-imagine-image-2.0 (20260801)SpaceXAI1426 ±548,559
7muse-imageMeta1403 ±4141,017
8mai-image-2.5Microsoft AI1401 ±4190,078
9seedream-5.0-proBytedance1394 ±3281,170
10grok-imagine-image-quality (20260519)SpaceXAI1391 ±638,195

As of Oct 5, 2026 · 30,962,085 votes · 58 models

Video generationLMArena ↗

#ModelDeveloperScoreVotes
1gemini-omni-1.1-flashGoogle1516 ±151,784
2gemini-omni-flashGoogle1513 ±926,576
3flux-3-videoPrelim.Black Forest Labs1493 ±171,302
4grok-imagine-video-1.5-agentPrelim.SpaceXAI1492 ±181,218
5dreamina-seedance-2.0-720pBytedance1479 ±856,318
6wan3.0Alibaba1476 ±132,598
7dreamina-seedance-2.5-720pBytedance1474 ±98,001
8minimax-h3OpenMiniMax1460 ±911,417
9muse-videoMeta1456 ±152,188
10happyhorse-1.0Alibaba-ATH1427 ±1322,264

As of Sep 21, 2026 · 718,577 votes · 48 models

Model releases

Figures from developers' official announcements or the cited reports. Scores are as published by the developer at launch. Test conditions differ between companies, so compare with care.

DateModelDeveloperAPI priceReported results (selected)
7 Oct 2026Claude Haiku 5.5 cuts prices 90% to $0.10 per million tokens and beats the previous Sonnet on office workAnthropic$0.10 / $0.50 per million input / output tokens (prompts up to 100k; $0.50 / $2.50 above)—
30 Sept 2026Google announces Gemini 4 ArgonGoogle DeepMindIntroductory $2 / $10 per million input / output tokens, later $4 / $20
DeepSWE v1.1: 77.9%
AutomationBench: 51.3%
28 Sept 2026Claude Sonnet 5.5 nears Opus 5.5 at half the price, scoring 70.6% on a coding test where Sonnet 5 managed 10.3%Anthropic$2 / $10 per million input / output tokens—
22 Sept 2026Claude Opus 5.5Anthropic$4 / $20 per million input / output tokens; cache reads $0.20
Terminal-Bench 4.0: 66.4%
Humanity's Last Exam (with tools): 67.7%
3 Sept 2026OpenAI releases GPT-6 AstraOpenAI—
ARC-AGI-3: 99.9%
1 Sept 2026Claude Fable 5.1 and Mythos 5.1Anthropic$10 / $50 per million input / output tokens; cache reads $0.25 (75% lower)—
24 Jul 2026Claude Opus 5Anthropic$5 / $25 per million input / output tokens—
9 Jun 2026Claude Fable 5 and Mythos 5Anthropic$10 / $50 per million input / output tokens—
24 Apr 2026DeepSeek V4: open weights with a 1-million-token contextDeepSeek—V4-Pro: 1.6T total / 49B active parameters. V4-Flash: 284B total / 13B active. 1M-token context.
5 Feb 2026Claude Opus 4.6 with a 1-million-token contextAnthropic$5 / $25 per million input / output tokens1M-token context window (beta)
24 Nov 2025Claude Opus 4.5, at a lower priceAnthropic$5 / $25 per million input / output tokens—
18 Nov 2025Google releases Gemini 3Google DeepMind—
Humanity's Last Exam (no tools): 37.5%
GPQA Diamond: 91.9%
29 Sept 2025Claude Sonnet 4.5Anthropic—
SWE-bench Verified: 77.2%
OSWorld: 61.4%
7 Aug 2025OpenAI releases GPT-5OpenAI—
SWE-bench Verified: 74.9%
AIME 2025 (no tools): 94.6%
22 May 2025Anthropic releases Claude 4Anthropic—
SWE-bench Verified: 72.5%
Terminal-bench: 43.2%
5 Apr 2025Meta releases Llama 4Meta—Scout: 17B active parameters, 16 experts, 10M-token context. Maverick: 17B active, 128 experts.
25 Mar 2025Google releases Gemini 2.5, its first "thinking" modelGoogle DeepMind—
Humanity's Last Exam (no tools): 18.8%
24 Feb 2025Claude 3.7 Sonnet and Claude CodeAnthropic$3 / $15 per million input / output tokens
SWE-bench Verified: 63.7%
20 Jan 2025DeepSeek releases R1, an open reasoning modelDeepSeek—
AIME 2024 (pass@1): 79.8%
MATH-500 (pass@1): 97.3%