Which AI model is best?
These rankings come from LMArena: users ask two anonymous models the same question and pick the better answer, and scores are calculated from millions of votes. They reflect real users' preferences, not developers' own claims.
Overall (text)LMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | gemini-4-argon-highPrelim. | 1525 ±9 | 4,932 | |
| 2 | claude-opus-4-6-high | Anthropic | 1505 ±3 | 77,636 |
| 3 | claude-fable-5-high | Anthropic | 1504 ±4 | 38,387 |
| 4 | claude-opus-5.5-high | Anthropic | 1504 ±9 | 4,552 |
| 5 | claude-opus-4-7-high | Anthropic | 1501 ±4 | 64,946 |
| 6 | claude-fable-5.1-max | Anthropic | 1501 ±6 | 11,800 |
| 7 | claude-opus-4-6 | Anthropic | 1497 ±3 | 82,189 |
| 8 | gemini-3.8-flash-highPrelim. | 1495 ±5 | 26,298 | |
| 9 | muse-spark-1.3-max | Meta | 1494 ±6 | 12,343 |
| 10 | claude-opus-4-7 | Anthropic | 1494 ±4 | 66,058 |
Open-weight models (text)LMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | kimi-k3-maxOpen | Moonshot | 1488 ±5 | 28,280 |
| 2 | mimo-v2.6-proOpen | Xiaomi | 1480 ±9 | 4,056 |
| 3 | glm-5.3-maxOpen | Z.ai | 1478 ±6 | 17,857 |
| 4 | glm-5.2-maxOpen | Z.ai | 1476 ±4 | 45,399 |
| 5 | deepseek-v4.1-flash-maxOpen | DeepSeek | 1474 ±7 | 8,728 |
| 6 | glm-5.3-flashOpen | Z.ai | 1473 ±5 | 22,971 |
| 7 | mimo-v2.5-proOpen | Xiaomi | 1468 ±4 | 70,871 |
| 8 | glm-5.1Open | Z.ai | 1465 ±4 | 58,358 |
| 9 | deepseek-v4-pro-high-20260813Open | DeepSeek | 1464 ±7 | 9,780 |
| 10 | kimi-k2.6Open | Moonshot | 1461 ±4 | 39,936 |
CodingLMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | gemini-4-argon-highPrelim. | 1560 ±17 | 1,230 | |
| 2 | claude-fable-5-high | Anthropic | 1552 ±7 | 10,038 |
| 3 | claude-opus-4-6-high | Anthropic | 1551 ±6 | 20,235 |
| 4 | claude-opus-4-7-high | Anthropic | 1551 ±6 | 18,521 |
| 5 | claude-opus-4-6 | Anthropic | 1548 ±5 | 22,901 |
| 6 | claude-opus-4-7 | Anthropic | 1546 ±6 | 18,728 |
| 7 | gpt-6-astra-max | OpenAI | 1543 ±13 | 2,135 |
| 8 | gpt-6.1-sol-max | OpenAI | 1542 ±24 | 609 |
| 9 | kimi-k3-maxOpen | Moonshot | 1541 ±8 | 7,285 |
| 10 | mimo-v2.6-proOpen | Xiaomi | 1540 ±18 | 1,103 |
Building web appsLMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | claude-opus-5.5-max | Anthropic | 1815 +16/-16 | 2,062 |
| 2 | gpt-6-astra-max | OpenAI | 1788 +10/-10 | 6,123 |
| 3 | claude-sonnet-5.5-xhigh | Anthropic | 1786 +18/-18 | 1,531 |
| 4 | gpt-6.1-sol-max | OpenAI | 1758 +17/-17 | 1,620 |
| 5 | claude-fable-5.1-max | Anthropic | 1749 +10/-10 | 6,318 |
| 6 | claude-sonnet-5.5-high | Anthropic | 1715 +13/-13 | 2,527 |
| 7 | claude-opus-5-max | Anthropic | 1695 +6/-6 | 16,954 |
| 8 | gpt-6-sol-max | OpenAI | 1689 +11/-11 | 3,411 |
| 9 | gemini-4-argon-highPrelim. | 1680 +13/-13 | 2,422 | |
| 10 | qwen3.8-maxPrelim. | Alibaba | 1671 +12/-12 | 3,454 |
Understanding imagesLMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | claude-fable-5-high | Anthropic | 1309 ±7 | 13,709 |
| 2 | qwen3.8-max | Alibaba | 1301 ±7 | 10,040 |
| 3 | claude-opus-4-6-high | Anthropic | 1299 ±6 | 22,328 |
| 4 | claude-opus-4-7 | Anthropic | 1298 ±7 | 23,169 |
| 5 | claude-opus-4-7-high | Anthropic | 1298 ±7 | 22,755 |
| 6 | gemini-3.7-flash-highPrelim. | 1296 ±10 | 3,814 | |
| 7 | claude-opus-4-6 | Anthropic | 1294 ±6 | 26,869 |
| 8 | muse-spark | Meta | 1294 ±9 | 5,868 |
| 9 | muse-spark-1.2 (xHigh) | Meta | 1293 ±14 | 1,951 |
| 10 | gpt-6.1-sol-max | OpenAI | 1291 ±16 | 1,408 |
Web searchLMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | gpt-5.6-sol-xhigh | OpenAI | 1257 ±7 | 29,663 |
| 2 | claude-opus-4-6-search | Anthropic | 1253 ±5 | 134,699 |
| 3 | gpt-5.5-search | OpenAI | 1242 ±5 | 89,873 |
| 4 | claude-opus-4-7 | Anthropic | 1233 ±5 | 91,394 |
| 5 | claude-fable-5-high | Anthropic | 1230 ±8 | 41,795 |
| 6 | ernie-5.1 | Baidu | 1227 ±10 | 3,788 |
| 7 | claude-sonnet-4-6-search | Anthropic | 1221 ±5 | 134,905 |
| 8 | grok-4.5 | SpaceXAI | 1213 ±7 | 31,505 |
| 9 | gemini-3.1-pro-grounding | 1210 ±5 | 113,282 | |
| 10 | gemini-3-pro-grounding | 1207 ±6 | 37,024 |
Image generationLMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | gpt-image-2.5-sunburstPrelim. | OpenAI | 1425 ±7 | 16,652 |
| 2 | gpt-image-2.5-flarePrelim. | OpenAI | 1397 ±7 | 15,792 |
| 3 | gpt-image-2 (medium) | OpenAI | 1384 ±4 | 92,549 |
| 4 | grok-imagine-image-2.0 (canvas)Prelim. | SpaceXAI | 1336 ±12 | 3,070 |
| 5 | mai-image-2.6 | Microsoft AI | 1333 ±5 | 23,918 |
| 6 | reve-2.1 | Reve | 1302 ±8 | 8,250 |
| 7 | grok-imagine-image-2.0 (20260801) | SpaceXAI | 1297 ±6 | 13,278 |
| 8 | muse-image | Meta | 1274 ±5 | 41,071 |
| 9 | reve-2.0 | Reve | 1269 ±6 | 15,774 |
| 10 | gemini-3.1-flash-image (nano-banana-2) [web-search] | 1261 ±4 | 59,301 |
Image editingLMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | gpt-image-2.5-sunburstPrelim. | OpenAI | 1524 ±5 | 58,172 |
| 2 | gpt-image-2.5-flarePrelim. | OpenAI | 1481 ±5 | 55,745 |
| 3 | gpt-image-2 (medium) | OpenAI | 1462 ±3 | 304,260 |
| 4 | grok-imagine-image-2.0 (canvas)Prelim. | SpaceXAI | 1437 ±9 | 10,821 |
| 5 | mai-image-2.6 | Microsoft AI | 1427 ±5 | 50,147 |
| 6 | grok-imagine-image-2.0 (20260801) | SpaceXAI | 1426 ±5 | 48,559 |
| 7 | muse-image | Meta | 1403 ±4 | 141,017 |
| 8 | mai-image-2.5 | Microsoft AI | 1401 ±4 | 190,078 |
| 9 | seedream-5.0-pro | Bytedance | 1394 ±3 | 281,170 |
| 10 | grok-imagine-image-quality (20260519) | SpaceXAI | 1391 ±6 | 38,195 |
Video generationLMArena ↗
| # | Model | Developer | Score | Votes |
|---|---|---|---|---|
| 1 | gemini-omni-1.1-flash | 1516 ±15 | 1,784 | |
| 2 | gemini-omni-flash | 1513 ±9 | 26,576 | |
| 3 | flux-3-videoPrelim. | Black Forest Labs | 1493 ±17 | 1,302 |
| 4 | grok-imagine-video-1.5-agentPrelim. | SpaceXAI | 1492 ±18 | 1,218 |
| 5 | dreamina-seedance-2.0-720p | Bytedance | 1479 ±8 | 56,318 |
| 6 | wan3.0 | Alibaba | 1476 ±13 | 2,598 |
| 7 | dreamina-seedance-2.5-720p | Bytedance | 1474 ±9 | 8,001 |
| 8 | minimax-h3Open | MiniMax | 1460 ±9 | 11,417 |
| 9 | muse-video | Meta | 1456 ±15 | 2,188 |
| 10 | happyhorse-1.0 | Alibaba-ATH | 1427 ±13 | 22,264 |
Model releases
| Date | Model | Developer | API price | Reported results (selected) |
|---|---|---|---|---|
| 7 Oct 2026 | Claude Haiku 5.5 cuts prices 90% to $0.10 per million tokens and beats the previous Sonnet on office work | Anthropic | $0.10 / $0.50 per million input / output tokens (prompts up to 100k; $0.50 / $2.50 above) | — |
| 30 Sept 2026 | Google announces Gemini 4 Argon | Google DeepMind | Introductory $2 / $10 per million input / output tokens, later $4 / $20 | DeepSWE v1.1: 77.9% AutomationBench: 51.3% |
| 28 Sept 2026 | Claude Sonnet 5.5 nears Opus 5.5 at half the price, scoring 70.6% on a coding test where Sonnet 5 managed 10.3% | Anthropic | $2 / $10 per million input / output tokens | — |
| 22 Sept 2026 | Claude Opus 5.5 | Anthropic | $4 / $20 per million input / output tokens; cache reads $0.20 | Terminal-Bench 4.0: 66.4% Humanity's Last Exam (with tools): 67.7% |
| 3 Sept 2026 | OpenAI releases GPT-6 Astra | OpenAI | — | ARC-AGI-3: 99.9% |
| 1 Sept 2026 | Claude Fable 5.1 and Mythos 5.1 | Anthropic | $10 / $50 per million input / output tokens; cache reads $0.25 (75% lower) | — |
| 24 Jul 2026 | Claude Opus 5 | Anthropic | $5 / $25 per million input / output tokens | — |
| 9 Jun 2026 | Claude Fable 5 and Mythos 5 | Anthropic | $10 / $50 per million input / output tokens | — |
| 24 Apr 2026 | DeepSeek V4: open weights with a 1-million-token context | DeepSeek | — | V4-Pro: 1.6T total / 49B active parameters. V4-Flash: 284B total / 13B active. 1M-token context. |
| 5 Feb 2026 | Claude Opus 4.6 with a 1-million-token context | Anthropic | $5 / $25 per million input / output tokens | 1M-token context window (beta) |
| 24 Nov 2025 | Claude Opus 4.5, at a lower price | Anthropic | $5 / $25 per million input / output tokens | — |
| 18 Nov 2025 | Google releases Gemini 3 | Google DeepMind | — | Humanity's Last Exam (no tools): 37.5% GPQA Diamond: 91.9% |
| 29 Sept 2025 | Claude Sonnet 4.5 | Anthropic | — | SWE-bench Verified: 77.2% OSWorld: 61.4% |
| 7 Aug 2025 | OpenAI releases GPT-5 | OpenAI | — | SWE-bench Verified: 74.9% AIME 2025 (no tools): 94.6% |
| 22 May 2025 | Anthropic releases Claude 4 | Anthropic | — | SWE-bench Verified: 72.5% Terminal-bench: 43.2% |
| 5 Apr 2025 | Meta releases Llama 4 | Meta | — | Scout: 17B active parameters, 16 experts, 10M-token context. Maverick: 17B active, 128 experts. |
| 25 Mar 2025 | Google releases Gemini 2.5, its first "thinking" model | Google DeepMind | — | Humanity's Last Exam (no tools): 18.8% |
| 24 Feb 2025 | Claude 3.7 Sonnet and Claude Code | Anthropic | $3 / $15 per million input / output tokens | SWE-bench Verified: 63.7% |
| 20 Jan 2025 | DeepSeek releases R1, an open reasoning model | DeepSeek | — | AIME 2024 (pass@1): 79.8% MATH-500 (pass@1): 97.3% |