AI news, models and products, with sources中文
Guide Basics · Chapter 2

Models, products and companies

How does ChatGPT relate to GPT-6? Which is bigger, Sonnet or Opus? What are open-weight models for, and how should you read a leaderboard? The names in AI, sorted out.

A model is not a product

A model is the trained “brain”, in essence a huge file of numbers, such as Claude Opus 5.5 or GPT-6 Astra. A product is what you actually open, such as the Claude app or the ChatGPT website. Products add a lot around the model:

  • the chat interface, mobile apps and voice;
  • tools: web search, reading files, running code, making images, operating a computer;
  • memory, projects, and connections to your email and cloud storage;
  • plans, usage limits, privacy rules and regional availability.

So “Claude” is both the product and the model family; choosing “Opus 5.5” inside the Claude app is choosing a model. ChatGPT works the same way. According to OpenAI’s release notes for October 2026, everyday chat on paid ChatGPT plans (Plus, Pro, Business, Enterprise) uses GPT-6 Sol, Free and Go use GPT-6 Luna, and the Pro reasoning option uses the most capable model, GPT-6 Astra (see ChatGPT brings GPT-6 and interactive answers to more users). Two people both “using ChatGPT” may be using quite different models.

Why companies offer several tiers

Like small, medium and large coffees, AI companies sell models by capability and price. Small models are quick and cheap, good for lots of simple tasks; big ones are stronger but slower and costlier. In the names, the number usually marks the generation (5.5, 6) and the word marks the tier.

The main tiers from the three big US labs as of October 2026 (developer API prices in US dollars per million input / output tokens; subscribers don’t pay these directly):

Company Small and fast Everyday workhorse Flagship and above
Anthropic Haiku 5.5 (from 0.10 / 0.50) Sonnet 5.5 (2 / 10) Opus 5.5 (4 / 20); above it, Fable 5.1 (10 / 50)
OpenAI GPT-6 Luna (0.10 / 0.50) GPT-6.1 Sol (2 / 10) GPT-6 Astra (10 / 50)
Google Gemini 3.5 Flash-Lite (0.30 / 2.50) Gemini 3.8 Flash (0.75 / 3.75, doubling from 2027) Gemini 4 Argon (introductory 2 / 10, then 4 / 20; not yet public)

Some notes:

  • Anthropic goes Haiku, Sonnet, Opus from small to large. In 2026 it added a higher tier: Fable and Mythos are the same model, with Fable open to the public under extra safeguards and Mythos, with fewer restrictions, only for vetted cybersecurity and life-science organisations. Anthropic recommends starting with Opus 5.5 for most work.
  • OpenAI names its tiers Luna, Terra, Sol and Astra, from smallest to largest. As of October, its official models page lists three GPT-6 generation models, Luna, 6.1 Sol and Astra, and no Terra.
  • Google has long used Flash-Lite and Flash (fast) and Pro (strong), plus Nano, which runs on phones. Gemini 4 Argon, announced in September 2026, is the first model under a new naming scheme replacing names like “Pro”. As of early October only a small group of testers can use it; the Gemini app still runs on 3.x models.
  • Chinese companies do the same: DeepSeek V4 comes in Pro and Flash; Alibaba’s Qwen has open smaller models and non-public Max and Plus versions.
API output price per million tokens (US$)As of October 2026; shorter is cheaper
  • GPT-6 Luna$0.50
  • Claude Haiku 5.5from $0.50
  • Gemini 3.8 Flash$3.75
  • Claude Sonnet 5.5$10
  • GPT-6.1 Sol$10
  • Claude Opus 5.5$20
  • Claude Fable 5.1$50
  • GPT-6 Astra$50

Haiku 5.5 costs more for prompts over 100,000 tokens; the Gemini 3.8 Flash price applies until the end of 2026. Prices change often; check the official pages. Sources: Anthropic, OpenAI and Google official pricing and model pages, checked 8 Oct 2026

A higher tier isn’t automatically better. A new generation’s middle tier often catches the previous top tier: Anthropic says Opus 5.5 matches Fable 5.1 on most work at a much lower price, and in Artificial Analysis’s independent command-line test Sonnet 5.5 (64%) scored above Opus 5.5 (60%). Besides choosing a tier, many models let you adjust how hard they think, another way to trade speed for quality (see the previous chapter).

The main companies and products

Company Chat product Model family Country
OpenAI ChatGPT GPT (GPT-6 Astra / Sol / Luna) US
Anthropic Claude Claude (Haiku / Sonnet / Opus / Fable) US
Google Gemini Gemini (Flash family, Argon) US
Meta Meta AI Muse (previously Llama) US
SpaceXAI (formerly xAI) Grok Grok US
DeepSeek DeepSeek DeepSeek V4 / V4.1 China
Alibaba Qwen Qwen China
Moonshot AI Kimi Kimi K3 China
ByteDance Doubao Doubao China
Mistral AI Mistral Vibe (formerly Le Chat) Mistral France

They differ hugely in scale: ChatGPT reached 900 million weekly users in February 2026, while Doubao, China’s most popular chatbot, had 330 million users by May 2026. Many other products don’t train their own base models but build on others’, including many of the writing, coding and office tools on our Products page.

Open-weight and closed models

Closed models, such as GPT-6, Claude and Gemini, can only be used through the company’s own apps or API. You never see the model itself, and every request goes to that company’s servers.

Open-weight models have their trained model files published, so anyone can download them, run them on their own computers or servers, and modify them. When DeepSeek-R1 was released under the permissive MIT licence in January 2025, open models drew worldwide attention. The main ones as of October 2026:

  • DeepSeek V4 / V4.1: MIT licence, commercial use and modification allowed.
  • Kimi K3 (Moonshot AI, July 2026): 2.8 trillion parameters, the largest open-weight model. But its licence requires companies with annual revenue over $20 million to sign a contract with Moonshot before offering it to customers as a service.
  • Qwen3.8 (Alibaba): 2.4 trillion parameters. Its larger models require revenue sharing from providers earning over $50 million a year; a distilled 27-billion-parameter version uses the more permissive Apache licence.
  • Llama (Meta): once the best-known open model, but its licence restricted use by very large companies, and some researchers argued it wasn’t truly “open source”. Since April 2026 Meta’s chat products have used its new Muse models instead, which LMArena lists as proprietary.

What it means for you

What you care about Closed models Open-weight models
Capability The strongest models today are closed Close behind
Cost Pay the official price Run your own, or pick a cheaper provider
Privacy Your data goes to the company Run it yourself and data stays on your servers
Effort Sign up and go Big models need expensive hardware and skills
Safety limits Controlled by the company Others can modify or even remove them

Note that using a company’s official app is different from running its open model yourself. DeepSeek’s official app follows Chinese censorship rules, and the US Department of Defense and the Australian government, among others, have barred DeepSeek products from government devices. But if you download the model and run it yourself, no data goes to DeepSeek’s servers.

How big is the gap? On LMArena’s overall text leaderboard (data as of 2 October 2026), the top closed model, Gemini 4 Argon, scored 1525 (a preliminary result with few votes so far) and the top open model, Kimi K3, scored 1488. According to Fortune, DeepSeek’s own V4 technical report estimated that it trails the frontier by about three to six months.

How to read benchmarks and leaderboards

Almost every model launch comes with a table showing the company in first place. When you read these numbers, separate two kinds:

Scores published by the developer. The company that made the model ran the tests, often chose tests it does well on, and didn’t always use the same conditions as its rivals. Some examples we’ve reported:

  • Different companies report different scores for the same test: Anthropic gives Opus 5.5 40.0% on AutomationBench, while Google’s comparison table gives it 42.5%.
  • Tools and conditions matter enormously: OpenAI reported GPT-6 Astra at 99.9% on ARC-AGI-3 with an enhanced set of helper tools, but 66% under standard conditions.
  • How you count matters: when xAI published a comparison chart for Grok 3 in 2025, an OpenAI employee pointed out that Grok 3’s score was the most common answer from 64 attempts, while the rival was shown with single attempts.
  • Developers admit the limits: in its Opus 5.5 announcement, Anthropic said that at this level of capability, gaps in benchmark scores no longer reliably reflect real differences.

Independent tests. A third party with no stake in the model tests everything the same way, so results compare better. On Artificial Analysis’s overall Intelligence Index, for example, Gemini 4 Argon and GPT-6 Astra both scored 53, with GPT-6.1 Sol on 52.

For any score, ask four questions: Who ran the test? What task does it measure? Compared with whom, under what conditions? And how much is that task like what I need to do?

How to choose

There is no “best AI”, only one that suits a particular job better. Some practical rules:

  1. Start with what you already have. For everyday questions, writing and translation, the free versions of the main products are usually enough. Switch when a kind of task clearly isn’t working.
  2. Test with your own real tasks. Pick three to five things you do often (a tricky email, a report to digest, a spreadsheet), give the same instructions to two or three products, and see which needs the least rework. That tells you more than any leaderboard.
  3. Start in the middle tier and move up when needed. Use the default or mid-tier model day to day; switch to a flagship or raise the thinking effort for complex coding, long analysis or careful reasoning.
  4. Consider what you already use. If you live in Gmail and Google Docs, Gemini fits in smoothly; if your company uses Microsoft Office, Copilot may already be included.
  5. Check privacy and availability. Before uploading work documents, check your company’s rules and the product’s data policy (see Privacy and safety); some features aren’t available in every country.
  6. Do the maths. Subscription or pay-as-you-go, and whether the limits are enough: see How AI is priced.
What you want to do Where to start
Everyday questions and writing Whatever you already have: ChatGPT, Claude, Gemini, DeepSeek and others
Research with sources Products with web search and citations, such as Perplexity
Coding and websites Claude Code, Codex, Cursor, plus LMArena’s coding and web development boards
Data that must stay in-house Business plans, or an open-weight model on your own servers

For more products by use, see the Products page; for the latest model rankings, see the Models page.

Sources