AI news, models and products, with sources中文
Guide Basics · Chapter 1

How AI works

Starting from "predict the next word": how models are trained, how much they can read at once, why they get things wrong, and how chatbots became agents that act for you.

A large language model: a machine that keeps writing

ChatGPT, Claude, Gemini and DeepSeek all have a large language model (LLM) at their core. What it does sounds simple: given some text, it predicts the most likely next small piece, writes it, adds it to the text, and predicts again, until the answer is done.

Your phone keyboard suggests “you” after “see”. A language model does the same kind of thing at a vastly larger scale. It has seen so much text (large parts of the public web, books and code) that it can continue not just a sentence but an email, an explanation or a program.

Its fluency doesn’t mean it understands the world the way a person does. It has learned extremely complex patterns linking words, facts and steps of reasoning. Those patterns are enough for a lot of genuinely useful work, and they also explain why it sometimes makes things up: its job is to produce text that looks right, not to check facts.

Almost every major language model is built on the Transformer architecture, published by Google researchers in 2017.

How a model is made

Making a model has two big phases: first it reads (pre-training), then it learns how to behave (post-training).

How a large language model is made
  1. 1Collect textVast amounts of web pages, books and code. DeepSeek says V4 was pre-trained on more than 32 trillion tokens.
  2. 2Pre-trainingGuess the next token again and again, adjusting parameters when wrong. The costly part: OpenAI says GPT-6 Astra used over 100,000 GPUs.
  3. 3Learn to answerExamples written by people teach it to follow instructions; then people rank its answers and those preferences train it further (RLHF).
  4. 4Reasoning trainingOn tasks with checkable answers, like maths and code, reinforcement learning rewards working that gets it right.
  5. 5Test and releaseInternal and outside testers probe for dangerous abilities and bad behaviour; safeguards are added before release.

Sources: InstructGPT paper (2022); DeepSeek-R1 and V4 technical materials; Fortune on GPT-6 Astra

Post-training makes a big difference. OpenAI’s 2022 InstructGPT research found that people preferred answers from a 1.3-billion-parameter model tuned with human feedback over those from the untuned 175-billion-parameter GPT-3. A lot of what makes today’s assistants pleasant to use comes from this step.

Training and using are different things

Training takes months and huge amounts of computing power, and afterwards the model is fixed. Each time you ask a question you are using that fixed model, which is called inference. Three things follow that people often get wrong:

  • It doesn’t get smarter by talking to you. Correct it and it will follow your correction in that conversation, but the model itself hasn’t changed. “Memory” features save notes about you and add them to later chats; that isn’t retraining.
  • Its knowledge has a cut-off date. A model only knows what was in its training data. Anthropic gives Claude Opus 5.5’s reliable knowledge cut-off as June 2026; OpenAI gives GPT-6 Astra’s as 30 April 2026. Newer events need web search.
  • Your chats may be used to train future models. That depends on the product and your settings; see Privacy and safety.

What happens when you send a message

From pressing send to seeing an answer
  1. 1Assemble the contextThe app combines system instructions, memory, the conversation so far, uploaded files and your message, and feeds it to the model as tokens.
  2. 2Think (optional)A reasoning model first writes a working draft to break down the problem and check its approach.
  3. 3Use tools (optional)If needed it searches the web, runs code or reads a file, and the results go back into the context.
  4. 4Write token by tokenThat's why the answer seems to stream onto your screen.
  5. 5Start again next turnWhen you reply, the whole conversation, including its last answer, is sent in again.

Based on Anthropic's context window documentation

That last step matters: the model doesn’t remember your previous message. The app sends the whole conversation again every time. The longer the chat, the more there is to process, which brings us to the next idea.

The context window: how much AI can see at once

The context window is the most a model can handle in one go, counting everything you send and everything it writes back. Anthropic calls it the model’s “working memory”: it is not the vast data the model was trained on, but what is in front of it in this conversation.

Figures from official documentation as of October 2026 (for developers using the API):

Model Context window Maximum output per response
Claude Opus 5.5 / Sonnet 5.5 / Haiku 5.5 / Fable 5.1 1M tokens 128K tokens
GPT-6 Astra / GPT-6.1 Sol / GPT-6 Luna 1.05M tokens 128K tokens
Gemini 3.8 Flash about 1.05M tokens (1,048,576) about 66K tokens (65,536)
DeepSeek V4 family 1M tokens —

Google offers a way to picture 1 million tokens: about 50,000 lines of code, eight average-length English novels, or the transcripts of more than 200 podcast episodes.

Why it matters

  • Past the limit, it forgets. In a very long chat the app drops or condenses the earliest material, and you may find it no longer follows a rule you set at the start.
  • More isn’t automatically better. Anthropic’s documentation warns that accuracy and recall fall as the amount of context grows. Rather than pasting a whole book, tell it which chapters matter.
  • It affects cost. Developers pay per token on every request, so long conversations get more expensive with each turn. See How AI is priced.

In practice: start a new chat for a new topic, restate key requirements in long chats, and when you upload a long file, say what question you’re trying to answer.

Why AI makes things up

When AI invents facts, cites papers that don’t exist or gives wrong numbers, that’s called hallucination. There are three layers to it:

  1. Its job is plausible text. To the model, a believable-looking name, page number or figure is not fundamentally different from a real one.
  2. Training rewards guessing. A 2025 paper, Why Language Models Hallucinate, compares models to students on a hard exam: most tests give no credit for “I don’t know”, so guessing scores better, and models learn to answer even when unsure.
  3. Knowledge has a cut-off and gaps. Where it never learned something, it often fills in with whatever seems closest.

Newer models are improving, but they vary a lot, and making fewer things up often comes with answering fewer questions. The independent firm Artificial Analysis measured how often models invent an answer when they don’t know, in September 2026:

AA-Omniscience hallucination rate: inventing answers when unsureLower is better; each model at its highest reasoning setting
  • Gemini 4 Argon15%
  • GPT-6 Astra51%
  • GPT-6.1 Sol54%

On the same test Gemini 4 Argon answered 50% of questions correctly and GPT-6 Astra 63%: the model that invents less also declines more often. Source: Artificial Analysis, 30 Sep 2026

Vendors’ own figures show progress too: OpenAI said GPT-5 with thinking gave an answer containing a factual error on 4.8% of real ChatGPT questions, against 22% for the earlier o3. Neither number is zero.

How to reduce hallucinations

  • Allow it to say “I don’t know”. Say so in your request; Anthropic’s documentation says this simple step can drastically reduce false information.
  • Give it the source material. Provide the document, page or data and ask it to answer only from that. This is the idea behind retrieval-augmented generation (RAG).
  • Ask for sources and open them. Ask for a quote or link behind each claim, and drop claims it can’t support.
  • Turn on web search, especially for recent events.
  • Check names, numbers, dates and quotations, and anything to do with health, money or the law.
  • Ask again in different words, or ask another AI. If the answers disagree, it’s probably guessing.

Reasoning models: thinking before answering

Early chat models started writing an answer the moment they got a question. In September 2024 OpenAI released o1, which first writes out a long internal chain of reasoning, breaking the problem down, trying things and checking, before it answers. In January 2025 DeepSeek-R1 published its weights and training method, showing this ability could be developed through reinforcement learning alone. Google’s Gemini 2.5 and Anthropic’s Claude 3.7 Sonnet followed, and reasoning became standard for frontier models.

By 2026 most leading models decide for themselves how long to think, and offer an effort setting:

  • GPT-6 Astra: low, medium, high, xhigh and max;
  • Gemini 3.8 Flash: low, medium and high;
  • DeepSeek V4: “non-think”, “high” and “max”;
  • Current Claude models use “adaptive thinking”, adjusting how much they think to the task.

More thinking usually means better results in maths, coding and multi-step analysis, but it is slower and costs more: thinking uses tokens, which developers pay for as output and which use up subscribers’ allowances faster. For a quick fact or rewording a sentence, a fast mode is fine; for hard calculations, code or long analysis, let it think.

Multimodal: images, sound and video

“Multimodal” means a model can handle more than text. GPT-4o, in May 2024, was a turning point: one model handling text, audio and images, enabling spoken conversations with near-human response times.

Where things stand now:

  • Reading images is standard. Anthropic says all current Claude models accept images; OpenAI says all its latest models accept text and images. Photograph a table, a screenshot or a homework problem and it can read it.
  • Video and audio depend on the model. Google’s Gemini 3.8 Flash accepts text, images, video, audio and PDFs directly. In October 2026 ChatGPT added audio uploads for paid users, to transcribe and summarise recordings.
  • Seeing is not the same as making. Images, video and speech are usually generated by specialised models, such as Google’s Nano Banana and Gemini Omni or OpenAI’s GPT-Image. Chat apps call them behind the scenes, so it feels like one AI doing everything.

Images and sound are converted into tokens too and take up context. Models also misread small print and complex charts, so check key figures.

From chatbot to agent

A chatbot can only tell you how to do something. An agent does it: it plans steps, uses tools, looks at the results and decides what to do next until the task is done.

The key steps:

  • Using tools. The model can decide “I need to search” or “run this code”; the app does it and hands back the result. Web search and deep research are built on this.
  • Connecting to your apps. In November 2024 Anthropic released the Model Context Protocol (MCP), an open standard for connecting AI to outside data and tools. Its documentation compares it to a USB-C port for AI applications; Claude, ChatGPT and others now support it.
  • Operating a computer. In October 2024 Claude learned to operate a computer: reading screenshots, moving the cursor, clicking and typing. OpenAI’s Operator followed, and in September 2026 GPT-6 Astra made computer use its headline feature.
  • Working in the background. Coding agents such as Codex and Claude Code read, change and test code on their own; OpenAI’s Dots, launched in September, has its own cloud computer and keeps working between conversations.

Progress is fast. In Anthropic’s OSWorld 2.1 test, which has AI use desktop software like a person, Claude Opus 5.5 completed 81.8% of tasks. That is the company’s own result; real life is messier.

When you use an agent: give it only the access and accounts the task needs; have it stop and ask you before paying, submitting a form, sending an email or deleting files; and keep an eye on what it’s doing.

Things to remember

  • AI writes the most likely next text. Fluent doesn’t mean correct.
  • Its knowledge stops at its training cut-off. New facts need search, and your corrections don’t change the model.
  • The context window has a limit and long chats forget. Start fresh for new topics and repeat key requirements.
  • Let it think on hard problems; use a fast mode for simple ones.
  • Before an agent acts, decide what it may access, and confirm the important steps yourself.

The next chapter covers models, products and companies: how ChatGPT relates to GPT-6, how companies split their models into tiers, and how to read the leaderboards.

Sources