Timeline

How we got here

33 milestones since 2017, chosen because they still explain how today's systems work. Each one links to its primary source where one exists.

2025

  1. 24 Feb 2025 · Product

    Claude Code preview

    An agent that works in the terminal across a whole codebase. Coding assistants move from autocomplete to delegation.

  2. 20 Jan 2025 · Model

    DeepSeek-R1

    An open-weight reasoning model trained largely with RL, competitive with closed models at a fraction of the reported cost.

2024

  1. 25 Nov 2024 · Standard

    Model Context Protocol

    An open protocol for connecting AI apps to tools and data. Integrations become reusable.

  2. 22 Oct 2024 · Product

    Computer use

    Claude operates a desktop through screenshots, mouse and keyboard: agents leave the API sandbox.

  3. 12 Sept 2024 · Model

    o1 and test-time compute

    A model trained with RL to reason at length before answering. More thinking at inference buys accuracy.

  4. 13 May 2024 · Model

    GPT-4o

    One model natively handles text, audio and images, enabling real-time voice conversation.

  5. 4 Mar 2024 · Model

    Claude 3

    A three-tier family (Haiku, Sonnet, Opus) with long context and vision.

  6. 15 Feb 2024 · Model

    Sora preview

    Minute-long coherent video from text. Video generation joins the frontier.

2023

  1. 11 Dec 2023 · Model

    Mixtral 8x7B

    An open sparse mixture-of-experts model. Only a fraction of parameters are active per token.

  2. 6 Dec 2023 · Model

    Gemini 1.0

    Google's first model family trained multimodal from the start.

  3. 28 Jul 2023 · Model

    RT-2

    A vision-language-action model transfers web knowledge into robot control.

  4. 18 Jul 2023 · Model

    Llama 2

    Open weights with a commercial-use licence make self-hosted LLMs a serious option for companies.

  5. 14 Mar 2023 · Model

    Claude launches

    Anthropic's first public assistant model, trained with Constitutional AI.

  6. 14 Mar 2023 · Model

    GPT-4

    A step change in capability on professional exams, with image input. The bar every lab now aims at.

  7. 24 Feb 2023 · Model

    LLaMA

    Strong models trained only on public data. The leaked weights ignite the open-model ecosystem.

  8. 9 Feb 2023 · Paper

    Toolformer

    A model teaches itself when to call APIs: an early sign of native tool use.

2022

  1. 15 Dec 2022 · Paper

    Constitutional AI

    Harmlessness trained from AI feedback against a written set of principles, reducing reliance on human labels.

  2. 30 Nov 2022 · Product

    ChatGPT

    A chat interface on an RLHF-tuned model reaches a mass audience within weeks.

  3. 6 Oct 2022 · Paper

    ReAct

    Reason, act, observe, repeat. The loop at the heart of nearly every agent framework.

  4. 22 Aug 2022 · Model

    Stable Diffusion released

    A capable image model ships as open weights and runs on consumer GPUs.

  5. 27 May 2022 · Paper

    FlashAttention

    Exact attention made memory-efficient by being IO-aware. A key enabler of long context windows.

  6. 29 Mar 2022 · Paper

    Chinchilla

    Most large models were undertrained. Compute-optimal training scales tokens with parameters.

  7. 4 Mar 2022 · Paper

    InstructGPT and RLHF

    Reinforcement learning from human feedback makes a smaller model preferred over a far larger raw one.

  8. 28 Jan 2022 · Paper

    Chain-of-thought prompting

    Asking models to show intermediate steps sharply improves multi-step reasoning: the seed of reasoning models.

2021

  1. 29 Jun 2021 · Product

    GitHub Copilot preview

    Code completion from a large model, inside the editor. AI coding becomes a daily tool.

  2. 5 Jan 2021 · Paper

    CLIP

    Contrastive training on image–text pairs gives a shared embedding space. It becomes the glue of multimodal systems.

2020

  1. 30 Nov 2020 · Result

    AlphaFold 2 at CASP14

    Protein structure prediction reaches near-experimental accuracy. The clearest case of AI solving a long-standing scientific problem.

  2. 28 May 2020 · Model

    GPT-3 and in-context learning

    At 175B parameters, a model learns new tasks from a few examples in the prompt, without retraining.

  3. 22 May 2020 · Paper

    Retrieval-Augmented Generation

    A generator conditioned on retrieved passages. The pattern behind most enterprise AI today.

  4. 23 Jan 2020 · Paper

    Scaling laws

    Loss falls as a smooth power law in parameters, data and compute, turning model building into a forecastable investment.

2019

  1. 14 Feb 2019 · Model

    GPT-2 and staged release

    A 1.5B-parameter model writes coherent paragraphs. OpenAI withholds full weights at first, starting the release-norms debate.

2018

  1. 11 Oct 2018 · Paper

    BERT

    Bidirectional pre-training then fine-tuning becomes the default recipe for language understanding.

2017

  1. 12 Jun 2017 · Paper

    The Transformer

    “Attention Is All You Need” replaces recurrence with self-attention. Almost every frontier model since is built on it.