AI news, models and products, with sources中文
NewsLandmark Models20 Jan 2025

DeepSeek releases R1, an open reasoning model

DeepSeek releases R1 and distilled reasoning models with public weights. Its published evaluations report strong maths and reasoning results; the release expands options for research and deployment.

Nvidia's headquarters building in Santa Clara, California
Nvidia's headquarters in Santa Clara, California, in 2018. A week after R1's release, Nvidia's shares fell nearly 17% in a single day. File photo. Photo: Coolcaesar / Wikimedia Commons, CC BY-SA 4.0

What happened

On 20 January 2025, DeepSeek released DeepSeek-R1. Its release note summed it up in a few lines: performance on par with OpenAI-o1; the model and technical report fully open; code and models under the MIT licence, free to “distill & commercialize”; the website and API live the same day, where users could switch on the “DeepThink” button at chat.deepseek.com.

Two other things came out at the same time:

  • R1-Zero, an experimental version trained with reinforcement learning alone (explained below).
  • Six “distilled” smaller models. DeepSeek used about 800,000 samples generated by R1 to train open models from Alibaba’s Qwen2.5 and Meta’s Llama 3 families, ranging from 1.5 billion to 70 billion parameters. DeepSeek said the 32-billion and 70-billion versions were on par with OpenAI’s o1-mini.

R1 itself has 671 billion parameters in total, but only about 37 billion are used for any given answer. This design, called “mixture of experts”, works like a team of specialists where only the relevant ones are called on each time, which saves computing power. R1 was built on top of V3, a base model DeepSeek had released a month earlier.

Background: what is a reasoning model?

R1 is a “reasoning model”: before answering, it writes out a long chain of thought, which users can read, and then gives its answer. Models like this are good at maths, coding and other problems that need step-by-step working. OpenAI’s o1, released in September 2024, is the best-known example of the type.

One reason R1 mattered is that DeepSeek published how it was trained. In its paper, DeepSeek says R1-Zero is the first open research to show that a large model’s reasoning ability can be developed purely through reinforcement learning, without first fine-tuning it on human-written examples.

Diagram of DeepSeek-R1's multi-stage training pipeline, from the V3 base model through reinforcement learning to R1-Zero, then through fine-tuning and reinforcement learning to R1
R1's training pipeline as published by DeepSeek in Nature. On the left, R1-Zero is trained with reinforcement learning (RL) only; the later stages alternate supervised fine-tuning (SFT) and RL to produce R1. Figure: DeepSeek-AI (Guo et al.), Nature 2025 / Wikimedia Commons, CC BY 4.0

How good is it?

DeepSeek’s model card compares R1 with OpenAI’s o1 (the 17 December 2024 version) and other models. The most quoted result is AIME 2024, the American Invitational Mathematics Examination, a competition for top US secondary-school maths students that is far harder than ordinary exams.

AIME 2024: hard competition mathsShare answered correctly on the first attempt (pass@1); higher is better
  • DeepSeek-R179.8%
  • OpenAI o179.2%
  • OpenAI o1-mini63.6%
  • DeepSeek-V339.2%
  • Claude 3.5 Sonnet16.0%
  • GPT-4o9.3%

All scores as published by DeepSeek in its model card; o1 is the 17 Dec 2024 version. Source: DeepSeek-R1 model card (Hugging Face), 20 Jan 2025

The same table shows R1 did not lead everywhere:

Test What it measures R1 o1
MATH-500 Maths problems 97.3% 96.4%
Codeforces Competitive programming (share of human contestants beaten) 96.3% 96.6%
SWE-bench Verified Fixing bugs in real software projects 49.2% 48.9%
GPQA Diamond Graduate-level science questions 71.5% 75.7%
SimpleQA Short factual questions 30.1% 47.0%

Roughly: the two were level in maths and coding, while o1 did better on science knowledge and factual questions.

Why it shook the markets

What made R1 big news was not just its scores, but its apparent cheapness.

A month before R1, DeepSeek had released its V3 base model and said its training cost only about $5.6 million (the computing cost of the final training run). According to Wikipedia’s summary of reporting, DeepSeek said V3 needed about 2,000 Nvidia H800 chips and around 55 days, while the world’s leading AI companies trained their models on as many as 16,000 chips. OpenAI’s chief executive Sam Altman had said in 2023 that training foundation models cost “much more” than $100 million.

R1 was trained on top of V3. In September 2025, in a peer-reviewed paper in the journal Nature, DeepSeek gave its first figures for R1 itself:

R1 training cost (excluding V3)
$294,000
Chips used
512 × H800
R1 training time
80 hours
V3's earlier self-reported cost
~$5.6m

Source: Reuters, 18 Sep 2025, on DeepSeek's Nature paper; V3 figure self-reported by DeepSeek in late 2024.

The market’s reasoning: if a model this good could be built with far fewer chips, were the huge planned investments in chips and data centres really necessary?

The week after R1
  1. 120 JanuaryR1 is released; its weights are published under the MIT licence and the DeepSeek app and website are free to use.
  2. 227 January (Monday)The DeepSeek app overtakes ChatGPT as the most-downloaded free app on the US App Store.
  3. 3Same dayNvidia's shares fall just under 17%, wiping about $593 billion off its value, a record one-day loss for a Wall Street stock, according to Reuters. The Nasdaq falls 3.1%.
  4. 4Same dayDeepSeek says it is facing "large-scale malicious" cyberattacks and temporarily limits new sign-ups.

Sources: Reuters and AFP reports, 27–28 Jan 2025

(Reports differ slightly on the size of the loss: Reuters, citing LSEG data, put it at about $593 billion; AFP at about $589 billion.) For the full story, see DeepSeek tops the App Store; Nvidia loses $590 billion in a day.

According to Wikipedia’s summary, US President Donald Trump called DeepSeek a “wake-up call”, and Microsoft’s Satya Nadella and OpenAI’s Sam Altman both called it “super impressive”. Others were sceptical: Scale AI’s chief executive Alexandr Wang speculated that China had more Nvidia H100 chips than was thought.

Things to keep in mind

  • The scores come from DeepSeek. Every comparison above is from DeepSeek’s own model card, under test conditions it chose.
  • “Training cost” is a narrow measure. The $294,000 covers only R1’s extra training on top of V3, and V3’s $5.6 million covers only the computing for its final run, not staff, earlier experiments or buying chips. According to The Decoder, outside estimates of V3’s real total cost range from tens of millions to several hundred million dollars.
  • Where the chips came from is disputed. Reuters reported that US officials said DeepSeek had access to large volumes of H100 chips obtained after US export controls took effect; Nvidia said DeepSeek used lawfully acquired H800s. In supplementary material to the Nature paper, DeepSeek acknowledged for the first time that it owns A100 chips and used them in early preparation.
  • Distillation claims. OpenAI said DeepSeek may have “inappropriately” used its models’ outputs as training data, though there is no way to prove this conclusively. In Nature, DeepSeek said V3’s training data came from crawled web pages that contained a “significant number of OpenAI-model-generated answers”, which may have let it learn indirectly from other models. In February 2026, Anthropic accused DeepSeek of using thousands of fraudulent accounts to generate millions of conversations with Claude to train its own models.
  • Censorship and privacy. The official app and API run on servers in mainland China and censor politically sensitive topics. Regulators in several countries, including Italy, South Korea and Australia, investigated or restricted it over data collection.

What it means for ordinary users

  • A free model that “thinks”. At launch, anyone could use R1 for free by switching on DeepThink in the DeepSeek website or app. DeepSeek’s products have since moved on to the V4 family; see our DeepSeek product page and DeepSeek V4.
  • You can run it yourself. Open weights mean companies can run the model on their own servers, keeping data in-house. The distilled models, the smallest with just 1.5 billion parameters, need far less computing power.
  • Pressure on prices. At launch, R1’s API cost $0.55 per million input tokens ($0.14 for cached input) and $2.19 per million output tokens. According to Fortune, DeepSeek’s open approach encouraged more Chinese labs to release open models, and even pushed OpenAI to release an open-weight model of its own, gpt-oss.

Model summary

Developer
DeepSeek
Availability
Open weights (MIT licence); DeepSeek app and API
Specs
671B total parameters, 37B active (mixture of experts)

Reported benchmark results

AIME 2024 (pass@1)79.8%OpenAI o1: 79.2%
MATH-500 (pass@1)97.3%OpenAI o1: 96.4%

Data source: DeepSeek's published model card. Scores are as published by the developer at launch. Test conditions differ between companies, so compare with care.

Sources