AI news, models and products, with sources中文
NewsMajor Models24 Apr 2026

DeepSeek V4: open weights with a 1-million-token context

DeepSeek releases V4 in two open-weight sizes, saying the larger Pro model rivals top closed models in maths, science and coding.

What happened

On Friday 24 April 2026, DeepSeek announced that “DeepSeek-V4 Preview is officially live & open-sourced”, under the slogan “Welcome to the era of cost-effective 1M context length”. It was the company’s biggest release since R1 shook the world in early 2025.

There are two models:

  • V4-Pro: 1.6 trillion parameters in total, 49 billion used for each answer. DeepSeek says its performance is “rivaling the world’s top closed-source models”.
  • V4-Flash: 284 billion in total, 13 billion used each time. Smaller, faster and cheaper; DeepSeek says its reasoning comes close to V4-Pro and it matches V4-Pro on simple agent tasks.

The weights of both are on Hugging Face under the MIT licence. From launch day they could be used at chat.deepseek.com in “Expert Mode” and “Instant Mode”, and through the updated API.

What’s new

According to the V4 model card, the main technical changes are:

  • A new attention mechanism. DeepSeek designed a “hybrid attention” that combines two kinds of compression. At 1 million tokens, V4-Pro needs only 27% of the computation per token and 10% of the memory for holding context (the “KV cache”) compared with the previous V3.2. This is what lets DeepSeek make a 1-million-token context the default across its services.
  • More training data. Both models were pre-trained on more than 32 trillion tokens.
  • Three thinking levels. “Non-think” for quick answers to everyday tasks, “High” for slower but more accurate work, and “Max” to push reasoning as far as it goes.
  • Built for agents. DeepSeek says V4 plugs directly into coding agents such as Claude Code, OpenClaw and OpenCode, and already drives its own in-house coding.

How it compares

DeepSeek’s model card compares V4-Pro’s highest thinking level (V4-Pro-Max) with several frontier models. The two tests below stand for knowledge-based reasoning and agent work:

GPQA Diamond: graduate-level science questionsShare answered correctly on the first attempt; higher is better
  • Gemini 3.1 Pro94.3%
  • GPT-5.493.0%
  • Claude Opus 4.691.3%
  • Kimi K2.690.5%
  • DeepSeek-V4-Pro90.1%
  • GLM-5.186.2%

All scores published by DeepSeek; V4-Pro at its Max level, other models at their high-effort settings. Source: DeepSeek-V4-Pro model card (Hugging Face), Apr 2026

Terminal Bench 2.0: completing tasks in a command lineShare of tasks completed; higher is better
  • GPT-5.475.1%
  • Gemini 3.1 Pro68.5%
  • DeepSeek-V4-Pro67.9%
  • Kimi K2.666.7%
  • Claude Opus 4.665.4%
  • GLM-5.163.5%

All scores published by DeepSeek. Source: DeepSeek-V4-Pro model card (Hugging Face), Apr 2026

On these tests V4-Pro sits between the leading closed models and other open models. DeepSeek’s own wording is fairly modest: its release note says V4-Pro “leads all current open models” in world knowledge, “trailing only Gemini-3.1-Pro”. And according to Fortune, the technical report says V4 “falls marginally short of GPT-5.4 and Gemini 3.1 Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately three to six months”.

Price: the main attraction

V4-Pro, per million output tokens
$3.48
V4-Flash, per million output tokens
$0.28
OpenAI, same amount of work
$30
Anthropic, same amount of work
$25

Source: Fortune, 24 Apr 2026. Prices at launch; DeepSeek has since changed its line-up and pricing.

Fortune noted that this ran against the industry trend: OpenAI and Anthropic had both raised prices or tightened usage limits. DeepSeek also said it expected to cut V4-Pro’s price later in the year as Huawei scaled up production of its new Ascend 950 chips.

The Huawei connection

On launch day, Huawei announced that its Ascend AI chips would offer “full support” for DeepSeek’s models. According to Fortune, DeepSeek said it used Ascend processors to train the new model. Shares in SMIC, the Chinese chipmaker that makes Ascend chips, rose 10% in Hong Kong that day, while shares in two of DeepSeek’s Chinese rivals, MiniMax and Knowledge Atlas, fell by more than 9%.

Since launch: V4 was quickly replaced

DeepSeek moved fast, and what you get from DeepSeek today is no longer the April version:

The V4 family over time
  1. 124 AprilV4-Pro and V4-Flash released as an open-weight preview.
  2. 224 JulyThe old API names deepseek-chat and deepseek-reasoner are retired (they had already been routed to V4-Flash).
  3. 331 July / 13 AugustAccording to Wikipedia, the official versions of V4-Flash and V4-Pro are released.
  4. 410 SeptemberV4.1-Flash arrives: 552 billion parameters, native image understanding, scores DeepSeek says beat V4-Pro, and lower API prices.
  5. 5From 14 SeptemberAll V4-Pro requests are routed to V4.1-Flash at Flash prices, until V4.1-Pro launches.

Sources: DeepSeek API documentation and release notes; Wikipedia

Things to keep in mind

  • The scores and comparisons come from DeepSeek, including the scores for rival models, and test conditions may differ. The rivals are the models of April 2026 (such as Claude Opus 4.6 and GPT-5.4), not today’s latest.
  • It was a preview. The official versions came only in July and August, and V4-Pro was already being phased out in September.
  • The distillation dispute continues. According to Fortune, the day before V4’s release White House technology adviser Michael Kratsios accused Chinese AI developers of “industrial-scale campaigns” to copy US technology, and OpenAI and Anthropic have accused Chinese developers, including DeepSeek, of “illicit” distillation. China’s foreign ministry called the claims “groundless”.
  • Official channels only. In its release note, DeepSeek reminded readers that only its official accounts speak for the company.
  • Censorship and privacy. DeepSeek’s official app follows Chinese censorship rules, and the US Department of Defense and the Australian government, among others, have barred DeepSeek products from government devices. Running the downloaded weights yourself does not send data to DeepSeek’s servers.

How to try it

  • Website and apps. Free at chat.deepseek.com or in the DeepSeek app, in Expert or Instant mode. See our DeepSeek product page.
  • Developers. The DeepSeek API accepts the OpenAI and Anthropic formats, so it works with tools such as Claude Code. The recommended model name is now deepseek-flash (V4.1-Flash).
  • Running it yourself. Weights for the V4 family and V4.1-Flash are available on Hugging Face.

Model summary

Developer
DeepSeek
Availability
Open weights; DeepSeek app and API
Specs
V4-Pro: 1.6T total / 49B active parameters. V4-Flash: 284B total / 13B active. 1M-token context.

Data source: DeepSeek release note

Sources