DeepSeek's long awaited V4 arrived on 24 April 2026 as a preview, with weights anyone can download, an API that went live the same day and word from Huawei that its entire Ascend supernode product line now supports the new models.
The release came in two sizes, V4-Pro and V4-Flash, both able to read 1 million tokens of context, which DeepSeek made the default across all its services. The Hangzhou company said V4-Pro is the strongest open model for agentic coding and trails only Google's Gemini 3.1 Pro on world knowledge. The other headline was hardware. DeepSeek's earlier V3 and R1 models were trained on Nvidia chips, Reuters noted, while V4 was adapted to run on Huawei's Ascend processors at a time when US export controls limit China's access to advanced American chips.
- What: DeepSeek V4 Preview: V4-Pro (1.6 trillion parameters, 49 billion active) and V4-Flash (284 billion, 13 billion active), open weights under the MIT licence.
- Price: list prices at launch USD 1.74 per million input tokens and USD 3.48 per million output tokens for Pro, USD 0.14 and USD 0.28 for Flash.
- Why it matters: a frontier class open model with a 1 million token context as standard, adapted for Chinese chips.
Two models, one million tokens
Both models use a mixture of experts design, in which only part of the network works on each token. According to the model card, V4-Pro has 1.6 trillion parameters in total with 49 billion active per token, and V4-Flash has 284 billion with 13 billion active. DeepSeek trained both on more than 32 trillion tokens and published the weights on Hugging Face under the MIT licence.
The main engineering change is in attention, the part of a model that decides which earlier tokens matter. DeepSeek combined two compressed attention schemes and says that at a 1 million token context V4-Pro needs only 27 percent of the computing operations per generated token, and 10 percent of the memory cache, that its previous model V3.2 required. That is what lets a million tokens become a default rather than a premium extra.
The API launched the same day, accepting requests in both the OpenAI and the Anthropic formats, and DeepSeek said V4 works with agent tools such as Claude Code, OpenClaw and OpenCode. The old model names deepseek-chat and deepseek-reasoner were set to retire on 24 July.
Where it leads, by DeepSeek's numbers
DeepSeek's own comparison, with V4-Pro at its maximum reasoning setting, is vendor reported and uses the rival models available in April. V4-Pro scored 93.5 on LiveCodeBench and a Codeforces rating of 3,206, the best figures in its table, and 80.6 percent on SWE Verified, level with Gemini 3.1 Pro and just behind Claude Opus 4.6 at 80.8 percent. On knowledge tests the gap was plain: 57.9 percent on SimpleQA Verified against 75.6 percent for Gemini 3.1 Pro, and 37.7 percent on Humanity's Last Exam against 44.4 percent. Reuters, citing a DeepSeek paper, summed it up as ahead of all other open models but still behind closed systems such as Gemini 3.1 Pro and OpenAI's GPT-5.4 in some areas.
Huawei and the price of compute
In a post on Weibo, Huawei said close collaboration between the two companies on core model technology had brought V4 support to its entire line of Ascend supernodes, and that its Ascend 950 SuperPod reaches an inference latency of 20 milliseconds for V4-Pro and 10 milliseconds for V4-Flash. DeepSeek, for its part, tied its prices to compute. V4-Pro could cost up to 12 times as much as Flash because of "constraints in high-end compute capacity", DeepSeek said, according to Reuters, and it expected the Pro price to fall significantly once Huawei's Ascend 950 supernodes are mass produced and deployed in the second half of the year, Huawei Central reported. At launch DataCamp listed V4-Pro at USD 1.74 per million input tokens and USD 3.48 per million output tokens, and V4-Flash at USD 0.14 and USD 0.28.
A quieter reception
The launch did not repeat the shock of 2025, when the reception of V3 and R1 "triggered a global tech share selloff", as Reuters wrote on 27 April. This time the market reaction was "subdued". "This announcement followed a rather predictable path," said Lian Jye Su, chief analyst at Omdia. Reuters cited Artificial Analysis data showing V4-Pro among the leading open weight models rather than clearly ahead of them, with Chinese rivals such as Kimi and Qwen narrowing the gap.
What matters now is whether China can continue advancing on AI development, and potentially do so with its own chipsAlfredo Montufar-Helu, managing director at Ankura China Advisors, to Reuters, 27 April 2026
What happened next
DeepSeek moved V4-Pro out of preview on 13 August and announced peak and off peak API pricing from 16 August, with off peak rates half the peak rates. On 10 September it released V4.1-Flash, a 552 billion parameter model with a new encoder and decoder design that activates 8 billion parameters for input and 16 billion for output, with native image understanding. DeepSeek said tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost and speed, and announced it was phasing V4-Pro out, with all V4-Pro requests to be served by V4.1-Flash from 14 September until a V4.1-Pro launches. It then changed course: "In response to user demand," its API change log now says, V4-Pro stays available after 14 September with billing unchanged. Five months after the preview, DeepSeek itself says a smaller sibling beats the model that carried the April headlines.





Comments
Comments are reviewed before they appear.
No comments yet. Be the first to share your thoughts.