DeepSeek released a preview of V4 on Friday, its first flagship model since R1 launched in January 2025, with a 1 million-token context window and pricing that undercuts every closed-source rival on the leaderboard. V4-Pro costs $1.74 per million input tokens and $3.48 per million output tokens. V4-Flash, the smaller sibling, runs at $0.14 and $0.28 — among the cheapest top-tier rates on the market.
Both versions ship open source. DeepSeek says V4-Pro matches Anthropic's Claude-Opus-4.6, OpenAI's GPT-5.4 and Google's Gemini-3.1 on major benchmarks, and beats Alibaba's Qwen-3.5 and Z.ai's GLM-5.1 on coding, math and STEM. In an internal survey of 85 experienced developers published with the model, more than 90% put V4-Pro among their top picks for coding work.
The 1 million-token window is the headline capability — enough to ingest all three Lord of the Rings volumes and The Hobbit in a single prompt — and it now applies as the default across every DeepSeek service. That brings the firm to parity with the long-context tiers of Gemini and Claude, and the way DeepSeek got there is the more interesting story.
Key facts
- 01DeepSeek released V4 on Friday in two open-source variants, V4-Pro and V4-Flash, both with a 1 million-token context window.
- 02V4-Pro is priced at $1.74 per million input tokens and $3.48 per million output tokens; V4-Flash runs $0.14 in and $0.28 out.
- 03In a 1M-token context, V4-Pro uses 27% of V3.2's compute and 10% of its memory; V4-Flash uses 10% and 7%.
- 04In a DeepSeek survey of 85 developers, more than 90% put V4-Pro among their top model choices for coding tasks.
- 05V4 is DeepSeek's first model optimized for Chinese chips, with Huawei's Ascend 950 supernodes officially supporting it on day one.
V4 reworks the attention mechanism so the model is selective about what it carries forward. Older text gets compressed; nearby tokens stay in full resolution. In a 1-million-token context, V4-Pro consumes 27% of the compute V3.2 needed and 10% of the memory. V4-Flash drops further, to 10% of the compute and 7% of the memory.
“In a 1-million-token context, V4-Pro uses just 27% of the compute and 10% of the memory required by V3.2, with V4-Flash dropping to 10% compute and 7% memory.”— Jaeden Schafer
For builders, that is the difference between a long-context coding agent that can sweep a full repository on a reasonable budget and one that can't. DeepSeek says it has specifically optimized V4 for Claude Code, OpenClaw and CodeBuddy, the agent frameworks where coding workloads are concentrating.
Performance gains aside, the more strategic move is the silicon. V4 is DeepSeek's first model optimized for domestic Chinese chips, and Huawei confirmed on Friday that its Ascend 950-based supernodes will support the model. The Information reported earlier this month that DeepSeek withheld pre-release access from Nvidia and AMD, giving early access only to Chinese chipmakers — a reversal of the usual launch choreography.
Reuters previously reported that Chinese officials recommended DeepSeek integrate Huawei silicon into training. The pressure tracks with a broader policy push that began with US export controls in 2022 cutting Chinese firms off from Nvidia's top chips, then tightening to cover the China-market downgrades. Beijing has responded with sourcing quotas, reported foreign-chip bans in public computing projects, and requirements to pair Nvidia hardware with domestic alternatives from Huawei and Cambricon.
The catch is that swapping Nvidia is not a one-line config change. The hard part is the software ecosystem — CUDA, kernels, tooling — that developers have spent more than a decade building around Nvidia. Adapting model code to Ascend, rebuilding tooling, and proving stability at production scale takes time DeepSeek hasn't fully spent yet.
Liu Zhiyuan, a computer science professor at Tsinghua University, told MIT Technology Review that "DeepSeek appears to have adapted only part of V4's training process for Chinese chips." DeepSeek's own technical report confirms Chinese chips are running inference, but is silent on whether the long-context training was ported. Several sources told the publication that domestic chips remain better suited to inference than to training, where Nvidia's lead is widest.
DeepSeek is also tying V4's future economics to the hardware transition, telling customers that V4-Pro prices could fall once Huawei's Ascend 950 supernodes ship in volume. That is a bet that domestic supply will scale on schedule — a bet Chinese AI firms are increasingly being asked to make whether they want to or not.
The release won't reset the field the way R1 did 15 months ago. R1 was a shock because nobody expected a constrained Chinese lab to match frontier reasoning at a fraction of the cost. V4 arrives in a market where open-weight Chinese models from Alibaba, Z.ai and others have already normalized that pattern.
What V4 does mark is the first credible test of whether a top-tier open model can be served, and eventually trained, on a non-Nvidia stack. If Huawei's Ascend 950 holds up at scale, the cost curve for Chinese AI inference detaches from US export policy. If it doesn't, V4 will still be remembered as the cheapest frontier-class model anyone could download — and that, on its own, is enough to keep OpenAI and Anthropic's pricing teams busy.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




