DeepSeek Moves V4-Flash Out of Preview, Closing the Agent-Benchmark Gap With Its Own Pro Model
DeepSeek's official DeepSeek-V4-Flash-0731 API left public beta on July 31 with the same architecture as the preview but a retrained post-training pipeline that lifts agent and coding benchmarks well above V4-Pro-Preview, at unchanged pricing.
DeepSeek took its V4-Flash API out of preview on July 31, publishing DeepSeek-V4-Flash-0731 as the official release. The company said on X that the update keeps the exact same model architecture and size as the April preview — a sparse mixture-of-experts model with 284 billion total parameters and roughly 13 billion active per token — and applies only to the Flash API; the V4-Pro API and DeepSeek's app and web models are unchanged for now, with an official V4-Pro release still to come.
What changed
DeepSeek said the gains come entirely from re-post-training the existing checkpoint with a pipeline focused on coding, agentic tool use and reasoning, not from a new architecture. The clearest jump is on Terminal-Bench 2.1, a benchmark for complex command-line agent work, where 0731 scores 82.7 versus 61.8 for the Flash preview and 72.1 for V4-Pro-Preview — meaning the smaller, cheaper Flash model now outperforms DeepSeek's own flagship preview on agentic execution despite activating a fraction of its parameters per token. DeepSeek and third-party trackers reported similar jumps on internal agent benchmarks DeepSWE and DSBench-FullStack. The model retains a 1-million-token context window and three selectable reasoning-effort levels, and ships with a speculative-decoding draft module carried over from the preview checkpoint, so existing serving configurations continue to work.
Pricing unchanged, for now
API pricing holds at the preview's rates — $0.14 per million input tokens (cache-miss) and $0.28 per million output tokens, among the cheapest access to a top-tier open-weight coding and agent model. DeepSeek has separately flagged a planned peak-pricing window that would double billing for roughly seven hours a day, though no effective date has been set.
Why it matters
The release underscores how quickly open-weight labs are closing the gap with proprietary frontier models on agentic coding tasks, and it does so via cheaper post-training rather than costlier pretraining runs — a pattern that keeps pressure on rival labs' pricing for agent-oriented API traffic. With V4-Pro's official release still pending, DeepSeek's roadmap suggests further agent-benchmark gains are still to come across its full model line.
Sources
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
Related agents
More ai news
Demis Hassabis Steps Down as Google DeepMind CEO, Hands Day-to-Day Control to Koray Kavukcuoglu
Hassabis becomes chairman of Google DeepMind and Alphabet's chief scientist to focus on AGI research, while longtime DeepMind CTO Koray Kavukcuoglu takes over Gemini development and frontier research, reporting directly to Sundar Pichai.
White House Finalizes Voluntary AI Safety Framework, Won't Say What's In It
The Trump administration told about a dozen AI labs on August 4 that its voluntary framework for early government access to frontier models is final, capping a process ordered by a June executive order — but it is keeping the framework's contents, and who has seen them, confidential.
EU Begins Enforcing AI Act Transparency Rules as High-Risk Deadlines Slip to 2027-2028
The European Commission's AI Office started enforcing new EU AI Act transparency obligations on August 2, requiring chatbots to disclose they're AI and deepfakes to be labeled, even as high-risk system rules were pushed back under the Digital Omnibus on AI.
Anthropic Names Tino Cuéllar as Its First Chief Global Affairs Officer
Anthropic hired former California Supreme Court Justice Mariano-Florentino Cuéllar to lead policy and government relations worldwide, a new senior role created as the company navigates a Pentagon technology blacklist and export-control friction with Washington.