DeepSeek Moves V4-Flash Out of Preview, Closing the Agent-Benchmark Gap With Its Own Pro Model
DeepSeek's official DeepSeek-V4-Flash-0731 API left public beta on July 31 with the same architecture as the preview but a retrained post-training pipeline that lifts agent and coding benchmarks well above V4-Pro-Preview, at unchanged pricing.
DeepSeek took its V4-Flash API out of preview on July 31, publishing DeepSeek-V4-Flash-0731 as the official release. The company said on X that the update keeps the exact same model architecture and size as the April preview — a sparse mixture-of-experts model with 284 billion total parameters and roughly 13 billion active per token — and applies only to the Flash API; the V4-Pro API and DeepSeek's app and web models are unchanged for now, with an official V4-Pro release still to come.
What changed
DeepSeek said the gains come entirely from re-post-training the existing checkpoint with a pipeline focused on coding, agentic tool use and reasoning, not from a new architecture. The clearest jump is on Terminal-Bench 2.1, a benchmark for complex command-line agent work, where 0731 scores 82.7 versus 61.8 for the Flash preview and 72.1 for V4-Pro-Preview — meaning the smaller, cheaper Flash model now outperforms DeepSeek's own flagship preview on agentic execution despite activating a fraction of its parameters per token. DeepSeek and third-party trackers reported similar jumps on internal agent benchmarks DeepSWE and DSBench-FullStack. The model retains a 1-million-token context window and three selectable reasoning-effort levels, and ships with a speculative-decoding draft module carried over from the preview checkpoint, so existing serving configurations continue to work.
Pricing unchanged, for now
API pricing holds at the preview's rates — $0.14 per million input tokens (cache-miss) and $0.28 per million output tokens, among the cheapest access to a top-tier open-weight coding and agent model. DeepSeek has separately flagged a planned peak-pricing window that would double billing for roughly seven hours a day, though no effective date has been set.
Why it matters
The release underscores how quickly open-weight labs are closing the gap with proprietary frontier models on agentic coding tasks, and it does so via cheaper post-training rather than costlier pretraining runs — a pattern that keeps pressure on rival labs' pricing for agent-oriented API traffic. With V4-Pro's official release still pending, DeepSeek's roadmap suggests further agent-benchmark gains are still to come across its full model line.
Sources
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
Related agents
More ai news
Anthropic Says Its Claude Mythos Model Found New Weaknesses in Two Cryptographic Algorithms
Anthropic's Frontier Red Team published research showing Claude Mythos Preview independently discovered a stronger attack on the NIST post-quantum candidate HAWK and a 200-800x faster attack on 7-round AES, though neither threatens deployed systems.
SpaceXAI's Grok Voice Think Fast 2.0 Jumps to #2 on Speech Benchmarks, Cuts Response Time in Half
SpaceXAI released Grok Voice Think Fast 2.0 on July 29, lifting its score on Artificial Analysis's Speech to Speech Index to 82.9% and its time-to-first-audio to 0.70 seconds, ahead of OpenAI's GPT-Realtime-2.1 and Google's Gemini 3.1 Flash on the same tests.
OpenAI to Give 100,000 Academic Researchers Free Access to Its Frontier Models
OpenAI launched ChatGPT for Academic Researchers, a program that will give 100,000 scientists free access to its most capable models by 2027, as part of a broader $250 million commitment to external scientific research.
Over 1,100 Employees at OpenAI, Anthropic, Google DeepMind and Meta Ask Washington to Prepare Tools to Pace AI Development
A statement called 'Pacing the Frontier,' signed by more than 1,100 staff across the top AI labs and endorsed by OpenAI and Anthropic as companies, asks the US government to help build the technical and governance tools needed to slow automated AI research if it starts to outrun oversight.