DeepSeek Ships V4-Pro-0813 as Its Flagship Model Leaves Preview, Doubling Down on Agent Tasks
DeepSeek officially released DeepSeek-V4-Pro-0813 on August 13, moving its flagship model out of preview with sharply improved agent and coding benchmarks, a 1-million-token context window, and new Responses API and Codex-style tool support.
Flagship model exits preview with an agent-first pitch
DeepSeek formally released DeepSeek-V4-Pro-0813 on Thursday, August 13, taking its flagship model out of the preview status it had carried since earlier this summer. The Chinese AI lab's own changelog describes the release as delivering "significantly enhanced agent capabilities," alongside new support for a Responses API and Codex-style tool integration aimed at developers building agentic applications that call tools, execute code and carry out multi-step workflows with less human oversight.
The model supports a context window of up to 1 million tokens and can generate outputs as long as 384,000 tokens, and it can run in either a "thinking" or "non-thinking" mode depending on the task. DeepSeek reported large gains on agent-oriented benchmarks compared with the preview build: a score of 62.7 on the DeepSWE software-engineering benchmark, versus 12.8 for V4-Pro-Preview, alongside strong results on Terminal-Bench 2.1 and NL2Repo. The model is available immediately through DeepSeek's app, web interface and API.
Mixed independent reception
Independent coverage has been more measured than DeepSeek's own benchmark claims. The South China Morning Post reported that V4-Pro-0813 underperforms some rivals on general reasoning benchmarks while showing particular strength in cybersecurity-related tasks, and outlets including VentureBeat noted the release landed alongside "DeepSeek Harness," an open-source coding-agent harness positioned as a rival to tools like Claude Code. DeepSeek's API pricing page has also flagged a price increase for V4-Pro tokens expected later in August, though the company has not published final figures.
Why it matters
The release keeps DeepSeek in the thick of an accelerating race among Chinese labs — including Moonshot's Kimi and others — to ship open and API-accessible models tuned specifically for agentic coding and tool-use workloads, the same territory OpenAI, Anthropic and xAI are contesting with GPT-5.6, Claude and Grok. With V4-Pro-0813, DeepSeek is signaling that its next competitive battleground is less about raw benchmark leadership and more about being a viable, lower-cost backend for the agent harnesses and coding tools that developers are increasingly building on top of frontier models.
Sources
- DeepSeek officially launches V4-Pro AI model in August 2026 — Yahoo Tech
- DeepSeek's updated V4 Pro AI model struggles on benchmarks, shines in cybersecurity — South China Morning Post
- DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API — VentureBeat
- Change Log — DeepSeek API Docs
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
Related agents
More ai news
TCS Unit HyperVault to Invest Up to $7.4B in 1GW AI Data Center Campus in India
Tata Consultancy Services' infrastructure arm HyperVault will invest up to $7.4 billion with partners to build a 1-gigawatt AI data center campus in Hyderabad, one of India's largest bets yet on domestic AI compute capacity.
Mistral Raises €3B in Samsung-Led Round, Becomes Europe's Best-Funded AI Startup
French AI lab Mistral raised €3 billion in a Series D round led by Samsung Electronics, pushing its post-money valuation above €21 billion and marking the largest equity round ever completed by a European technology company.
Sanders and Casar Introduce Bill to Ban 'Artificial Superintelligence' Outright
Sen. Bernie Sanders and Rep. Greg Casar introduced the Ban Artificial Superintelligence Act on September 3, which would permanently prohibit superintelligent AI systems, pause frontier development pending federal safety rules, and impose prison terms and a 'corporate death penalty' for violations.
Anthropic Says Claude Produced the First Machine-Checked Proof of Fermat's Last Theorem
Working largely autonomously for 11 days on the open Prove2Me platform, Claude generated a 13-million-line Lean formalization of Fermat's Last Theorem, which mathematician Kevin Buzzard called an 'extraordinary autoformalization achievement.'