DeepSeek Releases V4.1 Flash, Cuts API Prices and Sets End Date for V4 Pro
DeepSeek officially released V4.1 Flash, a cheaper and faster multimodal successor to V4 Pro with a 1-million-token context window, and said it will reroute all V4 Pro API traffic to the new model from September 14.
DeepSeek officially released DeepSeek-V4.1-Flash on September 10, a lower-cost, faster successor to V4 Pro built to handle both text and images natively rather than through a bolted-on vision module.
What changed under the hood
V4.1 Flash is a sparse mixture-of-experts model with a 1,048,576-token context window and support for up to 384,000 tokens of output. According to MarkTechPost's technical write-up, the model brings native multimodal input into its base architecture and uses several efficiency techniques — including an FP4 key-value cache and cross-layer attention reuse — to cut the memory cost of holding long contexts in memory compared with V4 Flash. DeepSeek said internal and external testing showed V4.1 Flash beating V4 Pro on performance, cost, speed and total task time, despite Flash traditionally being positioned as the cheaper, lighter-weight tier.
Pricing moves and V4 Pro's retirement
Alongside the release, DeepSeek cut Flash's API pricing, with reporting from Intelligent Living and OpenRouter's own listings pegging the new uncached input rate at a fraction of what V4 Pro charged, plus a steep discount on cache-hit input tokens. DeepSeek also set a firm transition date: starting at 12:00 Beijing time on September 14, all API requests still targeting deepseek-v4-pro will be automatically routed to V4.1 Flash and billed at Flash's lower price, effectively retiring V4 Pro until a V4.1 Pro model arrives.
Why it matters
The release keeps DeepSeek in the middle of the pricing war among frontier and near-frontier labs: rather than shipping a flagship "Pro" upgrade, it pushed its efficiency gains into the cheaper Flash tier and is forcing existing Pro customers onto it. That mirrors a broader pattern this quarter of labs competing less on headline benchmark scores — which are converging across Claude, GPT-6 Astra, Gemini and DeepSeek's own models — and more on cost per token, context-window efficiency and how gracefully a model handles very long inputs. For developers building on open-weight or budget-tier models, it also lowers the cost of long-context workloads like large-codebase analysis or document review that previously required a pricier "Pro"-class model.
Sources
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
More ai news
OpenAI Says a Swarm of 10,000 AI Agents Solved the Navier-Stokes Millennium Problem
OpenAI published a claimed solution to the Navier-Stokes existence and smoothness problem, one of math's seven Millennium Prize Problems, produced by roughly 10,000 coordinated AI agents and formally verified in the Lean proof language.
TCS Unit HyperVault to Invest Up to $7.4B in 1GW AI Data Center Campus in India
Tata Consultancy Services' infrastructure arm HyperVault will invest up to $7.4 billion with partners to build a 1-gigawatt AI data center campus in Hyderabad, one of India's largest bets yet on domestic AI compute capacity.
Mistral Raises €3B in Samsung-Led Round, Becomes Europe's Best-Funded AI Startup
French AI lab Mistral raised €3 billion in a Series D round led by Samsung Electronics, pushing its post-money valuation above €21 billion and marking the largest equity round ever completed by a European technology company.
Sanders and Casar Introduce Bill to Ban 'Artificial Superintelligence' Outright
Sen. Bernie Sanders and Rep. Greg Casar introduced the Ban Artificial Superintelligence Act on September 3, which would permanently prohibit superintelligent AI systems, pause frontier development pending federal safety rules, and impose prison terms and a 'corporate death penalty' for violations.