Agents AI

Update
ai

DeepSeek Releases V4.1 Flash, Cuts API Prices and Sets End Date for V4 Pro

DeepSeek officially released V4.1 Flash, a cheaper and faster multimodal successor to V4 Pro with a 1-million-token context window, and said it will reroute all V4 Pro API traffic to the new model from September 14.

AgentsAI NewsroomSeptember 11, 20262 min read

DeepSeek officially released DeepSeek-V4.1-Flash on September 10, a lower-cost, faster successor to V4 Pro built to handle both text and images natively rather than through a bolted-on vision module.

What changed under the hood

V4.1 Flash is a sparse mixture-of-experts model with a 1,048,576-token context window and support for up to 384,000 tokens of output. According to MarkTechPost's technical write-up, the model brings native multimodal input into its base architecture and uses several efficiency techniques — including an FP4 key-value cache and cross-layer attention reuse — to cut the memory cost of holding long contexts in memory compared with V4 Flash. DeepSeek said internal and external testing showed V4.1 Flash beating V4 Pro on performance, cost, speed and total task time, despite Flash traditionally being positioned as the cheaper, lighter-weight tier.

Pricing moves and V4 Pro's retirement

Alongside the release, DeepSeek cut Flash's API pricing, with reporting from Intelligent Living and OpenRouter's own listings pegging the new uncached input rate at a fraction of what V4 Pro charged, plus a steep discount on cache-hit input tokens. DeepSeek also set a firm transition date: starting at 12:00 Beijing time on September 14, all API requests still targeting deepseek-v4-pro will be automatically routed to V4.1 Flash and billed at Flash's lower price, effectively retiring V4 Pro until a V4.1 Pro model arrives.

Why it matters

The release keeps DeepSeek in the middle of the pricing war among frontier and near-frontier labs: rather than shipping a flagship "Pro" upgrade, it pushed its efficiency gains into the cheaper Flash tier and is forcing existing Pro customers onto it. That mirrors a broader pattern this quarter of labs competing less on headline benchmark scores — which are converging across Claude, GPT-6 Astra, Gemini and DeepSeek's own models — and more on cost per token, context-window efficiency and how gracefully a model handles very long inputs. For developers building on open-weight or budget-tier models, it also lowers the cost of long-context workloads like large-codebase analysis or document review that previously required a pricier "Pro"-class model.

AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.