DeepSeek Finalizes Steep API Price Hikes, Ending Its Flat-Rate Era
DeepSeek confirmed a new peak/off-peak pricing structure for its V4 Pro and V4 Flash models effective August 16, raising some output-token rates by more than 1,100% as the low-cost Chinese lab moves away from the flat pricing that fueled its rise.
DeepSeek's API pricing page confirmed this week that its flat, famously cheap token rates are gone. Effective August 16 at 16:00 UTC, the Hangzhou-based lab replaced single flat prices for DeepSeek-V4-Pro and DeepSeek-V4-Flash with a two-tier peak and off-peak schedule, with increases ranging from roughly 50% to more than 1,100% depending on the model, token type and time of day.
What's changing
Peak hours run 01:00–04:00 UTC and 06:00–10:00 UTC daily, windows that overlap peak business hours in Asia and the European morning. Off-peak rates are set at half the peak price. For V4-Pro, output tokens that previously cost a flat $0.87 per million now cost $3.96 per million at peak and $1.98 per million off-peak — more than a fourfold increase even at the cheaper rate. V4-Flash output climbs from a flat $0.28 per million to $1.32 at peak and $0.66 off-peak. DeepSeek says the shift to dynamic pricing is meant to "allocate resources more reasonably" as demand for its models has strained serving capacity.
A pattern DeepSeek flagged but didn't finalize
The increase isn't a surprise: DeepSeek warned developers on August 6 that a "significant" price change was coming, and when the company took DeepSeek-V4-Pro-0813 out of preview on August 13, its own pricing page already flagged that a hike was imminent without publishing final numbers. This week's update fills in those figures and, notably, introduces time-of-day billing — a first for the company — rather than simply raising a flat rate.
Why it matters
DeepSeek's ultra-low API pricing was a core part of its appeal over the past two years, undercutting Western labs badly enough to trigger price cuts across the industry and driving rapid adoption among developers building cost-sensitive agents and coding tools. Ending flat-rate billing in favor of demand-based pricing brings DeepSeek's economics closer to how cloud compute is typically priced, and it signals the company is prioritizing margin and capacity management over being the cheapest option on the market. For teams that built agent pipelines around DeepSeek's near-zero token costs, the change means re-evaluating architecture and budgets — or shifting latency-tolerant workloads into DeepSeek's new off-peak windows to blunt the impact.
Sources
- DeepSeek raises API prices by up to 1,100% starting Aug. 16 — Quartz
- DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity — InfoWorld
- DeepSeek's AI models are about to cost four times more — Engadget
- DeepSeek Introduces Peak-Hour Pricing That Quadruples Current Levels — PYMNTS
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
Related agents
More ai news
EU Set to Propose Barring Under-15s From AI Chatbots and Social Media in 'Kids Act'
The European Commission is preparing to unveil an EU Kids Act that would bar unsupervised access to AI chatbots, social media, video platforms and online games for under-15s, with tiered rules and mandatory age verification for 13-14 year-olds.
Amodei's 'Pace the Frontier' Plan Draws Same-Day Backing From OpenAI, DeepMind and xAI
Anthropic CEO Dario Amodei published an essay arguing frontier AI labs should deliberately slow capability gains, and within hours Sam Altman, Demis Hassabis and Elon Musk publicly endorsed the idea, with Microsoft's Satya Nadella following a day later.
Positron Raises $875M to Build an HBM-Free AI Inference Chip
Chip startup Positron closed an $875 million Series C at a $5 billion post-money valuation to fund its Asimov inference accelerator, which pairs its compute architecture with up to 2,304GB of commodity LPDDR5X memory instead of scarce high-bandwidth memory.
DeepSeek Releases V4.1 Flash, Cuts API Prices and Sets End Date for V4 Pro
DeepSeek officially released V4.1 Flash, a cheaper and faster multimodal successor to V4 Pro with a 1-million-token context window, and said it will reroute all V4 Pro API traffic to the new model from September 14.