DeepSeek's Aug 16 price hike: US hours are off-peak
DeepSeek's new peak/off-peak API pricing lands August 16 at 16:00 UTC, with some rates up 12x. US business hours fall entirely in the off-peak window.
The cheap tier is over. Tomorrow at 16:00 UTC, DeepSeek's API price hike takes effect — a peak/off-peak schedule that raises every rate on V4-Flash and V4-Pro, with the worst line item going up more than twelvefold. If you have automations pointed at DeepSeek, your cost model expires in about a day. The useful detail nobody is leading with: for a US operator, the expensive window is the middle of the night.
What actually happened
Per DeepSeek's own pricing docs, the new schedule starts August 16, 2026 at 16:00 UTC. Peak hours run 01:00–04:00 and 06:00–10:00 UTC. Everything else is off-peak, billed at exactly half the peak rate.
The numbers, per million tokens:
- V4-Flash cache-miss input goes from $0.14 flat to $0.22 off-peak / $0.44 peak. Output goes from $0.28 to $0.66 / $1.32.
- V4-Pro cache-miss input goes from $0.435 to $0.66 / $1.32. Output goes from $0.87 to $1.98 / $3.96.
- The brutal one is cached input. V4-Pro cache hits go from $0.003625 to $0.044 at peak — a 12x increase, or roughly 1,100%.
Engadget pegs V4-Pro output as roughly a fourfold rise. DeepSeek's stated reason is to "allocate resources more reasonably" — inference congestion, priced.
Why the DeepSeek price hike matters for your business
Off-peak is not a discount. It's a smaller increase. V4-Flash output at the off-peak rate is still 2.4x what you pay today. Anyone reading "half price off-peak" as a way to hold costs flat is going to miss by a lot. Reprice the workload, not the schedule.
Those peak windows are Beijing business hours, not yours. 01:00–10:00 UTC is 09:00–18:00 in Beijing. In Phoenix, that's 18:00–03:00. Which means a US shop running agents from 9am to 5pm local — 16:00 to 00:00 UTC — sits entirely in off-peak. The batch job you thoughtfully scheduled to run overnight is the one that now costs double.
Cache economics just inverted. Prompt caching was DeepSeek's sharpest edge: cache hits cost ~2% of a cache miss. After Wednesday it's ~10%. Long system prompts you were reusing for nearly free are now a real line item.
If you can't answer "what does this workflow cost per run, on which model, at what hour" from a dashboard, this hike will find you in next month's invoice.
Key takeaways
- New DeepSeek peak/off-peak pricing starts August 16, 2026 at 16:00 UTC; off-peak is exactly half of peak
- Peak hours are 01:00–04:00 and 06:00–10:00 UTC — Beijing business hours, overnight in the US
- V4-Flash output: $0.28 flat becomes $0.66 off-peak / $1.32 peak; V4-Pro output becomes $1.98 / $3.96
- V4-Pro cached input rises ~12x, from $0.003625 to $0.044 at peak — the biggest single jump
- US daytime work lands in off-peak; move overnight batch jobs into your own business hours
- Every rate went up, so re-run the cost math rather than assuming off-peak keeps you whole
A model price change should be a config edit, not a rewrite. We build AI systems with the model layer behind one interface and per-run cost logged, so a vendor's Wednesday repricing is a routing decision instead of a fire drill. See how we build portable AI systems or put real numbers on your automation spend.
Sources: DeepSeek API docs, Engadget.
- #deepseek
- #llm-pricing
- #model-routing
- #ai-costs
- #inference
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Anthropic's IPO rests on a $200B 2028 revenue forecast
Reuters reports Anthropic's IPO valuation hinges on $190-200B in 2028 revenue, up from a $47B run rate. Read what that growth assumption implies for your bill.
Read itAgentic AI Foundation hits 247 members: Visa, Wells Fargo join
The Agentic AI Foundation added 57 members and now governs MCP, goose, AGENTS.md and agentgateway. Neutral governance is what makes agent work portable.
Read it