Skip to content
Rush Commerce
AI & Automation3 min read

DeepSeek's Aug 16 price hike: US hours are off-peak

DeepSeek's new peak/off-peak API pricing lands August 16 at 16:00 UTC, with some rates up 12x. US business hours fall entirely in the off-peak window.

The cheap tier is over. Tomorrow at 16:00 UTC, DeepSeek's API price hike takes effect — a peak/off-peak schedule that raises every rate on V4-Flash and V4-Pro, with the worst line item going up more than twelvefold. If you have automations pointed at DeepSeek, your cost model expires in about a day. The useful detail nobody is leading with: for a US operator, the expensive window is the middle of the night.

What actually happened

Per DeepSeek's own pricing docs, the new schedule starts August 16, 2026 at 16:00 UTC. Peak hours run 01:00–04:00 and 06:00–10:00 UTC. Everything else is off-peak, billed at exactly half the peak rate.

The numbers, per million tokens:

  • V4-Flash cache-miss input goes from $0.14 flat to $0.22 off-peak / $0.44 peak. Output goes from $0.28 to $0.66 / $1.32.
  • V4-Pro cache-miss input goes from $0.435 to $0.66 / $1.32. Output goes from $0.87 to $1.98 / $3.96.
  • The brutal one is cached input. V4-Pro cache hits go from $0.003625 to $0.044 at peak — a 12x increase, or roughly 1,100%.

Engadget pegs V4-Pro output as roughly a fourfold rise. DeepSeek's stated reason is to "allocate resources more reasonably" — inference congestion, priced.

Why the DeepSeek price hike matters for your business

Off-peak is not a discount. It's a smaller increase. V4-Flash output at the off-peak rate is still 2.4x what you pay today. Anyone reading "half price off-peak" as a way to hold costs flat is going to miss by a lot. Reprice the workload, not the schedule.

Those peak windows are Beijing business hours, not yours. 01:00–10:00 UTC is 09:00–18:00 in Beijing. In Phoenix, that's 18:00–03:00. Which means a US shop running agents from 9am to 5pm local — 16:00 to 00:00 UTC — sits entirely in off-peak. The batch job you thoughtfully scheduled to run overnight is the one that now costs double.

Cache economics just inverted. Prompt caching was DeepSeek's sharpest edge: cache hits cost ~2% of a cache miss. After Wednesday it's ~10%. Long system prompts you were reusing for nearly free are now a real line item.

If you can't answer "what does this workflow cost per run, on which model, at what hour" from a dashboard, this hike will find you in next month's invoice.

Key takeaways

  • New DeepSeek peak/off-peak pricing starts August 16, 2026 at 16:00 UTC; off-peak is exactly half of peak
  • Peak hours are 01:00–04:00 and 06:00–10:00 UTC — Beijing business hours, overnight in the US
  • V4-Flash output: $0.28 flat becomes $0.66 off-peak / $1.32 peak; V4-Pro output becomes $1.98 / $3.96
  • V4-Pro cached input rises ~12x, from $0.003625 to $0.044 at peak — the biggest single jump
  • US daytime work lands in off-peak; move overnight batch jobs into your own business hours
  • Every rate went up, so re-run the cost math rather than assuming off-peak keeps you whole

A model price change should be a config edit, not a rewrite. We build AI systems with the model layer behind one interface and per-run cost logged, so a vendor's Wednesday repricing is a routing decision instead of a fire drill. See how we build portable AI systems or put real numbers on your automation spend.

Sources: DeepSeek API docs, Engadget.

  • #deepseek
  • #llm-pricing
  • #model-routing
  • #ai-costs
  • #inference
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.