AI Tools
What is DeepSeek's new peak and off-peak API pricing?
DeepSeek API pricing changes on August 16, 2026 at 16:00 UTC, when peak and off-peak rates take effect: off-peak hours cost 50% less than peak, with deepseek-v4-flash output at $0.66 per 1M tokens off-peak and $1.32 at peak. Peak hours are 01:00-04:00 and 06:00-10:00 UTC.
What matters
- DeepSeek switches to peak/off-peak API billing on Aug 16, 2026 at 16:00 UTC; off-peak rates are exactly 50% of peak rates.
- Peak hours are 01:00-04:00 and 06:00-10:00 UTC, which is China business time; the other 17 hours are off-peak, including all US daytime.
- New off-peak output prices per 1M tokens: deepseek-v4-flash $0.66 and deepseek-v4-pro $1.98, versus $1.32 and $3.96 at peak.
- The new card is higher than the flat rates the same page listed before the change: V4-Flash output rises from $0.28 to $0.66 off-peak.
- Off-peak V4-Flash output undercuts OpenAI gpt-5.6-luna ($1.20) and Anthropic Haiku 4.5 ($5.00) on output tokens.
What changed in DeepSeek API pricing
DeepSeek is an AI lab whose API currently offers two models: deepseek-v4-flash and deepseek-v4-pro. On August 16, 2026 at 16:00 UTC, its billing moves from flat per-token rates to time-based peak and off-peak rates, with off-peak priced at half the peak rate. The change was announced on the official DeepSeek API docs on August 13, 2026, alongside the DeepSeek-V4-Pro general availability release.
The same announcement covers V4-Pro's production features: agent upgrades, flexible reasoning effort (low, high, and max settings), and native OpenAI Responses API support optimized for Codex with one-click setup. V4-Pro is available through the API with model names unchanged, and on the app and web via Expert Mode. The pricing change is the part that affects every existing API customer's bill.
The new DeepSeek rate card, per 1M tokens
The official pricing page lists rates in units of per 1M tokens, broken into three components: input tokens served from context cache (cache hit), input tokens that miss the cache (cache miss), and output tokens. Context caching is built in, so cache-hit input is far cheaper than cache-miss input. The table below is the complete new schedule, verified on the official page on August 15, 2026.
Model versions on the current rate card are DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813. Both models support a 1M-token context window with up to 384K output tokens, and both default to thinking mode with a non-thinking option. Concurrency limits are 2500 for flash and 500 for pro.
When is off-peak in your timezone?
Peak hours are 01:00-04:00 and 06:00-10:00 UTC, and every other hour is off-peak, per the official pricing page. In Beijing time (UTC+8), peak falls at 09:00-12:00 and 14:00-18:00: the peak window is essentially China's business day. The DST conversions below reflect August 2026.
- US Eastern (EDT, UTC-4): peak at 21:00-00:00 and 02:00-06:00, so the working day is off-peak.
- US Pacific (PDT, UTC-7): peak at 18:00-21:00 and 23:00-03:00, again outside working hours.
- Central Europe (CEST, UTC+2): peak at 03:00-06:00 and 08:00-12:00, which covers the late morning.
For teams in the Americas, scheduling batch jobs, embeddings, and nightly syncs during the local working day lands in the off-peak window automatically. China-based teams and anyone serving China daytime traffic will find peak hours unavoidable for real-time requests.
The honest catch: off-peak is 50% off peak, not 50% off the old price
The framing in the announcement is that off-peak rates are 50% lower than peak. What the announcement does not emphasize is that the new card is higher than the flat rates the same page listed before the change. Compare the old flat rates with the new off-peak rates, per 1M tokens:
- V4-Flash output: $0.28 before, $0.66 off-peak, $1.32 at peak (2.4x and 4.7x the old price).
- V4-Pro output: $0.87 before, $1.98 off-peak, $3.96 at peak (2.3x and 4.6x).
- V4-Flash cache-hit input: $0.0028 before, $0.007 off-peak, $0.014 at peak.
- V4-Pro cache-hit input: $0.003625 before, $0.022 off-peak, $0.044 at peak.
So an off-peak bill is still higher than the same workload cost before August 16. The old flat prices were $0.0028 / $0.14 / $0.28 for V4-Flash and $0.003625 / $0.435 / $0.87 for V4-Pro (cache hit, cache miss, output). The pricing page also notes that product prices may vary and DeepSeek reserves the right to adjust them, so treat the card as the current snapshot, captured August 15, 2026.
DeepSeek vs OpenAI vs Anthropic API pricing
DeepSeek's own announcement is one source; the comparison below is the synthesis across vendor pricing pages. Per 1M tokens, standard tiers:
- OpenAI gpt-5.6-sol: $5.00 input / $30.00 output; gpt-5.6-terra: $2.00 / $12.00; gpt-5.6-luna: $0.20 / $1.20.
- Anthropic Fable 5: $10.00 / $50.00; Opus 5: $5.00 / $25.00; Sonnet 5: $2.00 / $10.00; Haiku 4.5: $1.00 / $5.00.
- DeepSeek V4-Flash off-peak: $0.22 cache-miss input / $0.66 output; at peak: $0.44 / $1.32.
- DeepSeek V4-Pro off-peak: $0.66 / $1.98; at peak: $1.32 / $3.96.
On output tokens, off-peak V4-Flash at $0.66 undercuts OpenAI's cheapest flagship gpt-5.6-luna ($1.20) and Anthropic's Haiku 4.5 ($5.00). At peak, V4-Flash output at $1.32 edges slightly above luna's $1.20, which is the first case where DeepSeek's flash tier loses the price comparison. V4-Pro at peak ($1.32 input / $3.96 output) still undercuts gpt-5.6-terra ($2.00 / $12.00) and Sonnet 5 ($2.00 / $10.00). DeepSeek cache-hit input ($0.007 to $0.044) remains far cheaper than OpenAI cached input ($0.02 to $0.50) and Anthropic cache reads ($0.10 to $1.00).
Who should shift workloads to off-peak DeepSeek usage
If your traffic is queueable, the off-peak window is a real cost lever. The most obvious candidates are batch inference, nightly indexing and sync jobs, bulk classification, and RAG pipelines, where a few hours of latency does not matter. Agent workloads that can tolerate slower cadence fit too, especially with V4-Pro's low reasoning-effort setting for simple tasks.
Teams in North America and Europe get the off-peak card during their working day without any code changes. Teams whose traffic peaks during China business hours (09:00-18:00 Beijing time) will mostly pay peak rates unless they queue work. Real-time user-facing apps that cannot defer requests should budget at the peak card: V4-Flash output at $1.32 per 1M tokens or V4-Pro at $3.96.
Two operational details from the official page matter for cost planning: billing is deducted from your topped-up balance or granted balance (granted balance is consumed first), and the concurrency ceiling is 2500 for V4-Flash and 500 for V4-Pro, which caps how much parallel off-peak traffic one account can push.
At a glance
| DeepSeek V4 rate (per 1M tokens) | Cache hit input | Cache miss input | Output |
|---|---|---|---|
| deepseek-v4-flash, off-peak | $0.007 | $0.22 | $0.66 |
| deepseek-v4-flash, peak | $0.014 | $0.44 | $1.32 |
| deepseek-v4-pro, off-peak | $0.022 | $0.66 | $1.98 |
| deepseek-v4-pro, peak | $0.044 | $1.32 | $3.96 |
FAQ
Is the DeepSeek API free?
No. The official pricing page bills per token in units of per 1M tokens, and fees are deducted from your topped-up balance or granted balance. No free tier is listed. New peak/off-peak rates apply from August 16, 2026 at 16:00 UTC.
When are DeepSeek off-peak hours?
Peak hours are 01:00-04:00 and 06:00-10:00 UTC, per the official pricing page. All other hours are off-peak, which is 17 hours per day, and off-peak rates are exactly 50% of peak rates.
Is DeepSeek cheaper than OpenAI for API use?
At off-peak rates, mostly yes: V4-Flash output at $0.66 per 1M tokens beats OpenAI's cheapest flagship gpt-5.6-luna at $1.20. At peak, V4-Flash output at $1.32 slightly exceeds luna's $1.20. The new DeepSeek card is also higher than its own previous flat rates.