
·
16 min read
DeepSeek API pricing depends on the selected model, cache-hit input tokens, cache-miss input tokens, output tokens, and whether requests run during peak or off-peak hours. As of October 1, 2026, V4.1 Flash starts at $0.003 per million cached input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens during off-peak hours.
Key Takeaways
DeepSeek bills cache-hit input tokens, cache-miss input tokens, and output tokens separately.
Off-peak prices are currently half the peak rates.
V4.1 Flash is substantially less expensive than V4 Pro per token.
OpenRouter prices vary by model checkpoint, provider, routing mode, and promotions.
The lowest price per token does not always yield the lowest cost per successful task.
Selective memory and retrieval, for example via Mem0, can reduce repeated context and uncached input usage.
DeepSeek API pricing at a glance
This section uses pricing from the official DeepSeek Models and Pricing page, verified on 2026-10-01 UTC.
All prices in this table are USD per one million tokens.
Model | Period | Cached input | Uncached input | Output |
|---|---|---|---|---|
DeepSeek V4.1 Flash | Off-peak | $0.003 / M | $0.15 / M | $0.60 / M |
DeepSeek V4.1 Flash | Peak | $0.006 / M | $0.30 / M | $1.20 / M |
DeepSeek V4 Pro | Off-peak | $0.022 / M | $0.66 / M | $1.98 / M |
DeepSeek V4 Pro | Peak | $0.044 / M | $1.32 / M | $3.96 / M |
Key facts for production builders:
deepseek-flash currently points to DeepSeek V4.1 Flash.
The direct DeepSeek model ID for Pro is
deepseek-v4-pro.Both models currently have a 1,000,000-token context window.
Maximum output is currently 384,000 tokens.
Always confirm details on the official pricing page before deploying.
Below are suggested actions for readers:
Calculate monthly cost
Compare models
Compare OpenRouter pricing
View official prices
How to calculate the DeepSeek API Cost?
Cost formula
For a single request in one rate period:
If usage spans both peak and off-peak, calculate each portion separately and sum:
Example Python cost helper
The following code uses the table above and can be adapted into a CLI, dashboard, or cost-guardrail service.
Calculator outputs
An implementation should show:
Cost per request
Daily cost
Monthly cost
Cached-input cost portion
Uncached-input cost portion
Output cost portion
Peak versus off-peak difference
Flash versus Pro difference
Effective blended cost per million tokens
This gives AI engineers a clear way to estimate DeepSeek API cost for different workloads.
How DeepSeek API pricing works
DeepSeek API pricing has four primary billing variables: model, input tokens, cache status, and output tokens, plus a rate period flag.
Model selection
DeepSeek currently promotes two main models for general-purpose use:
DeepSeek V4.1 Flash via
deepseek-flashDeepSeek V4 Pro via
deepseek-v4-pro
V4.1 Flash is designed for high-volume, cost-sensitive workloads with strong quality. V4 Pro targets more demanding tasks that may benefit from extra reasoning or multi-step outputs. Pro is more expensive per token, so teams should compare cost per successful task rather than assuming Pro always outperforms Flash.
Input tokens
Input tokens include everything sent to the API:
System instructions
Conversation history
Retrieved documents
Tool outputs and traces
Memory payloads and user profile details
Agents that naively replay the entire conversation plus all previous tool results on every call will grow uncached input rapidly, which increases cost. Mem0 and similar memory layers address this by selectively injecting only relevant context.
Cached versus uncached input
DeepSeek applies context caching to matching prompt prefixes. When a prefix has been processed recently, it may qualify for the lower cache-hit price. Tokens in new or non-matching prefixes are billed at cache-miss rates, which are significantly higher.
Key points:
Cache creation and reuse are best-effort and not guaranteed.
Only matching input prefixes benefit from cache-hit pricing.
Application changes to the early part of the prompt can reduce cache reuse.
Reordering content can affect which tokens are recognized as reusable.
Output tokens
Every generated token is billed as output, including:
Plain text replies
Long code segments
Detailed reasoning explanations
Tool traces if generated as part of the output
In tasks like content generation or large code diffs, output tokens can dominate the invoice even when inputs are modest.
Peak and off-peak rates
DeepSeek assigns each request to peak or off-peak periods:
Token prices are twice as high during peak as during off-peak.
Identical token usage may cost significantly more if executed in peak windows.
The same workloads can be scheduled differently for background tasks.
Scheduling batch jobs off-peak is often the simplest lever to reduce cost without changing prompts or models.
DeepSeek V4 Flash API pricing
Current V4 Flash pricing generally refers to DeepSeek V4.1 Flash, accessed through the deepseek-flash API model ID. Older V4 Flash names have either been retired or redirected to this model. Engineers should not assume that legacy model names represent separate price tiers.
V4.1 Flash price table
Pricing category | Off-peak | Peak |
|---|---|---|
Cache-hit input | $0.003 / M | $0.006 / M |
Cache-miss input | $0.15 / M | $0.30 / M |
Output | $0.60 / M | $1.20 / M |
V4.1 Flash capabilities and implications
Key configuration, based on DeepSeek documentation:
API model identifier:
deepseek-flashContext window: 1,000,000 tokens
Maximum output: 384,000 tokens
Modes: supports standard and reasoning modes where applicable
Vision: currently listed as supported for Flash
Tool calling: supports tools and structured output
Typical workloads:
Customer support chatbots
Coding assistants with moderate context
Agent planners with frequent but short calls
High-volume batch summarization
Cost implications:
Long outputs, such as full-page articles or long code, can significantly increase the output token bill.
The large context window allows big prompts, but repeated long prefixes without caching or memory discipline can raise costs.
Migration from older Flash model IDs should treat
deepseek-flashas the canonical target for evaluation.
DeepSeek V4 Pro API pricing
DeepSeek V4 Pro is available under the direct API model ID deepseek-v4-pro.
V4 Pro price table
Pricing category | Off-peak | Peak |
|---|---|---|
Cache-hit input | $0.022 / M | $0.044 / M |
Cache-miss input | $0.66 / M | $1.32 / M |
Output | $1.98 / M | $3.96 / M |
Pro offers higher prices across all token categories compared to Flash. It should not be assumed to outperform Flash on every task, especially short or simple ones.
Teams should:
Benchmark cost per successful task rather than focusing on raw token price.
Evaluate quality, tool usage, latency, and typical output length.
Use Pro only when task-level evaluation shows meaningful value over Flash that justifies the extra cost.
DeepSeek V4 Flash vs. V4 Pro pricing
The table below compares key factors for pricing and selection.
Factor | V4.1 Flash | V4 Pro |
|---|---|---|
Direct model ID |
|
|
Relative price | Lower | Higher |
Cache-hit cost | Lower | Higher |
Cache-miss cost | Lower | Higher |
Output cost | Lower | Higher |
Vision | Supported | Not currently listed |
Recommended use | High-volume production and cost-sensitive tasks | Tasks where testing demonstrates additional value |
Selection method | Default evaluation candidate | Benchmark against Flash before adoption |
Worked comparison example
Assumptions:
2,000 cached input tokens
3,000 uncached input tokens
1,000 output tokens
100,000 requests per month
All requests off-peak
Cost per request:
Flash off-peak:
Pro off-peak:
Monthly cost:
Flash: 100,000 × 0.001056 ≈ $105.60
Pro: 100,000 × 0.004004 ≈ $400.40
The same token usage costs nearly 3.8x more on Pro in this configuration, so real task-level benefits must justify that difference.
When are DeepSeek’s peak and off-peak hours?
DeepSeek’s current official schedule (checked 2026-10-01):
Peak: 01:00-04:00 UTC, Monday through Friday
Peak: 06:00-10:00 UTC, Monday through Friday
Off-peak: all remaining hours
Off-peak rates: 50 percent of peak rates
A local-time converter should be embedded in internal tools rather than hardcoding offsets. Time zones and daylight saving changes make static examples brittle.
Recommendations:
Display UTC times plus a clear timezone label, for example: “01:00-04:00 UTC (your local time: converted dynamically)”.
Avoid scheduling latency-sensitive tasks based only on price, because off-peak times for one region may still have high demand or different user expectations.
Schedule batch evaluation, summarization, and background processing off-peak where possible to reduce cost.
How DeepSeek context caching changes your bill
DeepSeek enables context caching automatically. When a prompt shares a matching prefix with a previously processed request, the matching portion may be served as cache-hit tokens at a lower price.
Key behaviors:
Prefix caching is automatic and best-effort.
Cache hits are not guaranteed even for repeated inputs.
Only matching prefixes benefit from cache-hit pricing.
Newly generated output is still billed at the output rate.
An example prompt structure:
Keeping reusable content, such as system prompts and static reference material, at the beginning of the prompt can improve opportunities for prefix reuse. This does not guarantee a specific cache-hit ratio, but it aligns with DeepSeek’s prefix-based caching design.
Applications can inspect token breakdown metrics to track cache-hit and cache-miss tokens and feed that into cost dashboards.
DeepSeek V4 Pro OpenRouter pricing
DeepSeek V4 Pro pricing on OpenRouter differs from direct DeepSeek API pricing because OpenRouter can route requests across multiple inference providers. The final listed rate can depend on the model checkpoint, selected provider, routing mode, cache support, batch mode, and temporary discounts.
This section uses OpenRouter model pages checked on 2026-10-01 UTC. Readers should always confirm the latest figures.
Available V4 Pro model IDs on OpenRouter
OpenRouter model | Meaning |
|---|---|
| Newer dated V4 Pro 0813 checkpoint |
| Earlier V4 Pro 0423 checkpoint |
| Alias pointing to the latest Pro-family model |
| Batch variant, when available |
Engineers should treat each model ID as potentially having different prices, providers, and capabilities.
What pricing information to display
For each OpenRouter DeepSeek V4 Pro model, it is useful to surface:
Model ID
Model checkpoint date
Lowest available input rate (USD / 1M)
Lowest available output rate (USD / 1M)
Cache-read rate if supported
Default or displayed route
Batch rate, if present
Context length
Number of providers for this model
Date and time checked
OpenRouter’s V4 Pro pages expose provider-specific rates and routing options, so UI tools should link directly to the live V4 Pro 0813 page. Prices can change with provider promotions or capacity constraints.
Why OpenRouter rates change
OpenRouter introduces more pricing variability:
Multiple providers may serve the same model ID.
Different providers set different token prices.
Temporary promotions can change headline rates.
A cheapest-provider route can have a different price from a fastest-provider route.
Pinning a provider can trade cost for latency or reliability.
Batch endpoints may have separate rates and minimums.
Cache-read support can vary between providers.
Applications should avoid hardcoding OpenRouter prices and instead treat them as dynamic configuration.
Editorial rule for OpenRouter prices
OpenRouter pricing references should follow this pattern:
OpenRouter rates checked on [date and UTC time]. Provider-level prices and promotional discounts can change independently.
Undated OpenRouter price claims should not appear in introductions or meta descriptions.
Direct DeepSeek API vs. OpenRouter pricing
The table below summarizes structural differences.
Factor | Direct DeepSeek API | OpenRouter |
|---|---|---|
Provider model | Served directly by DeepSeek | Multiple inference providers |
Price structure | Peak and off-peak | Provider- and route-dependent |
Model selection | Current official IDs | Multiple checkpoints and aliases |
Cache pricing | Official DeepSeek cache-hit rates | Model- and provider-dependent |
Routing control | Direct provider | Cheapest, fastest, balanced, pinned, or other routes |
Failover | Managed by DeepSeek | Potential cross-provider routing |
Billing relationship | DeepSeek | OpenRouter |
Best for | Direct access and official billing | Multi-model access and routing flexibility |
Answering a common question:
Is OpenRouter cheaper than direct DeepSeek?
Sometimes yes, sometimes no. The final cost depends on model checkpoint, chosen provider, routing policy, cache support, promotions, and whether the direct alternative would have run during peak or off-peak hours.
DeepSeek vs every major API: the full comparison
This section positions DeepSeek in the broader LLM API pricing landscape. Prices here must be rechecked at publication time and normalized to USD per one million tokens. Reasoning tokens, if billed separately, should be classified as input or output per provider documentation.
A typical comparison table might include:
Provider | Flagship model | Input / 1M tokens | Output / 1M tokens | Context window |
|---|---|---|---|---|
DeepSeek | V4.1 Flash | $0.15 off-peak / $0.30 peak | $0.60 off-peak / $1.20 peak | 1M |
DeepSeek | V4 Pro | $0.66 off-peak / $1.32 peak | $1.98 off-peak / $3.96 peak | 1M |
OpenAI | GPT-6 Astra | $10.00 | $50.00 | 1.05M |
OpenAI | GPT-6.1 Sol | $2.00 | $10.00 | 1.05M |
Anthropic | Claude Fable 5.1 | $10.00 | $50.00 | 1M |
Anthropic | Claude Opus 5.5 | $4.00 | $20.00 | 1M |
Gemini 3.1 Pro Preview | $2.00 ≤200K / $4.00 >200K | $12.00 ≤200K / $18.00 >200K | 1M | |
Gemini 3.8 Flash | $0.75* | $3.75* | 1M | |
xAI | Grok 4.7 | $2.00 | $6.00 | 500K |
xAI | Grok 4.3 | $1.25 ≤200K / $2.50 >200K | $2.50 ≤200K / $5.00 >200K | 1M |
Mistral | Mistral Medium 3.5 | $1.50 | $7.50 | 256K |
Mistral | Mistral Large 3 | $0.50 | $1.50 | 256K |
Alibaba | Qwen3.8 Max | $2.00 | $6.00 | 1M |
Alibaba | Qwen3.8 Flash | $0.15 | $0.47 | 1M |
Meta via Together AI | Llama 4 Maverick | $0.27 | $0.85 | ~1M |
Meta via Together AI | Llama 4 Scout | $0.18 | $0.59 | ~1M |
Note: Mentioned pricing is as of 1 Oct, 2026, sourced from respective official websites.
Recommended comparison strategy:
Match V4.1 Flash against fast/low-cost models from each provider.
Match V4 Pro against higher-capability models from each provider.
Separate standard and batch tiers.
Include reasoning tokens where applicable.
Describe cache pricing explicitly when providers expose it.
From this comparison, engineers can distinguish:
Cheapest input per million tokens.
Cheapest output per million tokens.
Cheapest repeated-context workloads (where caching matters).
Cheapest batch workloads.
Lowest estimated cost per completed task, which is more operationally meaningful.
Worked DeepSeek API cost examples
This section uses the cost formulas from earlier to ground token pricing in realistic agent workloads. All assumptions are explicitly stated.
Customer-support chatbot
Assumptions:
Model:
deepseek-flashPeriod: off-peak
Stable system prompt and policy: 1,500 tokens
Recent messages included: 1,000 tokens
Retrieved knowledge: 1,500 tokens
Output size: 600 tokens
Cache-hit percentage on system + policy: 80 percent
Requests per month: 1,000,000
Token breakdown (average per request):
Cached input: 1,500 × 0.8 = 1,200 tokens
Uncached input: 1,000 (chat) + 1,500 (retrieved) + 300 uncached portion of system ≈ 2,800 tokens
Output: 600 tokens
Cost per request:
Monthly total:
1,000,000 requests × 0.0007836 ≈ $783.60
If 90 percent of tickets resolve successfully, cost per successful task ≈ $783.60 / 900,000 ≈ $0.00087.
Coding assistant
Assumptions:
Model:
deepseek-v4-proPeriod: peak
Repository context and file snippets: 5,000 tokens
Tool results and error logs: 3,000 tokens
Conversation history: 2,000 tokens
Output code: 2,500 tokens
Minimal caching; assume everything is uncached for the worst case
Calls per task: 4
Tasks per month: 20,000
Per-call tokens:
Cached input: 0
Uncached input: 10,000 tokens
Output: 2,500 tokens
Cost per call:
Per task (4 calls): 4 × 0.0231 ≈ $0.0924
Monthly total:
20,000 tasks × 0.0924 ≈ $1,848
If 85 percent of tasks succeed, cost per successful task ≈ $1,848 / 17,000 ≈ $0.1088.
Document analysis
Assumptions:
Model:
deepseek-flashPeriod: off-peak
Long input document: 40,000 tokens
Short analytical output: 800 tokens
Minimal caching, single-pass per document
Calls per task: 1
Tasks per month: 50,000
Tokens per request:
Cached input: 0
Uncached input: 40,000
Output: 800
Cost per request:
Monthly total:
50,000 × 0.00648 ≈ $324
If 95 percent success, cost per successful task ≈ $324 / 47,500 ≈ $0.00682.
Content-generation workflow
Assumptions:
Model:
deepseek-flashPeriod: off-peak
Prompt: 3,000 tokens
Output article: 4,000 tokens
Minimal caching
Calls per task: 1
Tasks per month: 30,000
Tokens:
Cached input: 0
Uncached input: 3,000
Output: 4,000
Cost per request:
Monthly total:
30,000 × 0.00285 ≈ $85.50
Output tokens dominate, about 84 percent of the cost.
Long-running AI agent
Assumptions:
Model:
deepseek-flashPeriod: mixed; assume half peak, half off-peak for simplicity
Conversation history per step: 8,000 tokens
Memory retrieval per step: 2,000 tokens
Tool use per step: 3,000 tokens
Output per step: 1,500 tokens
10 steps per task
Tasks per month: 10,000
Some stable prefix cached; assume 3,000 cached tokens per step
Per step tokens:
Cached input: 3,000
Uncached input: 10,000 (remaining)
Output: 1,500
Off-peak cost per step:
Peak cost per step:
Average per step (half peak, half off-peak):
≈ (0.002409 + 0.004818) / 2 ≈ $0.0036135
Per task (10 steps): ≈ $0.036135
Monthly total:
10,000 tasks × 0.036135 ≈ $361.35
If 80 percent success, cost per successful task ≈ $361.35 / 8,000 ≈ $0.045.
These examples show how input vs output balance, cache usage, and peak scheduling affect DeepSeek API pricing.
The costs not shown on DeepSeek’s pricing page
DeepSeek invoices only cover LLM token usage. Production agents incur additional costs that are not reflected in the LLM pricing table, including:
Embedding models and vector storage
Search and reranking services
Memory layers and user profile storage
Tool and third-party API calls
Orchestration frameworks and job queues
Retries and failed generations
Logging, metrics, and auditing
Evaluation and red-teaming pipelines
Security controls and compliance tooling
Networking, bandwidth, and egress
Data processing and ETL pipelines
Engineering time for prompt and agent design
Self-hosting infrastructure for complementary components
Understanding total cost of ownership requires tracking both DeepSeek tokens and these surrounding elements.
How to reduce DeepSeek API costs
Practical steps for AI engineers:
Schedule flexible jobs, such as batch summarization and offline evaluation, during off-peak hours.
Use V4.1 Flash by default, and upgrade to V4 Pro only when evaluations justify it.
Keep reusable prompt prefixes stable and early in the prompt to improve cache reuse.
Monitor actual cache-hit and cache-miss token counts via DeepSeek metrics.
Limit unnecessary output and avoid verbose free-form reasoning where not needed.
Retrieve
Useful resources

Aashi Dutt
She is a senior technical content writer at Mem0. She covers agent memory architecture and the engineering decisions behind building agents that actually remember. She experiments with new features and turns research into posts developers can put straight to use.
Start building with memory
Free tier, no card









