/

/

/

DeepSeek API Pricing in 2026: V4.1 Flash, V4 Pro, and How to Save

Share

DeepSeek API Pricing in 2026: V4.1 Flash, V4 Pro, and How to Save

DeepSeek API Pricing in 2026: V4.1 Flash, V4 Pro, and How to Save

DeepSeek API Pricing in 2026: V4.1 Flash, V4 Pro, and How to Save

aashi dutt
Published on Oct 1, 2026

·

16 min read

On this page

DeepSeek API pricing depends on the selected model, cache-hit input tokens, cache-miss input tokens, output tokens, and whether requests run during peak or off-peak hours. As of October 1, 2026, V4.1 Flash starts at $0.003 per million cached input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens during off-peak hours.

Key Takeaways

  • DeepSeek bills cache-hit input tokens, cache-miss input tokens, and output tokens separately.

  • Off-peak prices are currently half the peak rates.

  • V4.1 Flash is substantially less expensive than V4 Pro per token.

  • OpenRouter prices vary by model checkpoint, provider, routing mode, and promotions.

  • The lowest price per token does not always yield the lowest cost per successful task.

  • Selective memory and retrieval, for example via Mem0, can reduce repeated context and uncached input usage.

DeepSeek API pricing at a glance

This section uses pricing from the official DeepSeek Models and Pricing page, verified on 2026-10-01 UTC.

All prices in this table are USD per one million tokens.

Model

Period

Cached input

Uncached input

Output

DeepSeek V4.1 Flash
deepseek-flash

Off-peak

$0.003 / M

$0.15 / M

$0.60 / M

DeepSeek V4.1 Flash
deepseek-flash

Peak

$0.006 / M

$0.30 / M

$1.20 / M

DeepSeek V4 Pro
deepseek-v4-pro

Off-peak

$0.022 / M

$0.66 / M

$1.98 / M

DeepSeek V4 Pro
deepseek-v4-pro

Peak

$0.044 / M

$1.32 / M

$3.96 / M

Key facts for production builders:

  • deepseek-flash currently points to DeepSeek V4.1 Flash.

  • The direct DeepSeek model ID for Pro is deepseek-v4-pro.

  • Both models currently have a 1,000,000-token context window.

  • Maximum output is currently 384,000 tokens.

  • Always confirm details on the official pricing page before deploying.

Below are suggested actions for readers:

  • Calculate monthly cost

  • Compare models

  • Compare OpenRouter pricing

  • View official prices

How to calculate the DeepSeek API Cost?

Cost formula

For a single request in one rate period:

Total cost =
(cached_input_tokens / 1_000_000) * cache_hit_price
+
(uncached_input_tokens / 1_000_000) * cache_miss_price
+
(output_tokens / 1_000_000) * output_price
Total cost =
(cached_input_tokens / 1_000_000) * cache_hit_price
+
(uncached_input_tokens / 1_000_000) * cache_miss_price
+
(output_tokens / 1_000_000) * output_price
Total cost =
(cached_input_tokens / 1_000_000) * cache_hit_price
+
(uncached_input_tokens / 1_000_000) * cache_miss_price
+
(output_tokens / 1_000_000) * output_price

If usage spans both peak and off-peak, calculate each portion separately and sum:

Total monthly cost = off_peak_cost + peak_cost
Total monthly cost = off_peak_cost + peak_cost
Total monthly cost = off_peak_cost + peak_cost

Example Python cost helper

The following code uses the table above and can be adapted into a CLI, dashboard, or cost-guardrail service.

from dataclasses import dataclass
from typing import Literal

ModelName = Literal["deepseek-flash", "deepseek-v4-pro"]
Period = Literal["peak", "off_peak"]

@dataclass
class DeepSeekRate:
    cached_input: float   # USD per 1M tokens
    uncached_input: float
    output: float

RATES = {
    ("deepseek-flash", "off_peak"): DeepSeekRate(0.003, 0.15, 0.60),
    ("deepseek-flash", "peak"):     DeepSeekRate(0.006, 0.30, 1.20),
    ("deepseek-v4-pro", "off_peak"):DeepSeekRate(0.022, 0.66, 1.98),
    ("deepseek-v4-pro", "peak"):    DeepSeekRate(0.044, 1.32, 3.96),
}

TOKENS_PER_MILLION = 1_000_000.0

def cost_per_request(
    model: ModelName,
    period: Period,
    cached_input_tokens: int,
    uncached_input_tokens: int,
    output_tokens: int,
) -> float:
    rate = RATES[(model, period)]
    return (
        (cached_input_tokens / TOKENS_PER_MILLION) * rate.cached_input
        + (uncached_input_tokens / TOKENS_PER_MILLION) * rate.uncached_input
        + (output_tokens / TOKENS_PER_MILLION) * rate.output
    )

def monthly_cost(
    model: ModelName,
    period: Period,
    cached_input_tokens: int,
    uncached_input_tokens: int,
    output_tokens: int,
    requests_per_day: int,
    active_days_per_month: int,
) -> float:
    cpr = cost_per_request(
        model,
        period,
        cached_input_tokens,
        uncached_input_tokens,
        output_tokens,
    )
    return cpr * requests_per_day * active_days_per_month

if __name__ == "__main__":
    # Example: 3k cached, 2k uncached, 1k output tokens, 50k req/day, 30 days, Flash off-peak
    cpr = cost_per_request("deepseek-flash", "off_peak", 3000, 2000, 1000)
    monthly = monthly_cost("deepseek-flash", "off_peak", 3000, 2000, 1000, 50_000, 30)
    print(f"Cost per request: ${cpr:.6f}")
    print(f"Monthly cost: ${monthly:,.2f}")
from dataclasses import dataclass
from typing import Literal

ModelName = Literal["deepseek-flash", "deepseek-v4-pro"]
Period = Literal["peak", "off_peak"]

@dataclass
class DeepSeekRate:
    cached_input: float   # USD per 1M tokens
    uncached_input: float
    output: float

RATES = {
    ("deepseek-flash", "off_peak"): DeepSeekRate(0.003, 0.15, 0.60),
    ("deepseek-flash", "peak"):     DeepSeekRate(0.006, 0.30, 1.20),
    ("deepseek-v4-pro", "off_peak"):DeepSeekRate(0.022, 0.66, 1.98),
    ("deepseek-v4-pro", "peak"):    DeepSeekRate(0.044, 1.32, 3.96),
}

TOKENS_PER_MILLION = 1_000_000.0

def cost_per_request(
    model: ModelName,
    period: Period,
    cached_input_tokens: int,
    uncached_input_tokens: int,
    output_tokens: int,
) -> float:
    rate = RATES[(model, period)]
    return (
        (cached_input_tokens / TOKENS_PER_MILLION) * rate.cached_input
        + (uncached_input_tokens / TOKENS_PER_MILLION) * rate.uncached_input
        + (output_tokens / TOKENS_PER_MILLION) * rate.output
    )

def monthly_cost(
    model: ModelName,
    period: Period,
    cached_input_tokens: int,
    uncached_input_tokens: int,
    output_tokens: int,
    requests_per_day: int,
    active_days_per_month: int,
) -> float:
    cpr = cost_per_request(
        model,
        period,
        cached_input_tokens,
        uncached_input_tokens,
        output_tokens,
    )
    return cpr * requests_per_day * active_days_per_month

if __name__ == "__main__":
    # Example: 3k cached, 2k uncached, 1k output tokens, 50k req/day, 30 days, Flash off-peak
    cpr = cost_per_request("deepseek-flash", "off_peak", 3000, 2000, 1000)
    monthly = monthly_cost("deepseek-flash", "off_peak", 3000, 2000, 1000, 50_000, 30)
    print(f"Cost per request: ${cpr:.6f}")
    print(f"Monthly cost: ${monthly:,.2f}")
from dataclasses import dataclass
from typing import Literal

ModelName = Literal["deepseek-flash", "deepseek-v4-pro"]
Period = Literal["peak", "off_peak"]

@dataclass
class DeepSeekRate:
    cached_input: float   # USD per 1M tokens
    uncached_input: float
    output: float

RATES = {
    ("deepseek-flash", "off_peak"): DeepSeekRate(0.003, 0.15, 0.60),
    ("deepseek-flash", "peak"):     DeepSeekRate(0.006, 0.30, 1.20),
    ("deepseek-v4-pro", "off_peak"):DeepSeekRate(0.022, 0.66, 1.98),
    ("deepseek-v4-pro", "peak"):    DeepSeekRate(0.044, 1.32, 3.96),
}

TOKENS_PER_MILLION = 1_000_000.0

def cost_per_request(
    model: ModelName,
    period: Period,
    cached_input_tokens: int,
    uncached_input_tokens: int,
    output_tokens: int,
) -> float:
    rate = RATES[(model, period)]
    return (
        (cached_input_tokens / TOKENS_PER_MILLION) * rate.cached_input
        + (uncached_input_tokens / TOKENS_PER_MILLION) * rate.uncached_input
        + (output_tokens / TOKENS_PER_MILLION) * rate.output
    )

def monthly_cost(
    model: ModelName,
    period: Period,
    cached_input_tokens: int,
    uncached_input_tokens: int,
    output_tokens: int,
    requests_per_day: int,
    active_days_per_month: int,
) -> float:
    cpr = cost_per_request(
        model,
        period,
        cached_input_tokens,
        uncached_input_tokens,
        output_tokens,
    )
    return cpr * requests_per_day * active_days_per_month

if __name__ == "__main__":
    # Example: 3k cached, 2k uncached, 1k output tokens, 50k req/day, 30 days, Flash off-peak
    cpr = cost_per_request("deepseek-flash", "off_peak", 3000, 2000, 1000)
    monthly = monthly_cost("deepseek-flash", "off_peak", 3000, 2000, 1000, 50_000, 30)
    print(f"Cost per request: ${cpr:.6f}")
    print(f"Monthly cost: ${monthly:,.2f}")

Calculator outputs

An implementation should show:

  • Cost per request

  • Daily cost

  • Monthly cost

  • Cached-input cost portion

  • Uncached-input cost portion

  • Output cost portion

  • Peak versus off-peak difference

  • Flash versus Pro difference

  • Effective blended cost per million tokens

This gives AI engineers a clear way to estimate DeepSeek API cost for different workloads.

How DeepSeek API pricing works

DeepSeek API pricing has four primary billing variables: model, input tokens, cache status, and output tokens, plus a rate period flag.

Model selection

DeepSeek currently promotes two main models for general-purpose use:

  • DeepSeek V4.1 Flash via deepseek-flash

  • DeepSeek V4 Pro via deepseek-v4-pro

V4.1 Flash is designed for high-volume, cost-sensitive workloads with strong quality. V4 Pro targets more demanding tasks that may benefit from extra reasoning or multi-step outputs. Pro is more expensive per token, so teams should compare cost per successful task rather than assuming Pro always outperforms Flash.

Input tokens

Input tokens include everything sent to the API:

  • System instructions

  • Conversation history

  • Retrieved documents

  • Tool outputs and traces

  • Memory payloads and user profile details

Agents that naively replay the entire conversation plus all previous tool results on every call will grow uncached input rapidly, which increases cost. Mem0 and similar memory layers address this by selectively injecting only relevant context.

Cached versus uncached input

DeepSeek applies context caching to matching prompt prefixes. When a prefix has been processed recently, it may qualify for the lower cache-hit price. Tokens in new or non-matching prefixes are billed at cache-miss rates, which are significantly higher.

Key points:

  • Cache creation and reuse are best-effort and not guaranteed.

  • Only matching input prefixes benefit from cache-hit pricing.

  • Application changes to the early part of the prompt can reduce cache reuse.

  • Reordering content can affect which tokens are recognized as reusable.

Output tokens

Every generated token is billed as output, including:

  • Plain text replies

  • Long code segments

  • Detailed reasoning explanations

  • Tool traces if generated as part of the output

In tasks like content generation or large code diffs, output tokens can dominate the invoice even when inputs are modest.

Peak and off-peak rates

DeepSeek assigns each request to peak or off-peak periods:

  • Token prices are twice as high during peak as during off-peak.

  • Identical token usage may cost significantly more if executed in peak windows.

  • The same workloads can be scheduled differently for background tasks.

Scheduling batch jobs off-peak is often the simplest lever to reduce cost without changing prompts or models.

DeepSeek V4 Flash API pricing

Current V4 Flash pricing generally refers to DeepSeek V4.1 Flash, accessed through the deepseek-flash API model ID. Older V4 Flash names have either been retired or redirected to this model. Engineers should not assume that legacy model names represent separate price tiers.

V4.1 Flash price table

Pricing category

Off-peak

Peak

Cache-hit input

$0.003 / M

$0.006 / M

Cache-miss input

$0.15 / M

$0.30 / M

Output

$0.60 / M

$1.20 / M

V4.1 Flash capabilities and implications

Key configuration, based on DeepSeek documentation:

  • API model identifier: deepseek-flash

  • Context window: 1,000,000 tokens

  • Maximum output: 384,000 tokens

  • Modes: supports standard and reasoning modes where applicable

  • Vision: currently listed as supported for Flash

  • Tool calling: supports tools and structured output

Typical workloads:

  • Customer support chatbots

  • Coding assistants with moderate context

  • Agent planners with frequent but short calls

  • High-volume batch summarization

Cost implications:

  • Long outputs, such as full-page articles or long code, can significantly increase the output token bill.

  • The large context window allows big prompts, but repeated long prefixes without caching or memory discipline can raise costs.

  • Migration from older Flash model IDs should treat deepseek-flash as the canonical target for evaluation.

DeepSeek V4 Pro API pricing

DeepSeek V4 Pro is available under the direct API model ID deepseek-v4-pro.

V4 Pro price table

Pricing category

Off-peak

Peak

Cache-hit input

$0.022 / M

$0.044 / M

Cache-miss input

$0.66 / M

$1.32 / M

Output

$1.98 / M

$3.96 / M

Pro offers higher prices across all token categories compared to Flash. It should not be assumed to outperform Flash on every task, especially short or simple ones.

Teams should:

  • Benchmark cost per successful task rather than focusing on raw token price.

  • Evaluate quality, tool usage, latency, and typical output length.

  • Use Pro only when task-level evaluation shows meaningful value over Flash that justifies the extra cost.

DeepSeek V4 Flash vs. V4 Pro pricing

The table below compares key factors for pricing and selection.

Factor

V4.1 Flash

V4 Pro

Direct model ID

deepseek-flash

deepseek-v4-pro

Relative price

Lower

Higher

Cache-hit cost

Lower

Higher

Cache-miss cost

Lower

Higher

Output cost

Lower

Higher

Vision

Supported

Not currently listed

Recommended use

High-volume production and cost-sensitive tasks

Tasks where testing demonstrates additional value

Selection method

Default evaluation candidate

Benchmark against Flash before adoption

Worked comparison example

Assumptions:

  • 2,000 cached input tokens

  • 3,000 uncached input tokens

  • 1,000 output tokens

  • 100,000 requests per month

  • All requests off-peak

Cost per request:

  • Flash off-peak:

    cost =
      (2,000 / 1e6) * 0.003
    + (3,000 / 1e6) * 0.15
    + (1,000 / 1e6) * 0.60
    = 0.000006 + 0.00045 + 0.0006
    = $0.001056
    cost =
      (2,000 / 1e6) * 0.003
    + (3,000 / 1e6) * 0.15
    + (1,000 / 1e6) * 0.60
    = 0.000006 + 0.00045 + 0.0006
    = $0.001056
    cost =
      (2,000 / 1e6) * 0.003
    + (3,000 / 1e6) * 0.15
    + (1,000 / 1e6) * 0.60
    = 0.000006 + 0.00045 + 0.0006
    = $0.001056
  • Pro off-peak:

    cost =
      (2,000 / 1e6) * 0.022
    + (3,000 / 1e6) * 0.66
    + (1,000 / 1e6) * 1.98
    = 0.000044 + 0.00198 + 0.00198
    = $0.004004
    cost =
      (2,000 / 1e6) * 0.022
    + (3,000 / 1e6) * 0.66
    + (1,000 / 1e6) * 1.98
    = 0.000044 + 0.00198 + 0.00198
    = $0.004004
    cost =
      (2,000 / 1e6) * 0.022
    + (3,000 / 1e6) * 0.66
    + (1,000 / 1e6) * 1.98
    = 0.000044 + 0.00198 + 0.00198
    = $0.004004

Monthly cost:

  • Flash: 100,000 × 0.001056 ≈ $105.60

  • Pro: 100,000 × 0.004004 ≈ $400.40

The same token usage costs nearly 3.8x more on Pro in this configuration, so real task-level benefits must justify that difference.

When are DeepSeek’s peak and off-peak hours?

DeepSeek’s current official schedule (checked 2026-10-01):

  • Peak: 01:00-04:00 UTC, Monday through Friday

  • Peak: 06:00-10:00 UTC, Monday through Friday

  • Off-peak: all remaining hours

  • Off-peak rates: 50 percent of peak rates

A local-time converter should be embedded in internal tools rather than hardcoding offsets. Time zones and daylight saving changes make static examples brittle.

Recommendations:

  • Display UTC times plus a clear timezone label, for example: “01:00-04:00 UTC (your local time: converted dynamically)”.

  • Avoid scheduling latency-sensitive tasks based only on price, because off-peak times for one region may still have high demand or different user expectations.

  • Schedule batch evaluation, summarization, and background processing off-peak where possible to reduce cost.

How DeepSeek context caching changes your bill

DeepSeek enables context caching automatically. When a prompt shares a matching prefix with a previously processed request, the matching portion may be served as cache-hit tokens at a lower price.

Key behaviors:

  • Prefix caching is automatic and best-effort.

  • Cache hits are not guaranteed even for repeated inputs.

  • Only matching prefixes benefit from cache-hit pricing.

  • Newly generated output is still billed at the output rate.

An example prompt structure:

[Stable system instructions]
+ [Stable reference material]
+ [Changing user question]
[Stable system instructions]
+ [Stable reference material]
+ [Changing user question]
[Stable system instructions]
+ [Stable reference material]
+ [Changing user question]

Keeping reusable content, such as system prompts and static reference material, at the beginning of the prompt can improve opportunities for prefix reuse. This does not guarantee a specific cache-hit ratio, but it aligns with DeepSeek’s prefix-based caching design.

Applications can inspect token breakdown metrics to track cache-hit and cache-miss tokens and feed that into cost dashboards.

DeepSeek V4 Pro OpenRouter pricing

DeepSeek V4 Pro pricing on OpenRouter differs from direct DeepSeek API pricing because OpenRouter can route requests across multiple inference providers. The final listed rate can depend on the model checkpoint, selected provider, routing mode, cache support, batch mode, and temporary discounts.

This section uses OpenRouter model pages checked on 2026-10-01 UTC. Readers should always confirm the latest figures.

Available V4 Pro model IDs on OpenRouter

OpenRouter model

Meaning

deepseek/deepseek-v4-pro-0813

Newer dated V4 Pro 0813 checkpoint

deepseek/deepseek-v4-pro

Earlier V4 Pro 0423 checkpoint

~deepseek/deepseek-pro-latest

Alias pointing to the latest Pro-family model

deepseek/deepseek-v4-pro-0813:batch

Batch variant, when available

Engineers should treat each model ID as potentially having different prices, providers, and capabilities.

What pricing information to display

For each OpenRouter DeepSeek V4 Pro model, it is useful to surface:

  • Model ID

  • Model checkpoint date

  • Lowest available input rate (USD / 1M)

  • Lowest available output rate (USD / 1M)

  • Cache-read rate if supported

  • Default or displayed route

  • Batch rate, if present

  • Context length

  • Number of providers for this model

  • Date and time checked

OpenRouter’s V4 Pro pages expose provider-specific rates and routing options, so UI tools should link directly to the live V4 Pro 0813 page. Prices can change with provider promotions or capacity constraints.

Why OpenRouter rates change

OpenRouter introduces more pricing variability:

  • Multiple providers may serve the same model ID.

  • Different providers set different token prices.

  • Temporary promotions can change headline rates.

  • A cheapest-provider route can have a different price from a fastest-provider route.

  • Pinning a provider can trade cost for latency or reliability.

  • Batch endpoints may have separate rates and minimums.

  • Cache-read support can vary between providers.

Applications should avoid hardcoding OpenRouter prices and instead treat them as dynamic configuration.

Editorial rule for OpenRouter prices

OpenRouter pricing references should follow this pattern:

OpenRouter rates checked on [date and UTC time]. Provider-level prices and promotional discounts can change independently.

Undated OpenRouter price claims should not appear in introductions or meta descriptions.

Direct DeepSeek API vs. OpenRouter pricing

The table below summarizes structural differences.

Factor

Direct DeepSeek API

OpenRouter

Provider model

Served directly by DeepSeek

Multiple inference providers

Price structure

Peak and off-peak

Provider- and route-dependent

Model selection

Current official IDs

Multiple checkpoints and aliases

Cache pricing

Official DeepSeek cache-hit rates

Model- and provider-dependent

Routing control

Direct provider

Cheapest, fastest, balanced, pinned, or other routes

Failover

Managed by DeepSeek

Potential cross-provider routing

Billing relationship

DeepSeek

OpenRouter

Best for

Direct access and official billing

Multi-model access and routing flexibility

Answering a common question:

  • Is OpenRouter cheaper than direct DeepSeek?
    Sometimes yes, sometimes no. The final cost depends on model checkpoint, chosen provider, routing policy, cache support, promotions, and whether the direct alternative would have run during peak or off-peak hours.

DeepSeek vs every major API: the full comparison

This section positions DeepSeek in the broader LLM API pricing landscape. Prices here must be rechecked at publication time and normalized to USD per one million tokens. Reasoning tokens, if billed separately, should be classified as input or output per provider documentation.

A typical comparison table might include:

Provider

Flagship model

Input / 1M tokens

Output / 1M tokens

Context window

DeepSeek

V4.1 Flash

$0.15 off-peak / $0.30 peak

$0.60 off-peak / $1.20 peak

1M

DeepSeek

V4 Pro

$0.66 off-peak / $1.32 peak

$1.98 off-peak / $3.96 peak

1M

OpenAI

GPT-6 Astra

$10.00

$50.00

1.05M

OpenAI

GPT-6.1 Sol

$2.00

$10.00

1.05M

Anthropic

Claude Fable 5.1

$10.00

$50.00

1M

Anthropic

Claude Opus 5.5

$4.00

$20.00

1M

Google

Gemini 3.1 Pro Preview

$2.00 ≤200K / $4.00 >200K

$12.00 ≤200K / $18.00 >200K

1M

Google

Gemini 3.8 Flash

$0.75*

$3.75*

1M

xAI

Grok 4.7

$2.00

$6.00

500K

xAI

Grok 4.3

$1.25 ≤200K / $2.50 >200K

$2.50 ≤200K / $5.00 >200K

1M

Mistral

Mistral Medium 3.5

$1.50

$7.50

256K

Mistral

Mistral Large 3

$0.50

$1.50

256K

Alibaba

Qwen3.8 Max

$2.00

$6.00

1M

Alibaba

Qwen3.8 Flash

$0.15

$0.47

1M

Meta via Together AI

Llama 4 Maverick

$0.27

$0.85

~1M

Meta via Together AI

Llama 4 Scout

$0.18

$0.59

~1M

Note: Mentioned pricing is as of 1 Oct, 2026, sourced from respective official websites.

Recommended comparison strategy:

  • Match V4.1 Flash against fast/low-cost models from each provider.

  • Match V4 Pro against higher-capability models from each provider.

  • Separate standard and batch tiers.

  • Include reasoning tokens where applicable.

  • Describe cache pricing explicitly when providers expose it.

From this comparison, engineers can distinguish:

  • Cheapest input per million tokens.

  • Cheapest output per million tokens.

  • Cheapest repeated-context workloads (where caching matters).

  • Cheapest batch workloads.

  • Lowest estimated cost per completed task, which is more operationally meaningful.

Worked DeepSeek API cost examples

This section uses the cost formulas from earlier to ground token pricing in realistic agent workloads. All assumptions are explicitly stated.

Customer-support chatbot

Assumptions:

  • Model: deepseek-flash

  • Period: off-peak

  • Stable system prompt and policy: 1,500 tokens

  • Recent messages included: 1,000 tokens

  • Retrieved knowledge: 1,500 tokens

  • Output size: 600 tokens

  • Cache-hit percentage on system + policy: 80 percent

  • Requests per month: 1,000,000

Token breakdown (average per request):

  • Cached input: 1,500 × 0.8 = 1,200 tokens

  • Uncached input: 1,000 (chat) + 1,500 (retrieved) + 300 uncached portion of system ≈ 2,800 tokens

  • Output: 600 tokens

Cost per request:

cached cost = (1,200 / 1e6) * 0.003 ≈ $0.0000036
uncached cost = (2,800 / 1e6) * 0.15 ≈ $0.00042
output cost = (600 / 1e6) * 0.60 ≈ $0.00036
total ≈ $0.0007836
cached cost = (1,200 / 1e6) * 0.003 ≈ $0.0000036
uncached cost = (2,800 / 1e6) * 0.15 ≈ $0.00042
output cost = (600 / 1e6) * 0.60 ≈ $0.00036
total ≈ $0.0007836
cached cost = (1,200 / 1e6) * 0.003 ≈ $0.0000036
uncached cost = (2,800 / 1e6) * 0.15 ≈ $0.00042
output cost = (600 / 1e6) * 0.60 ≈ $0.00036
total ≈ $0.0007836

Monthly total:

  • 1,000,000 requests × 0.0007836 ≈ $783.60

  • If 90 percent of tickets resolve successfully, cost per successful task ≈ $783.60 / 900,000 ≈ $0.00087.

Coding assistant

Assumptions:

  • Model: deepseek-v4-pro

  • Period: peak

  • Repository context and file snippets: 5,000 tokens

  • Tool results and error logs: 3,000 tokens

  • Conversation history: 2,000 tokens

  • Output code: 2,500 tokens

  • Minimal caching; assume everything is uncached for the worst case

  • Calls per task: 4

  • Tasks per month: 20,000

Per-call tokens:

  • Cached input: 0

  • Uncached input: 10,000 tokens

  • Output: 2,500 tokens

Cost per call:

uncached cost = (10,000 / 1e6) * 1.32 ≈ $0.0132
output cost = (2,500 / 1e6) * 3.96 ≈ $0.0099
total ≈ $0.0231
uncached cost = (10,000 / 1e6) * 1.32 ≈ $0.0132
output cost = (2,500 / 1e6) * 3.96 ≈ $0.0099
total ≈ $0.0231
uncached cost = (10,000 / 1e6) * 1.32 ≈ $0.0132
output cost = (2,500 / 1e6) * 3.96 ≈ $0.0099
total ≈ $0.0231

Per task (4 calls): 4 × 0.0231 ≈ $0.0924

Monthly total:

  • 20,000 tasks × 0.0924 ≈ $1,848

  • If 85 percent of tasks succeed, cost per successful task ≈ $1,848 / 17,000 ≈ $0.1088.

Document analysis

Assumptions:

  • Model: deepseek-flash

  • Period: off-peak

  • Long input document: 40,000 tokens

  • Short analytical output: 800 tokens

  • Minimal caching, single-pass per document

  • Calls per task: 1

  • Tasks per month: 50,000

Tokens per request:

  • Cached input: 0

  • Uncached input: 40,000

  • Output: 800

Cost per request:

uncached cost = (40,000 / 1e6) * 0.15 ≈ $0.006
output cost = (800 / 1e6) * 0.60 ≈ $0.00048
total ≈ $0.00648
uncached cost = (40,000 / 1e6) * 0.15 ≈ $0.006
output cost = (800 / 1e6) * 0.60 ≈ $0.00048
total ≈ $0.00648
uncached cost = (40,000 / 1e6) * 0.15 ≈ $0.006
output cost = (800 / 1e6) * 0.60 ≈ $0.00048
total ≈ $0.00648

Monthly total:

  • 50,000 × 0.00648 ≈ $324

  • If 95 percent success, cost per successful task ≈ $324 / 47,500 ≈ $0.00682.

Content-generation workflow

Assumptions:

  • Model: deepseek-flash

  • Period: off-peak

  • Prompt: 3,000 tokens

  • Output article: 4,000 tokens

  • Minimal caching

  • Calls per task: 1

  • Tasks per month: 30,000

Tokens:

  • Cached input: 0

  • Uncached input: 3,000

  • Output: 4,000

Cost per request:

uncached cost = (3,000 / 1e6) * 0.15 ≈ $0.00045
output cost = (4,000 / 1e6) * 0.60 ≈ $0.0024
total ≈ $0.00285
uncached cost = (3,000 / 1e6) * 0.15 ≈ $0.00045
output cost = (4,000 / 1e6) * 0.60 ≈ $0.0024
total ≈ $0.00285
uncached cost = (3,000 / 1e6) * 0.15 ≈ $0.00045
output cost = (4,000 / 1e6) * 0.60 ≈ $0.0024
total ≈ $0.00285

Monthly total:

  • 30,000 × 0.00285 ≈ $85.50

  • Output tokens dominate, about 84 percent of the cost.

Long-running AI agent

Assumptions:

  • Model: deepseek-flash

  • Period: mixed; assume half peak, half off-peak for simplicity

  • Conversation history per step: 8,000 tokens

  • Memory retrieval per step: 2,000 tokens

  • Tool use per step: 3,000 tokens

  • Output per step: 1,500 tokens

  • 10 steps per task

  • Tasks per month: 10,000

  • Some stable prefix cached; assume 3,000 cached tokens per step

Per step tokens:

  • Cached input: 3,000

  • Uncached input: 10,000 (remaining)

  • Output: 1,500

Off-peak cost per step:

cached = (3,000 / 1e6) * 0.003 ≈ $0.000009
uncached = (10,000 / 1e6) * 0.15 ≈ $0.0015
output = (1,500 / 1e6) * 0.60 ≈ $0.0009
off_peak_step ≈ $0.002409
cached = (3,000 / 1e6) * 0.003 ≈ $0.000009
uncached = (10,000 / 1e6) * 0.15 ≈ $0.0015
output = (1,500 / 1e6) * 0.60 ≈ $0.0009
off_peak_step ≈ $0.002409
cached = (3,000 / 1e6) * 0.003 ≈ $0.000009
uncached = (10,000 / 1e6) * 0.15 ≈ $0.0015
output = (1,500 / 1e6) * 0.60 ≈ $0.0009
off_peak_step ≈ $0.002409

Peak cost per step:

cached = (3,000 / 1e6) * 0.006 ≈ $0.000018
uncached = (10,000 / 1e6) * 0.30 ≈ $0.003
output = (1,500 / 1e6) * 1.20 ≈ $0.0018
peak_step ≈ $0.004818
cached = (3,000 / 1e6) * 0.006 ≈ $0.000018
uncached = (10,000 / 1e6) * 0.30 ≈ $0.003
output = (1,500 / 1e6) * 1.20 ≈ $0.0018
peak_step ≈ $0.004818
cached = (3,000 / 1e6) * 0.006 ≈ $0.000018
uncached = (10,000 / 1e6) * 0.30 ≈ $0.003
output = (1,500 / 1e6) * 1.20 ≈ $0.0018
peak_step ≈ $0.004818

Average per step (half peak, half off-peak):

  • ≈ (0.002409 + 0.004818) / 2 ≈ $0.0036135

Per task (10 steps): ≈ $0.036135

Monthly total:

  • 10,000 tasks × 0.036135 ≈ $361.35

  • If 80 percent success, cost per successful task ≈ $361.35 / 8,000 ≈ $0.045.

These examples show how input vs output balance, cache usage, and peak scheduling affect DeepSeek API pricing.

The costs not shown on DeepSeek’s pricing page

DeepSeek invoices only cover LLM token usage. Production agents incur additional costs that are not reflected in the LLM pricing table, including:

  • Embedding models and vector storage

  • Search and reranking services

  • Memory layers and user profile storage

  • Tool and third-party API calls

  • Orchestration frameworks and job queues

  • Retries and failed generations

  • Logging, metrics, and auditing

  • Evaluation and red-teaming pipelines

  • Security controls and compliance tooling

  • Networking, bandwidth, and egress

  • Data processing and ETL pipelines

  • Engineering time for prompt and agent design

  • Self-hosting infrastructure for complementary components

Understanding total cost of ownership requires tracking both DeepSeek tokens and these surrounding elements.

How to reduce DeepSeek API costs

Practical steps for AI engineers:

  1. Schedule flexible jobs, such as batch summarization and offline evaluation, during off-peak hours.

  2. Use V4.1 Flash by default, and upgrade to V4 Pro only when evaluations justify it.

  3. Keep reusable prompt prefixes stable and early in the prompt to improve cache reuse.

  4. Monitor actual cache-hit and cache-miss token counts via DeepSeek metrics.

  5. Limit unnecessary output and avoid verbose free-form reasoning where not needed.

  6. Retrieve

Useful resources

Start building with memory

Wire persistent memory into your own agent in about 15 minutes. Free tier, no credit card.

Share on:

aashi dutt

Aashi Dutt

She is a senior technical content writer at Mem0. She covers agent memory architecture and the engineering decisions behind building agents that actually remember. She experiments with new features and turns research into posts developers can put straight to use.

Start building with memory

Free tier, no card