LLM workload cost calculator

Describe monthly requests, tokens, attempts, cache reuse, required features, and budget. The buying-decision engine excludes incompatible models and recommends the lowest-cost tracked official API. A separate OpenRouter market explorer remains available below.

Describe the workload

LLM API buying decision

6 official offers
Cheapest compatible official API
Official API · DeepSeek

DeepSeek V4 Flash

$2.24
within monthly budget
$0.0002 per successful request
10,000
attempts
10,000,000
input tokens
0 cached
3,000,000
output tokens
Why this route ranks first

It satisfies the selected model class, context, output, vision, and tool constraints, then has the lowest modeled token cost for the monthly workload.

10,000 requests × 1 attempts = 10,000,000 input × $0.14/1M + 3,000,000 output × $0.28/1M
Rates checked 2026-08-26 · 4 incompatible offers excluded
Next compatible official APIsSame selected constraints
GPT-5.6 Luna via OpenAI
fast · 1,050,000 context

Official standard API token rates only. Cached input means cache-hit reads; cache writes or storage, batch discounts, separately billed tools, regional taxes, rate limits, third-party routes, and subscriptions are excluded. Model class is a workload filter, not a quality score; test candidates on the real task.

AI API cost FAQ

How is API cost calculated?

Complete cost = successful requests × attempts × input and output tokens at the applicable official rates. Documented cached-input rates are applied to the selected cache share. The broader table below is a separate OpenRouter route explorer.

Why do prices vary between providers?

Each provider sets its own rates based on model capability, compute cost, and competitive positioning. Premium reasoning models (Claude Opus, GPT high-effort) cost more than fast variants (nano, mini, flash) because they spend more compute per token.

Which model is cheapest for production?

For most workloads, look at the fast tier: Claude Haiku, GPT-5 mini, Gemini Flash, DeepSeek V3. They're typically 10–30× cheaper than flagship models while staying good enough for chat, classification, and routing. Reserve flagship models for reasoning-heavy tasks.

Does this include caching and batch discounts?

The official-API decision applies a cached-input rate only when that rate is present in the sourced model data. Batch discounts remain excluded because eligibility and processing constraints are not yet normalized.

Optional third-party API access

These platforms provide additional multi-model routes. They are not part of the official-API recommendation above unless their complete route cost is explicitly compared. Some links earn us a small commission at no extra cost to you.

See our affiliate disclosure for the full list and how we choose what to recommend.

Supporting market explorer

Live OpenRouter model-route prices

Use this table to inspect a third-party route across the broader model market. It is not mixed into the audited official-API recommendation because route capabilities and billing adjustments are not normalized here.

425 models
★ Best
inclusionai
Free
Ling 3.0 Flash Sante (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 262K
inclusionai
Free
Ling 3.0 Flash Fin (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 262K
dots-studio
Free
Dots3-Note Preview (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 512K
liquid
Free
LFM2.5-2.6B (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 66K
nvidia
Free
Nemotron 3.5 Lightning (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 1.0M
thinkingmachines
Free
Inkling Small (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 1.0M
poolside
Free
Laguna S 2.1 (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 262K
thinkingmachines
Free
Inkling (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 1.0M
poolside
Free
Laguna XS 2.1 (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 262K
Cohere
Free
North Mini Code (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 256K
nvidia
Free
Nemotron 3.5 Content Safety (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 128K
nvidia
Free
Nemotron 3 Ultra (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 1.0M
minimax
Free
MiniMax M3 (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 1.0M
nvidia
Free
Nemotron 3 Nano Omni (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 256K
Google
Free
Gemma 4 26B A4B (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 262K
Google
Free
Gemma 4 31B (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 262K
Google
Flagship
Lyria 3 Pro Preview
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 1.0M
Google
Flagship
Lyria 3 Clip Preview
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 1.0M
minimax
Free
MiniMax M2.7 (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 197K
nvidia
Free
Nemotron 3 Super (free)
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 262K
openrouter
Free
Free Models Router
Input
Free
Output
Free
Est. Cost
<$0.0001
Context: 200K
Mistral
Flagship
Mistral Nemo
Input
$0.02
Output
$0.03
Est. Cost
<$0.0001
Context: 131K
inclusionai
Fast
Ling 3.0 Flash
Input
$0.02
Output
$0.06
Est. Cost
<$0.0001
Context: 262K
sao10k
Reasoning
Llama 3 8B Lunaris
Input
$0.04
Output
$0.05
Est. Cost
<$0.0001
Context: 8K
ibm-granite
Flagship
Granite 4.0 Micro
Input
$0.02
Output
$0.11
Est. Cost
<$0.0001
Context: 131K
nex-agi
Fast
Nex-N2-Mini
Input
$0.02
Output
$0.10
Est. Cost
<$0.0001
Context: 262K
Mistral
Flagship
Mistral Small 3
Input
$0.05
Output
$0.08
Est. Cost
<$0.0001
Context: 33K
meta-llama
Flagship
Llama 3.1 8B Instruct
Input
$0.05
Output
$0.08
Est. Cost
<$0.0001
Context: 131K
gryphe
Flagship
MythoMax 13B
Input
$0.06
Output
$0.06
Est. Cost
<$0.0001
Context: 8K
~deepseek
Fast
DeepSeek V4 Flash Latest
Input
$0.04
Output
$0.09
Est. Cost
<$0.0001
Context: 1.3M
upstage
Reasoning
Solar Pro 4
Input
$0.03
Output
$0.12
Est. Cost
<$0.0001
Context: 524K
Qwen
Fast
Qwen3.7 Flash
Input
$0.03
Output
$0.13
Est. Cost
<$0.0001
Context: 1.0M
OpenAI
Flagship
gpt-oss-20b
Input
$0.03
Output
$0.13
Est. Cost
<$0.0001
Context: 131K
DeepSeek
Fast
DeepSeek V4 Flash 0731
Input
$0.05
Output
$0.10
Est. Cost
<$0.0001
Context: 1.3M
Google
Flagship
Gemma 3 4B
Input
$0.05
Output
$0.10
Est. Cost
<$0.0001
Context: 131K
amazon
Flagship
Nova Micro 1.0
Input
$0.04
Output
$0.14
Est. Cost
$0.0001
Context: 128K
Cohere
Flagship
Command R7B (12-2024)
Input
$0.04
Output
$0.15
Est. Cost
$0.0001
Context: 128K
inception
Flagship
Mercury 2.5 Preview
Input
$0.04
Output
$0.15
Est. Cost
$0.0001
Context: 260K
poolside
Flagship
Laguna XS 2.1
Input
$0.06
Output
$0.12
Est. Cost
$0.0001
Context: 262K
OpenAI
Flagship
gpt-oss-120b
Input
$0.04
Output
$0.17
Est. Cost
$0.0001
Context: 131K
Google
Flagship
Gemma 3 12B
Input
$0.05
Output
$0.15
Est. Cost
$0.0001
Context: 131K
OpenAI
Fast
GPT-5 Nano (batch)
Input
$0.02
Output
$0.20
Est. Cost
$0.0001
Context: 400K
meta-llama
Flagship
Llama 3.2 1B Instruct
Input
$0.03
Output
$0.20
Est. Cost
$0.0001
Context: 60K
tencent
Flagship
Hy-MT2-1.8B
Input
$0.04
Output
$0.18
Est. Cost
$0.0001
Context: 8K
microsoft
Flagship
Phi 4
Input
$0.07
Output
$0.14
Est. Cost
$0.0001
Context: 16K
Qwen
Flagship
Qwen3 30B A3B Instruct 2507
Input
$0.05
Output
$0.19
Est. Cost
$0.0001
Context: 262K
rekaai
Flagship
Reka Edge
Input
$0.10
Output
$0.10
Est. Cost
$0.0001
Context: 16K
Mistral
Fast
Ministral 3 3B 2512
Input
$0.10
Output
$0.10
Est. Cost
$0.0001
Context: 131K
nvidia
Fast
Nemotron 3 Nano 30B A3B
Input
$0.05
Output
$0.20
Est. Cost
$0.0001
Context: 262K
OpenAI
Flagship
gpt-oss-20b (batch)
Input
$0.05
Output
$0.20
Est. Cost
$0.0001
Context: 131K
Google
Fast
Gemini 2.5 Flash Lite (batch)
Input
$0.05
Output
$0.20
Est. Cost
$0.0001
Context: 1.0M
OpenAI
Fast
GPT-4.1 Nano (batch)
Input
$0.05
Output
$0.20
Est. Cost
$0.0001
Context: 1.0M
inclusionai
Fast
Ling 3.0 Flash Fin
Input
$0.06
Output
$0.18
Est. Cost
$0.0002
Context: 262K
DeepSeek
Fast
DeepSeek V4 Flash 0423
Input
$0.08
Output
$0.16
Est. Cost
$0.0002
Context: 1.0M
ibm-granite
Flagship
Granite 4.2 8B
Input
$0.10
Output
$0.15
Est. Cost
$0.0002
Context: 131K
Qwen
Flagship
Qwen3.5-9B
Input
$0.10
Output
$0.15
Est. Cost
$0.0002
Context: 262K
Mistral
Flagship
Mistral Small 3.2 24B
Input
$0.07
Output
$0.20
Est. Cost
$0.0002
Context: 131K
nvidia
Flagship
Nemotron 3.5 Lightning
Input
$0.08
Output
$0.20
Est. Cost
$0.0002
Context: 262K
poolside
Flagship
Laguna S 2.1
Input
$0.09
Output
$0.18
Est. Cost
$0.0002
Context: 1.0M
amazon
Flagship
Nova Lite 1.0
Input
$0.06
Output
$0.24
Est. Cost
$0.0002
Context: 300K
Qwen
Fast
Qwen3.5-Flash
Input
$0.07
Output
$0.26
Est. Cost
$0.0002
Context: 1.0M
Meta
Flagship
Muse Spark 1.3 Contributor
Input
$0.10
Output
$0.20
Est. Cost
$0.0002
Context: 1.0M
Meta
Flagship
Muse Spark 1.2 Contributor
Input
$0.10
Output
$0.20
Est. Cost
$0.0002
Context: 1.0M
bytedance
Flagship
UI-TARS 7B
Input
$0.10
Output
$0.20
Est. Cost
$0.0002
Context: 128K
rekaai
Fast
Reka Flash 3
Input
$0.10
Output
$0.20
Est. Cost
$0.0002
Context: 66K
Qwen
Flagship
Qwen2.5 7B Instruct
Input
$0.10
Output
$0.20
Est. Cost
$0.0002
Context: 33K
~z-ai
Fast
GLM Flash Latest
Input
$0.07
Output
$0.25
Est. Cost
$0.0002
Context: 1.3M
z-ai
Fast
GLM 5.3 Flash
Input
$0.07
Output
$0.25
Est. Cost
$0.0002
Context: 1.3M
Qwen
Flagship
Qwen3 Coder 30B A3B Instruct
Input
$0.07
Output
$0.28
Est. Cost
$0.0002
Context: 262K
meta-llama
Flagship
Llama 3.2 3B Instruct
Input
$0.05
Output
$0.33
Est. Cost
$0.0002
Context: 131K
Qwen
Flagship
Qwen3 32B
Input
$0.08
Output
$0.28
Est. Cost
$0.0002
Context: 131K
tencent
Flagship
Hy-MT2-30B-A3B
Input
$0.07
Output
$0.29
Est. Cost
$0.0002
Context: 8K
tencent
Flagship
Hy-MT2-7B
Input
$0.07
Output
$0.29
Est. Cost
$0.0002
Context: 8K
Mistral
Fast
Ministral 3 8B 2512
Input
$0.15
Output
$0.15
Est. Cost
$0.0002
Context: 262K
bytedance-seed
Fast
Seed 1.6 Flash
Input
$0.07
Output
$0.30
Est. Cost
$0.0002
Context: 262K
OpenAI
Flagship
gpt-oss-safeguard-20b
Input
$0.07
Output
$0.30
Est. Cost
$0.0002
Context: 131K
OpenAI
Fast
GPT-4o-mini (batch)
Input
$0.07
Output
$0.30
Est. Cost
$0.0002
Context: 128K
Google
Flagship
Gemma 4 26B A4B
Input
$0.07
Output
$0.34
Est. Cost
$0.0002
Context: 262K
Qwen
Flagship
Qwen3 14B
Input
$0.12
Output
$0.24
Est. Cost
$0.0002
Context: 131K
stepfun
Fast
Step 3.5 Flash
Input
$0.10
Output
$0.30
Est. Cost
$0.0003
Context: 262K
Mistral
Flagship
Voxtral Small 24B 2507
Input
$0.10
Output
$0.30
Est. Cost
$0.0003
Context: 33K
meta-llama
Flagship
Llama 4 Scout
Input
$0.10
Output
$0.30
Est. Cost
$0.0003
Context: 1.3M
OpenAI
Fast
GPT-5 Nano
Input
$0.05
Output
$0.40
Est. Cost
$0.0003
Context: 400K
Google
Flagship
Gemma 4 31B
Input
$0.09
Output
$0.34
Est. Cost
$0.0003
Context: 262K
z-ai
Fast
GLM 4.7 Flash
Input
$0.06
Output
$0.40
Est. Cost
$0.0003
Context: 203K
meta-llama
Flagship
Llama 3.3 70B Instruct
Input
$0.10
Output
$0.32
Est. Cost
$0.0003
Context: 131K
meta-llama
Flagship
Llama Guard 4 12B
Input
$0.18
Output
$0.18
Est. Cost
$0.0003
Context: 164K
DeepSeek
Fast
DeepSeek V4 Flash 0731 (batch)
Input
$0.14
Output
$0.28
Est. Cost
$0.0003
Context: 1.0M
xiaomi
Flagship
MiMo-V2.5
Input
$0.14
Output
$0.28
Est. Cost
$0.0003
Context: 1.1M
nvidia
Flagship
Nemotron 3 Super
Input
$0.08
Output
$0.40
Est. Cost
$0.0003
Context: 1.0M
Qwen
Flagship
Qwen3.5-9B (batch)
Input
$0.17
Output
$0.25
Est. Cost
$0.0003
Context: 262K
nvidia
Flagship
Nemotron 3.5 Content Safety
Input
$0.20
Output
$0.20
Est. Cost
$0.0003
Context: 131K
Mistral
Fast
Ministral 3 14B 2512
Input
$0.20
Output
$0.20
Est. Cost
$0.0003
Context: 262K
bytedance-seed
Fast
Seed-2.0-Mini
Input
$0.10
Output
$0.40
Est. Cost
$0.0003
Context: 262K
Google
Fast
Gemini 2.5 Flash Lite
Input
$0.10
Output
$0.40
Est. Cost
$0.0003
Context: 1.0M
OpenAI
Fast
GPT-4.1 Nano
Input
$0.10
Output
$0.40
Est. Cost
$0.0003
Context: 1.0M
Google
Flagship
Gemma 3 27B
Input
$0.08
Output
$0.45
Est. Cost
$0.0003
Context: 131K
Qwen
Flagship
Qwen3 VL 32B Instruct
Input
$0.10
Output
$0.42
Est. Cost
$0.0003
Context: 131K
Nous
Flagship
Hermes 4 70B
Input
$0.13
Output
$0.40
Est. Cost
$0.0003
Context: 131K
Qwen
Flagship
Qwen3 VL 8B Instruct
Input
$0.12
Output
$0.45
Est. Cost
$0.0003
Context: 262K