Describe monthly requests, tokens, attempts, cache reuse, required features, and budget. The buying-decision engine excludes incompatible models and recommends the lowest-cost tracked official API. A separate OpenRouter market explorer remains available below.
It satisfies the selected model class, context, output, vision, and tool constraints, then has the lowest modeled token cost for the monthly workload.
Official standard API token rates only. Cached input means cache-hit reads; cache writes or storage, batch discounts, separately billed tools, regional taxes, rate limits, third-party routes, and subscriptions are excluded. Model class is a workload filter, not a quality score; test candidates on the real task.
Complete cost = successful requests × attempts × input and output tokens at the applicable official rates. Documented cached-input rates are applied to the selected cache share. The broader table below is a separate OpenRouter route explorer.
Each provider sets its own rates based on model capability, compute cost, and competitive positioning. Premium reasoning models (Claude Opus, GPT high-effort) cost more than fast variants (nano, mini, flash) because they spend more compute per token.
For most workloads, look at the fast tier: Claude Haiku, GPT-5 mini, Gemini Flash, DeepSeek V3. They're typically 10–30× cheaper than flagship models while staying good enough for chat, classification, and routing. Reserve flagship models for reasoning-heavy tasks.
The official-API decision applies a cached-input rate only when that rate is present in the sourced model data. Batch discounts remain excluded because eligibility and processing constraints are not yet normalized.
These platforms provide additional multi-model routes. They are not part of the official-API recommendation above unless their complete route cost is explicitly compared. Some links earn us a small commission at no extra cost to you.
See our affiliate disclosure for the full list and how we choose what to recommend.
Use this table to inspect a third-party route across the broader model market. It is not mixed into the audited official-API recommendation because route capabilities and billing adjustments are not normalized here.