Fastest AI Models 2026 — Live Speed Leaderboard (Celeris-1, 2,121 tok/s) Compare real-world speed for 329+ AI models: response latency, time to first token (TTFT) , and throughput (tokens/sec) . We pull live performance data from Artificial Analysis and pair it with per-token pricing so you can spot the fastest and most cost-efficient models for production APIs, real-time chat, and code completion.
For raw benchmark scores (GPQA, AIME, MMLU-Pro, HLE, SWE-bench) see the Benchmarks leaderboard . For pricing-only comparison try the LLM cost calculator .
Price the latency your workload needs Model speed is only one constraint. Add your request volume, token mix, retries, feature requirements, caching, and budget to find the lowest-cost compatible tracked official API.
Complete monthly cost Compatibility first Primary price sourcesCalculate my workload Top 5 Fastest AI Models Right Now 1. Celeris-1 1621.3 tok/s 2. Mercury 2 804.1 tok/s 3. Gemini 3.5 Flash-Lite 390.6 tok/s 4. Gemini 2.5 Flash-Lite (Reasoning) 367.1 tok/s 5. LFM2.5-VL-1.6B 357.7 tok/s
Live speed data from Artificial Analysis API
AI Model Speed Rankings Compare 329+ AI models by response speed, latency, and throughput. Find the fastest models for your use case.
329 models · click headers to sort
#
Model
Throughput ↓
TTFT
$/1M
$/Speed
Price×TTFT
Showing top 100 of 329 models. Use search/filter to narrow down.
Speed Metrics Guide Throughput (tokens/s)
Output generation speed in tokens per second. Higher is better.
Good: >50 t/s · Excellent: >100 t/s
Time to First Token (TTFT)
Delay before the first token appears. Lower is better.
Good: <500ms · Excellent: <200ms
Price/Performance
Cost efficiency ratios. Lower values indicate better value.
$/Speed: price per t/s · Price×TTFT: latency penalty