Fastest AI Models 2026 — Live Speed Leaderboard (Celeris-1, 2,121 tok/s)

Compare real-world speed for 329+ AI models: response latency, time to first token (TTFT), and throughput (tokens/sec). We pull live performance data from Artificial Analysis and pair it with per-token pricing so you can spot the fastest and most cost-efficient models for production APIs, real-time chat, and code completion.

For raw benchmark scores (GPQA, AIME, MMLU-Pro, HLE, SWE-bench) see the Benchmarks leaderboard. For pricing-only comparison try the LLM cost calculator.

Top 5 Fastest AI Models Right Now

  1. 1.Celeris-11621.3 tok/s
  2. 2.Mercury 2804.1 tok/s
  3. 3.Gemini 3.5 Flash-Lite390.6 tok/s
  4. 4.Gemini 2.5 Flash-Lite (Reasoning)367.1 tok/s
  5. 5.LFM2.5-VL-1.6B357.7 tok/s
Live speed data from Artificial Analysis API

AI Model Speed Rankings

Compare 329+ AI models by response speed, latency, and throughput. Find the fastest models for your use case.

329 models · click headers to sort
#
Model
Throughput
TTFT
$/1M
$/Speed
Price×TTFT
1
Celeris-1
Celeris
1621 t/s
630ms
$0.33
$0.000
204.8
2
Mercury 2
Inception
804 t/s
5.10s
$0.38
$0.000
1912.5
3
391 t/s
6.98s
$0.85
$0.002
5933.0
4
367 t/s
18.98s
$0.17
$0.000
3321.5
5
358 t/s
1.49s
$0.00
6
350 t/s
790ms
$0.07
$0.000
51.4
7
Ling 3.0 Flash
InclusionAI
346 t/s
2.58s
$0.11
$0.000
278.6
8
LFM2.5-8B-A1B
Liquid AI
334 t/s
1.65s
$0.00
9
328 t/s
1.09s
$0.13
$0.000
139.5
10
327 t/s
12.91s
$1.5
$0.005
19365.0
11
314 t/s
1.15s
$0.41
$0.001
474.9
12
313 t/s
700ms
$1.5
$0.005
1050.0
13
312 t/s
6.44s
$1.5
$0.005
9660.0
14
310 t/s
12.01s
$1.5
$0.005
18015.0
15
286 t/s
1.00s
$0.11
$0.000
108.0
16
280 t/s
5.77s
$1.5
$0.005
8655.0
17
279 t/s
5.22s
$0.56
$0.002
2938.9
18
276 t/s
310ms
$0.17
$0.001
54.3
19
267 t/s
870ms
$0.07
$0.000
56.6
20
258 t/s
980ms
$1.5
$0.006
1470.0
21
248 t/s
390ms
$0.00
22
237 t/s
20.32s
$1.9
$0.008
39116.0
23
234 t/s
15.76s
$2.0
$0.009
31520.0
24
228 t/s
1.81s
$0.28
$0.001
497.8
25
227 t/s
1.19s
$0.09
$0.000
104.7
26
225 t/s
17.48s
$0.85
$0.004
14858.0
27
217 t/s
430ms
$0.85
$0.004
365.5
28
216 t/s
610ms
$0.10
$0.000
61.0
29
214 t/s
440ms
$0.05
$0.000
23.3
30
214 t/s
1.05s
$0.15
$0.001
157.5
31
210 t/s
13.85s
$3.4
$0.016
46743.8
32
205 t/s
1.06s
$0.15
$0.001
159.0
33
205 t/s
15.26s
$3.4
$0.016
51502.5
34
204 t/s
2.27s
$3.8
$0.018
8512.5
35
o3-mini
OpenAI
203 t/s
5.10s
$1.9
$0.009
9817.5
36
203 t/s
950ms
$0.09
$0.000
83.6
37
202 t/s
2.86s
$0.56
$0.003
1610.2
38
200 t/s
7.94s
$1.7
$0.008
13402.7
39
195 t/s
920ms
$3.4
$0.017
3105.0
40
193 t/s
1.05s
$0.90
$0.005
945.0
41
190 t/s
6.33s
$1.1
$0.006
7121.3
42
188 t/s
2.27s
$0.41
$0.002
937.5
43
187 t/s
910ms
$1.1
$0.006
1023.7
44
186 t/s
42.55s
$2.0
$0.011
85100.0
45
182 t/s
15.68s
$1.5
$0.008
23520.0
46
180 t/s
5.62s
$1.7
$0.009
9486.6
47
179 t/s
1.45s
$2.0
$0.011
2900.0
48
Nova Lite
Amazon
173 t/s
930ms
$0.10
$0.001
97.7
49
Ling 3.0 Tiny
InclusionAI
172 t/s
2.68s
$0.00
50
170 t/s
2.19s
$0.41
$0.002
904.5
51
170 t/s
1.98s
$0.30
$0.002
594.0
52
168 t/s
2.00s
$0.15
$0.001
300.0
53
GPT-4.1
OpenAI
168 t/s
910ms
$3.5
$0.021
3185.0
54
166 t/s
710ms
$0.46
$0.003
328.7
55
164 t/s
2.15s
$0.69
$0.004
1479.2
56
164 t/s
3.41s
$0.46
$0.003
1578.8
57
163 t/s
7.60s
$0.85
$0.005
6460.0
58
163 t/s
830ms
$0.26
$0.002
217.5
59
162 t/s
3.64s
$0.46
$0.003
1685.3
60
160 t/s
1.57s
$0.09
$0.001
138.2
61
159 t/s
2.72s
$1.2
$0.007
3204.2
62
157 t/s
17.55s
$0.85
$0.005
14917.5
63
155 t/s
840ms
$1.7
$0.011
1417.9
64
155 t/s
35.38s
$0.14
$0.001
4882.4
65
154 t/s
1.05s
$0.85
$0.006
892.5
66
154 t/s
870ms
$0.26
$0.002
226.2
67
153 t/s
850ms
$0.26
$0.002
222.7
68
153 t/s
14.41s
$0.85
$0.006
12248.5
69
152 t/s
1.09s
$0.10
$0.001
112.3
70
152 t/s
2.05s
$0.69
$0.005
1410.4
71
146 t/s
95.00s
$0.14
$0.001
13110.0
72
145 t/s
2.10s
$0.85
$0.006
1780.8
73
145 t/s
22.85s
$1.9
$0.013
43986.3
74
144 t/s
2.32s
$0.75
$0.005
1740.0
75
144 t/s
2.32s
$1.1
$0.008
2552.0
76
144 t/s
620ms
$0.11
$0.001
67.0
77
144 t/s
2.44s
$0.26
$0.002
639.3
78
144 t/s
920ms
$0.14
$0.001
127.0
79
144 t/s
780ms
$0.30
$0.002
234.0
80
142 t/s
1.97s
$0.35
$0.002
689.5
81
142 t/s
830ms
$0.26
$0.002
217.5
82
141 t/s
820ms
$0.03
$0.000
23.0
83
141 t/s
790ms
$0.15
$0.001
118.5
84
140 t/s
830ms
$0.15
$0.001
124.5
85
140 t/s
1.19s
$0.66
$0.005
785.4
86
138 t/s
880ms
$0.15
$0.001
132.0
87
135 t/s
1.99s
$0.35
$0.003
696.5
88
Nex-N2-Pro
Nex AGI
135 t/s
1.71s
$1.0
$0.007
1710.0
89
135 t/s
18.95s
$1.6
$0.012
29618.8
90
135 t/s
880ms
$0.26
$0.002
230.6
91
134 t/s
132.71s
$5.6
$0.042
746493.8
92
133 t/s
760ms
$0.17
$0.001
133.0
93
133 t/s
2.26s
$0.80
$0.006
1808.0
94
132 t/s
2.25s
$0.00
95
132 t/s
2.28s
$0.80
$0.006
1824.0
96
132 t/s
2.37s
$3.0
$0.023
7110.0
97
132 t/s
2.34s
$1.1
$0.008
2574.0
98
129 t/s
1.04s
$11.3
$0.087
11700.0
99
129 t/s
1.09s
$0.56
$0.004
613.7
100
128 t/s
940ms
$0.09
$0.001
87.4
Showing top 100 of 329 models. Use search/filter to narrow down.

Speed Metrics Guide

Throughput (tokens/s)

Output generation speed in tokens per second. Higher is better.

Good: >50 t/s · Excellent: >100 t/s
Time to First Token (TTFT)

Delay before the first token appears. Lower is better.

Good: <500ms · Excellent: <200ms
Price/Performance

Cost efficiency ratios. Lower values indicate better value.

$/Speed: price per t/s · Price×TTFT: latency penalty

Compare pricing for all models side by side

Open AI API Cost Calculator →