What PinchBench Tests

PinchBench evaluates AI models on 6 real-world agent scenarios: coding agent (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA, AIME 2025, MATH-500, HLE), instruction following (IFBench, MMLU-Pro), research & analysis, and tool use & agentic (τ²-bench, LCR). For raw academic scores, see our full LLM benchmark leaderboard, and for raw speed numbers, the AI model speed rankings.

What is PinchBench? (AI Agent Benchmark Leaderboard 2026)

PinchBench is a real-world AI agent benchmark that scores 400+ models on production-style agent tasks — coding agent scenarios (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA Diamond, AIME 2025, HLE, MATH-500), instruction following (IFBench, MMLU-Pro), tool use & agentic (τ²-bench, LCR), and research workflows. Each model receives two composite scores: a Coding Index weighted toward agentic coding performance and an Intelligence Index weighted toward general reasoning.

Live data · Last updated · Refreshed hourly from Artificial Analysis

Top agents today by Coding Index: GPT-5.6 Sol (xhigh) 78.3, Claude Opus 5 (Max) 78.0, GPT-5.6 Sol (max) 77.4. Top agents by Intelligence Index: Claude Opus 5 (Max) II 60.7, Claude Opus 5 (Xhigh) II 60.1, Claude Fable 5 II 59.9. The leaderboard is updated hourly from Artificial Analysis data — no signup required.

How is PinchBench different from GPQA, HLE, or other benchmarks?

Most benchmarks measure single-turn academic performance (GPQA, MMLU-Pro, AIME 2025). PinchBench measures end-to-end agent behavior: can the model write a function, run it, see it fail, and iterate? That requires tool use (τ²-bench, LCR), instruction following (IFBench), and multi-step coding (LiveCodeBench, TerminalBench, SciCode). It's the difference between "can solve a math problem" and "can build a working app from a spec."

Top 3 on PinchBench Today (August 2026)

  1. 🥇Claude Opus 5 (Adaptive Reasoning, Max Effort) — Intelligence Index 60.7, Coding Index 78.0View →
  2. 🥈GPT-5.6 Sol (xhigh) — Coding Index 78.3, Intelligence Index 57.7View →
  3. 🥉Claude Fable 5 — Intelligence Index 59.9, HLE 53.3% #1View →
Live data · Updated hourly

PinchBench — Real-World AI Agent Benchmarks

How do AI models perform on real agent tasks? PinchBench scores 630+ models across coding, reasoning, tool use, and instruction following — with live pricing data.

Models Tested
630
Scenarios
6
Avg Score
24.2
Best Value
Ling 3.0 Flash
Overall

Balanced score across all agent capabilities

intelligence index (15%)coding index (15%)math index (10%)gpqa (10%)livecodebench (10%)ifbench (10%)tau2 (10%)terminalbench hard (10%)hle (10%)
🥇#173.6
Anthropic

Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)

Price
$20.00
Speed
66
Efficiency
3.7
🥈#272.8
Anthropic

Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)

Price
$20.00
Speed
69
Efficiency
3.6
🥉#370.8
Anthropic

Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)

Price
$20.00
Speed
56
Efficiency
3.5
#ModelScoreBarInput $/MOutput $/MSpeedTTFTEfficiency
1
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Anthropic
73.6
$10.00$50.0066282.34s3.7
2
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Anthropic
72.8
$10.00$50.0069111.68s3.6
3
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)
Anthropic
70.8
$10.00$50.005630.36s3.5
4
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Anthropic
70.5
$5.00$25.005977.25s7.1
5
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
Anthropic
69.8
$5.00$25.005532.99s7.0
6
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Anthropic
69.3
$10.00$50.0070104.62s3.5
7
Muse Spark 1.3 (max)
Meta
69.2
8
GPT-5.6 Sol (max)
OpenAI
69.2
$4.00$20.007798.87s8.6
9
GPT-6 Astra (max)
OpenAI
69.0
$10.00$50.003.5
10
Claude Opus 5 (Adaptive Reasoning, High Effort)
Anthropic
69.0
$5.00$25.005617.14s6.9
11
Grok 4.6 (high)
SpaceXAI
68.9
$2.00$6.006147.23s23.0
12
Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)
Anthropic
68.8
$10.00$50.00558.71s3.4
13
GPT-6 Astra (high)
OpenAI
68.7
$10.00$50.003.4
14
Muse Spark 1.3 (xhigh)
Meta
68.6
$1.25$4.2518642.55s34.3
15
GPT-5.6 Sol (xhigh)
OpenAI
68.6
$4.00$20.007733.68s8.6
16
GPT-6 Astra (xhigh)
OpenAI
68.5
$10.00$50.003.4
17
Grok 4.6 (xhigh)
SpaceXAI
68.0
$2.00$6.005937.66s22.7
18
GPT-6 Astra (medium)
OpenAI
68.0
$10.00$50.003.4
19
Kimi K3 (max)
Kimi
68.0
$3.00$15.00385.33s11.3
20
Gemini 3.8 Flash (high)
Google
67.5
$0.75$3.7532712.91s45.0

💰 Best Cost Efficiency — Overall

Score per dollar (higher = better value). Only models with pricing data.

1
Ling 3.0 Flash
409.3$0.11
2
Agnes 2.5 Pro Beta
371.3$0.15
3
Qwen3.5 4B (Reasoning)
358.3$0.06
4
Hy3-preview (Reasoning)
351.0$0.10
5
HyperNova 60B 2605 (high, based on gpt-oss-120b)
319.2$0.07
6
Qwen3.5 4B (Non-reasoning)
303.3$0.06
7
Granite 4.2 3B
300.0$0.05
8
DeepSeek V4 Flash (Reasoning, Max Effort)
292.6$0.17
9
Qwen3.8-Flash-Next
280.2$0.23
10
Hy3-preview (Non-reasoning)
271.4$0.10

⚡ Score vs Speed — Overall

Models in the top-right are both fast and capable.

Celeris
Celeris-1
Score
13.4
Speed
1621
Inception
Mercury 2
Score
26.5
Speed
804
Google
Gemini 3.5 Flash-Lite
Score
43.4
Speed
391
Google
Gemini 3.8 Flash (high)
Score
67.5
Speed
327
Google
Gemini 3.8 Flash (medium)
Score
65.4
Speed
312
Google
Gemini 3.7 Flash (high)
Score
66.0
Speed
310
Google
Gemini 3.8 Flash (low)
Score
62.6
Speed
313
InclusionAI
Ling 3.0 Flash
Score
44.2
Speed
346
Google
Gemini 3.7 Flash (medium)
Score
62.5
Speed
280
Google
Gemini 3.7 Flash (low)
Score
61.0
Speed
258

Frequently Asked Questions

What is PinchBench and how does it differ from traditional benchmarks?

PinchBench evaluates AI models on real-world agent tasks spanning coding, reasoning, tool use, and instruction following. Unlike academic benchmarks that test isolated capabilities, PinchBench combines multiple benchmark dimensions to reflect how models perform as autonomous agents in practical workflows.

Which scenarios does PinchBench test?

PinchBench covers 6 scenarios: Coding Agent (code generation, debugging, terminal use), Reasoning & Logic (math, science, multi-step problems), Instruction Following (format compliance, structured output), Research & Analysis (scientific reasoning, knowledge), Tool Use & Agentic (multi-turn orchestration, planning), and an Overall balanced score.

How are scores calculated?

Each scenario uses a weighted combination of relevant benchmarks. For example, Coding Agent combines LiveCodeBench, TerminalBench, SciCode, and the Artificial Analysis Coding Index. Scores are normalized to 0-100. Cost efficiency is calculated as score divided by price per million tokens.

Why do real-world results differ from academic benchmarks?

Academic benchmarks test specific skills in controlled conditions. Real agent tasks require combining multiple skills — a model might score well on individual benchmarks but struggle when tasks require coding + tool use + instruction following simultaneously. PinchBench's weighted scenario scores better approximate this combined performance.

How often is the data updated?

PinchBench data refreshes hourly from the Artificial Analysis API, ensuring you see the latest benchmark scores and pricing for all models.