LLM Evaluation · Release 2026-06-25
The Seerist LLM Leaderboard Contamination-free benchmarks with objective ground-truth scoring, refreshed on a rolling release cadence. No LLM judges, no leaked test sets.
45 models· 550 ground-truth questions· 1 release· leader Claude Fable 5 Max Effort 83.0 →
Open weights only Release 2026-06-25 · latest
Category All Reasoning Coding Agentic Coding Mathematics Data Analysis Language Instruction Following
LLM leaderboard for release 2026-06-25, sortable by global average, category scores # Model Avg $/MTok Reasoning Coding Agentic Coding Mathematics Data Analysis Language Instruction Following 1 83.0 $20.0 89.7 86.0 62.2 96.0 80.5 90.7 75.8 2 81.0 $4.00 91.7 83.9 56.2 96.2 79.8 87.7 71.8 3 80.2 $11.3 89.7 82.2 54.0 95.9 81.6 87.4 70.7 4 80.1 $10.0 91.2 81.5 65.2 95.7 74.5 88.7 63.8 5 79.5 – 90.3 82.5 64.7 83.9 79.9 84.4 71.0 6 79.2 $6.00 90.7 81.5 62.2 84.4 78.7 85.5 71.4 7 78.8 $0.75 87.8 78.9 58.3 93.5 68.0 85.5 79.9 8 78.5 $3.00 88.2 72.9 64.7 91.3 78.4 79.7 74.1 9 78.0 $3.00 90.5 76.8 57.0 92.6 73.9 83.7 71.9 10 78.0 $5.63 88.1 77.5 53.8 94.2 79.3 82.6 70.2 11 78.0 $2.00 90.0 77.5 57.6 91.2 76.5 78.6 74.3 12 77.9 $4.50 90.6 78.3 55.0 94.9 79.3 82.9 64.6 13 77.4 $1.78 85.8 77.2 55.0 95.1 79.2 82.1 67.7 14 77.0 $4.50 84.0 76.5 44.1 91.0 78.5 85.4 79.1 15 76.5 $10.0 87.2 82.1 50.7 92.8 78.3 77.9 66.7 16 76.2 $10.0 89.2 81.8 50.5 94.3 66.0 79.7 72.0 17 76.0 $4.00 88.7 80.7 59.4 92.9 71.7 75.0 63.9 18 75.8 $3.00 87.2 68.6 56.5 90.8 73.0 82.8 71.5 19 75.3 $2.00 87.7 77.2 58.5 87.1 72.5 74.3 69.6 20 75.3 $1.14 80.0 75.7 61.4 86.2 76.6 74.3 72.7
Show all 45 modelsScores are the mean percentage correct on ground-truth graded tasks. Tint = position within each column: min max Click ▸ on a row for its per-task breakdown. Methodology →
Signal analysis
The efficiency frontier What a point of capability costs. Amber models are Pareto-optimal — nothing scores higher for less. Everything below the dashed line is paying for capability it does not get.
60 70 80 $0.10 $0.30 $1.00 $3.00 $10.0 $30.0 blended list price $/MTok · log scale Claude Fable 5 Max Effort — 83.0 at $20.0/MTok GPT 5.6 Sol Max — 81.0 at $4.00/MTok GPT 5.5 Xhigh — 80.2 at $11.3/MTok Claude Opus 5 Max Effort — 80.1 at $10.0/MTok Kimi K3 — 79.2 at $6.00/MTok Gemini 3.7 Flash High — 78.8 at $0.75/MTok Qwen3.8 Max — 78.5 at $3.00/MTok Grok 4.6 — 78.0 at $3.00/MTok GPT 5.4 Xhigh — 78.0 at $5.63/MTok Muse Spark 1.2 Xhigh — 78.0 at $2.00/MTok GPT 5.6 Terra Max — 77.9 at $4.50/MTok Deepseek V4 Pro 0813 — 77.4 at $1.78/MTok Gemini 3.1 Pro Preview High — 77.0 at $4.50/MTok Claude Opus 4 7 Xhigh Effort — 76.5 at $10.0/MTok Claude Opus 4 8 Max Effort — 76.2 at $10.0/MTok Claude Sonnet 5 Xhigh Effort — 76.0 at $4.00/MTok Grok 4.5 — 75.8 at $3.00/MTok Muse Spark 1.1 Xhigh — 75.3 at $2.00/MTok Qwen3.8 27b — 75.3 at $1.14/MTok Gemini 3.5 Flash High — 74.6 at $3.38/MTok GPT 5.2 2025 12 11 High — 74.6 at $4.81/MTok Claude Opus 4 6 Thinking Auto High Effort — 74.5 at $10.0/MTok Deepseek V4 Flash 0731 — 74.2 at $0.10/MTok GPT 5.2 Codex — 74.0 at $4.81/MTok Gemini 3.6 Flash High — 73.6 at $1.50/MTok GPT 5.6 Luna Max — 73.6 at $0.45/MTok GLM 5.2 — 73.2 at $1.48/MTok Qwen3.7 Max — 73.1 at $2.21/MTok Claude Sonnet 4 6 Thinking Auto Medium Effort — 73.0 at $6.00/MTok Claude Opus 4 5 20251101 Thinking 64k High Effort — 72.6 at $10.0/MTok Inkling Xhigh — 71.9 at $1.72/MTok Deepseek V4 Pro — 71.6 at $0.60/MTok Kimi K2.6 Thinking — 70.5 at $0.98/MTok GPT 5.4 Nano Xhigh — 69.6 at $0.46/MTok Qwen3.6 Plus — 68.9 at $0.73/MTok Kimi K2.7 Code — 68.4 at $1.35/MTok Grok Build 0.1 — 67.8 at $1.25/MTok Minimax M3 — 67.3 at $0.52/MTok GPT 5.4 Mini Xhigh — 66.4 at $1.69/MTok Deepseek V4 Flash — 65.5 at $0.09/MTok Qwen3.6 27b — 64.0 at $1.35/MTok Gemini 3.5 Flash Lite High — 63.9 at $0.85/MTok Grok 4.3 — 62.3 at $1.56/MTok Deepseek V4 Flash Deepseek V4 Flash 0731 Gemini 3.7 Flash High GPT 5.6 Sol Max Claude Fable 5 Max Effort on the frontier dominated — a frontier model is better and cheaper43 of 45 models priced (OpenRouter list, 3:1 in:out blend) · release 2026-06-25