Skip to content

LLM Evaluation · Release 2026-06-25

The Seerist LLM Leaderboard

Contamination-free benchmarks with objective ground-truth scoring, refreshed on a rolling release cadence. No LLM judges, no leaked test sets.

45 models550 ground-truth questions1 releaseleader Claude Fable 5 Max Effort 83.0

Release
Category
LLM leaderboard for release 2026-06-25, sortable by global average, category scores
#
183.0$20.089.786.062.296.080.590.775.8
281.0$4.0091.783.956.296.279.887.771.8
380.2$11.389.782.254.095.981.687.470.7
480.1$10.091.281.565.295.774.588.763.8
579.590.382.564.783.979.984.471.0
6
Kimi K3
79.2$6.0090.781.562.284.478.785.571.4
778.8$0.7587.878.958.393.568.085.579.9
878.5$3.0088.272.964.791.378.479.774.1
978.0$3.0090.576.857.092.673.983.771.9
1078.0$5.6388.177.553.894.279.382.670.2
1178.0$2.0090.077.557.691.276.578.674.3
1277.9$4.5090.678.355.094.979.382.964.6
1377.4$1.7885.877.255.095.179.282.167.7
1477.0$4.5084.076.544.191.078.585.479.1
1576.5$10.087.282.150.792.878.377.966.7
1676.2$10.089.281.850.594.366.079.772.0
1776.0$4.0088.780.759.492.971.775.063.9
1875.8$3.0087.268.656.590.873.082.871.5
1975.3$2.0087.777.258.587.172.574.369.6
2075.3$1.1480.075.761.486.276.674.372.7

Scores are the mean percentage correct on ground-truth graded tasks. Tint = position within each column:Click ▸ on a row for its per-task breakdown. Methodology →

Signal analysis

The efficiency frontier

What a point of capability costs. Amber models are Pareto-optimal — nothing scores higher for less. Everything below the dashed line is paying for capability it does not get.

607080$0.10$0.30$1.00$3.00$10.0$30.0blended list price $/MTok · log scaleClaude Fable 5 Max Effort — 83.0 at $20.0/MTokGPT 5.6 Sol Max — 81.0 at $4.00/MTokGPT 5.5 Xhigh — 80.2 at $11.3/MTokClaude Opus 5 Max Effort — 80.1 at $10.0/MTokKimi K3 — 79.2 at $6.00/MTokGemini 3.7 Flash High — 78.8 at $0.75/MTokQwen3.8 Max — 78.5 at $3.00/MTokGrok 4.6 — 78.0 at $3.00/MTokGPT 5.4 Xhigh — 78.0 at $5.63/MTokMuse Spark 1.2 Xhigh — 78.0 at $2.00/MTokGPT 5.6 Terra Max — 77.9 at $4.50/MTokDeepseek V4 Pro 0813 — 77.4 at $1.78/MTokGemini 3.1 Pro Preview High — 77.0 at $4.50/MTokClaude Opus 4 7 Xhigh Effort — 76.5 at $10.0/MTokClaude Opus 4 8 Max Effort — 76.2 at $10.0/MTokClaude Sonnet 5 Xhigh Effort — 76.0 at $4.00/MTokGrok 4.5 — 75.8 at $3.00/MTokMuse Spark 1.1 Xhigh — 75.3 at $2.00/MTokQwen3.8 27b — 75.3 at $1.14/MTokGemini 3.5 Flash High — 74.6 at $3.38/MTokGPT 5.2 2025 12 11 High — 74.6 at $4.81/MTokClaude Opus 4 6 Thinking Auto High Effort — 74.5 at $10.0/MTokDeepseek V4 Flash 0731 — 74.2 at $0.10/MTokGPT 5.2 Codex — 74.0 at $4.81/MTokGemini 3.6 Flash High — 73.6 at $1.50/MTokGPT 5.6 Luna Max — 73.6 at $0.45/MTokGLM 5.2 — 73.2 at $1.48/MTokQwen3.7 Max — 73.1 at $2.21/MTokClaude Sonnet 4 6 Thinking Auto Medium Effort — 73.0 at $6.00/MTokClaude Opus 4 5 20251101 Thinking 64k High Effort — 72.6 at $10.0/MTokInkling Xhigh — 71.9 at $1.72/MTokDeepseek V4 Pro — 71.6 at $0.60/MTokKimi K2.6 Thinking — 70.5 at $0.98/MTokGPT 5.4 Nano Xhigh — 69.6 at $0.46/MTokQwen3.6 Plus — 68.9 at $0.73/MTokKimi K2.7 Code — 68.4 at $1.35/MTokGrok Build 0.1 — 67.8 at $1.25/MTokMinimax M3 — 67.3 at $0.52/MTokGPT 5.4 Mini Xhigh — 66.4 at $1.69/MTokDeepseek V4 Flash — 65.5 at $0.09/MTokQwen3.6 27b — 64.0 at $1.35/MTokGemini 3.5 Flash Lite High — 63.9 at $0.85/MTokGrok 4.3 — 62.3 at $1.56/MTokDeepseek V4 FlashDeepseek V4 Flash 0731Gemini 3.7 Flash HighGPT 5.6 Sol MaxClaude Fable 5 Max Effort
on the frontier dominated — a frontier model is better and cheaper43 of 45 models priced (OpenRouter list, 3:1 in:out blend) · release 2026-06-25