About Pelican
A leaderboard you can trust is a leaderboard you can verify.
Pelican is Seerist's live evaluation platform for large language models. It exists to answer one question with evidence: which models are actually best at the work that matters — and where is each one weak? Every score on the board traces back to a concrete question with a verifiable answer.
Principle 01
Contamination-free by design
Test questions rot. Once a benchmark circulates, it leaks into training data and stops measuring ability. Pelican follows the LiveBench discipline: questions are released in dated batches, retired on a rolling cadence, and new questions are drawn from sources newer than any model's training cutoff — recent competitions, fresh datasets, current documents.
Principle 02
Ground truth, not vibes
Every question ships with a machine-verifiable answer. Scoring is objective — no LLM judges grading other LLMs, no preference panels. A model's score is the percentage of questions it answers correctly, and anyone can re-run the grader and get the same number.
Principle 03
Seven categories, one signal
The global average is the headline, but the breakdown is the story. Each category aggregates several tasks:
- Reasoning
- Logic puzzles, spatial reasoning, and multi-step deduction with a single verifiable answer.
- Coding
- Code completion and generation graded by executing the result against hidden tests.
- Agentic Coding
- Multi-file repository tasks: navigation, coordinated edits, and test repair.
- Mathematics
- Competition problems, olympiad proofs, and applied linear algebra.
- Data Analysis
- Table joins, column type annotation, and reformatting over real datasets.
- Language
- Word puzzles, typo correction, and plot unscrambling — precision language work.
- Instruction Following
- Paraphrase, simplification, and summarization under strict constraints.
Reading the board
How to interpret a score
Scores are percentages of questions answered correctly, averaged per category. A bar like this one reads as "87.42% of Reasoning questions correct in this release":
Reasoning
87.4
Comparisons are only meaningful within a release tag — the question pool changes between releases, so use the score history on each model page to track trend, not raw cross-release deltas.
Cadence
Releases and commentary
New models are added as they ship and receive a "new" badge for their first release. Each model page carries an annotation thread where analysts record what changed and why it matters — the qualitative record behind the quantitative one.
Pelican is built and operated by Seerist, the risk and threat intelligence company. Methodology adapted from LiveBench (White et al.), extended with analyst commentary and an open question pipeline. Contribute questions →