Skip to content

About Pelican

A leaderboard you can trust is a leaderboard you can verify.

Pelican is Seerist's live evaluation platform for large language models. It exists to answer one question with evidence: which models are actually best at the work that matters — and where is each one weak? Every score on the board traces back to a concrete question with a verifiable answer.

Principle 01

Contamination-free by design

Test questions rot. Once a benchmark circulates, it leaks into training data and stops measuring ability. Pelican follows the LiveBench discipline: questions are released in dated batches, retired on a rolling cadence, and new questions are drawn from sources newer than any model's training cutoff — recent competitions, fresh datasets, current documents.

Principle 02

Ground truth, not vibes

Every question ships with a machine-verifiable answer. Scoring is objective — no LLM judges grading other LLMs, no preference panels. A model's score is the percentage of questions it answers correctly, and anyone can re-run the grader and get the same number.

Principle 03

Seven categories, one signal

The global average is the headline, but the breakdown is the story. Each category aggregates several tasks:

Reasoning
Logic puzzles, spatial reasoning, and multi-step deduction with a single verifiable answer.
Coding
Code completion and generation graded by executing the result against hidden tests.
Agentic Coding
Multi-file repository tasks: navigation, coordinated edits, and test repair.
Mathematics
Competition problems, olympiad proofs, and applied linear algebra.
Data Analysis
Table joins, column type annotation, and reformatting over real datasets.
Language
Word puzzles, typo correction, and plot unscrambling — precision language work.
Instruction Following
Paraphrase, simplification, and summarization under strict constraints.

Reading the board

How to interpret a score

Scores are percentages of questions answered correctly, averaged per category. A bar like this one reads as "87.42% of Reasoning questions correct in this release":

Reasoning

87.4

Comparisons are only meaningful within a release tag — the question pool changes between releases, so use the score history on each model page to track trend, not raw cross-release deltas.

Cadence

Releases and commentary

New models are added as they ship and receive a "new" badge for their first release. Each model page carries an annotation thread where analysts record what changed and why it matters — the qualitative record behind the quantitative one.

Pelican is built and operated by Seerist, the risk and threat intelligence company. Methodology adapted from LiveBench (White et al.), extended with analyst commentary and an open question pipeline. Contribute questions →