Every AI benchmark, as a graph you can read.

What each benchmark measures, who publishes it, whether it has gone stale, and which models it covers, with the date on every score. A place to learn what the numbers mean before you trust them.

Today the graph holds 164 benchmarks across 549 scored models, 10,887 scores in all, each carrying the date it was taken.

What a benchmark page holds

One page per benchmark, kept current by daily research and merged by a person. This preview uses the scores already in the ModelSpec cards.

SWE-bench Verifiedbenchgraph.dev/b/swe-bench-verified
Measures
Whether a model can resolve real GitHub issues end to end: read the repository, write a patch, and pass the hidden tests. Coding, agentic.
Publisher
OpenAI, as a human-validated subset of Princeton's SWE-bench.
Dataset
500 tasks drawn from 12 open-source Python repositories.
Lineage
SWE-bench, then Verified; successors in the graph: SWE-bench Pro, Multilingual, Multimodal.
Saturation
Watch. Leading scores crossed 70 percent in 2025; the harder successors are where the spread now lives.
Freshness
Scores carry the date they were recorded, and the page shows when a score was last confirmed.

Models covered

112 models carry a score today. The top of the table, from the cards:

modelproviderscoreas of
Claude Mythos Previewanthropic93.92026-04
Claude Opus 4.6anthropic80.82026-04
Gemini 3.1 Pro Previewgoogle80.62026-04
Claude Haiku 4.5 (latest)anthropic73.32026-04
o3-proopenai73.22026-04
Claude Opus 4anthropic72.72026-04

The families in the graph today

One node per benchmark, grouped by what it measures. The graph also links each benchmark to the models that report it and to the benchmark it succeeded.

knowledge 66 benchmarks coding 28 benchmarks multimodal 13 benchmarks embeddings 10 benchmarks preference 9 benchmarks reasoning 9 benchmarks domain 8 benchmarks agentic 6 benchmarks math 5 benchmarks translation 4 benchmarks safety 4 benchmarks long context 2 benchmarks

For learning

A plain explanation of each benchmark: the task, the dataset, the licence, the publisher, and what a good score does and does not tell you.

For choosing

Start from the job you need done and see which benchmarks actually measure it, then which models report a recent score and how the spread looks.

For agents

The same graph feeds ModelSpec, so an agent picking a model gets the benchmark evidence and its date, not a leaderboard rank with no context.

Coming online. The benchmark data lives in the open ModelSpec cards today and moves into pages here as the daily research job comes up. Both projects are free and open source, built by the team behind Dark Product Factories.