- Measures
- Whether a model can resolve real GitHub issues end to end: read the repository, write a patch, and pass the hidden tests. Coding, agentic.
- Publisher
- OpenAI, as a human-validated subset of Princeton's SWE-bench.
- Dataset
- 500 tasks drawn from 12 open-source Python repositories.
- Lineage
- SWE-bench, then Verified; successors in the graph: SWE-bench Pro, Multilingual, Multimodal.
- Saturation
- Watch. Leading scores crossed 70 percent in 2025; the harder successors are where the spread now lives.
- Freshness
- Scores carry the date they were recorded, and the page shows when a score was last confirmed.
Models covered
112 models carry a score today. The top of the table, from the cards:
| model | provider | score | as of |
|---|---|---|---|
| Claude Mythos Preview | anthropic | 93.9 | 2026-04 |
| Claude Opus 4.6 | anthropic | 80.8 | 2026-04 |
| Gemini 3.1 Pro Preview | 80.6 | 2026-04 | |
| Claude Haiku 4.5 (latest) | anthropic | 73.3 | 2026-04 |
| o3-pro | openai | 73.2 | 2026-04 |
| Claude Opus 4 | anthropic | 72.7 | 2026-04 |