🏆 AI Benchmarks Index
What each major benchmark measures and who leads it — every score served from our own tracking pipeline and refreshed daily. JSON is available at /api/benchmarks.
Graduate-level science questions designed to be Google-proof — tests deep reasoning.
- 1Gemini 3.7 Flashgoogle94.8%
- 2GPT-5.4 Proopenai94.6%
- 3Gemini 3.6 Flashgoogle94.1%
- 4Grok 4.6x-ai94.0%
- 5GPT-5.5openai94.0%
- 6GPT-5.5 Proopenai93.9%
- 7Claude Opus 5anthropic93.9%
- 8GPT-5.6 Solopenai93.5%
- 9Grok 4.5x-ai93.4%
- 10GPT-5.6 Terraopenai93.3%
Real GitHub issues an agent must fix end-to-end — the standard for coding agents.
- 1Claude Opus 4.7anthropic83.5%
- 2Claude Opus 4.7 (Fast)anthropic83.5%
- 3GPT-5.5openai80.6%
- 4Gemini 3.5 Flashgoogle79.3%
- 5Claude Opus 4.6anthropic78.7%
- 6GLM 5.2z-ai78.7%
- 7DeepSeek V4 Prodeepseek77.6%
- 8Qwen3.7 Maxqwen77.3%
- 9GPT-5.4openai76.9%
- 10Kimi K2.6moonshotai76.7%
The hardest tier of competition math problems from the MATH benchmark.
- 1GPT-5openai98.1%
- 2GPT-5 Miniopenai97.8%
- 3o4 Miniopenai97.8%
- 4o3openai97.8%
- 5Claude Sonnet 4.5anthropic97.7%
- 6Qwen3 Maxqwen97.1%
- 7R1deepseek96.6%
- 8o3 Miniopenai96.5%
- 9Claude Haiku 4.5anthropic96.4%
- 10Gemini 2.5 Progoogle95.9%
Olympiad-style mock AIME exam (2024–25) — elite high-school competition math.
- 1Claude Fable 5anthropic100.0%
- 2GPT-5.5 Proopenai100.0%
- 3GPT-5.5openai100.0%
- 4GPT-5.6 Solopenai100.0%
- 5GPT-5.6 Terraopenai99.7%
- 6Grok 4.6x-ai99.2%
- 7Claude Opus 5anthropic98.9%
- 8DeepSeek V4 Pro 0813deepseek98.6%
- 9GPT-5.6 Lunaopenai98.3%
- 10Claude Opus 4.8anthropic98.3%
Research-level math problems written by mathematicians — the hardest math benchmark.
- 1GPT-5.5 Proopenai52.4%
- 2GPT-5.5openai51.7%
- 3GPT-5.4 Proopenai50.0%
- 4GPT-5.4openai47.6%
- 5Claude Opus 4.8anthropic47.2%
- 6Claude Opus 4.8 (Fast)anthropic47.2%
- 7Claude Opus 4.7 (Fast)anthropic43.8%
- 8Claude Opus 4.7anthropic43.8%
- 9GPT-5.2openai40.7%
- 10Claude Opus 4.6anthropic40.7%
Short factual questions measuring hallucination resistance — can it just be right?
- 1GPT-5.6 Solopenai71.6%
- 2Gemini 3.7 Flashgoogle71.2%
- 3Gemini 3.6 Flashgoogle68.7%
- 4Gemini 3.5 Flashgoogle68.4%
- 5Claude Fable 5anthropic68.3%
- 6Qwen3 Maxqwen67.5%
- 7GPT-5.5 Proopenai64.5%
- 8GPT-5.5openai63.1%
- 9Qwen3.7 Maxqwen58.5%
- 10DeepSeek V4 Prodeepseek57.0%
Human preference Elo from blind head-to-head chat battles — the de-facto vibes benchmark.
- 1Claude Opus 4.6anthropic1497
- 2Claude Fable 5anthropic1495
- 3Claude Opus 4.7anthropic1483
- 4Gemini 3.5 Flashgoogle1482
- 5Qwen3.8 Maxqwen1482
- 6Gemini 3.1 Pro Previewgoogle1480
- 7Gemini 3.6 Flashgoogle1479
- 8Muse Spark 1.1meta1478
- 9Kimi K3moonshotai1473
- 10ernie-5.1baidu1468
Human preference Elo for text-to-image generation battles.
- 1gpt-image-2 (medium)openai1381
- 2reve-2.1reve1302
- 3muse-imagemeta1283
- 4gemini-3.1-flash-image-preview (nano-banana-2) [web-search]google1270
- 5reve-2.0reve1270
- 6gemini-3.1-flash-image (nano-banana-2) [web-search]google1263
- 7qwen-image-3.0-proalibaba1258
- 8seedream-5.0-probytedance1257
- 9mai-image-2.5microsoft-ai1256
- 10gemini-3.1-flash-lite-image (nano-banana-2-lite)google1251
Human preference Elo for image-to-video generation battles.
- 1minimax-h3minimax1489
- 2dreamina-seedance-2.5-720pbytedance1484
- 3dreamina-seedance-2.0-720pbytedance1478
- 4grok-imagine-video-1.5-preview-720pxai1466
- 5gemini-omni-flashgoogle1462
- 6grok-imagine-video-1.5-720pxai1460
- 7flux-3-video-20260811bfl1449
- 8happyhorse-1.0aorizon1442
- 9wan2.7-i2vwan1428
- 10grok-imagine-video-720pxai1415
Score data: Epoch AI Benchmarking Hub (CC-BY) and LMArena (CC-BY), ingested daily and matched to our model catalog. Rankings are never for sale — see the methodology and full model leaderboard.