AINews Portal

🏆 AI Benchmarks Index

What each major benchmark measures and who leads it — every score served from our own tracking pipeline and refreshed daily. JSON is available at /api/benchmarks.

🔬 GPQA Diamond

Graduate-level science questions designed to be Google-proof — tests deep reasoning.

  1. 1Gemini 3.7 Flash94.8%
  2. 2GPT-5.4 Pro94.6%
  3. 3Gemini 3.6 Flash94.1%
  4. 4Grok 4.694.0%
  5. 5GPT-5.594.0%
  6. 6GPT-5.5 Pro93.9%
  7. 7Claude Opus 593.9%
  8. 8GPT-5.6 Sol93.5%
  9. 9Grok 4.593.4%
  10. 10GPT-5.6 Terra93.3%
💻 SWE-Bench Verified

Real GitHub issues an agent must fix end-to-end — the standard for coding agents.

  1. 1Claude Opus 4.783.5%
  2. 2Claude Opus 4.7 (Fast)83.5%
  3. 3GPT-5.580.6%
  4. 4Gemini 3.5 Flash79.3%
  5. 5Claude Opus 4.678.7%
  6. 6GLM 5.278.7%
  7. 7DeepSeek V4 Pro77.6%
  8. 8Qwen3.7 Max77.3%
  9. 9GPT-5.476.9%
  10. 10Kimi K2.676.7%
MATH Level 5

The hardest tier of competition math problems from the MATH benchmark.

  1. 1GPT-598.1%
  2. 2GPT-5 Mini97.8%
  3. 3o4 Mini97.8%
  4. 4o397.8%
  5. 5Claude Sonnet 4.597.7%
  6. 6Qwen3 Max97.1%
  7. 7R196.6%
  8. 8o3 Mini96.5%
  9. 9Claude Haiku 4.596.4%
  10. 10Gemini 2.5 Pro95.9%
🧮 OTIS Mock AIME

Olympiad-style mock AIME exam (2024–25) — elite high-school competition math.

  1. 1Claude Fable 5100.0%
  2. 2GPT-5.5 Pro100.0%
  3. 3GPT-5.5100.0%
  4. 4GPT-5.6 Sol100.0%
  5. 5GPT-5.6 Terra99.7%
  6. 6Grok 4.699.2%
  7. 7Claude Opus 598.9%
  8. 8DeepSeek V4 Pro 081398.6%
  9. 9GPT-5.6 Luna98.3%
  10. 10Claude Opus 4.898.3%
♾️ FrontierMath

Research-level math problems written by mathematicians — the hardest math benchmark.

  1. 1GPT-5.5 Pro52.4%
  2. 2GPT-5.551.7%
  3. 3GPT-5.4 Pro50.0%
  4. 4GPT-5.447.6%
  5. 5Claude Opus 4.847.2%
  6. 6Claude Opus 4.8 (Fast)47.2%
  7. 7Claude Opus 4.7 (Fast)43.8%
  8. 8Claude Opus 4.743.8%
  9. 9GPT-5.240.7%
  10. 10Claude Opus 4.640.7%
SimpleQA Verified

Short factual questions measuring hallucination resistance — can it just be right?

  1. 1GPT-5.6 Sol71.6%
  2. 2Gemini 3.7 Flash71.2%
  3. 3Gemini 3.6 Flash68.7%
  4. 4Gemini 3.5 Flash68.4%
  5. 5Claude Fable 568.3%
  6. 6Qwen3 Max67.5%
  7. 7GPT-5.5 Pro64.5%
  8. 8GPT-5.563.1%
  9. 9Qwen3.7 Max58.5%
  10. 10DeepSeek V4 Pro57.0%
💬 LMArena (Text)

Human preference Elo from blind head-to-head chat battles — the de-facto vibes benchmark.

  1. 1Claude Opus 4.61497
  2. 2Claude Fable 51495
  3. 3Claude Opus 4.71483
  4. 4Gemini 3.5 Flash1482
  5. 5Qwen3.8 Max1482
  6. 6Gemini 3.1 Pro Preview1480
  7. 7Gemini 3.6 Flash1479
  8. 8Muse Spark 1.11478
  9. 9Kimi K31473
  10. 10ernie-5.11468
🖼️ LMArena (Image)

Human preference Elo for text-to-image generation battles.

  1. 1gpt-image-2 (medium)1381
  2. 2reve-2.11302
  3. 3muse-image1283
  4. 4gemini-3.1-flash-image-preview (nano-banana-2) [web-search]1270
  5. 5reve-2.01270
  6. 6gemini-3.1-flash-image (nano-banana-2) [web-search]1263
  7. 7qwen-image-3.0-pro1258
  8. 8seedream-5.0-pro1257
  9. 9mai-image-2.51256
  10. 10gemini-3.1-flash-lite-image (nano-banana-2-lite)1251
🎬 LMArena (Video)

Human preference Elo for image-to-video generation battles.

  1. 1minimax-h31489
  2. 2dreamina-seedance-2.5-720p1484
  3. 3dreamina-seedance-2.0-720p1478
  4. 4grok-imagine-video-1.5-preview-720p1466
  5. 5gemini-omni-flash1462
  6. 6grok-imagine-video-1.5-720p1460
  7. 7flux-3-video-202608111449
  8. 8happyhorse-1.01442
  9. 9wan2.7-i2v1428
  10. 10grok-imagine-video-720p1415

Score data: Epoch AI Benchmarking Hub (CC-BY) and LMArena (CC-BY), ingested daily and matched to our model catalog. Rankings are never for sale — see the methodology and full model leaderboard.