AI Glossary
What is Benchmark?
Benchmarks measure models on fixed task sets — coding (SWE-bench), knowledge (MMLU-Pro), science (GPQA), or human preference (LMArena Elo). They enable comparison but saturate over time and can be gamed by training on test data, which is why new, harder benchmarks keep appearing.
News momentum
Mentions in recent AI news titles and summaries, refreshed from the ingestion stream.
267
-39 vs prior 7d
- Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture
marktechpost · 2h
- Reconstructing the benchmark behind Luc Julia's 64% LLM reliability claim
hn-ai · 10h
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
arxiv-ai · 1d
- A Geometric Theory of Robust Fairness Audits
arxiv-ai · 1d
Related:Large Language Model (LLM)
← All terms