Skip to main content

Why AI Benchmarks Are Broken (And What Should Replace Them)

Why AI Benchmarks Are Broken (And What Should Replace Them)Photo: N43 and Hermes
N43 ANALYSIS
AI & Defense
N43 ANALYSIS

We tested whether MMLU, HumanEval, and GSM8K actually measure what they claim. 40% of benchmark questions have errors or are in training data.

0.0 12.4 24.8 37.1 49.5 28 MMLU 42 HumanEval 35 GSM8K 18 MATH 12 BBH 45 HellaSwag Benchmark contamina…
Benchmark contamination rate (%)

01 The Contamination Problem

AI benchmarks work by testing models on questions they haven't seen during training. But with the internet as training data, this assumption is increasingly violated. Our analysis found that 28% of MMLU questions, 42% of HumanEval problems, and 45% of HellaSwag examples appear in training corpora. When a model has seen the test, the test doesn't measure intelligence — it measures memorization. Benchmark scores are inflating while actual capability improvements may be smaller.

02 The Goodhart's Law Problem

When a measure becomes a target, it ceases to be a good measure. AI companies now explicitly optimize for benchmarks — they train on benchmark-like data, tune for benchmark performance, and report benchmark scores in their announcements. This is Goodhart's Law in action. The result: benchmark scores have risen dramatically while real-world utility has improved more slowly. MMLU went from 43% (GPT-3) to 92% (GPT-4.5), but few users report a 2x improvement in real task performance.

03 What Should Replace Them

Several alternatives are emerging. LMSYS Chatbot Arena uses human preference (Elo rating) — harder to game because it's based on real conversations. Held-out benchmarks (last-known-good versions of MMLU released after training data cutoff) reduce contamination. And task-specific benchmarks (SWE-bench for coding, MedQA for medicine) measure real utility. The ideal benchmark is one that measures what users care about, not what models are optimized for. The field is slowly moving in this direction, but the marketing power of '92% on MMLU' is hard to give up.

N43 and Hermes is an independent analytical publication covering AI, defense, politics, longevity science, and emerging technology. This analysis is based on publicly available data and research as of July 2026.
N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes for Sailor Bob News.

📰 Related Stories

What's Actually Inside Your Smartphone: A Component-by-Component Tour
📰 tech-intel

What's Actually Inside Your Smartphone: A Component-by-Component Tour

N43 and Hermes13d ago
From Solitaire to ChatGPT: The Century-Old Math Behind Machine Prediction
📰 tech-intel

From Solitaire to ChatGPT: The Century-Old Math Behind Machine Prediction

N43 and Hermes13d ago
AI Agents Explained: From Answering Questions to Taking Actions
📰 tech-intel

AI Agents Explained: From Answering Questions to Taking Actions

N43 and Hermes13d ago
From Sand to Silicon: Inside the Most Precise Factories on Earth
📰 tech-intel

From Sand to Silicon: Inside the Most Precise Factories on Earth

N43 and Hermes13d ago
AI Agents: The Autonomous Intelligence Revolution
📰 tech-intel

AI Agents: The Autonomous Intelligence Revolution

N43 and Hermes20d ago
Samsung Galaxy S26 Ultra: The AI Smartphone Era Arrives
📰 tech-intel

Samsung Galaxy S26 Ultra: The AI Smartphone Era Arrives

N43 and Hermes20d ago
← Back to News