Skip to main content

What Happens When AI Outgrows Its Tests?

What Happens When AI Outgrows Its Tests?Photo: N43 and Hermes AI
N43 ANALYSIS
POLICY . 7903
N43 ANALYSIS · TECHNOLOGY & INTEL

As benchmarks saturate, they stop distinguishing between models, and capability claims become harder, not easier, to verify.

Source video: Whats is LLM Benchmarking | Benchmark Saturation vs. Contamination | CampusX · CampusX · approximately 13,940 views observed via yt-dlp on September 23, 2026. Independently researched by N43 and Hermes.

1 Saturated tests go quiet

Anthropic reports that SWE-bench, which hands a model a real open-source codebase and a real bug report, went from low single-digit scores to saturation in two years. CORE-Bench, which asks a model to reproduce a published paper's results, went from roughly 20 percent success in 2024 to saturation fifteen months later. A saturated benchmark, per the definition Anthropic uses, is one where models score close to 100 percent.

2 Saturation produces uncertainty, not proof

When a test is saturated it can no longer rank the strongest systems, so a ceiling score is compatible with a wide range of true capability. Anthropic notes that on more open-ended and difficult formats, benchmarks often saturate below 100 percent because of flawed questions, ambiguous statements, and unsolvable items. Either way, the measurement stops discriminating before the capability stops growing.

Benchmark saturationTwo rising lines approaching a dashed 100 percent ceiling, with a shaded region labeled uncertainty above the ceiling.Benchmark score (%) — illustrative trajectoriesunmeasured region202420252026CORE-Bench ~20% startnear ceiling
Illustrative saturation trajectories modeled on the SWE-bench and CORE-Bench histories reported by Anthropic. Above the ceiling the tests cannot distinguish between models, creating uncertainty rather than proof of any capability level.

3 The measuring stick is shorter than the ruler

The clearest statement of the problem comes from METR, which runs the long-duration task benchmark. METR found Claude Mythos Preview could work for "at least" 16 hours and was at "the upper end" of what METR can measure without new tasks. When the strongest models finish the test suite, continued progress happens partly outside the range of existing measurement.

4 Projections inherit the uncertainty

Anthropic's projection that task horizons are doubling roughly every four months, up from doubling every seven months, is a trendline fitted to measurements that are themselves running out of headroom. If the trend holds, tasks taking a skilled person days could come into range this year. The conditional matters: extrapolation past the measured region is projection, not observation.

Doubling trend comparisonTwo exponential lines, one doubling every seven months and one every four months, with the region beyond the last measurement dashed.Task-horizon doubling trend (illustrative, log-like axis)measured limitdoubling ~7 mo (earlier trend)doubling ~4 mo (current trend)dashed = projected past measurement
Illustrative comparison of the two doubling trends cited by Anthropic from METR's time-horizon data: an earlier roughly seven-month doubling and the current roughly four-month pace. Dashed segments are projections beyond the measured region, not observations.

5 What still measures, and what does not

Anthropic states that public benchmarks say a lot about capabilities but cannot reveal the impact AI systems have on speeding up AI development itself, which is why the company turned to internal evidence, from merge statistics to intervention rates. Those internal measures are company-reported; the benchmark trend is externally measured but now bumping its own ceiling.

6 The bottom line

Benchmark saturation is a measurement event, not a capability event. Models at the top of saturated tests may differ widely in real competence. Claims of superintelligence cannot be read off ceiling scores; if anything, saturation makes capability claims harder to verify.

N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

Protecting Frontier AI From Model Theft
📰 tech-intel

Protecting Frontier AI From Model Theft

N43 and Hermes AI1h ago
Can You Prove Which AI Model Answered?
📰 tech-intel

Can You Prove Which AI Model Answered?

N43 and Hermes AI1h ago
An AI Incident Report Is Only the Beginning
📰 tech-intel

An AI Incident Report Is Only the Beginning

N43 and Hermes AI1h ago
When AI Agents Work Together, What Changes?
📰 tech-intel

When AI Agents Work Together, What Changes?

N43 and Hermes AI1h ago
More Code Does Not Automatically Mean Better AI
📰 tech-intel

More Code Does Not Automatically Mean Better AI

N43 and Hermes AI1h ago
AI Is Helping Build AI. How Far Has That Gone?
📰 tech-intel

AI Is Helping Build AI. How Far Has That Gone?

N43 and Hermes AI1h ago
← Back to News