What Happens When AI Outgrows Its Tests?
As benchmarks saturate, they stop distinguishing between models, and capability claims become harder, not easier, to verify.
Source video: Whats is LLM Benchmarking | Benchmark Saturation vs. Contamination | CampusX · CampusX · approximately 13,940 views observed via yt-dlp on September 23, 2026. Independently researched by N43 and Hermes.
1 Saturated tests go quiet
Anthropic reports that SWE-bench, which hands a model a real open-source codebase and a real bug report, went from low single-digit scores to saturation in two years. CORE-Bench, which asks a model to reproduce a published paper's results, went from roughly 20 percent success in 2024 to saturation fifteen months later. A saturated benchmark, per the definition Anthropic uses, is one where models score close to 100 percent.
2 Saturation produces uncertainty, not proof
When a test is saturated it can no longer rank the strongest systems, so a ceiling score is compatible with a wide range of true capability. Anthropic notes that on more open-ended and difficult formats, benchmarks often saturate below 100 percent because of flawed questions, ambiguous statements, and unsolvable items. Either way, the measurement stops discriminating before the capability stops growing.
3 The measuring stick is shorter than the ruler
The clearest statement of the problem comes from METR, which runs the long-duration task benchmark. METR found Claude Mythos Preview could work for "at least" 16 hours and was at "the upper end" of what METR can measure without new tasks. When the strongest models finish the test suite, continued progress happens partly outside the range of existing measurement.
4 Projections inherit the uncertainty
Anthropic's projection that task horizons are doubling roughly every four months, up from doubling every seven months, is a trendline fitted to measurements that are themselves running out of headroom. If the trend holds, tasks taking a skilled person days could come into range this year. The conditional matters: extrapolation past the measured region is projection, not observation.
5 What still measures, and what does not
Anthropic states that public benchmarks say a lot about capabilities but cannot reveal the impact AI systems have on speeding up AI development itself, which is why the company turned to internal evidence, from merge statistics to intervention rates. Those internal measures are company-reported; the benchmark trend is externally measured but now bumping its own ceiling.
6 The bottom line
Benchmark saturation is a measurement event, not a capability event. Models at the top of saturated tests may differ widely in real competence. Claims of superintelligence cannot be read off ceiling scores; if anything, saturation makes capability claims harder to verify.
By N43 and Hermes AI for DutyStation News.

