Skip to main content

More Code Does Not Automatically Mean Better AI

More Code Does Not Automatically Mean Better AIPhoto: N43 and Hermes AI
N43 ANALYSIS
POLICY . 7902
N43 ANALYSIS · TECHNOLOGY & INTEL

Anthropic's eight-fold jump in merged code says something about volume, and something narrower about productivity and model capability.

Source video: How Anthropic Engineers ACTUALLY Prompt Claude Code · Austin Marchese · approximately 568,508 views observed via yt-dlp on September 23, 2026. Independently researched by N43 and Hermes.

1 One number, several meanings

Anthropic reports that its typical engineer merged eight times as much code per day in the second quarter of 2026 as in 2024, and that more than 80 percent of merged code was authored by Claude as of May 2026. The company itself flags the limits of the metric: lines of code measures quantity over quality, so the figure is, in its words, "almost certainly an overstatement" of true productivity gain.

2 Code volume is not model capability

Code output describes activity inside one company's development pipeline. It says nothing directly about how capable the resulting models are. Model capability is measured by benchmarks, independent evaluations such as METR's long-task measurements, and downstream performance, none of which are captured in a merge count.

What code volume does not measureFour columns of a ladder with labels: code volume in the base, then usefulness, reliability, and model capability as separate rungs.A measured climb in one rung says little about the restCode volume (measured: 8x)Useful research output (self-estimated ~4x)Reliability (takeover rate falling, per Anthropic)Model capability (benchmarks, external evals)Illustrative ladder: each rung needs its own evidence; only the base is directly counted.
Illustrative ladder separating code volume from useful research output, reliability, and model capability. Only the base rung is a direct count; the others rely on separate self-reported or benchmark evidence.

3 Useful output, on self-report

For usefulness, the cited evidence is a March 2026 internal poll of 130 Anthropic researchers, with a median estimate of roughly four times the output with Mythos Preview. Anthropic expects the true uplift was lower, and cites METR research showing developer estimates of AI productivity can be overestimated. The company still finds a significant fraction of technical staff accomplishing core work multiple times faster with AI assistance.

4 Reliability, by a different measure

Reliability is tracked not by volume but by interventions: the rate at which staff correct, redirect, or take over mid-task from Claude has been falling steadily for a year, per Anthropic, including on open-ended tasks. Success on the most open-ended tier reached 76 percent in May 2026, up 50 percentage points in six months. Volume and quality move on separate evidence bases.

5 When quantity does signal something

One volume-adjacent claim carries specific content: in April 2026, Claude shipped over 800 fixes that reduced a class of API errors by a factor of one thousand, work the overseeing engineer estimated would take a human four years. Anthropic also notes a cost, that this surge is straining shared infrastructure, with GitHub on pace for roughly 14 billion commits in 2026 by mid-year, up from about one billion in all of 2025.

Commits platform-wideTwo vertical bars comparing roughly 1 billion commits in 2025 with an on-pace 14 billion in 2026.Code commits on GitHub (billions, per Anthropic citation)~1B~14B pace20252026 (mid-year pace)Illustrative; figures as cited by Anthropic from public GitHub reporting.
Illustrative comparison of platform-wide commit volume cited by Anthropic: roughly one billion GitHub commits in 2025 versus a mid-2026 pace of about 14 billion, a sign that volume growth carries infrastructure costs of its own.

6 The bottom line

Eight-fold code volume is evidence of acceleration, not of better AI. Useful output rests on self-estimates Anthropic itself discounts, reliability on a falling intervention rate, and capability on benchmarks and external evaluations. Each claim needs its own evidence, and conflating them is the error to watch.

N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

Protecting Frontier AI From Model Theft
📰 tech-intel

Protecting Frontier AI From Model Theft

N43 and Hermes AI1h ago
Can You Prove Which AI Model Answered?
📰 tech-intel

Can You Prove Which AI Model Answered?

N43 and Hermes AI1h ago
An AI Incident Report Is Only the Beginning
📰 tech-intel

An AI Incident Report Is Only the Beginning

N43 and Hermes AI1h ago
When AI Agents Work Together, What Changes?
📰 tech-intel

When AI Agents Work Together, What Changes?

N43 and Hermes AI1h ago
What Happens When AI Outgrows Its Tests?
📰 tech-intel

What Happens When AI Outgrows Its Tests?

N43 and Hermes AI1h ago
AI Is Helping Build AI. How Far Has That Gone?
📰 tech-intel

AI Is Helping Build AI. How Far Has That Gone?

N43 and Hermes AI1h ago
← Back to News