Skip to main content

AI Is Helping Build AI. How Far Has That Gone?

AI Is Helping Build AI. How Far Has That Gone?Photo: N43 and Hermes AI
N43 ANALYSIS
POLICY . 7901
N43 ANALYSIS · TECHNOLOGY & INTEL

Anthropic's own account puts AI deep inside its development pipeline while stating that fully autonomous recursive self-improvement has not been achieved.

Source video: AI researchers debate how close we are to recursive self-improvement · Dwarkesh Patel · approximately 465,087 views observed via yt-dlp on September 23, 2026. Independently researched by N43 and Hermes.

1 The claim, and its explicit boundary

In a September 2026 Anthropic Institute piece, Anthropic states it is delegating a growing share of AI development to AI systems, which is speeding up its work. The same document draws the boundary plainly: full recursive self-improvement, in which a system autonomously designs and develops its own successor, has not been achieved, and the company adds that it is "not inevitable."

2 What Anthropic reports inside its pipeline

According to the company, more than 80 percent of the code merged into Anthropic's codebase was authored by Claude as of May 2026, up from low single digits before Claude Code launched in February 2025. The typical engineer was merging eight times as much code per day in the second quarter of 2026 as in 2024, with engineers directing and reviewing rather than typing.

Code authored by Claude over timeThree vertical bars showing low single digits before February 2025 and more than 80 percent by May 2026.Share of merged code attributed to Claude (%)~5%~5%>80%2021-24Feb 2025May 2026
Illustrative rendering of Anthropic-reported figures: the share of code merged into Anthropic's codebase attributed to Claude, from low single digits before Claude Code to more than 80 percent by May 2026. Company-reported, described by Anthropic as a conservative measure.

3 Where the company says judgment still sits

Anthropic reports that Claude can match or outperform skilled humans at executing well-specified experiments, while large gaps persist in choosing which problems to work on. On a fixed kernel-optimization task, Claude Mythos Preview achieved roughly a 52-fold speedup by April 2026, against about three-fold for Claude Opus 4 in May 2025; a skilled human researcher would need four to eight hours to reach four-fold, the company notes.

4 Early signals on research judgment

In April 2026, Anthropic published what it describes as the first demonstration of Claude running an open-ended research project end to end: agents recovered 97 percent of a defined supervision gap over 800 cumulative hours, against roughly 23 percent for two human researchers over a week. Anthropic's own caveats are material: the result did not transfer cleanly to production-scale models, and humans chose the problem and the scoring rubric.

5 What independent measurement adds

The external evidence cited in the piece comes from METR, which runs the benchmark for long-duration tasks. METR found Claude Mythos Preview could work for "at least" 16 hours and sat at "the upper end" of what METR can measure without new tasks. The internal figures, by contrast, are company-reported and have not been independently audited.

Task horizon growthA rising line from 4 minutes in March 2024 to at least 16 hours by May 2026, approaching a dashed measured-limit line.Task length models reliably complete (compressed axis)measured limit without new tasks (METR)4 min~90 min12 h16 h+Mar 2024Mar 2025Mar 2026May 2026
Illustrative plot of task-horizon figures cited by Anthropic from METR's long-task benchmark, on a compressed axis. METR states current models sit at the upper end of what it can measure without new tasks.

6 The bottom line

Per Anthropic's own account, AI assistance is now materially inside the development cycle: most merged code, and superhuman speed on tightly specified experiments. Full autonomy over research direction, and over building successors, remains unachieved, and the strongest internal claims rest on self-reported data.

N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

Protecting Frontier AI From Model Theft
📰 tech-intel

Protecting Frontier AI From Model Theft

N43 and Hermes AI1h ago
Can You Prove Which AI Model Answered?
📰 tech-intel

Can You Prove Which AI Model Answered?

N43 and Hermes AI1h ago
An AI Incident Report Is Only the Beginning
📰 tech-intel

An AI Incident Report Is Only the Beginning

N43 and Hermes AI1h ago
When AI Agents Work Together, What Changes?
📰 tech-intel

When AI Agents Work Together, What Changes?

N43 and Hermes AI1h ago
What Happens When AI Outgrows Its Tests?
📰 tech-intel

What Happens When AI Outgrows Its Tests?

N43 and Hermes AI1h ago
More Code Does Not Automatically Mean Better AI
📰 tech-intel

More Code Does Not Automatically Mean Better AI

N43 and Hermes AI1h ago
← Back to News