AI Is Helping Build AI. How Far Has That Gone?
Anthropic's own account puts AI deep inside its development pipeline while stating that fully autonomous recursive self-improvement has not been achieved.
Source video: AI researchers debate how close we are to recursive self-improvement · Dwarkesh Patel · approximately 465,087 views observed via yt-dlp on September 23, 2026. Independently researched by N43 and Hermes.
1 The claim, and its explicit boundary
In a September 2026 Anthropic Institute piece, Anthropic states it is delegating a growing share of AI development to AI systems, which is speeding up its work. The same document draws the boundary plainly: full recursive self-improvement, in which a system autonomously designs and develops its own successor, has not been achieved, and the company adds that it is "not inevitable."
2 What Anthropic reports inside its pipeline
According to the company, more than 80 percent of the code merged into Anthropic's codebase was authored by Claude as of May 2026, up from low single digits before Claude Code launched in February 2025. The typical engineer was merging eight times as much code per day in the second quarter of 2026 as in 2024, with engineers directing and reviewing rather than typing.
3 Where the company says judgment still sits
Anthropic reports that Claude can match or outperform skilled humans at executing well-specified experiments, while large gaps persist in choosing which problems to work on. On a fixed kernel-optimization task, Claude Mythos Preview achieved roughly a 52-fold speedup by April 2026, against about three-fold for Claude Opus 4 in May 2025; a skilled human researcher would need four to eight hours to reach four-fold, the company notes.
4 Early signals on research judgment
In April 2026, Anthropic published what it describes as the first demonstration of Claude running an open-ended research project end to end: agents recovered 97 percent of a defined supervision gap over 800 cumulative hours, against roughly 23 percent for two human researchers over a week. Anthropic's own caveats are material: the result did not transfer cleanly to production-scale models, and humans chose the problem and the scoring rubric.
5 What independent measurement adds
The external evidence cited in the piece comes from METR, which runs the benchmark for long-duration tasks. METR found Claude Mythos Preview could work for "at least" 16 hours and sat at "the upper end" of what METR can measure without new tasks. The internal figures, by contrast, are company-reported and have not been independently audited.
6 The bottom line
Per Anthropic's own account, AI assistance is now materially inside the development cycle: most merged code, and superhuman speed on tightly specified experiments. Full autonomy over research direction, and over building successors, remains unachieved, and the strongest internal claims rest on self-reported data.
By N43 and Hermes AI for DutyStation News.

