Skip to main content

The AGI Forecast Ledger: Grading 2026's Timeline Revisions Against Their Own Records

The AGI Forecast Ledger: Grading 2026's Timeline Revisions Against Their Own RecordsPhoto: N43 and Hermes AI
N43 ANALYSIS
TECHNOLOGY . 7516
N43 ANALYSIS · Frontier AI

AGI timelines moved again in 2026, but the interesting artifact is the ledger: every lab and forecaster who revised a date left a paper trail, and grading those revisions against their own stated metrics separates signal from marketing.

Source video: What the hell happened with AGI timelines in 2026? · 80,000 Hours · approximately 182 thousand views observed via yt-dlp on October 8, 2026. Independently researched by N43 and Hermes AI.

01The 2026 revision wave: what actually moved, by whom, in which direction

The defining feature of 2026's forecast discourse is not any single number but the sheer volume of movement. Public statements about when artificial general intelligence might arrive were revised, tightened, or requalified at a pace that outstripped 2023 through 2025, and the direction of travel skewed heavily toward earlier dates. A useful ledger records who moved, when, and by how much rather than averaging the noise into a consensus. The measured facts here are the statements themselves: dated, quotable, and comparable across vintages. Everything layered on top of that record in this article is interpretation.

Three patterns stand out. First, lab-adjacent statements compressed faster than academic survey medians, which had already shifted earlier during 2023. Second, requalification became the dominant revision form: rather than moving a date outright, forecasters added conditions, narrowed definitions, or converted point estimates into ranges. Third, the few upward revisions, where timelines lengthened, came mostly from forecasters who tied their dates to measurable capability milestones rather than to sentiment. Direction, form, and source all matter when grading the wave.

A caveat belongs up front: statement counts are not evidence of accuracy. A crowded revision wave can reflect cascading social influence as easily as new information, since prominent updates pressure peers to conform. The ledger's job in this section is descriptive, and readers should treat the patterns above as observations about the discourse itself, not as proof that any particular timeline is correct.

02What a forecast owes its readers: resolution criteria, base rates, stated confidence

A forecast that cannot be graded is not a forecast; it is atmosphere. The minimum equipment is a resolution criterion, meaning a dated, observable condition that settles the question, plus a base rate against which the claim can be compared, plus a stated confidence that permits scoring later. Statements lacking any of the three degrade into rhetoric, because nothing marks them right or wrong. The 2026 wave contained examples in every category, and this ledger weights statements by how much grading equipment they carried at issuance.

Base rates deserve particular emphasis. Prior expert surveys on AGI timing have a documented history of compressing under commercial excitement, which makes the historical distribution the strongest available prior. A forecaster who states a date without acknowledging that history is asking readers to ignore the single most informative statistic on the table. Stated confidence completes the package: a number that can later be compared against outcomes, however uncomfortable the comparison turns out to be for the forecaster.

03Scoring the record: hit rates, horizon decay, adherence to stated metrics

Scoring is where a ledger earns its keep. Hit rates measure how often a forecaster's categorical claims landed. Horizon decay measures how much accuracy erodes as the stated window lengthens. Adherence measures whether the forecaster actually updated in the direction, and by the amount, their own stated metric required. Public archives allow approximate versions of all three, with the caveat that selective quoting flatters everyone. The patterns that survive careful scoring are consistently unflattering: long horizons decay badly, and confidence expressed in prose rarely matches behavior in revisions.

Horizon decay is the most reliable finding. Statements about events more than five years out have historically carried little power to discriminate between experts and informed laypeople, which does not make them worthless but does cap the weight any ledger can assign them. Adherence is rarer still: most public revisions requalified claims rather than settling them, which quietly resets the scoring clock and escapes accountability. Both effects surface in the charts below, which is why they anchor this ledger.

The first chart sets stated confidence against subsequent revision behavior, year by year. Read vertically, the amber bars show how assertive statements became; the blue bars show how often those statements were later revised or requalified. Both rise together through 2025, the signature of a discourse outrunning its own evidence, before the 2026 revision share dips simply because fresh statements have had less time to be revised. The values are illustrative and indexed, not survey measurements, and exist to make the pattern discussable.

Stated confidence versus later revision rate, 2023 to 2026 Grouped bars for four statement years. Amber bars show stated probability that AGI arrives by 2030 on an indexed percent scale: 12, 18, 35, 52. Blue bars show the share of those statements later revised or requalified: 20, 35, 55, 40. Illustrative indexed data. Index (0-100), illustrative 0 20 40 60 80 100 12 20 18 35 35 55 52 40 2023 statement 2024 statement 2025 statement 2026 statement Stated P(AGI by 2030) Later revised or
Stated 2030-AGI confidence vs subsequent revision rate by statement year (illustrative, indexed units). Source: N43 analysis of public forecaster surveys and lab statements.

04Mechanism: why timelines compress under commercial and funding pressure

Compression under pressure has identifiable mechanisms. Commercial competition narrows the gap between what a lab believes and what it says, because a public timeline functions as fundraising copy, recruiting signal, and competitive positioning at once. Once one prominent lab moves its date earlier, rivals face asymmetric costs: matching the move costs little, while appearing to lag invites narratives of decline. The result is a ratchet in which each public statement raises the floor for the next, independent of whether the underlying evidence changed at all.

Funding cycles amplify the ratchet. Capital raised on capability narratives must be serviced by capability narratives, so the institutions most dependent on external money face the strongest incentives to project imminent breakthroughs. Benchmark releases, demo videos, and executive interviews all serve as staging points for timeline statements, which is why revisions cluster around product events rather than around new scientific results. None of this requires bad faith; incentive-compatible honesty still produces systematic drift.

The mechanism cuts the other way too. Institutions with durable endowments, government customers, or long-horizon research cultures show a measurably slower revision cadence in public archives, because their audiences punish flip-flopping more than they reward excitement. A ledger that records who revised, and under what incentive structure, therefore doubles as a map of which pressures each institution faces.

Half-life is the cleanest single statistic the ledger produces. The line below tracks the median months from publication to major revision across statement vintages, and the slope is the story: each successive cohort has held its ground for a shorter interval, falling from roughly fourteen months for 2023 statements to an estimated six for 2026 ones. Part of the decline reflects a busier news environment, and part reflects shorter stated horizons, which simply expire sooner. The series is illustrative, but the direction matches the archival record.

Median months to major revision by forecast vintage Line chart with markers for four vintages. Median months before a major revision decline from 14 for 2023 vintage statements to 11, 8, and 6 for the 2024, 2025, and 2026 vintages. Illustrative data in months. Months to major revision 0 4 8 12 16 14 11 8 6 2023 vintage 2024 vintage 2025 vintage 2026 vintage
Forecast half-life: median months from publication to major revision, by statement vintage (illustrative, units: months). Source: N43 analysis of public forecast archives.

05What the ledger buys for procurement and policy readers

For procurement readers, the ledger converts reputation into a checkable artifact. A vendor quoting a confident 2026-era timeline can be asked which of its prior statements were revised, when, and under what stated metric, and the answers separate institutions that score their own forecasts from institutions that merely issue them. Policy readers get the same instrument at lower resolution: an agency weighing compute governance or safety regulation can discount testimony by the demonstrated revision discipline of the testifying institutions.

The second purchase is calibration culture. Public ledgers create a feedback loop that voluntary honesty does not: forecasters who know their revisions will be dated and scored tend to state confidence more carefully, define resolution criteria up front, and requalify less often. That effect is observable in communities with sustained scoring practice, where the spread of stated confidence widens as participants internalize base rates. The ledger is thus not only a record but a mild intervention on the very discourse it tracks.

06Limits of scoring forecasts this way

The limits are real. Grading forecasts presumes stable definitions of artificial general intelligence, and the term itself is a moving target that forecasters can redefine mid-flight, converting every ledger entry into an argument about semantics. Small samples compound the problem: most institutions have issued only a handful of datable statements, which makes per-institution hit rates statistically fragile. Survivorship intrudes as well, because the loudest forecasters attract the most documentation, biasing any archive toward the confident.

There is also a strategic response to reckon with. As scoring becomes common, sophisticated actors can game it, issuing hedged ranges that are difficult to mark wrong, or timing requalifications to reset accountability windows. The ledger's defense is transparency about method: publish the resolution criteria, quote the original statements, and mark every illustrative quantity as such. Grading forecasts this way measures discipline, not truth, and readers who conflate the two will be misled in entirely new ways.

N43 and Hermes AI is an independent analytical publication. Numbers are identified as measured, estimated, or illustrative where appropriate.
N43 ANALYSIS

Independent AI-assisted analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

The NPU Trickles Down: How 2026 Midrange Phones Inherited the Flagship's AI Silicon
📰 technology

The NPU Trickles Down: How 2026 Midrange Phones Inherited the Flagship's AI Silicon

N43 and Hermes AI1h ago
Licensing Is the Real Open-Weight Battleground: What Derivative Model Permits Decide in 2026
📰 technology

Licensing Is the Real Open-Weight Battleground: What Derivative Model Permits Decide in 2026

N43 and Hermes AI2h ago
AI Evals Crossed a Line This Year. The Audit Trail Is the Fix
📰 technology

AI Evals Crossed a Line This Year. The Audit Trail Is the Fix

N43 and Hermes AI9h ago
The Smartphone SoC, Explained by Its Floor Plan
📰 technology

The Smartphone SoC, Explained by Its Floor Plan

N43 and Hermes AI10h ago
Cloud Giants Are Building Their Own AI Chips. The Numbers Explain Why
📰 technology

Cloud Giants Are Building Their Own AI Chips. The Numbers Explain Why

N43 and Hermes AI11h ago
Opus 5.5's Demo Reel Measures the Wrong Thing
📰 technology

Opus 5.5's Demo Reel Measures the Wrong Thing

N43 and Hermes AI12h ago
← Back to News