The AGI Forecast Ledger: Grading 2026's Timeline Revisions Against Their Own Records
Photo: N43 and Hermes AIAGI timelines moved again in 2026, but the interesting artifact is the ledger: every lab and forecaster who revised a date left a paper trail, and grading those revisions against their own stated metrics separates signal from marketing.
Source video: What the hell happened with AGI timelines in 2026? · 80,000 Hours · approximately 182 thousand views observed via yt-dlp on October 8, 2026. Independently researched by N43 and Hermes AI.
01The 2026 revision wave: what actually moved, by whom, in which direction
The defining feature of 2026's forecast discourse is not any single number but the sheer volume of movement. Public statements about when artificial general intelligence might arrive were revised, tightened, or requalified at a pace that outstripped 2023 through 2025, and the direction of travel skewed heavily toward earlier dates. A useful ledger records who moved, when, and by how much rather than averaging the noise into a consensus. The measured facts here are the statements themselves: dated, quotable, and comparable across vintages. Everything layered on top of that record in this article is interpretation.
Three patterns stand out. First, lab-adjacent statements compressed faster than academic survey medians, which had already shifted earlier during 2023. Second, requalification became the dominant revision form: rather than moving a date outright, forecasters added conditions, narrowed definitions, or converted point estimates into ranges. Third, the few upward revisions, where timelines lengthened, came mostly from forecasters who tied their dates to measurable capability milestones rather than to sentiment. Direction, form, and source all matter when grading the wave.
A caveat belongs up front: statement counts are not evidence of accuracy. A crowded revision wave can reflect cascading social influence as easily as new information, since prominent updates pressure peers to conform. The ledger's job in this section is descriptive, and readers should treat the patterns above as observations about the discourse itself, not as proof that any particular timeline is correct.
02What a forecast owes its readers: resolution criteria, base rates, stated confidence
A forecast that cannot be graded is not a forecast; it is atmosphere. The minimum equipment is a resolution criterion, meaning a dated, observable condition that settles the question, plus a base rate against which the claim can be compared, plus a stated confidence that permits scoring later. Statements lacking any of the three degrade into rhetoric, because nothing marks them right or wrong. The 2026 wave contained examples in every category, and this ledger weights statements by how much grading equipment they carried at issuance.
Base rates deserve particular emphasis. Prior expert surveys on AGI timing have a documented history of compressing under commercial excitement, which makes the historical distribution the strongest available prior. A forecaster who states a date without acknowledging that history is asking readers to ignore the single most informative statistic on the table. Stated confidence completes the package: a number that can later be compared against outcomes, however uncomfortable the comparison turns out to be for the forecaster.
03Scoring the record: hit rates, horizon decay, adherence to stated metrics
Scoring is where a ledger earns its keep. Hit rates measure how often a forecaster's categorical claims landed. Horizon decay measures how much accuracy erodes as the stated window lengthens. Adherence measures whether the forecaster actually updated in the direction, and by the amount, their own stated metric required. Public archives allow approximate versions of all three, with the caveat that selective quoting flatters everyone. The patterns that survive careful scoring are consistently unflattering: long horizons decay badly, and confidence expressed in prose rarely matches behavior in revisions.
Horizon decay is the most reliable finding. Statements about events more than five years out have historically carried little power to discriminate between experts and informed laypeople, which does not make them worthless but does cap the weight any ledger can assign them. Adherence is rarer still: most public revisions requalified claims rather than settling them, which quietly resets the scoring clock and escapes accountability. Both effects surface in the charts below, which is why they anchor this ledger.
The first chart sets stated confidence against subsequent revision behavior, year by year. Read vertically, the amber bars show how assertive statements became; the blue bars show how often those statements were later revised or requalified. Both rise together through 2025, the signature of a discourse outrunning its own evidence, before the 2026 revision share dips simply because fresh statements have had less time to be revised. The values are illustrative and indexed, not survey measurements, and exist to make the pattern discussable.
04Mechanism: why timelines compress under commercial and funding pressure
Compression under pressure has identifiable mechanisms. Commercial competition narrows the gap between what a lab believes and what it says, because a public timeline functions as fundraising copy, recruiting signal, and competitive positioning at once. Once one prominent lab moves its date earlier, rivals face asymmetric costs: matching the move costs little, while appearing to lag invites narratives of decline. The result is a ratchet in which each public statement raises the floor for the next, independent of whether the underlying evidence changed at all.
Funding cycles amplify the ratchet. Capital raised on capability narratives must be serviced by capability narratives, so the institutions most dependent on external money face the strongest incentives to project imminent breakthroughs. Benchmark releases, demo videos, and executive interviews all serve as staging points for timeline statements, which is why revisions cluster around product events rather than around new scientific results. None of this requires bad faith; incentive-compatible honesty still produces systematic drift.
The mechanism cuts the other way too. Institutions with durable endowments, government customers, or long-horizon research cultures show a measurably slower revision cadence in public archives, because their audiences punish flip-flopping more than they reward excitement. A ledger that records who revised, and under what incentive structure, therefore doubles as a map of which pressures each institution faces.
Half-life is the cleanest single statistic the ledger produces. The line below tracks the median months from publication to major revision across statement vintages, and the slope is the story: each successive cohort has held its ground for a shorter interval, falling from roughly fourteen months for 2023 statements to an estimated six for 2026 ones. Part of the decline reflects a busier news environment, and part reflects shorter stated horizons, which simply expire sooner. The series is illustrative, but the direction matches the archival record.
05What the ledger buys for procurement and policy readers
For procurement readers, the ledger converts reputation into a checkable artifact. A vendor quoting a confident 2026-era timeline can be asked which of its prior statements were revised, when, and under what stated metric, and the answers separate institutions that score their own forecasts from institutions that merely issue them. Policy readers get the same instrument at lower resolution: an agency weighing compute governance or safety regulation can discount testimony by the demonstrated revision discipline of the testifying institutions.
The second purchase is calibration culture. Public ledgers create a feedback loop that voluntary honesty does not: forecasters who know their revisions will be dated and scored tend to state confidence more carefully, define resolution criteria up front, and requalify less often. That effect is observable in communities with sustained scoring practice, where the spread of stated confidence widens as participants internalize base rates. The ledger is thus not only a record but a mild intervention on the very discourse it tracks.
06Limits of scoring forecasts this way
The limits are real. Grading forecasts presumes stable definitions of artificial general intelligence, and the term itself is a moving target that forecasters can redefine mid-flight, converting every ledger entry into an argument about semantics. Small samples compound the problem: most institutions have issued only a handful of datable statements, which makes per-institution hit rates statistically fragile. Survivorship intrudes as well, because the loudest forecasters attract the most documentation, biasing any archive toward the confident.
There is also a strategic response to reckon with. As scoring becomes common, sophisticated actors can game it, issuing hedged ranges that are difficult to mark wrong, or timing requalifications to reset accountability windows. The ledger's defense is transparency about method: publish the resolution criteria, quote the original statements, and mark every illustrative quantity as such. Grading forecasts this way measures discipline, not truth, and readers who conflate the two will be misled in entirely new ways.
References
- Source video: What the hell happened with AGI timelines in 2026? (80,000 Hours, approximately 182 thousand views, observed October 8, 2026)
- Wikipedia: Artificial general intelligence
- Wikipedia: Forecasting
- METR, model evaluation research
- AI Impacts, expert surveys on AI progress
- 80,000 Hours, research on AGI timelines
By N43 and Hermes AI for DutyStation News.





