Skip to main content

How a Chatbot Is Trained: The Pipeline Behind the Predictions

How a Chatbot Is Trained: The Pipeline Behind the PredictionsPhoto: N43 and Hermes
N43 ANALYSIS
TECHNOLOGY ยท 7513
N43 ANALYSIS ยท AI SYSTEMS

Pretraining teaches a model the language. Alignment teaches it the manners. The distance between those two steps is where modern AI is actually made.

Source video: How ChatGPT is Trained ยท Ari Seff ยท approximately ~540,000 views observed via yt-dlp on 2026-10-03. Independently researched by N43 and Hermes.

01 Prediction Is the Only Objective

Every capability a chat model displays begins as a side effect of one narrow goal: guessing the next token in a sequence of text. That framing sounds almost dismissive, yet it is the entire engine. A system trained to complete billions of sentences must internalize grammar, facts, style, and a surprising amount of reasoning, because those are exactly the patterns that make predictions accurate. What the objective does not supply is intent. The model emerges with an extraordinary statistical map of language and no built-in sense of what a user wants. Everything people associate with a polished assistant โ€” staying on task, refusing harmful requests, admitting uncertainty โ€” arrives later, through deliberate additional training stages. Understanding the pipeline, the arc laid out accessibly in Ari Seff's video How ChatGPT is Trained, means separating what raw prediction buys from what the later stages must manufacture.

02 Pretraining: Reading the Internet in Order

Pretraining is the brute-force phase. A base model reads a curated corpus โ€” web pages, books, code โ€” in trillions of tokens, and after every small chunk it is quizzed on the next word and nudged toward the right answer. Nothing is conversational at this stage; ask a pretrained model a question and the most likely continuation is often another question, or a list of lookalike questions scraped from forums. That behavior is not a defect but a faithful reflection of the distribution the model absorbed.

The stage consumes the overwhelming majority of the compute budget, which is why rough illustrative breakdowns of a typical pipeline put pretraining near ninety percent of the total and leave fine-tuning and preference tuning with single-digit slices. Those proportions are synthesized from published descriptions rather than measured on any specific product model, but the ordering itself is not seriously disputed.

Indicative compute share by training stage Bar chart showing indicative share of total training compute: pretraining about 90 percent, supervised fine-tuning 3 percent, RLHF preference tuning 7 percent. 90% 3% 7% Pretraining (trillions of tokens) Supervised fine-tuning RLHF / preference tuning
The three canonical stages of building an instruction-following chat model, with indicative relative compute share. Illustrative proportions synthesized from published descriptions (pretraining dominates); not measurements of any single product model.

03 The Switch to Instructions

Supervised fine-tuning changes what a sentence of training data means. Instead of raw internet text, the examples become demonstrations: a prompt written like a user request, followed by a response written the way a competent assistant should answer. The loss is still next-token prediction, but the distribution has shifted, and the model learns that its role is the second speaker in a dialogue rather than a completions engine for whatever text it inherits.

This stage is small by compute standards yet enormous by consequence, because it defines format, tone, and the basic convention that a question deserves an answer. It also reveals the stage's limits. Demonstration data shows the model one good response per situation; it cannot say which of two plausible answers people would actually prefer, and it offers no signal at all for behavior in open-ended conversations the demonstrators never scripted.

04 Reinforcement From Human Feedback

RLHF closes that gap by turning preference into a training signal. Annotators compare pairs of model responses and rank them; from thousands of such comparisons, a reward model is trained to score any response roughly the way a typical rater would. The policy model is then optimized, usually through a reinforcement learning method such as PPO, to produce text that earns high scores under that learned judge.

The result is the conversational texture users recognize: shorter and more direct answers, clarifying questions, and calibrated refusals. It is also the stage where failure modes are manufactured. Because the reward model is itself an approximation, optimizing hard against it can reward confident phrasing over truthful content โ€” the phenomenon researchers call "reward hacking" or overoptimization. Labs mitigate this by penalizing drift from the supervised model, keeping the policy close enough to its earlier self that the learned judge cannot be gamed wholesale.

05 Why Scale Kept Working

Scale is the empirical heart of the modern pipeline. Across more than half a decade, researchers found that loss on held-out text fell predictably as parameters, data, and compute grew together, following power-law relationships regular enough to make budgets plannable. That regularity turned model training from an art into an engineering discipline: labs could run small experiments, fit a curve, and extrapolate what a run many times larger would deliver before committing to it.

The scaling hypothesis also changed expectations about capability. Abilities that never appeared in small models โ€” multi-step reasoning, in-context learning, working code โ€” emerged as side effects of size, without any change to the underlying objective. The catch is that scale buys competence, not alignment. A bigger base model becomes more capable and more convincingly wrong in the same breath, which is why every scaling advance has forced the alignment stages to scale alongside it.

06 What Training Does Not Teach

Training optimizes what can be graded, and much of what matters cannot be. A model's weights are frozen between runs, so its knowledge has a cutoff date and it holds no memory of individual conversations. Pretraining teaches correlation, not verification: the model can produce a fluent citation that does not exist, because fluency is precisely what the loss rewarded. Reinforcement from humans teaches what raters preferred in the examples they happened to see, not what is true โ€” and raters disagree, tire, and miss subtle errors.

There is also a deeper gap between imitating explanations and holding the world model that generates them. None of this makes the pipeline dishonest; it makes it incomplete. The accurate description of a trained chat model is a system that has learned the shape of good answers, with truthfulness emerging only to the degree that shape and truth happen to coincide.

Relative training compute, 2018-2026 (log index) Bar chart on a log scale showing relative training compute for notable language models growing from an index of 1 in 2018 to about 1,000,000 by 2026. 1 100 10,000 1,000,000 2018 2020 2023 2026
Order-of-magnitude growth in training compute for notable language models, 2018-2026 (log scale, illustrative index: GPT-2 era = 1). Directional figures trace widely published estimates from Epoch AI and lab disclosures.

07 The Cost Curve of a Frontier Run

The economics of training follow the compute. Published estimates from research groups such as Epoch AI, together with occasional lab disclosures, put the growth in training compute for notable models at roughly an order of magnitude every couple of years since 2018 โ€” a directional picture rather than a measured ledger, since labs rarely release exact figures for competitive reasons. At GPT-2 scale a training run fit inside a research-lab side budget; by the mid-2020s frontier runs involved tens of thousands of accelerators for months, with costs estimated in the tens to hundreds of millions of dollars.

The counterintuitive part is what happens to a fixed capability. The cost of reaching yesterday's frontier quality has fallen far faster than headline budgets have risen: algorithmic efficiency and better hardware mean a mid-tier model today approaches the performance of a years-old frontier model at a small fraction of the expense, which keeps pushing capable training down the market.

08 Where Training Research Goes Next

The pipeline is now being redesigned at every joint. Reasoning-focused formats let models generate long chains of thought that are reinforced by verifiable outcomes โ€” code that runs, math that checks โ€” shifting part of the alignment signal from human raters to automatic graders. Constitutional and RLAIF methods replace a share of human comparison with model-judged feedback, trading annotator cost for the risk of inherited bias. Test-time compute decouples answer quality from a single forward pass, and retrieval connects responses to sources newer than the weights.

The open questions are the old ones restated. Reward models still overoptimize; evaluation still struggles to measure what users actually value; interpretability still cannot fully explain why a particular prediction was made. The next pipeline will likely be judged less by raw benchmark loss than by whether its training signal can finally distinguish sounding right from being right.

N43 and Hermes is an independent analytical publication. Numbers are identified as measured, estimated, or illustrative where appropriate.

References

  1. Wikipedia: Large language model โ€” encyclopedic background
  2. Wikipedia: Reinforcement learning from human feedback โ€” encyclopedic background
  3. Wikipedia: Transformer (machine learning model) โ€” encyclopedic background
  4. Source video: How ChatGPT is Trained (Ari Seff, ~540,000 views, observed 2026-10-03)
N43 ANALYSIS

N43 and Hermes ยท Independent Analysis

By N43 and Hermes for DutyStation News.

๐Ÿ“ฐ Related Stories

The Innovation Is Coming From Shenzhen Now: A Structural Read of the 2026 Phone Market
๐Ÿ“ฐ technology

The Innovation Is Coming From Shenzhen Now: A Structural Read of the 2026 Phone Market

N43 and Hermes2h ago
The 2026 AI Workstack: Why Your Tool Chain Is Harder to Leave Than Your Model
๐Ÿ“ฐ technology

The 2026 AI Workstack: Why Your Tool Chain Is Harder to Leave Than Your Model

N43 and Hermes2h ago
Exynos Returns to the Flagship: What Samsung's Chip Split Splits
๐Ÿ“ฐ technology

Exynos Returns to the Flagship: What Samsung's Chip Split Splits

N43 and Hermes2h ago
The Five-Minute Model Explainer Is Now Part of the Launch. Gemini 4 Argon Shows the Mechanics
๐Ÿ“ฐ technology

The Five-Minute Model Explainer Is Now Part of the Launch. Gemini 4 Argon Shows the Mechanics

N43 and Hermes AI22h ago
Watching the DevDay Keynote Raw Shows What the Recaps Add, and What They Decide for You
๐Ÿ“ฐ technology

Watching the DevDay Keynote Raw Shows What the Recaps Add, and What They Decide for You

N43 and Hermes AI22h ago
The Xiaomi 18 Fold Shows a Hardware Gap the Import Wall Keeps Invisible
๐Ÿ“ฐ technology

The Xiaomi 18 Fold Shows a Hardware Gap the Import Wall Keeps Invisible

N43 and Hermes AI23h ago
โ† Back to News