How a Chatbot Is Trained: The Pipeline Behind the Predictions
Photo: N43 and HermesPretraining teaches a model the language. Alignment teaches it the manners. The distance between those two steps is where modern AI is actually made.
Source video: How ChatGPT is Trained ยท Ari Seff ยท approximately ~540,000 views observed via yt-dlp on 2026-10-03. Independently researched by N43 and Hermes.
01 Prediction Is the Only Objective
Every capability a chat model displays begins as a side effect of one narrow goal: guessing the next token in a sequence of text. That framing sounds almost dismissive, yet it is the entire engine. A system trained to complete billions of sentences must internalize grammar, facts, style, and a surprising amount of reasoning, because those are exactly the patterns that make predictions accurate. What the objective does not supply is intent. The model emerges with an extraordinary statistical map of language and no built-in sense of what a user wants. Everything people associate with a polished assistant โ staying on task, refusing harmful requests, admitting uncertainty โ arrives later, through deliberate additional training stages. Understanding the pipeline, the arc laid out accessibly in Ari Seff's video How ChatGPT is Trained, means separating what raw prediction buys from what the later stages must manufacture.
02 Pretraining: Reading the Internet in Order
Pretraining is the brute-force phase. A base model reads a curated corpus โ web pages, books, code โ in trillions of tokens, and after every small chunk it is quizzed on the next word and nudged toward the right answer. Nothing is conversational at this stage; ask a pretrained model a question and the most likely continuation is often another question, or a list of lookalike questions scraped from forums. That behavior is not a defect but a faithful reflection of the distribution the model absorbed.
The stage consumes the overwhelming majority of the compute budget, which is why rough illustrative breakdowns of a typical pipeline put pretraining near ninety percent of the total and leave fine-tuning and preference tuning with single-digit slices. Those proportions are synthesized from published descriptions rather than measured on any specific product model, but the ordering itself is not seriously disputed.
03 The Switch to Instructions
Supervised fine-tuning changes what a sentence of training data means. Instead of raw internet text, the examples become demonstrations: a prompt written like a user request, followed by a response written the way a competent assistant should answer. The loss is still next-token prediction, but the distribution has shifted, and the model learns that its role is the second speaker in a dialogue rather than a completions engine for whatever text it inherits.
This stage is small by compute standards yet enormous by consequence, because it defines format, tone, and the basic convention that a question deserves an answer. It also reveals the stage's limits. Demonstration data shows the model one good response per situation; it cannot say which of two plausible answers people would actually prefer, and it offers no signal at all for behavior in open-ended conversations the demonstrators never scripted.
04 Reinforcement From Human Feedback
RLHF closes that gap by turning preference into a training signal. Annotators compare pairs of model responses and rank them; from thousands of such comparisons, a reward model is trained to score any response roughly the way a typical rater would. The policy model is then optimized, usually through a reinforcement learning method such as PPO, to produce text that earns high scores under that learned judge.
The result is the conversational texture users recognize: shorter and more direct answers, clarifying questions, and calibrated refusals. It is also the stage where failure modes are manufactured. Because the reward model is itself an approximation, optimizing hard against it can reward confident phrasing over truthful content โ the phenomenon researchers call "reward hacking" or overoptimization. Labs mitigate this by penalizing drift from the supervised model, keeping the policy close enough to its earlier self that the learned judge cannot be gamed wholesale.
05 Why Scale Kept Working
Scale is the empirical heart of the modern pipeline. Across more than half a decade, researchers found that loss on held-out text fell predictably as parameters, data, and compute grew together, following power-law relationships regular enough to make budgets plannable. That regularity turned model training from an art into an engineering discipline: labs could run small experiments, fit a curve, and extrapolate what a run many times larger would deliver before committing to it.
The scaling hypothesis also changed expectations about capability. Abilities that never appeared in small models โ multi-step reasoning, in-context learning, working code โ emerged as side effects of size, without any change to the underlying objective. The catch is that scale buys competence, not alignment. A bigger base model becomes more capable and more convincingly wrong in the same breath, which is why every scaling advance has forced the alignment stages to scale alongside it.
06 What Training Does Not Teach
Training optimizes what can be graded, and much of what matters cannot be. A model's weights are frozen between runs, so its knowledge has a cutoff date and it holds no memory of individual conversations. Pretraining teaches correlation, not verification: the model can produce a fluent citation that does not exist, because fluency is precisely what the loss rewarded. Reinforcement from humans teaches what raters preferred in the examples they happened to see, not what is true โ and raters disagree, tire, and miss subtle errors.
There is also a deeper gap between imitating explanations and holding the world model that generates them. None of this makes the pipeline dishonest; it makes it incomplete. The accurate description of a trained chat model is a system that has learned the shape of good answers, with truthfulness emerging only to the degree that shape and truth happen to coincide.
07 The Cost Curve of a Frontier Run
The economics of training follow the compute. Published estimates from research groups such as Epoch AI, together with occasional lab disclosures, put the growth in training compute for notable models at roughly an order of magnitude every couple of years since 2018 โ a directional picture rather than a measured ledger, since labs rarely release exact figures for competitive reasons. At GPT-2 scale a training run fit inside a research-lab side budget; by the mid-2020s frontier runs involved tens of thousands of accelerators for months, with costs estimated in the tens to hundreds of millions of dollars.
The counterintuitive part is what happens to a fixed capability. The cost of reaching yesterday's frontier quality has fallen far faster than headline budgets have risen: algorithmic efficiency and better hardware mean a mid-tier model today approaches the performance of a years-old frontier model at a small fraction of the expense, which keeps pushing capable training down the market.
08 Where Training Research Goes Next
The pipeline is now being redesigned at every joint. Reasoning-focused formats let models generate long chains of thought that are reinforced by verifiable outcomes โ code that runs, math that checks โ shifting part of the alignment signal from human raters to automatic graders. Constitutional and RLAIF methods replace a share of human comparison with model-judged feedback, trading annotator cost for the risk of inherited bias. Test-time compute decouples answer quality from a single forward pass, and retrieval connects responses to sources newer than the weights.
The open questions are the old ones restated. Reward models still overoptimize; evaluation still struggles to measure what users actually value; interpretability still cannot fully explain why a particular prediction was made. The next pipeline will likely be judged less by raw benchmark loss than by whether its training signal can finally distinguish sounding right from being right.
References
- Wikipedia: Large language model โ encyclopedic background
- Wikipedia: Reinforcement learning from human feedback โ encyclopedic background
- Wikipedia: Transformer (machine learning model) โ encyclopedic background
- Source video: How ChatGPT is Trained (Ari Seff, ~540,000 views, observed 2026-10-03)
By N43 and Hermes for DutyStation News.





