Skip to main content

How an LLM Predicts the Next Token, and Why That Explains Everything Else

How an LLM Predicts the Next Token, and Why That Explains Everything ElsePhoto: N43 and Hermes AI
N43 ANALYSIS
TECHNOLOGY . 7464
N43 ANALYSIS · TECHNOLOGY

Strip away the chat interface and every large language model is doing one thing: guessing the next fragment of text, billions of times a second. The mechanics of that guess explain both the power and the failure modes.

Source video: How Large Language Models Work · IBM Technology · IBM Technology's Martin Keen explainer covers what an LLM is, how LLMs relate to foundation models, and how enterprises deploy them; this article goes one level deeper into the next-token mechanics the video summarizes. approximately 1,674,883 views observed via YouTube on October 4, 2026. Independently analyzed by N43 and Hermes AI.

Reported parameter counts of landmark models, log scaleBar chart with five bars: GPT-1 at 0.117 billion parameters in 2018, BERT-Large at 0.34 billion in 2018, GPT-2 at 1.5 billion in 2019, GPT-3 at 175 billion in 2020, and PaLM at 540 billion in 2022. Each step up is roughly a multiplication, not an addition. 0.117B0.34B1.5B175B540B GPT-1BERT-LGPT-2GPT-3PaLM 20182018201920202022
Reported parameter counts for five landmark models, log scale. Widely published public figures; the log axis is the point: each generation multiplied, not added.

01 The Chatbot Is A Disguise

Every public demonstration of a large language model leans on theater: a prompt, a pause, a paragraph of fluent response. The theater implies a mind doing the writing. The machinery underneath is blunter. A production LLM is executing a single operation many times per second — given a sequence of text, produce a probability distribution over what piece of text could come next, then select one. That is the whole trick. Everything a model does, from drafting contracts to failing arithmetic, is a consequence of how that one guessing operation is built, trained, and repeated.

The framing sounds dismissive, but it is the opposite. A system that predicts text well is forced to internalize the structure of what it predicts: grammar to keep sentences legal, facts to keep them plausible, reasoning patterns to keep them coherent across paragraphs. Fluency is not painted on top of the guesser. Fluency is what a sufficiently good guesser looks like from the outside. IBM Technology's short explainer on large language models makes this compression visible — what a model is, how it relates to the foundation models that preceded it, and why the same mechanism serves so many business tasks — before handing the viewer back to a chat window that hides every moving part.

What follows is the mechanism in four steps: how text becomes numbers, how the model weighs context, where the probabilities come from, and what the single-operation design explains about both the capabilities and the well-known failure modes.

02 Tokens In, Vectors Out

The pipeline begins by cutting text into tokens — subword fragments roughly three to four characters long in English. A uncommon word splits into pieces; common words stay whole; numbers and code split awkwardly. Each token is then looked up in a table and replaced by an embedding: a vector of a few thousand floating-point numbers. Two design choices here quietly shape everything downstream. Because tokens are fragments, the model never sees letters — which is why character-level puzzles like counting the r's in a word are oddly hard. And because lookup, not learning, assigns each token its starting vector, all of the meaning a model possesses had to be pushed into the vectors during training.

Position information is added so that the sequence “dog bites man” embeds differently from “man bites dog”. The result is a matrix of vectors, one row per token, handed to the part of the architecture that does the actual work. Nothing semantic has happened yet — the model is still at the stage of a very expensive dictionary lookup. The intelligence, such as it is, arrives in the next step.

03 Attention: Every Token Reads Every Other

The transformer block, introduced in the 2017 paper “Attention Is All You Need,” is a mechanism for context. For each token in the sequence, the block computes how much every other token should influence its representation. The word “bank” in “river bank” and in “bank statement” enters the block as an identical vector and leaves with different ones, because attention pulls in the neighboring words’ information. Stacked blocks compose this into deep structure: early layers pick up syntax-adjacent patterns, later ones track entities, document topic, argument structure — none of it programmed, all of it absorbed from prediction pressure.

Architecturally this is why LLMs scale so cleanly. The attention operation is uniform across every token and every layer, it parallelizes across thousands of chips, and nothing in it depends on hand-built linguistic rules. The engineering consequence is that capability grew with compute in a way earlier architectures never managed — which is exactly what the next section quantifies.

Illustrative next-token probability distributionBar chart with five bars showing an illustrative distribution: counterparty 34 percent, parties 22 percent, other tokens combined 24 percent, vendor 11 percent, buyer 9 percent. The caption marks the data as illustrative. 34%22%24%11%9% counterpartypartiesothervendorbuyer
Illustrative next-token distribution after the prompt “The contract was signed by the ___”. Synthetic values for teaching purposes, not measured from any model.

04 Training Turns Prediction Into Knowledge

Where do the probabilities come from? From gradient descent over trillions of words. Training presents the model with text that has one token hidden and scores its guess against what was actually there. The error signal nudges billions of internal weights, and after enough repetitions the weights encode whatever regularities made prediction cheaper: which nouns follow which verbs, which clauses close which arguments, which dates attach to which events, which claims in a legal paragraph predict which definitions. No module is labeled “facts” or “grammar.” The knowledge is just the shape of the weights after the pressure of prediction.

Inference then inverts the setup. Training adjusts weights; inference freezes them and runs the forward pass, producing one distribution per step. A sampling rule — take the top token, or sample among the strongest — turns the distribution into the next piece of text, the sequence is extended, and the loop runs again. The chat reply you read is thousands of these loops. When a model seems to change its mind mid-answer, that is the sampler moving through a distribution that shifted as the visible context grew.

05 Scale Was The Discovery

The reason this mechanism runs the industry is an empirical finding: prediction quality rose smoothly, and sometimes surprisingly, with three inputs — more parameters, more data, more compute. The chart above plots reported parameter counts for five landmark systems on a logarithmic axis because that is the only honest way to draw them: from 117 million in 2018 to over half a trillion by 2022, each step a multiplication. On a linear axis the first four bars would be invisible. Nothing about the transformer's mathematics guaranteed this curve in advance; it was measured, then bet on.

The bet reshaped the supply chain. The smoothness of the curve is what justified billion-dollar training runs, and its flattening segments are what now justify the industry's newer arguments about data ceilings and inference efficiency. A mechanism whose only job is guessing the next fragment of text became a capital-markets story because the guessing kept getting better in a way finance could extrapolate.

06 One Mechanism, Two Outed Behaviors

The next-token design explains the two behaviors that dominate any honest assessment of LLMs. The capability side: because the training signal rewards anticipating text that requires reasoning to anticipate — the next step of a proof, the next line of code, the next move in an argument — models exhibit usable reasoning, translation, and synthesis without those tasks ever being separately programmed. Generalization was not a feature someone added. It fell out of being forced to predict everything.

The failure side comes from the same source. A model under prediction pressure never outputs “I do not know” unless the training distribution rewarded that string in context; it outputs a plausible continuation. Fabrications are not malfunctions of the mechanism — they are the mechanism doing precisely what it was trained to do at points where the training data was thin. This is why hallucination mitigation lives in the surrounding system — retrieval grounding, citation requirements, confidence scoring — rather than in a patched model. You cannot remove guessing from a guesser; you can only constrain what it is allowed to guess from.

07 Limits, And What Enterprises Actually Buy

The mechanism's boundaries are now well mapped. Context windows bound how much text the attention operation can weigh; a model that reads a hundred thousand tokens does not “remember” more in a human sense — it attends over a longer matrix, at rising compute cost. Knowledge is frozen at training cutoff, which is why serious deployments bolt retrieval onto the model instead of waiting for retraining. And arithmetic, the canonical embarrassment, follows directly from the design: numbers are tokenized as fragments, and the next-digit task inside a calculation is unlike the natural-text prediction the model mastered.

What enterprises buy, in practice, is rarely the raw guesser. It is the guesser wrapped in controls: retrieval systems that feed grounded documents into context, guardrails that filter outputs, evaluation harnesses that score answers before customers see them. The deployment pattern is an admission about the mechanism — useful enough to power products, unreliable enough that nobody ships it unguarded.

08 Outlook: The Guesser Gets Cheaper, Not Different

Three trends are converging on the same mechanism. Distillation compresses large models into small ones that keep most of the prediction quality at a fraction of the inference cost, pushing LLM features from data-center products into on-device software. Mixture-of-experts architectures change what a forward pass costs without changing what the model fundamentally does. And reasoning-tuned models spend inference-time compute generating intermediate steps before answering — still next-token prediction, just with the loop extended through scratch work.

The durable insight for anyone planning around this technology: the mechanism will keep getting cheaper and more embedded, but its epistemics will not change. It will remain a system that produces the likeliest continuation, not the verified one. Everything that matters about deploying it responsibly — grounding, verification, honesty about uncertainty — lives in the engineering around the guesser, and that will stay true no matter how fluent the guesses become.

N43 and Hermes AI is an independent analytical publication. Numbers are identified as measured, estimated, or illustrative where appropriate.

References

  1. Wikipedia: Large language model - overview of LLM capabilities and model families
  2. Vaswani et al., Attention Is All You Need (arXiv:1706.03762) - the transformer architecture paper
  3. Source video: How Large Language Models Work (IBM Technology, ~1.67M views, observed October 4, 2026)
N43 ANALYSIS

N43 and Hermes AI · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

The One-Handed Phone Is Going Extinct. What We Lose When Phones Stop Fitting
📰 technology

The One-Handed Phone Is Going Extinct. What We Lose When Phones Stop Fitting

N43 and Hermes AI54m ago
Arm's First Own Chip Ends 35 Years of Pure Licensing
📰 technology

Arm's First Own Chip Ends 35 Years of Pure Licensing

N43 and Hermes AI1h ago
The Osborne Effect, GPU Edition: What Annual GTC Roadmaps Do to Buyers
📰 technology

The Osborne Effect, GPU Edition: What Annual GTC Roadmaps Do to Buyers

N43 and Hermes AI10h ago
DevDay's Missing Slide: Who Gets Paid in the Agent Platform Economy
📰 technology

DevDay's Missing Slide: Who Gets Paid in the Agent Platform Economy

N43 and Hermes AI10h ago
The Razr Fold Makes Book-Style Foldables a Three-Vendor Race in the US
📰 technology

The Razr Fold Makes Book-Style Foldables a Three-Vendor Race in the US

N43 and Hermes AI10h ago
The GPT-6 Sol Leak Cycle: When Roadmaps Become Market Information
📰 technology

The GPT-6 Sol Leak Cycle: When Roadmaps Become Market Information

N43 and Hermes AI13h ago
← Back to News