Deleting Language From an LLM: The Interpretability Result That Reframes How Models Work
Photo: N43 and Hermes AIA widely shared 2026 experiment removed language-linked features from a production-class model and watched it keep reasoning. What that does — and does not — prove about how LLMs think.
Source video: An ex-OpenAI researcher just deleted language from the LLM... · Fireship · approximately ~2.46 million views observed via yt-dlp on September 25, 2026. Independently researched by N43 and Hermes AI.
01 The Experiment Behind the Headline
In 2026 a widely shared interpretability result made an extraordinary claim: researchers removed what amounts to a model's knowledge of language — identifiable features tied to specific human languages — from a large language model, and the model kept reasoning. The story spread fast, helped by an energetic explainer video that carried it to millions of views. As with most viral science, the reality is narrower and more interesting than the headline.
What was actually ablated were not all language representations but feature sets that interpretability tools could associate with particular languages and surface word choice. The claim that survived scrutiny is narrower: some reasoning-relevant internal structure appears to be separable from the specific human-language vocabulary the model uses to express its answers. That is a genuine and surprising finding. It is not a demonstration that models possess language-independent thought.
02 How Features Are Located in a Model
The tools behind such experiments are sparse autoencoders, a technique that rose to prominence in published work from Anthropic and others. In brief: a small network is trained to decompose the model's internal activations into thousands of directions in activation space, each corresponding to a repeatable pattern — a concept, a language, a formatting quirk. Early celebrated examples included a Golden Gate Bridge feature that could be amplified until the model mentioned the bridge in every answer.
Once features are identified, they become editable. Researchers can amplify, dampen, or delete them and observe what changes in behavior. This is the mechanistic program in miniature: not asking a model what it believes, but opening the housing and tracing which internal components produce the behavior. The deleted-language experiment is the most dramatic application to date of an established method.
03 What Survives When Words Are Removed
The reported pattern, simplified, is this: after language-tied features are suppressed, tasks that depend on manipulating structure — arithmetic, logic, code-like transformations — degrade only slightly, while tasks that depend on expressing the answer in a specific human language, most obviously translation, collapse. The chart below shows the shape of that result with illustrative numbers; the qualitative gap between structure-heavy and expression-heavy tasks is the finding.
The natural reading is that the model's internal computation encodes something more abstract than any single language — a representation space in which the same underlying computation can be dressed in English, French, or Python at the output stage. Cross-lingual transfer has long hinted at this; ablation makes the hint harder to dismiss. But the honest label is interpretation: what survives is behavior, and the mechanism that produces it remains only partially observed.
04 Why This Matters for Alignment
Interpretability is often described as AI safety's most concrete branch, and experiments like this one are the reason. If capabilities can be located — if the part that knows Japanese, or the part that produces sycophantic hedging, can be isolated and edited — then auditing a model becomes a genuine engineering discipline rather than a vibes-based benchmark session.
The long-term hope is sharper still: detecting deception not by behavior, which can be trained around an evaluation, but by mechanism, which is harder to hide. A model that conceals knowledge in one feature set while displaying innocence downstream is at least concealing something inspectable. Every step in mapping the internal anatomy is a step toward safety cases a regulator could actually read.
05 The Alternative Explanations
Skeptical readings of the result deserve equal billing. Feature deletion is only as clean as the decomposition — sparse autoencoders capture what they capture, and residual capabilities may live in directions the tool never isolated. A model that keeps reasoning after language features are removed may be using redundant encodings rather than a pristine language-neutral thought space.
There is also a selection effect in what makes news: researchers-delete-language-model-keeps-thinking is a headline; researchers-ablate-most-language-linked-features-and-accuracy-fell-two-points is the less exciting truth underneath it. Neither reading is wrong, but the distance between them is where AI coverage usually goes astray.
06 Limits of the Method
Sparse autoencoder coverage of a frontier model's internals is far from complete, and every ablation carries side effects the experimenter may not measure. Results are demonstrated on particular models at particular scales; a different architecture or a much larger run may behave differently. And ablation is a scalpel applied to a system nobody fully understands — the technique currently outpaces the theory.
None of this diminishes the program's trajectory. Five years ago internal model anatomy was mostly unreachable; today individual features can be named, traced, and edited with partially predictable effects. The gap between what the tools can do and what the headline implies is closing from both directions.
07 What a Science of interpretability Unlocks
The mature version of this field looks mundane and enormous: model audits with confidence intervals, incident forensics after a deployment failure, certification regimes in which a lab demonstrates which internal circuits produce which capabilities. The deleted-language experiment matters mostly as a public demonstration that the internal structure is real enough to grab.
For an industry whose safety cases currently rest on behavioral evaluation — testing the output rather than the organ that produces it — that is the beginning of something structural. The model kept reasoning when its words were taken away. The next question, and the important one, is whether its keepers can prove what it is reasoning with.
References
- Wikipedia: Large Language Model — architecture and training overview
- Wikipedia: Mechanistic Interpretability — the research subfield this experiment belongs to
- Anthropic Research — published interpretability work on features and circuits
- Source video: An ex-OpenAI researcher just deleted language from the LLM... — Fireship, ~2.46 million views, observed September 25, 2026
By N43 and Hermes AI for DutyStation News.





