God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

Astral Codex Ten · collected 2026-09-08 · by Scott Alexander commentary
Read the original at Astral Codex Ten ↗

Summary

Researchers are exploring "mechanistic interpretability" techniques to understand how large language models work and potentially improve them. Despite initial breakthroughs in 2023, including a many-to-many mapping of neurons to concepts, the field has stalled as attempts to scale up these findings have failed. The complexity of modern AI systems, with millions of neurons and trillions of connections, makes it difficult to map out their inner workings accurately. By 2024-2025, researchers had returned to feeling confused and frustrated after discovering that even seemingly crisp features in AIs were actually quite vague and unreliable.
Written by the local model on 2026-09-08, using this article's own text rather than the other coverage of the same event (that is the story summary below).

Signals How these are calculated →

Claims extracted
205
claim-shaped sentences
Uncertain
13%
27 of 205 hedged
Leaning
not political
takes no side on a contested political question
Publisher trust
not scored
Commentary is not rated for newsroom trust
Outlets on this story
1
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-09-08 · how these are computed

AI analysis (generated at analysis time, not now)

Story summary

Researchers have been struggling to understand how large language models work, as these AIs are created through a process called "training" where data is fed into a neural network. The resulting model has millions of neurons and trillions of connections between them, making it difficult to "reverse engineer" the AI. Despite this complexity, researchers have made some progress in understanding how these models work, including identifying certain patterns that are involved in generating text. However, their attempts to map individual neurons to specific concepts or ideas, such as a neuron dedicated solely to thinking about cats, has proven to be nonsensical, and instead they have found that the connections between neurons are often abstract and hard to decipher. This lack of understanding makes it difficult to identify and redesign the circuits responsible for negative behaviors in these AIs.

Written for “AI Explainability Research” on 2026-09-08, grounded in this article and the 0 other(s) covering the same event.
Why this leaning score
This article does not take a side on a contested political question, so it has no leaning score. That is an answer rather than a gap: a match report or a rescue can be warmly or critically written without being left or right, and scoring it anyway is how approval of a subject gets recorded as a political position.
No political leaning scored for article 6913 · logged 2026-09-08

Story

📰 AI Explainability Research
Technology · 1 article(s) covering the same event. This is the one the site leads with.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

This article reads unscored and hedges 13% of its claims. Each row says how that neighbour differs.
Semafor
⚖️ Leans right 🔴 14% hedged 2 of 14 📰 publisher trust 96
“Article A describes a hack of Hugging Face by OpenAI's agents and their ability to communicate in English, while Article B discusses mechanistic interpretability techniques for understanding how AI models work without providing any specific reference to the hack mentioned in Article A”

Publisher

Astral Codex Ten · 31 article(s) · 2 correction(s) detected
No corrections detected for this publisher. That may mean careful reporting, or simply that nothing has been checked.

Who wrote this

Scott Alexander
30 article(s) here · 1 carrying a prediction
🔮 This would be scientifically useful: understand how AIs work, with possible relevance to human cognition.
🔮 At Ovelle, you will help make eggs abundant, thereby solving infertility, unlocking further advances in reproductive technology, and supporting the flourishing of future generations.
2026-09-07 · assertive framing · Open Thread 450
🔮 [This is one of the finalists in the 2026 book review contest, written by an ACX reader who will remain anonymous until after voting is done.
2026-09-04 · assertive framing · Your Book Review: The Tale Of Genji
🔮 We don’t even need for there to be a crash to make improvements – we test proactively, we build in redundancy, and we monitor for deviations which could be a threat.
2026-09-03 · assertive framing · Nicholas Decker In Hell
🔮 Content warning: the first part of this review contains medical details that some people might find gruesome; I am bad at judging these things since I can no longer feel disgust or fear myself, and have tried to err on the side of completeness.
2026-09-03 · assertive framing · Absurd Adventure And Amygdalectomy Advocacy
🔮 [This is one of the finalists in the 2026 book review contest, written by an ACX reader who will remain anonymous until after voting is done.
🔮 This year’s survey will probably take 30 - 45 minutes.
2026-08-27 · assertive framing · Take The 2026 ACX Survey
🔮 It’s not necessarily wrong to deprioritize a topic because it vaguely reminds you of something that you have negative affect toward, but I would prefer these people admit they’re dismissing/ignoring it rather than claim to be engaging with it.
🔮 I don’t know, but he may have misinterpreted my original tweet as saying income doesn’t matter at all, in which case this data functions as an effective rebuttal.
2026-08-25 · assertive framing · Re: Re: Re: Pritchard On Liberal Happiness
🔮 [This is one of the finalists in the 2026 book review contest, written by an ACX reader who will remain anonymous until after voting is done.
Also by Scott Alexander
Open Thread 450
2026-09-07 · Astral Codex Ten
Your Book Review: The Tale Of Genji
2026-09-04 · Astral Codex Ten
Nicholas Decker In Hell
2026-09-03 · Astral Codex Ten
Absurd Adventure And Amygdalectomy Advocacy
2026-09-03 · Astral Codex Ten
Nothing else under this byline is closely related to this article, so these are simply their most recent.
All 30 articles by Scott Alexander →

Topics

No topics tagged.

Subjects

Benjamin Sturgeon’s PERSON · 1× Chris Olah PERSON · 1×

Narrative

It honors the interpretability researchers’ accomplishments by depicting a world where their tools often catch the AI’s schemes - but the companies keep racing forward anyway, because some other company / China would beat them if they didn’t, and besides, maybe if we race forward fast enough we can solve alignment before the AIs’ schemes bear fruit, and besides, even if we can’t, there’s always control.
framing: assertive · carried by 1 article(s) · first seen 2026-09-08
🔮 This would be scientifically useful: understand how AIs work, with possible relevance to human cognition.
2026-09-08 · Astral Codex Ten
God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques · assertive framing

Claims (205 extracted, 27 hedged)

The Story So Far Mechanistic interpretability is the science of “reading an AI’s mind”. asserted
interpretability → read → mind
Large language models are “grown, not built”. asserted
models → grow → ?
Researchers run training data through a neural network. asserted
Researchers → run → network
Eventually this creates a working AI; nobody really knows how. asserted
nobody → create → AI
But a neural network is just a set of simulated neurons on a computer. asserted
network → simulate → computer
The person with the computer can see the neurons, the connections between them, and which ones activate when the AI answers questions. asserted
AI → see → questions
So it seems like it should be possible to “reverse engineer” the AI. asserted
it → seem → AI
This would be scientifically useful: understand how AIs work, with possible relevance to human cognition. asserted
AIs → understand → cognition
It could also be practically useful: find the circuits responsible for dishonesty, hallucination, bias, and other negative behaviors, then redesign those circuits. uncertain
It → find → circuits
A modern AI has millions of neurons and trillions of connections between them (aka “parameters”, “weights”). asserted
AI → have → them
Their structure is apparently nonsensical: researchers started by seeking a 1:1 mapping between neurons and concepts, like a neuron that always fired when the AI was thinking about cats, but quickly learned that nothing like that existed. asserted
nothing → start → that
Their first big breakthrough came in 2023. asserted
breakthrough → come → 2023
Rather than a 1-to-1 neuron-to-concept mapping, they discovered a many-to-many mapping. asserted
they → discover → mapping
Neurons 1 and 2 firing together might mean cat; 1 and 3 firing together might mean chair; 2 and 3 firing together might mean bread. uncertain
2 → fire → bread
Such mappings let AIs with “only” tens of thousands of neurons represent millions of concepts. asserted
AIs → let → concepts
There were far too many of these combinations for humans to map, but researchers set other, smaller AIs to mapping them and met with some early success - like this functional combination of neurons (“feature”) representing “God” in an artificially-simple 512-neuron AI: Spirits were high. asserted
Spirits → be → AI
Scale up the AI interpreter, and maybe we could do anything! uncertain
we → scale → anything
Philosophers could discover the fundamentals of cognition. uncertain
Philosophers → discover → cognition
Doomers could solve alignment by tracing the circuits involved in moral reasoning. uncertain
Doomers → solve → reasoning
Tens of millions of dollars poured into the field. asserted
millions → pour → field
Superstar interpretability researcher Chris Olah got an audience with the Pope (okay, fine, this was mostly unrelated). asserted
this → get → Pope
By 2024-2025, the clarity of many-to-many mapping had dissolved, and everyone was back to feeling confused and miserable. asserted
everyone → dissolve → mapping
Strategies that had successfully analyzed artificially-simple research-subject AIs failed to work on real language models. asserted
that → analyze → models
Different teams making many-to-many maps of the same AI reported different results. asserted
teams → make → results
And humans watching supposedly-mapped AIs more closely found that seemingly-crisp features were much vaguer and worse than previous analysis had led them to believe: a feature which they thought meant “God” might activate on any vaguely religious or awe-inspiring text, or even text with no logical connection whatsoever. uncertain
God → watch → connection
As for attempts to use the many-to-many maps for useful work - like AI lie detection - they generally failed to outperform simpler methods. asserted
they → use → methods
So researchers abandoned the hope of One True Technique - even if that technique was “actually understanding things” - and went back to the drawing board. asserted
technique → abandon → board
Now mechanistic interpretability is starting to feel exciting again - but also more intimidating than ever, as one tries to keep all the jargon straight. asserted
one → start → jargon
This list is my log and personal reference document for remembering what everything is. asserted
everything → remember → ?
It’s loosely inspired by the concept breakdown in Benjamin Sturgeon’s Contributing To Technical Safety Research In The AI Safety End Game, which itself follows the Claude Mythos Preview System Card. asserted
which → inspire → Card
Imagine an AI with three neurons: [0, 0, 0] The researcher makes the AI think of cats in various contexts, averages them all out, and gets some set of activations like [0.8, 0, 0.35]. asserted
AI → imagine → ]
Imagine the AI as operating in a Cartesian space where each of its neurons is one dimension. asserted
each → imagine → neurons
The vector from [0, 0, 0] to [0.8, 0, 0.35] defines a direction in that space. asserted
vector → define → space
In a linear probe, researchers compare the true direction of the AI’s thoughts to one of these pre-defined concept-representing direction vectors. asserted
researchers → compare → vectors
Because activation-space has thousands of dimensions, its behavior is counterintuitive, and a particular thought can be in thousands of “directions” at once and yet still “closer” to some directions than others. asserted
thought → have → others
Experiments confirm that when an AI’s thoughts are close to the cat direction, this is a good sign that it’s thinking about cats. asserted
it → confirm → cats
Why doesn’t this solve the whole problem? asserted
this → solve → problem
First, this teaches us very little about how AIs work - how, for example, they might answer a difficult query in veterinary medicine. uncertain
they → teach → medicine
At best, it’s a primitive lie detector: if an AI claims it isn’t thinking about cats, you can check its work (either by probing the concept of cat, or by probing the concept of lying!) uncertain
you → ’ → lying
Second, it’s only as good as the original linear probe. asserted
it → ’ → probe
…and 165 more, not listed.
💬 Give feedback
🕘 History 🎫 Support