What Are Transformer Circuits in Neural Networks?

Transformer circuits are the internal computational pathways that researchers have identified inside transformer-based AI models, tracing how specific groups of components work together to perform identifiable tasks like completing a pattern, identifying a grammatical subject, or recalling a fact. The field of “mechanistic interpretability” treats a transformer not as an opaque black box but as a system of interacting parts whose behavior can, in principle, be reverse-engineered the way you might trace wiring in an electronic device. Over the past few years, this work has moved from toy demonstrations to discoveries inside real language models, revealing surprisingly structured internal mechanisms and raising hard questions about whether we can ever fully map what large AI systems are doing.

How Information Moves Through a Transformer

A transformer processes text by passing it through a series of layers. Each layer reads from and writes to a shared workspace called the residual stream, which you can think of as a running tally of everything the model has figured out so far about the input. The key insight from circuit-level analysis is that individual layers do not each do a little bit of everything. Instead, they write information into specific, narrow slices of this workspace, and later layers read from those same slices. These narrow slices form what researchers call low-rank communication channels between layers.1NeurIPS. Transformer Circuits

This matters because it means the model’s computation is not a tangled mess. There is structure to how one part of the network talks to another. An attention head in layer three might write a signal that is specifically picked up by a feedforward layer in layer eight, and nowhere else. That selective communication is what makes it possible to talk about “circuits” at all. Without it, you would just have a blur of numbers flowing forward with no identifiable pathways.

Induction Heads and the Birth of In-Context Learning

One of the most celebrated circuit discoveries involves induction heads, a type of attention head that performs a deceptively simple operation: if the model has already seen the sequence [A][B] earlier in the text, and it now encounters [A] again, the induction head boosts the prediction that [B] comes next. It is pattern completion at its most basic, yet it appears to be the engine behind something much bigger.

Researchers at Anthropic found that induction heads develop at a specific, identifiable moment during training. Early in the process, there is a visible bump in the training loss, a brief period where the model’s performance dips before improving sharply. This “phase change” coincides with the formation of induction heads, and it marks the point where the model acquires most of its ability to learn from context within a single prompt.2Transformer Circuits Thread. In-context Learning and Induction Heads – Section: Argument 1: Transformer language models undergo a “phase change” during training, during which induction heads form and simultaneously in-context learning improves dramatically. The evidence suggests this is not a coincidence. Six complementary lines of evidence point toward induction heads being the primary mechanism behind in-context learning across models of all sizes.3arXiv. In-context Learning and Induction Heads

What makes this striking is that the induction head’s underlying operation, match and copy, is simple enough to describe in one sentence.4arXiv. What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation Yet these heads implement “fairly abstract and fuzzy” versions of pattern completion, meaning they are not limited to exact token matches. They can generalize the matching operation in ways that let the model adapt its behavior to new instructions, new formats, and new tasks it encounters mid-prompt. A single, identifiable circuit element appears to underpin one of the most useful capabilities of modern language models.

Feedforward Layers as Lookup Tables

Attention heads get most of the popular attention, but transformers have another major component: feedforward layers, which sit between the attention operations in every transformer block. Research has shown that these layers function like key-value memory banks. Each neuron in the feedforward layer acts as a key that activates when it detects particular textual patterns, and the corresponding value pushes the model’s output distribution toward specific words or tokens.5arXiv. Transformer Feed-Forward Layers Are Key-Value Memories

This means that when a feedforward neuron fires strongly, it is essentially voting for certain words to appear next, based on patterns it learned during training. Some neurons respond to factual associations (detecting patterns related to, say, capital cities), while others respond to syntactic structures. The feedforward layers are where much of the model’s stored knowledge lives, as opposed to the attention layers, which are more about routing and combining information from different positions in the input.

Why Individual Neurons Are Hard to Interpret

If every neuron in a transformer responded to one clean, identifiable concept, reverse-engineering circuits would be straightforward. In practice, most neurons are polysemantic: a single neuron fires in response to multiple, seemingly unrelated inputs.6arXiv. Understanding polysemanticity in neural networks through coding theory A neuron might activate for both academic citations and certain types of list formatting, with no obvious connection between the two. This makes it extremely difficult to assign a human-readable label to what any given neuron “means.”

The leading explanation for why this happens is called superposition. When a model needs to represent more distinct concepts than it has neurons, it packs multiple sparse features into overlapping patterns of neuron activations. Think of it like storing several different signals on the same set of wires by giving each signal a slightly different encoding. As long as the features are rarely active at the same time, the interference stays manageable. Toy model experiments have confirmed that networks do this deliberately when they have more features to track than dimensions to store them in.7arXiv. Toy Models of Superposition

Superposition is arguably the single biggest obstacle to fully understanding transformer circuits. Even when you can identify a circuit at the level of attention heads and layers, the features flowing through that circuit may be tangled together inside individual neurons, making the fine-grained picture murky.

Untangling Features with Sparse Autoencoders

To deal with superposition, researchers have turned to a technique called dictionary learning, most commonly implemented through sparse autoencoders. The idea is to train a secondary, much wider network to decompose a transformer’s internal activations into a large set of individual features, each of which ideally responds to one coherent concept. Anthropic demonstrated this approach on a one-layer transformer and extracted a large number of interpretable, relatively monosemantic features, meaning each feature corresponded to a recognizable concept rather than a jumble of unrelated ones.8Anthropic. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Sparse autoencoders have quickly become one of the most active tools in the field. They allow researchers to peer past the polysemanticity barrier and work with cleaner units of analysis. Rather than asking “what does neuron 4,237 do?” you can ask “what does feature 18,902 in the sparse autoencoder do?” and get a more coherent answer. These features are not perfect, and the decomposition involves choices that affect the results, but they represent the best current method for getting interpretable building blocks out of a model’s internals.

A Circuit in the Wild

One of the most detailed circuit analyses to date dissected how GPT-2 Small handles a natural language task called indirect object identification. Given a sentence like “When Mary and John went to the store, John gave a drink to,” the model needs to predict “Mary” rather than “John.” Researchers traced this ability to a circuit involving 26 attention heads organized into seven functional classes, each with a distinct role: some heads identify the repeated name, others suppress it, and others promote the correct alternative.9arXiv. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

This study was a landmark because it showed that a real model performing a real language task uses an internally organized circuit that humans can understand in functional terms. It is not just one or two attention heads doing vaguely interpretable things. It is a coordinated set of components with different jobs that work together. The methods used, including causal interventions where researchers surgically alter activations and observe how outputs change, have since become standard in the field.

Automating the Search

Manually tracing circuits head by head is labor-intensive and does not scale well to models with billions of parameters. This has pushed researchers toward automating the circuit discovery process. One notable approach, the ACDC algorithm, takes a model and a specified behavior and systematically prunes the model’s computational graph to find the minimal subnetwork responsible. In validation tests, ACDC successfully rediscovered all five component types in a previously known GPT-2 Small circuit that performs greater-than comparisons.10arXiv. Towards Automated Circuit Discovery for Mechanistic Interpretability

Other methods use edge pruning to make automated discovery more scalable, reducing the search space by focusing on which connections between components matter rather than which components individually matter.11NeurIPS Proceedings. Edge Pruning for Scalable Circuit Discovery Automation is essential if circuit analysis is ever going to keep pace with the rapid growth in model size. Frontier models have hundreds of billions of parameters; finding circuits in them by hand is not realistic.

Steering Model Behavior by Manipulating Circuits

Understanding circuits is not just an academic exercise. If you know which internal features correspond to specific behaviors, you can intervene on them directly. Steering vectors are one practical application: by adding a calculated vector to a model’s internal activations at a specific layer, you can push its outputs in a desired direction without retraining the model. This approach is easier than fine-tuning and may be more reliable than prompt engineering.12arXiv. Improving Steering Vectors by Targeting Sparse Autoencoder Features

Recent work has demonstrated remarkably precise control using sparse autoencoder features. In one experiment, modifying a single feature at one transformer layer was enough to shift the language a model generates, achieving controlled language switches with up to 90% success while preserving the meaning of the output.13arXiv.org. Causal Language Control in Multilingual Transformers via Sparse Feature Steering That level of surgical control, changing one feature and watching a specific, measurable behavior flip, is strong evidence that the features being extracted by sparse autoencoders correspond to real functional units inside the model.

Do Different Models Develop the Same Circuits?

A natural question is whether the circuits researchers find in one model also show up in others. The universality hypothesis suggests that different neural networks trained on similar tasks may converge on similar internal algorithms, the way evolution has independently produced eyes in unrelated species.14arXiv. Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures If this were true, understanding circuits in a small model could teach us about large ones.

The reality is more complicated. A study examining numerical comparison circuits across several model families and sizes found a split. Models within the same family, specifically the Qwen family spanning from about 1.7 billion to 9 billion parameters, showed highly consistent circuit structures with localized attention heads performing the same role at similar positions. But models from different families implemented the same task in qualitatively different ways, distributing the relevant computation across many more components earlier in the network rather than concentrating it in a few heads.15ACL Anthology. Mechanistic Analysis Of Universality: Numerical Comparison Circuits Across Transformer Architectures The takeaway is sobering: two models can produce the same correct answers through genuinely different internal mechanisms. Getting the same output does not mean they are wired the same way inside.

The Measurement Problem

Even when circuits are identified, there is an underappreciated problem with how researchers verify them. The standard approach involves ablation: you knock out part of the model’s computation (by zeroing out activations, replacing them with averages, or scrambling them) and measure how much the model’s performance drops on the target task. A big drop means the ablated component was important to the circuit. But the choice of how you ablate turns out to matter a lot.

A systematic investigation found that existing circuit faithfulness scores are highly sensitive to seemingly minor methodological decisions, things like which baseline value you replace the ablated activation with, or exactly how you define the task the circuit is supposed to perform. The same circuit can look essential or irrelevant depending on these researcher-side choices.16arXiv. Transformer Circuit Faithfulness Metrics are not Robust This does not mean the circuits are not real, but it does mean the field lacks a gold standard for measuring how well a proposed circuit actually explains the model’s behavior. Two research groups studying the same model could reach different conclusions about its circuitry based on methodological differences they might not even realize are important.

Circuits for Chain-of-Thought Reasoning

More recent work has extended circuit analysis to chain-of-thought reasoning, where a model “thinks out loud” by generating intermediate steps before producing a final answer. Researchers have developed sequential activation patching techniques that trace how attention heads contribute across token positions during multi-step reasoning, identifying distributed sub-circuits responsible for maintaining the reasoning trajectory, anchoring the answer, and separating examples from targets.17arXiv. Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching When the identified heads were ablated, the model’s ability to generate correct answers degraded, confirming that these components are functionally important rather than just correlated with the task.

This is a frontier area because chain-of-thought is one of the main techniques used to improve model performance on hard problems. Understanding the circuits behind it could eventually help distinguish genuine multi-step reasoning from superficial pattern-matching that mimics reasoning.

Connections to AI Safety

The interest in transformer circuits is not purely scientific curiosity. For the AI safety community, understanding a model’s internal circuitry is one of the most promising paths toward ensuring that powerful AI systems behave as intended. If you can identify the circuit responsible for, say, deceptive behavior or sycophantic agreement, you might be able to detect and suppress it. Surveys of the field identify mechanistic interpretability as directly relevant to alignment strategies, while also acknowledging that key challenges remain, including the difficulty of interpreting emergent behaviors in very large models.18PubMed Central. Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions – Section: 4.2 Detecting and Mitigating Deception

The tension is between the precision of current circuit-level findings, which mostly come from small or medium models, and the scale of the models that actually raise safety concerns. Scaling interpretability analysis to frontier models is one of the field’s central open problems.19ACM Computing Surveys. Bridging the Black Box: A Survey on Mechanistic Interpretability in AI

Parallels with the Brain

An intriguing side thread in transformer circuit research involves comparing how transformers and biological brains represent information. In one study, researchers aligned transformer circuit mechanisms to neural recordings from humans performing a relational reasoning task. They found that positional encoding in the transformer, which captures the spatial structure of the input, aligned with representations in the visual cortex. Meanwhile, the transformer’s attention mechanism, which captures relational structure, mapped onto frontoparietal and default-mode networks in the brain.20bioRxiv. Aligning transformer circuit mechanisms to neural representations in relational reasoning

This does not mean transformers and brains work the same way. But it suggests that certain computational strategies, like separating “where things are” from “how things relate,” may be convergent solutions to similar problems. As both mechanistic interpretability and neuroscience develop better tools for decomposing their respective systems into functional components, the comparisons are likely to get more precise and more interesting.

When Models Get Compressed

A practical question for deploying interpretability tools is whether they survive model compression. Companies routinely prune and quantize large models to make them cheaper to run. If a sparse autoencoder trained on the original model becomes useless after pruning, then interpretability analysis would need to be redone from scratch every time a model is compressed for deployment. Encouragingly, experiments on GPT-2 Small and Gemma-2-2B found that sparse autoencoders pre-trained on uncompressed models retain high interpretability and reconstruction quality even when applied directly to pruned versions. Simply pruning the sparse autoencoder itself produced results nearly as good as retraining from scratch on the compressed model.21arXiv. On the transferability of Sparse Autoencoders for interpreting compressed models This is a practical win: it means interpretability tools may not need to be rebuilt every time a model gets slimmed down for production use.