Skip to content

Source. Interpretability in the Wild (Wang et al., 2023).

Description. Path patching over GPT-2 small identifies 26 attention heads in 7 classes—Previous Token, Duplicate Token, Induction, S-Inhibition, Name Mover, Negative Name Mover, Backup Name Mover—for sentences of the form “When Mary and John went to the store, John gave a drink to ___”. Name Mover roles are supported at the weight level (OV copy score >>95%), the rest at the activation level. Component-level redundancy (Backup Name Movers) is recorded under minimality (I3) rather than necessity: a circuit whose parts compensate for one another is still necessary as a circuit. Description mode: implementational-functional.

IOI circuit. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. A circuit comprising these 26 heads performs IOI in GPT-2 small on the origin templatesCausally SuggestiveIndependent evidence (C3), separation from other circuits (I6), and robustness to method and prompt (M3, M6)
These 26 heads are the mechanism, to the exclusion of othersUnderdeterminedA test that discriminates among rival head sets; naïve and greedy-search circuits reach comparable faithfulness (I5)
The mechanism generalizes across prompts and modelsDisconfirmedNothing — tested and failed: model and circuit diverge at 10610^6 prompt pairs (E2), and faithfulness varies across the origin’s own templates (M6)
The seven classes compose in the stated order, each performing the role its name assertsUnderdeterminedWeight-level accounts for two of seven classes (C2); role semantics currently exceed the evidence (V2)

IOI circuit. Verdict: Causally Suggestive. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Reached.
Mechanistically Supportedinternal (I2, I4); external (E1)Blocked. E1 intervention reach — One ablation value at origin, three granularities; methods disagree; I4 specificity — Head overlap and task effect give opposite verdicts (Merullo et al., 2024)
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked at Mechanistically Supported. C4 discriminant validity — 78% overlap (Merullo et al., 2024); faithfulness opposite (Hanna et al., 2024); E2 prompt generalization — At 10610^6 clean/corrupted pairs, model and circuit diverge (uit de Bos & Garriga-Alonso, 2024); I5 rival mechanism exclusion — Greedy search finds knockout sets carrying 87% of the behavior; I6 double dissociation — No IOI test; the nearest design finds none across tasks (Li & Subramani, 2026)
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Mechanistically Supported. M5 sensitivity — Unattempted; no planted circuit of known extent

IOI circuit. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityC26 heads named with roles; criteria stated as failable quantities
C2Structural plausibilityPCWeight-level for two of seven classes; three omissions the authors name
C3Convergent validityPCFive evidence types; only three share the mean-ablation primitive (Conmy et al., 2023)
C4Discriminant validityI78% overlap (Merullo et al., 2024); faithfulness opposite (Hanna et al., 2024)
C5Nomological validityPCInduction-head theory applied operationally; no formal derivation
C6Complementation validityPCSeven classes knocked out singly; two later crossed (Gong et al., 2026)
Measurement Validity
M1ReliabilityISDs plotted for attention probabilities; none on faithfulness
M2Baseline separationPCNaïve circuit 0.1 against the full 0.46; null supplied later (Shi et al., 2024)
M3StabilityDSix choices move the same quantity below 0% to over 100% (Miller et al., 2024)
M4CalibrationPCLogit difference has sign, null and baseline; faithfulness does not
M5SensitivityUUnattempted; no planted circuit of known extent
M6InvarianceDFaithfulness differs across the origin’s own templates (Miller et al., 2024)
M7Selection correctionUN is unstated for the post-knockout sweep, the selection step itself
Internal Validity
I1NecessityPCName Mover knockout costs only 5%; the paper says so itself
I2SufficiencyPCMean-ablating outside the circuit leaves 87% of the logit difference
I3MinimalityIEvery node clears 1%; contested under other ablations (Li & Janson, 2024)
I4SpecificityIHead overlap and task effect give opposite verdicts (Merullo et al., 2024)
I5Rival mechanism exclusionIGreedy search finds knockout sets carrying 87% of the behavior
I6Double dissociationUNo IOI test; the nearest design finds none across tasks (Li & Subramani, 2026)
I7Confound controlPCSequence length controlled twice; name frequency is not
I8Confounding sensitivityUUnattempted; no E-value or sensitivity analysis
I9Epistatic interactionPCBackup Name Movers are a discovered non-additivity
I10Rescue reversibilityPCDirection reversed at origin; corrupt-then-restore run later
I11Onset couplingN/AUntestable in GPT-2 small, a released checkpoint with no training run
I12Offset couplingPCA Pythia name-mover head loses the behavior late in training (Tigges et al., 2024)
External Validity
E1Intervention reachIOne ablation value at origin, three granularities; methods disagree
E2Prompt generalizationDAt 10610^6 clean/corrupted pairs, model and circuit diverge (uit de Bos & Garriga-Alonso, 2024)
E3Cross-task generalizationPCRepeated-random-token at origin; transfer shown later (Merullo et al., 2024)
E4Cross-model recurrencePCGPT-2 medium at origin; reproduced across Pythia (Tigges et al., 2024; Mueller et al., 2025)
E5Graded responsePCGraded in circuit extent, never in intervention strength
E6Novel predictionCDuplicated-name prediction, with a matched control
Interpretive Validity
V1Level declarationPCCircuits formally defined in §2.1; description level not declared
V2Level-evidence matchPCTopography and information flow supported; role semantics exceed it
V3Alternative levelIThe naïve circuit is simpler and not clearly worse
V4Unlicensed labelingPC“Hedge” for the Negative Name Movers, marked by the authors as speculation
V5Scope declarationCLimits declared in the abstract, §1, §3 and §5
Total: 3 Confirmed, 18 Partially confirmed, 7 Inconclusive, 4 Untested, 3 Disconfirmed, 1 Not applicable
Verdict: Causally Suggestive

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at IOI Circuit — exploratory lens analysis.