Skip to content

Source. Sparse Autoencoders Find Highly Interpretable Features (Cunningham et al., 2024).

Description. Sparse autoencoders are trained on residual-stream activations to recover an overcomplete dictionary of directions, evaluated two ways: automated interpretability scoring of 150 features per method against six baselines (default basis, random directions, PCA, ICA, top-KK PCA, top-KK ICA), and activation patching on 50 IOI data points. Evaluated on: Pythia-70M and 410M at origin, GPT-2 small and Gemma 2 2B in the follow-up. Being a method-level claim, these are the systems the evidence was gathered on rather than one audited model. Description mode: representational.

SAE features. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. Sparse dictionary learning recovers directions more interpretable than PCA, ICA and the neuron basis, and causally usableProposedBaseline separation (M2), where the origin’s within-model controls and the later random-network controls disagree; then specificity (I4), independent convergence (C3), separation from other decompositions (I6), and calibration of the scoring instrument (M1, M3, M5, M7)
SAE features are atomic, canonical units of the modelDisconfirmedNothing — tested and failed: meta-SAEs decompose latents further (C4), and the instrument cannot distinguish a trained transformer from a random one (Heap et al., 2025)
SAE features localize causation better than neuronsDisconfirmedNothing — tested and failed: (Mueller et al., 2025) report no causal-localization advantage over the neuron basis

SAE features. Verdict: Proposed. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Blocked. M2 baseline separation — Separates from random directions; not from a random network when that control is run
Mechanistically Supportedinternal (I2, I4); external (E1)Blocked at Causally Suggestive. I4 specificity — Single-feature ablation moves 12,000 logits; left unanalyzed
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked at Causally Suggestive. I6 double dissociation — Unattempted; no second decomposition shown intact
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Causally Suggestive. C6 complementation validity — Feature splitting is the open question; no test of atomicity is run; E3 cross-task generalization — Unattempted; generalization stated as an expectation; I10 rescue reversibility — Unattempted, though patching is reversible by construction; I11 onset coupling — Unattempted, and cheap: Pythia ships pretraining checkpoints (Biderman et al., 2023); I12 offset coupling — Unattempted; the nearest result is at an extreme, not a trajectory (Heap et al., 2025)

SAE features. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityCDirectional prediction against four named baselines; partly fails in layer 4
C2Structural plausibilityPCPredicted by superposition; MLP training fails, attention not attempted
C3Convergent validityPCTwo evidence types; the concurrent-origin pairing comes later (Leask et al., 2025)
C4Discriminant validityPCScored against the origin’s operational construct, not the field’s
C5Nomological validityPCFits superposition; the theory does not underwrite the method
C6Complementation validityUFeature splitting is the open question; no test of atomicity is run
Measurement Validity
M1ReliabilityIOne run at origin; later seed studies disagree (Paulo & Belrose, 2025)
M2Baseline separationISeparates from random directions; not from a random network when that control is run
M3StabilityISparsity sweep gives a smooth tradeoff, not a stable optimum
M4CalibrationPCInterpretability score is a correlation with a meaningful zero
M5SensitivityDTested later on planted features and the recovery is poor (Korznikov et al., 2026)
M6InvariancePCAll layers of Pythia-70M, five expansion ratios, two scoring regimes
M7Selection correctionUA feature scoring fewer than 20 varying fragments is skipped; the count is unreported
Internal Validity
I1NecessityPCOne feature ablated for behavior; a full previous-layer sweep reads out to features
I2SufficiencyPCInterchange intervention on IOI; counterfactual features written in
I3MinimalityPCACDC orders features; the next feature usually adds little
I4SpecificityISingle-feature ablation moves 12,000 logits; left unanalyzed
I5Rival mechanism exclusionPCRival methods excluded: PCA, ICA, top-KK, non-sparse dictionary
I6Double dissociationUUnattempted; no second decomposition shown intact
I7Confound controlPCThree designed controls, including an =0=0 dictionary
I8Confounding sensitivityUUnattempted; no E-value or sensitivity analysis
I9Epistatic interactionPCAssumed away explicitly; ACDC treats features as a flat graph
I10Rescue reversibilityUUnattempted, though patching is reversible by construction
I11Onset couplingUUnattempted, and cheap: Pythia ships pretraining checkpoints (Biderman et al., 2023)
I12Offset couplingUUnattempted; the nearest result is at an extreme, not a trajectory (Heap et al., 2025)
External Validity
E1Intervention reachPCTwo forms: activation patching and less-than-rank-one ablation
E2Prompt generalizationPCCross-corpus by construction: trained on the Pile, scored on OpenWebText
E3Cross-task generalizationUUnattempted; generalization stated as an expectation
E4Cross-model recurrenceCTransfers everywhere tried: Pythia, GPT-2 small, Gemma 2 2B (Leask et al., 2025)
E5Graded responsePCTwo graded curves: KL against features patched, edit size against KL
E6Novel predictionPCThe IOI experiment is a prediction in the weak sense
Interpretive Validity
V1Level declarationPCFormal object declared precisely; “feature” carries more than it
V2Level-evidence matchPCEvidence supports the decomposition claim the origin makes
V3Alternative levelPCRaised at origin and not pursued: no single correct decomposition
V4Unlicensed labelingPC“Monosemantic” projects semantics; the origin hedges consistently
V5Scope declarationCBoth papers declare limits and quantify them (Leask et al., 2025)
Total: 3 Confirmed, 20 Partially confirmed, 4 Inconclusive, 8 Untested, 1 Disconfirmed
Verdict: Proposed

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at SAE Features — exploratory lens analysis.