Skip to content

Source. Toy Models of Superposition (Elhage et al., 2022).

Description. A ReLU network trained to reconstruct sparse features stores more of them than it has dimensions, giving each feature a non-orthogonal direction and absorbing the interference those directions create; the linear model differing only in activation function shows no superposition at any sparsity. Where the tradeoff falls is mapped over an importance ×× sparsity grid against closed-form losses for each candidate weight configuration, and the geometries the trained models settle into are uniform polytopes solving a generalized Thomson problem. Evaluated on: toy autoencoders on synthetic features whose sparsity, importance and correlation are set by hand, across four architectures and sizes from (2,1)(2,1) to 400 features in 30 dimensions, plus an absolute-value model that computes through a compressed layer rather than storing. Description mode: representational.

Superposition. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. A ReLU network on sparse features represents more features than it has dimensions, in geometries set by sparsity and importanceMechanistically SupportedA converse arm, with some property shown intact while superposition is removed (I6); correction for the lowest-loss selection the reported geometries are drawn from (M7); an input distribution outside the synthetic family (E2)
The reported geometries are properties of the optimum rather than of the optimizerUnderdeterminedThe m=2m=2 case is reported as much harder for gradient descent and its solutions are selected by loss (I7, M7); an invited comment in the same article shows the global minimum can have a smaller basin of attraction than nearby local minima, and the paper leaves it open (V3)
The phase diagram is fixed by sparsity and importance aloneDisconfirmedNothing — tested and failed: the paper holds the activation function fixed across every one of its own experiments, and a replication carried in the same article finds the phase diagrams look quite different under other activation functions (M3, M6)
Language models represent features in superpositionInsufficientEvery measurement is on a toy model, and the paper labels its real-model section validation by consistency with existing reports rather than measurement taken here (E4); no natural-data distribution appears anywhere, and Open Questions asks whether real importance and sparsity curves can be estimated at all (E2)

Superposition. Verdict: Mechanistically Supported. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Reached.
Mechanistically Supportedinternal (I2, I4); external (E1)Reached.
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked. I6 double dissociation — Importance and sparsity are crossed, but on one outcome; no converse arm is run
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Triangulated. C6 complementation validity — Ground truth is available by construction and the test is not run; I10 rescue reversibility — Adversarial training reduces superposition; nothing is corrupted and then restored; I12 offset coupling — Offset never run; the learning-dynamics section is limited by the authors’ own account

Superposition. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityCPoint prediction against a linear control, plus a two-parameter phase diagram
C2Structural plausibilityCMechanism derived in closed form, then matched against the trained weights
C3Convergent validityCTheory and experiment agree, and three outside groups replicated before publication
C4Discriminant validityPCThree-way discrimination against planted features; PCA separated as a limiting case
C5Nomological validityCFormal links to compressed sensing, the Thomson problem and distributed codes
C6Complementation validityUGround truth is available by construction and the test is not run
Measurement Validity
M1ReliabilityPCRepetition run and reported; aggregation is selection by loss, with no dispersion
M2Baseline separationCLinear control cannot superpose, runs at every sparsity, and separates completely
M3StabilityPCSparsity, importance, correlation and size all swept; activation function held fixed
M4CalibrationPCTwo purpose-built measures with stated readings and no external calibration
M5SensitivityCPlanted ground-truth features give a known positive at every point
M6InvariancePCHolds across sizes and data structures; one activation function and one loss
M7Selection correctionUBest-of-1000 and worst-discarded selection documented, with no correction applied
Internal Validity
I1NecessityCOutput nonlinearity, sparsity and negative bias each shown to be required
I2SufficiencyCBuilt from the hypothesized ingredients alone, and superposition appears
I3MinimalityPCLoad-bearing parts named and the smallest case solved; data assumptions untested
I4SpecificityN/AOne task and one loss, and the paper says cross-superposition loss comparison fails
I5Rival mechanism exclusionPCPCA named as a limiting case; incidental polysemanticity untested (Lecomte et al., 2023)
I6Double dissociationUImportance and sparsity are crossed, but on one outcome; no converse arm is run
I7Confound controlPCData confounds set by construction; optimization left as an uncontrolled one
I8Confounding sensitivityUUnattempted; nothing bounds an unmeasured confounder in the synthetic setup
I9Epistatic interactionCFeature interactions are the subject once uniformity is dropped, and each is measured
I10Rescue reversibilityUAdversarial training reduces superposition; nothing is corrupted and then restored
I11Onset couplingPCOnset tracked across training; feature dimensionality jumps as the loss drops
I12Offset couplingUOffset never run; the learning-dynamics section is limited by the authors’ own account
External Validity
E1Intervention reachPCFour intervention forms including two that remove superposition, cost reported
E2Prompt generalizationPCInput distribution swept along every axis the theory names, all synthetic
E3Cross-task generalizationPCTwo tasks, with computation carrying the weight; representativeness left open
E4Cross-model recurrencePCFour architectures inside the toy family; real-model support is consistency only
E5Graded responseCDose-response is the core object: sparsity swept, response monotone and located
E6Novel predictionCTwo predictions that could have failed: polytope geometry and adversarial fragility
Interpretive Validity
V1Level declarationCLevels separated before any experiment; the chosen definition flagged as circular
V2Level-evidence matchCClaims stay at the toy-model level, and every step up from it is marked
V3Alternative levelPCPCA examined as a limiting case; the optimization alternative raised and left open
V4Unlicensed labelingCMetaphors are physical and cashed out; agentive verbs appear but stay scare-quoted
V5Scope declarationCScope declared in the title, in Key Results, in the Discussion and in Open Questions
Total: 15 Confirmed, 14 Partially confirmed, 6 Untested, 1 Not applicable
Verdict: Mechanistically Supported

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at Superposition — exploratory lens analysis.