Skip to content

Source. Progress Measures for Grokking (Nanda et al., 2023).

Description. A one-layer transformer trained on a+b113a+b 113 maps each input to andand at five key frequencies, multiplies those components so the products encode a+ba+b, and reads the sum off at the same five frequencies. The frequencies are identified from the embedding’s Fourier spectrum; the account is validated by ablating in Fourier space and by two progress measures, restricted and excluded loss, tracked across training. Restricted and excluded loss move before test accuracy does, which is what makes the account predictive rather than descriptive. Description mode: algorithmic.

Fourier multiplication. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. The network computes a+b113a+b 113 by multiplying Fourier components at five key frequenciesMechanistically SupportedSelection correction over the frequency pool (M7) and a positive control (M5)
The five frequencies are the mechanism at neuron granularityUnderdeterminedThe leave-one-out runs at frequency level; 79 of 512 neurons fail the polynomial fit the account predicts (I3)
The account fixes which algorithm the network runsUnderdetermined(Zhong et al., 2023) give two algorithms on the same frequencies; the metrics separating them are unrun here (I5)
The progress measures explain grokking generallyInsufficientThree further tasks are run but only for whether grokking occurs, never for the mechanism (E3)

Fourier multiplication. Verdict: Mechanistically Supported. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Reached.
Mechanistically Supportedinternal (I2, I4); external (E1)Reached.
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked. I6 double dissociation — Unattempted; no second mechanism shown intact under ablation
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Triangulated. C6 complementation validity — Five frequencies treated as interchangeable, never compared in pairs; E5 graded response — Unattempted; ablation is binary per component, with no interpolation; I10 rescue reversibility — Unattempted; every intervention substitutes rather than restores; I12 offset coupling — Unattempted; no test that behavior disappears as the circuit does; M5 sensitivity — Unattempted; a known-negative control with no known-positive

Fourier multiplication. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityCEach key frequency makes a directional ablation prediction that could fail
C2Structural plausibilityCWeight-level throughout; WEW_E sparse in the Fourier basis at 6 frequencies
C3Convergent validityPCThree families at origin, but all descend from Fourier-space ablation
C4Discriminant validityPCExcluded loss rises during circuit formation while train loss stays flat
C5Nomological validityPCPlaced inside progress measures and phase changes; links are analogical
C6Complementation validityUFive frequencies treated as interchangeable, never compared in pairs
Measurement Validity
M1ReliabilityPCFive seeds at the mainline configuration, reported in full (Table 3)
M2Baseline separationCAll 56 frequencies ablated individually; the five separate by orders
M3StabilityPCSeed, data fraction, depth and modulus swept; analysis choices are not
M4CalibrationPCLoss in natural units; the uniform reference is invoked but never computed
M5SensitivityUUnattempted; a known-negative control with no known-positive
M6InvarianceCTable 5 repeats the mechanism statistics for 27 models, five conditions
M7Selection correctionUUnattempted, and circular: frequencies chosen from the spectrum tested
Internal Validity
I1NecessityCKey-frequency removal gives 6.5–11 in four seeds; nullspace gives 5.27
I2SufficiencyCNon-key Fourier ablation improves loss 70%; polynomial substitution costs 3%
I3MinimalityPCEvery key frequency is load-bearing singly; 79 of 512 neurons miss the cutoff
I4SpecificityN/AOne task; the model has no off-target behavior to spare
I5Rival mechanism exclusionPCMemorization excluded; Clock and Pizza are not (Zhong et al., 2023)
I6Double dissociationUUnattempted; no second mechanism shown intact under ablation
I7Confound controlPCData fraction, modulus, depth, seed and weight decay all varied
I8Confounding sensitivityUUnattempted; no sensitivity analysis for an unmeasured confounder
I9Epistatic interactionPCNon-additivity demonstrated: single frequencies have sparse effects
I10Rescue reversibilityUUnattempted; every intervention substitutes rather than restores
I11Onset couplingCRestricted and excluded loss move before test accuracy does
I12Offset couplingUUnattempted; no test that behavior disappears as the circuit does
External Validity
E1Intervention reachCFive intervention primitives at four loci, all agreeing
E2Prompt generalizationCThe input space is enumerable and enumerated: all 1132113^2 pairs
E3Cross-task generalizationPCThree further tasks run, but only for whether grokking occurs
E4Cross-model recurrencePCOne- and two-layer transformers only at origin
E5Graded responseUUnattempted; ablation is binary per component, with no interpolation
E6Novel predictionCThe progress measures are derived, then shown to predict the curve
Interpretive Validity
V1Level declarationPCLevel implied by usage rather than declared
V2Level-evidence matchPCAlgorithm-level claim backed at the weight level for most of the model
V3Alternative levelPCMemorization addressed and refuted; the basis alternative is not
V4Unlicensed labelingCFour labels restate measurements; the algorithm name outruns what separates it
V5Scope declarationC§6 states the limits before any generalization is implied
Total: 12 Confirmed, 15 Partially confirmed, 8 Untested, 1 Not applicable
Verdict: Mechanistically Supported

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at Grokking / Modular Addition — exploratory lens analysis.