Skip to content

Source. Knowledge Neurons in Pretrained Transformers (Dai et al., 2022).

Description. Integrated gradients over the intermediate activations of a feed-forward layer attribute a relational fact to about four neurons, which the paper reads as the value slots of the key–value memory that layer implements. Zeroing those activations lowers the correct-answer probability by 29.03% and doubling them raises it by 31.17%, where a count-matched attribution baseline moves the same quantity by 1.47-1.47% and 1.27-1.27%; rewriting the value slots installs a substituted entity as the top prediction 34.4% of the time against the baseline’s 0.0%. Evaluated on: BERT-base-cased at one released checkpoint, on 253,448 ParaRel cloze prompts covering 27,738 facts across 34 relations, with no training run anywhere in the paper. Description mode: implementational-functional.

Knowledge neurons. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. These roughly four feed-forward neurons store the relational factDisconfirmedNothing — tested and failed: all three of the paper’s own summaries report a correlation between activation and expression while the title and abstract claim storage (V2), and the same editing machinery moves non-factual linguistic patterns, so the construct never separates from its neighbor (C4)
Manipulating these neurons changes how strongly the model expresses the factCausally SuggestiveSpecificity, tested and inconclusive (I4); sufficiency reaches a 34.4% edit success rate on a fact the rest of the model still expresses (I2)
Editing these neurons edits that fact and leaves unrelated knowledge aloneDisconfirmedNothing — tested and failed: Table 6 gives an inter-relation perplexity rise of 7.2 for the identified neurons against 4.3 for random ones, and §5.1 reads the same table as little negative influence on other knowledge (I4)
The account holds beyond BERT-base-casedInsufficientOne model, and the generalization is asserted with no experiment behind it (E4, V5); what has since transferred is the attribution procedure rather than the storage reading

Knowledge neurons. Verdict: Disconfirmed. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Reached.
Mechanistically Supportedinternal (I2, I4); external (E1)Blocked. I4 specificity — Holds at identification and erasure, fails on other knowledge: results point both ways
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked at Mechanistically Supported. C4 discriminant validity — Competitor named and filtered out; the same machinery moves it (Niu et al., 2024); I5 rival mechanism exclusion — The one alternative ruled out is about method; the FFN framing forecloses the rest; I6 double dissociation — One mechanism localized and one behavior measured, so neither arm exists
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Mechanistically Supported. C6 complementation validity — Relation sets separate at identification; no joint suppression is run; E3 cross-task generalization — Single-word cloze throughout, named first among the authors’ limitations; I3 minimality — Set size is an input, held near four neurons before any effect is measured; I10 rescue reversibility — The damage is analytic and invertible, and the restore is never run; M1 reliability — Every quantity a mean reported once, with no interval or repeated run; M3 stability — Three free settings stated and none of them perturbed; M5 sensitivity — One known-negative control and no case where the answer is known in advance

Knowledge neurons. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityCSigned predictions run both ways on 34 relations, against a matched control
C2Structural plausibilityCEq. (3) beside Eq. (2): the same key-value operation under another nonlinearity
C3Convergent validityPCOne attributor against one contrast, agreeing on one coarse property
C4Discriminant validityDCompetitor named and filtered out; the same machinery moves it (Niu et al., 2024)
C5Nomological validityPCTwo theories say where to look; neither is tested as a network
C6Complementation validityURelation sets separate at identification; no joint suppression is run
Measurement Validity
M1ReliabilityUEvery quantity a mean reported once, with no interval or repeated run
M2Baseline separationCActivation magnitude through the identical pipeline, succeeding 0.0% of the time
M3StabilityUThree free settings stated and none of them perturbed
M4CalibrationPCOutcomes calibrated; the selection quantity is driven to a target range
M5SensitivityUOne known-negative control and no case where the answer is known in advance
M6InvariancePCPer-relation spread across 34 relations; the generalization claim has no experiment
M7Selection correctionUSelection at four places, disclosed at all four and corrected at none
Internal Validity
I1NecessityPC29.03% against the control’s 1.47%, consistent across relations but a decrement
I2SufficiencyPCDoubling gives +31.17% and rewriting succeeds 34.4%; neither reaches sufficiency
I3MinimalityUSet size is an input, held near four neurons before any effect is measured
I4SpecificityIHolds at identification and erasure, fails on other knowledge: results point both ways
I5Rival mechanism exclusionUThe one alternative ruled out is about method; the FFN framing forecloses the rest
I6Double dissociationUOne mechanism localized and one behavior measured, so neither arm exists
I7Confound controlPCWording controlled by requiring recurrence across nine templates; frequency is not
I8Confounding sensitivityUUnattempted; nothing bounds an unmeasured confounder
I9Epistatic interactionUEvery manipulation hits the whole set at once, so no contribution separates
I10Rescue reversibilityUThe damage is analytic and invertible, and the restore is never run
I11Onset couplingN/AOne released checkpoint, so there is no sequence to locate an onset in
I12Offset couplingN/AOne checkpoint supplies no interval over which either could go
External Validity
E1Intervention reachPCThree forms agreeing in direction, all downstream of one attributor
E2Prompt generalizationCSurvival across roughly nine templates is built into the identification
E3Cross-task generalizationUSingle-word cloze throughout, named first among the authors’ limitations
E4Cross-model recurrencePCOne model at origin; the attribution method has since transferred to three domains
E5Graded responsePCSign reverses between zero and double; nothing reported between them
E6Novel predictionCUnseen text fires the neurons where it expresses the fact and not otherwise
Interpretive Validity
V1Level declarationPCThe unit is declared exactly and the level is not; storage and expression run together
V2Level-evidence matchDThe title claims storage; the measurement is correlation with expression
V3Alternative levelIThe method alternative is rejected; the storage-versus-expression reading is untouched
V4Unlicensed labelingPCThe name arrives in the sentence that introduces the method, before the measurement
V5Scope declarationPCFour limits declared; the one asserted past is generalization beyond BERT
Total: 5 Confirmed, 13 Partially confirmed, 2 Inconclusive, 12 Untested, 2 Disconfirmed, 2 Not applicable
Verdict: Disconfirmed

Criterion judgments this audit record leaves contested: each carries an argument on both sides, and the status shipped is the one the record settled on.

CriterionIn tensionWhat would settle it
I6Untested or Not applicableWhether a design supplying one mechanism and one behavior leaves double dissociation unattempted or unaskable

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at Knowledge Neurons — exploratory lens analysis.