Skip to content

Each criterion is a specific, falsifiable condition that must be met for a validity type to be satisfied. The 36 criteria are grouped into five validity types following a dependency chain: Construct → Measurement → Internal → External → Interpretive. A failure early in the chain limits what later evidence can establish.

Each criterion receives one of six statuses: Confirmed, Partially confirmed, Inconclusive, Disconfirmed, Untested, or Not applicable.

Construct validity (C1–C6) — Is the target concept well-defined?

Section titled “Construct validity (C1–C6) — Is the target concept well-defined?”
#CriterionOne-line descriptionPage
C1FalsifiabilityCan the claim be refuted?falsifiability
C2Structural plausibilityIs the mechanism physically possible in the architecture?structural-plausibility
C3Convergent validityDo multiple independent methods agree?convergent-validity
C4Discriminant validityDoes the measure distinguish this from neighboring constructs?discriminant-validity
C5Nomological validityDoes the claim fit into a broader theory?nomological-validity
C6Complementation validityAre the construct’s labeled subdivisions functionally distinct?complementation-validity

Measurement validity (M1–M7) — Are the instruments trustworthy?

Section titled “Measurement validity (M1–M7) — Are the instruments trustworthy?”
#CriterionOne-line descriptionPage
M1ReliabilityDo repeated measurements give the same answer?reliability
M2Baseline separationIs the score distinguishable from random/untrained baselines?baseline-separation
M3StabilityIs the classification robust to perturbation?stability
M4CalibrationAre the numbers meaningful?calibration
M5SensitivityCan the instrument detect known-true effects?sensitivity
M6InvarianceDoes the metric behave consistently across conditions?invariance
M7Selection correctionWhen k findings are selected from N candidates, is N reported and multiplicity controlled?selection-correction

Internal validity (I1–I12) — Does the evidence support the causal claim?

Section titled “Internal validity (I1–I12) — Does the evidence support the causal claim?”

The twelve internal criteria fall into five blocks:

  • I1–I3: Properties of the set as a whole (necessity, sufficiency, minimality)
  • I4–I6: Discrimination across tasks, rival circuits, and both (specificity, rival exclusion, double dissociation)
  • I7–I8: Measured and unmeasured confounders
  • I9–I10: Internal structure probes (epistatic interaction, rescue reversibility)
  • I11–I12: Developmental coupling (onset, offset)
#CriterionOne-line descriptionPage
I1NecessityIs the circuit required for the behavior?necessity
I2SufficiencyIs the circuit enough to produce the behavior?sufficiency
I3MinimalityDoes every component earn its place?minimality
I4SpecificityDoes intervening on the circuit affect this task more than matched control tasks?specificity
I5Rival mechanism exclusionIs this the mechanism, or a mechanism?rival-mechanism-exclusion
I6Double dissociationDo two interventions cross, each breaking what the other spares?double-dissociation
I7Confound controlAre alternative explanations ruled out?confound-control
I8Confounding sensitivityHow strong must an unmeasured confounder be to explain the result?confounding-sensitivity
I9Epistatic interactionDo circuit components interact non-additively, and does the direction of interaction distinguish shared pathways from mutual compensation?epistatic-interaction
I10Rescue reversibilityDoes restoring a corrupted component recover behavior?rescue-reversibility
I11Onset couplingDoes the mechanism appear when the capability appears?onset-coupling
I12Offset couplingDoes the mechanism go when the capability is removed?offset-coupling

I6 (double dissociation) caps every claim that reaches Mechanistically Supported in the sixteen audited case studies. I8 (confounding sensitivity) is untested in all sixteen. No claim reaches Validated.

External validity (E1–E6) — Does the mechanism generalize?

Section titled “External validity (E1–E6) — Does the mechanism generalize?”
#CriterionOne-line descriptionPage
E1Intervention reachHas the result been reproduced under at least two intervention families, and do they agree?intervention-reach
E2Prompt generalizationDoes it work on diverse prompts?prompt-generalization
E3Cross-task generalizationDoes the mechanism transfer to related tasks?cross-task-generalization
E4Cross-model generalizationDoes the mechanism appear in other models?cross-model-generalization
E5Graded responseDoes partial ablation produce partial effects?graded-response
E6Novel predictionDoes the mechanism predict new, untested behaviors?novel-prediction

Interpretive validity (V1–V5) — Is the interpretation correct?

Section titled “Interpretive validity (V1–V5) — Is the interpretation correct?”
#CriterionOne-line descriptionPage
V1Level declarationAt what description mode is the claim made?level-declaration
V2Level-evidence matchDoes the evidence support claims at that level?level-evidence-match
V3Alternative levelCould the evidence be explained at a different level?alternative-level
V4Unlicensed labelingDoes a name import a property that was not measured?unlicensed-labeling
V5Scope declarationWhat does the claim explicitly not cover?scope-declaration

V3, V4, and V5 have no counterpart in the validity frameworks surveyed from other fields. They address failure modes specific to mechanistic interpretability: claiming an algorithm when only an implementation was shown (V3), calling a representation a “world model” when only a state summary was demonstrated (V4), and silently generalizing beyond the tested system (V5).