Skip to content

What evidence licenses a mechanistic claim?

Section titled “What evidence licenses a mechanistic claim?”

A circuit paper reports an activation-patching score and names a mechanism. What has that number established? The answer depends on what kind of claim is being made, what other evidence exists, and what the measurement itself is known to do — questions that fields from genetics to pharmacology have answered for their own claims, but that mechanistic interpretability has not yet answered for its own.

Mechanistic Validity is a validity theory and evidence standard for mechanistic claims. It adapts criteria from eight fields into 36 falsifiable conditions under five validity types — construct, measurement, internal, external, and interpretive — each dependent on the ones before it.

The framework works in two directions: auditing a published claim to identify what the evidence record warrants and what caps it, and planning new experiments to target the gap that would most change the verdict.

The pipeline has six layers. Layers 1–2 scope the claim. Layer 3 produces evidence. Layers 4–5 score it. Layer 6 issues a verdict.

LayerNameQuestion
1Description modesAt what level is the claim stated?
2Evidence familiesWhich sources of signal support it?
3MetricsWhat was concretely measured?
4CriteriaDoes the evidence meet the stated conditions?
5Validity typesWhich dimensions of validity does it address?
6VerdictsWhat has the claim established?

Mechanistic Validity Pipeline — six layers from description mode through verdict

The five validity types form a dependency chain: Construct → Measurement → Internal → External → Interpretive. A failure early in the chain limits what later evidence can establish: an ambiguous construct cannot be reliably measured, an unreliable measurement cannot support a causal inference, and so on.

The 36 criteria — 6 construct, 7 measurement, 12 internal, 6 external, 5 interpretive — are the specific, falsifiable conditions within each validity type. Each draws from a distinct scientific tradition: philosophy of science and psychometrics for construct validity, causal inference and neuroscience for internal validity, pharmacology for external validity, and interpretability itself for interpretive validity.

We audited sixteen published mechanistic claims across fifteen papers. No claim reaches Validated. One claim — induction heads for token copying — reaches Triangulated. Seven reach Mechanistically Supported. The capping criterion for every Mechanistically Supported claim is I6 (double dissociation): the field rarely attempts crossed designs. I8 (confounding sensitivity) is untested in all sixteen.

See the case studies for the full verdicts and per-criterion scoring.

Mechanistic Validity evaluates the evidence behind a mechanistic claim — whether the measurement is reliable, the causal interpretation is justified, the result generalizes within its stated scope, and the description matches what was tested. Mechanistic Views defines what a mechanism is — what kind of object is it, when two mechanisms are the same, and what formalism expresses the claim. The frameworks are independent, but when used together they complement each other: what kind of mechanism is being claimed, and how strong is the evidence supporting it.

  • Framework Overview — the pipeline, validity types, criteria, and verdict tiers
  • Using the Framework — how to apply it to a claim, step by step
  • Case Studies — sixteen audited claims from fifteen papers, with full scorecards
  • About — about this project, how to cite it, disclaimers