Skip to content
LabelDiagnostic (replaces the tier)
What it meansEvidence actively contradicts the claimed mechanism — a specific prediction has failed or the finding is shown to be artifactual
When to assignA prediction of the mechanism has been tested and refuted, OR the mechanism is demonstrated to be a measurement artifact
Relationship to progressive tiersAny claim at any progressive tier can be moved to Disconfirmed when contradicting evidence emerges
Scientific valueHigh — disconfirmation narrows the hypothesis space and is informative

Disconfirmed is not “failed research.” It is a positive scientific conclusion: the evidence actively contradicts the mechanistic claim. A field that never disconfirms is not doing science. The lateral position (rather than placing it below Proposed) reflects this: disconfirmation is a different kind of conclusion, not a worse one.

Disconfirmation can take three forms: prediction failure (the mechanism predicts X, the model does not-X), artifact demonstration (the finding disappears under improved methodology), or construct dissolution (the named entity is not a coherent construct separable from other processing).

Each form is informative. Prediction failure narrows the space of viable mechanisms. Artifact demonstration improves methodology for the whole field. Construct dissolution reveals that the question was ill-posed, redirecting inquiry.

Verdict: Disconfirmed — [implementational-topographic] Claim: The IOI circuit is sufficient for indirect object identification under distribution-respecting ablation. Disconfirming evidence: Miller et al. (2024) demonstrated that sufficiency (R=0.87R = 0.87 under mean ablation) drops to R<0.50R < 0.50 under resample ablation. The original sufficiency claim is an artifact of mean ablation’s distributional assumptions. Type: Artifact demonstration — the finding is method-conditional, not mechanism-intrinsic. Remaining valid claims: Necessity of the circuit components remains established. Sufficiency under mean ablation remains a true statement (with method qualification). Scope: GPT-2 Small, IOI task, sufficiency specifically (not the full circuit claim)

TypeDefinitionExample
Prediction failureMechanism predicts behavior XX; model produces ¬X\neg XA claimed “gender circuit” predicts male bias; model shows no gender preference on the test distribution
Artifact demonstrationFinding disappears under improved methodologyPatching result vanishes when mean ablation is replaced by resample ablation
Construct dissolutionNamed entity is not separable from other processing”The bias circuit” is indistinguishable from “the gender knowledge circuit” — the construct has no independent existence
  • The original claim stated precisely (what was predicted)
  • The disconfirming evidence (what was observed instead)
  • The type of disconfirmation (prediction failure, artifact, or dissolution)
  • What remains valid from the original work (disconfirmation is usually partial)
  • Whether the disconfirmation is total (mechanism is wrong) or scoped (mechanism is method-conditional or distribution-limited)
TransitionMeaning
Any tier → DisconfirmedNew evidence contradicts the claim
Disconfirmed → Proposed (rare)The disconfirming evidence is itself shown to be flawed; the original claim is reopened
Disconfirmed → refined claim at Tier 1+The original claim is revised to accommodate the disconfirming evidence — the revised claim is a new entity

Two of the sixteen audited claims reach this tier, both on evidence that was collected rather than missing.

  • Knowledge neurons (Dai et al., 2022) — the claim that roughly four feed-forward neurons store a relational fact. All three of the paper’s own summaries report a correlation between activation and expression while the title claims storage (V2), and the same editing machinery moves non-factual linguistic patterns, so the construct never separates from its neighbor (C4).
  • Induction heads as the source of general in-context learning (Olsson et al., 2022) — the broad reading. No ablation runs above the twelve small models, so nothing at scale separates induction heads from whatever else forms alongside them (E4, I5), and the adopted measure does not separate general in-context learning from the few-shot accuracy the field reads it as (C4). The narrow reading, that induction heads implement prefix-matching token copying, reaches Triangulated.

Individual criteria are Disconfirmed more often than whole claims. IOI’s stability (M3) and invariance (M6) are both Disconfirmed by Miller et al. (2024), who move the same faithfulness quantity from below 0% to over 100% across six methodological choices, and the greater-than circuit’s cross-model recurrence (E4) is Disconfirmed post-origin — without either claim as a whole reaching this tier.