Criterion V4 — Unlicensed Labeling
Section titled “Criterion V4 — Unlicensed Labeling”| Validity type | Interpretive |
| Pass condition | The name assigned to a component does not import a property beyond what was measured |
| Evidence family | N/A (criterion is about naming, not experiments) |
| Minimum reporting | For each named component: the measurement that motivated the name, and the property the name additionally implies |
| Common failure mode | Naming a head after one measurement, then treating the name’s connotations as an established finding |
What this criterion requires
Section titled “What this criterion requires”A name is a claim. Calling a head a “name mover” after observing its direct logit attribution to name tokens implies an intentional function — that the head’s job is to move names — when the measurement shows only that the head’s output correlates with the logit. Calling a probed representation a “world model” imports a construct with commitments (structure, use in planning, counterfactual reasoning) that were never contrasted against a simpler alternative such as “state summary.”
Satisfied when:
- Names are traced to measurements. Each label is linked to the specific result that motivated it.
- The import is stated. The study names what the label implies beyond that result.
- The gap is closed or acknowledged. Additional evidence closes it, or the verdict states the label exceeds what was measured.
MI example
Section titled “MI example”Othello: “world model” is never contrasted against “causally-used state summary” — the weaker label the probing and patching evidence directly supports. Knowledge neurons: calling the causal effect “storage” imports a retrieval mechanism — the network looks a fact up from where it is kept — that was never demonstrated; the evidence shows a causal effect of a component on an output, not a storage-and-retrieval architecture.
The failure mode is interpretive inflation: the label does more work than the evidence licenses, and once published, the name is what gets cited rather than the measurement behind it.
Relation to other interpretive criteria
Section titled “Relation to other interpretive criteria”Unlicensed labeling (V4) is where an alternative-level (V3) failure gets fixed into prose. A name chosen before the alternative-level check is run tends to encode the higher-mode reading by default. V4 also bounds scope declaration (V5): a label that imports unmeasured properties makes every downstream scope statement about the claim too broad, because the scope ends up stated in the label’s terms rather than the measurement’s.