Skip to content

Source. Probing Classifiers: Promises, Shortcomings, and Advances (Belinkov, 2022).

Description. A classifier gg is trained on a model’s intermediate output fl(x)f_l(x) to predict a property zz, and its accuracy is reported as evidence that the model has learned information relevant to zz. The audited object is the method rather than one experiment: the anchor is a review that runs no experiments of its own, so the evidence sits in the studies it formalizes, and its Figure 1 separates the nine components a probing result depends on from the ten controls the literature has added to them. Evaluated on: every architecture family the method has been applied to, from static word embeddings through recurrent and recursive networks to transformers, and outward to speech recognition and computer vision. Description mode: representational.

Probing classifiers. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. A probing result licenses that property zz is extractable from the representation by a classifier of the stated capacityProposedNecessity: the method performs no removal of its own, and where removal was performed the studies the anchor reports disagree four ways (I1, E1)
A good probe score shows that the model represents zzDisconfirmedNothing — tested and failed: control tasks attribute most of a nonlinear probe’s accuracy to the probe itself, and even random features decode the property (C4)
A good probe score shows that the model uses zz in producing its outputDisconfirmedNothing — tested and failed: a control dataset holds zz non-discriminative for the original task and the probe recovers it anyway, so decodability does not localize to the task (I4)
Interventions on a probe-identified direction show which features the model usesUnderdeterminedA result that decides the four-way disagreement among the intervention studies of §4.3 (I1, E1). This is the broad scope, under which I2 and E5 are unattempted; under probing as §2 defines it, a read-out with no path writing back in, they are not applicable

Probing classifiers. Verdict: Proposed. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Blocked. I1 necessity — Standard probing performs no removal; removal studies disagree four ways
Mechanistically Supportedinternal (I2, I4); external (E1)Blocked at Causally Suggestive. E1 intervention reach — Probing has no intervention of its own; those applied to it disagree; I4 specificity — A control dataset holds the property constant and probing fails on it
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked at Causally Suggestive. C4 discriminant validity — Two neighbors the instrument must separate: probe memorization, random features; I6 double dissociation — Needs two properties and two behaviors; the framework supplies one of each
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Causally Suggestive. I3 minimality — The unit is a whole intermediate output; no operation removes part of one; I10 rescue reversibility — Every catalogued intervention runs one way, projecting or training out; I11 onset coupling — The original model is taken as given and trained once; I12 offset coupling — The converse needs the training sequence onset coupling also lacks; M1 reliability — A review reports no variance; the anchor states what it requires of studies

Probing classifiers. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityCThe prediction can fail and did (Giulianelli et al., 2018; Elazar et al., 2021)
C2Structural plausibilityPCEvery symbol a probing result depends on is named and typed
C3Convergent validityPCProbe families agree on accuracy, reverse under selectivity (Hewitt & Liang, 2019)
C4Discriminant validityDTwo neighbors the instrument must separate: probe memorization, random features
C5Nomological validityPCTied to mutual information, description length and information gain
C6Complementation validityN/AThe audited object is a method, not a construct with named parts
Measurement Validity
M1ReliabilityUA review reports no variance; the anchor states what it requires of studies
M2Baseline separationCThe criterion the literature solved; three control families formalized
M3StabilityIOrdering holds under accuracy and inverts under selectivity (Hewitt & Liang, 2019)
M4CalibrationDThe shortcomings section opens by denying the measurement has a scale
M5SensitivityPCSkylines and floors exist; no planted property is ever recovered
M6InvariancePCInvariance is what probing measures, so the instrument is the question
M7Selection correctionUThree control families catalogued; the omission from the catalogue is the gap
Internal Validity
I1NecessityIStandard probing performs no removal; removal studies disagree four ways
I2SufficiencyN/AProbing is a map out of the representation, with no path writing back in
I3MinimalityUThe unit is a whole intermediate output; no operation removes part of one
I4SpecificityDA control dataset holds the property constant and probing fails on it
I5Rival mechanism exclusionPCControl tasks, solvable only by memorization, separate cleanly (Hewitt & Liang, 2019)
I6Double dissociationUNeeds two properties and two behaviors; the framework supplies one of each
I7Confound controlPCProbe capacity, information in a random baseline and task difficulty
I8Confounding sensitivityUThe anchor answers confounding with designed controls, never with a bound
I9Epistatic interactionUNon-additivity needs two manipulable components; the object is one output
I10Rescue reversibilityUEvery catalogued intervention runs one way, projecting or training out
I11Onset couplingUThe original model is taken as given and trained once
I12Offset couplingUThe converse needs the training sequence onset coupling also lacks
External Validity
E1Intervention reachIProbing has no intervention of its own; those applied to it disagree
E2Prompt generalizationPCOne line of work reports the same property across several datasets
E3Cross-task generalizationPCTransfers to any property with an annotated dataset, which is the appeal
E4Cross-model recurrenceCPortability is established without qualification, from static embeddings on
E5Graded responseN/ANo intervention, so no strength to dial and no response to grade
E6Novel predictionPCFour design changes followed probing results and then worked
Interpretive Validity
V1Level declarationPCThe formalism fixes which objects a probing result is about
V2Level-evidence matchDThe evidence supports extractability; the claim made is representation
V3Alternative levelCEvery rival reading of a good score is given a measure of its own
V4Unlicensed labelingPC“The model knows z” projects an epistemic relation onto decodability
V5Scope declarationCA scope declaration end to end, conceding limits before any content
Total: 5 Confirmed, 12 Partially confirmed, 3 Inconclusive, 9 Untested, 4 Disconfirmed, 3 Not applicable
Verdict: Proposed

Criterion judgments this audit record leaves contested: each carries an argument on both sides, and the status shipped is the one the record settled on.

CriterionIn tensionWhat would settle it
I6Untested or Not applicableWhether a design supplying one mechanism and one behavior leaves double dissociation unattempted or unaskable

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at Probing Classifiers — exploratory lens analysis.