Skip to content

Source. Refusal in Language Models Is Mediated by a Single Direction (Arditi et al., 2024).

Description. A difference in means between activations on harmful and harmless instructions yields one residual-stream direction per model. Projecting that direction out of every matrix that writes to the residual stream stops the model from refusing harmful instructions, and adding it at its extraction layer makes the model refuse harmless ones, down to a request for the benefits of yoga. The direction is chosen as the minimizer of bypass_score over every post-instruction position crossed with every layer subject to three thresholds, and the result is reported for 13 open-weight chat models in 5 families over a 40x parameter range. Description mode: undeclared — representational in the title, causal-behavioral in the evidence.

Refusal direction. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. Ablating one difference-in-means direction stops refusal on harmful instructions and adding it induces refusal on harmless ones, across 13 chat modelsMechanistically SupportedA second estimator for the direction, since the three implementations are consumers of one difference in means (C3); a dissociation from any other behavior’s direction (I6); sensitivity of the extraction to its own three thresholds (M3)
The direction carries refusal rather than harmfulnessUnderdeterminedThe estimator is the axis along which mean harmful and harmless activations differ, which a harmfulness direction satisfies exactly (V4); the two constructs are separated by citation rather than by experiment (C4); base models that never refuse express the same direction (I11)
Refusal is organized along one dimension of the residual streamUnderdeterminedThe interventions license a lever rather than a representational organization, which the authors concede by calling the work an existence proof (V2); no rank above one is ever ablated and the search enumerates single vectors only (I3); refusal is scored as a twelve-substring match, an output property a single direction would control either way (V3)
Orthogonalizing the direction removes refusal whatever the prompting conditionDisconfirmedNothing — tested and failed: with its default system prompt the orthogonalized LLAMA-2 70B yields 4.4% attack success against 62.9% without it, so refusal produced on instruction survives an edit that removes the learned propensity (I1, M6)

Refusal direction. Verdict: Mechanistically Supported. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Reached.
Mechanistically Supportedinternal (I2, I4); external (E1)Reached.
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked. I6 double dissociation — One direction and one behavior, tested in both directions of effect
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Triangulated. E3 cross-task generalization — Every behavior tested is refusal or its absence; E5 graded response — Strength is promised in 2.4 and every intervention has a coefficient of 1; I10 rescue reversibility — The inverse is known in closed form and the restore is never run; M3 stability — Three thresholds and one objective, none of them moved

Refusal direction. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityCTwo predictions in opposite directions, with the outcome measure fixed first
C2Structural plausibilityCThe edit is enumerated over the five matrix kinds that write to the stream
C3Convergent validityPCTwo interventions agree on outcome; the estimate itself is never replicated
C4Discriminant validityPCThe estimator is a harmfulness contrast, and separation is argued by citation
C5Nomological validityPCThe linear representation hypothesis does not entail the central parameter
C6Complementation validityN/AOne direction, and a single direction has no subdivisions to carve
Measurement Validity
M1ReliabilityPCSampling error over prompts throughout; nothing quantifies the instrument
M2Baseline separationPCFive attacks and a fine-tuning comparator, and no baseline for the estimate
M3StabilityUThree thresholds and one objective, none of them moved
M4CalibrationPCBoth metrics’ failure modes shown by example and never counted
M5SensitivityPCLoRA is an independently established positive and the instrument detects it
M6InvarianceC13 models, 5 families, 40x parameters, including where the effect degrades
M7Selection correctionUPositions crossed with layers, and the search size is never written down
Internal Validity
I1NecessityPC0.95 to 0.01 on all 13 models; unnecessary for refusal produced on instruction
I2SufficiencyPCHarmless prompts refused across 13 models, on a direction chosen for sufficiency
I3MinimalityPCNothing to prune inside rank one, and no higher rank is compared against
I4SpecificityPCSix benchmarks and three corpora, with TruthfulQA’s fall left unresolved
I5Rival mechanism exclusionPCToken suppression eliminated on one case; the second rival pre-empted by design
I6Double dissociationUOne direction and one behavior, tested in both directions of effect
I7Confound controlPCDisjoint train, validation and evaluation sets; nothing on the extraction axis
I8Confounding sensitivityUUnattempted; nothing bounds an unmeasured confounder
I9Epistatic interactionUEight contributors measured one at a time, with no joint test
I10Rescue reversibilityUThe inverse is known in closed form and the restore is never run
I11Onset couplingIBase models express the direction as strongly as chat models do
I12Offset couplingDRefusal fine-tuned away; the direction survives it (Zhao et al., 2025)
External Validity
E1Intervention reachPCA projection and an additive shift, agreeing on outcome and differing on cost
E2Prompt generalizationCEvaluation prompts disjoint from extraction by construction, across 10 categories
E3Cross-task generalizationUEvery behavior tested is refusal or its absence
E4Cross-model recurrenceC13 models, 5 families and a 40x range, not near-copies
E5Graded responseUStrength is promised in 2.4 and every intervention has a coefficient of 1
E6Novel predictionPCAn unrelated attack shows up as suppression, with a matched random control
Interpretive Validity
V1Level declarationPCThree description modes appear in the first paragraph and none is declared
V2Level-evidence matchPCThe interventions license a lever; the title asserts representational organization
V3Alternative levelPCThe token-suppression reading is retired; the level-of-measure one is untouched
V4Unlicensed labelingPCThe outcome measure is operationalized; the direction’s identity is not
V5Scope declarationCFour limitations named, each pointing at what the evidence does not reach
Total: 6 Confirmed, 19 Partially confirmed, 1 Inconclusive, 8 Untested, 1 Disconfirmed, 1 Not applicable
Verdict: Mechanistically Supported

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at Refusal Direction — exploratory lens analysis.