Skip to content

Source. Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task (Li et al., 2023).

Description. A GPT variant trained from random initialization on Othello transcripts carries the board state in its activations. Nonlinear probes decode all 64 tiles where linear probes never dip below 20% error, and 2,000 interventions move the top-NN predictions onto the counterfactual legal-move set (errors 0.12 and 0.06 against null baselines of 2.68 and 2.59), half of them on boards unreachable by legal play. Evaluated on: one 8-layer architecture trained twice, on championship games and on uniformly sampled legal moves, with both runs probed across all eight layers. Description mode: representational.

Othello world model. Readings of the claim, from the version the authors state to the version the field cites.

ReadingVerdictMissing
Primary. Othello-GPT carries a decodable representation of board state that its move predictions causally useCausally SuggestiveNecessity by removal (I1), a second intervention operator (E1), a dissociation (I6), and a known-positive control, which the origin does not contain (M5)
The representation is a model of the process producing the sequences, which is the paper’s own definition of a world modelUnderdeterminedAn experiment separating that from a decodable, causally-used state summary, which every result in the paper also satisfies (V4, V3); the rule composition that would distinguish them is measured once and fails on the OR across lines (V2, I9)
The board state is encoded nonlinearlyUnderdeterminedA perturbation of the probe target, held at black/white/empty throughout and the one configuration the stability sweep never varied (M3, I5)
The account travels beyond Othello and beyond this architectureInsufficientThe origin runs one task on one architecture trained twice, and lists other games and natural language as future work; E3 and E4 are scored on post-origin evidence

Othello world model. Verdict: Causally Suggestive. The claim stops at the rung marked Blocked; rungs below it are Reached and rungs above it are blocked by that rung. Criteria in bold are absent or adverse at that rung; only those the rung requires can block it.

TierRequiresMissing
Proposedconstruct (C1–C2)Reached.
Causally Suggestivemeasurement (M2); internal (I1)Reached.
Mechanistically Supportedinternal (I2, I4); external (E1)Blocked. E1 intervention reach — One intervention operator, reported under three outcome metrics
Triangulatedconstruct (C3–C4); internal (I5–I7); external (E2, E4)Blocked at Mechanistically Supported. I6 double dissociation — One representation, one behavior; no converse leg is available
Validatedconstruct (C5–C6); measurement (M1–M6); internal (I3, I10–I12); external (E2–E6); interpretive (V1–V5)Blocked at Mechanistically Supported. C6 complementation validity — Sixty-four tiles, intervened on one at a time throughout; I11 onset coupling — Training-step axis available and explicitly declined; I12 offset coupling — Two models differ in competence; the coupling is never analyzed

Othello world model. Full 36-criterion audit. Status: C = Confirmed, PC = Partially confirmed, U = Untested, I = Inconclusive, D = Disconfirmed, N/A = Not applicable.

IDCriterionStatusEvidence
Construct Validity
C1FalsifiabilityCIntervention fixes the failing outcome before it is measured
C2Structural plausibilityPCLayerwise argument and probe geometry; no weight-level account
C3Convergent validityCProbe, intervention and three metrics agree; two share one instrument
C4Discriminant validityPCSeparated from memorization and from a random net, not from a summary
C5Nomological validityPCNamed rival theory and a literature; no theory of what a world model is
C6Complementation validityUSixty-four tiles, intervened on one at a time throughout
Measurement Validity
M1ReliabilityCProbe accuracies re-run 100 times; deviations in Tables 4 and 5
M2Baseline separationCUntrained net, constant guess and null intervention floors all reported
M3StabilityPCOptimizer, $$ and probe capacity swept; the probe target never varied
M4CalibrationPCIntervention error has a null floor; probe accuracy has no task scale
M5SensitivityPCKnown-negative at origin; the known-positive is later (Vafa et al., 2024)
M6InvarianceCAll eight layers, both datasets, both benchmarks, three metrics
M7Selection correctionUHeadline numbers are the maximum over swept Ls and $$, uncorrected
Internal Validity
I1NecessityPCSubstitution to a counterfactual state; nothing is removed
I2SufficiencyPCEditing the representation alone moves predictions, including off-distribution
I3MinimalityN/AOne activation vector holds all 64 tiles; no components to prune
I4SpecificityPCOff-target flips raised, mitigation swept, effects never counted
I5Rival mechanism exclusionPCMemorization and off-distribution rivals excluded; basis choice was not
I6Double dissociationUOne representation, one behavior; no converse leg is available
I7Confound controlPCMemorization, distribution and probe capacity controlled; target is not
I8Confounding sensitivityUNo sensitivity analysis, and no quantity that could serve as a bound
I9Epistatic interactionUSingle-tile edits only; rule composition named as future work
I10Rescue reversibilityPCRestoration observed as an obstacle, not designed as a rescue test
I11Onset couplingUTraining-step axis available and explicitly declined
I12Offset couplingUTwo models differ in competence; the coupling is never analyzed
External Validity
E1Intervention reachUOne intervention operator, reported under three outcome metrics
E2Prompt generalizationC1000 natural and 1000 unnatural cases; both training distributions
E3Cross-task generalizationPCOthello only at origin; chess and rule variants come later (Karvonen, 2024)
E4Cross-model recurrenceCAbsent at origin; seven architectures tested later (Yuan & Sogaard, 2025)
E5Graded responsePCTwo knobs swept with degradation outside the optimum; no dose-response
E6Novel predictionCUnnatural boards predicted to work and did; saliency maps split two models
Interpretive Validity
V1Level declarationPCLevel named by the phrase used, and the phrase changes between sections
V2Level-evidence matchPCEvidence matches decodable plus causal; the definition asks for more
V3Alternative levelPCMemorization alternative posed and refuted; state summary never posed
V4Unlicensed labelingPC“World model” is defined, and only a state summary is ever measured
V5Scope declarationCSynthetic scope declared in the title, §1.1 and the conclusion
Total: 9 Confirmed, 18 Partially confirmed, 8 Untested, 1 Not applicable
Verdict: Causally Suggestive

Criterion judgments this audit record leaves contested: each carries an argument on both sides, and the status shipped is the one the record settled on.

CriterionIn tensionWhat would settle it
M5Partially confirmed or UntestedWhether one instrument’s known-positive licenses a sensitivity verdict for another
I6Untested or Not applicableWhether a design supplying one mechanism and one behavior leaves double dissociation unattempted or unaskable

An exploratory reading of this claim through the framework’s five lenses, written for this site and not part of the paper, is at Othello Board State — exploratory lens analysis.