Analysis · July 14, 2026
The scoreboard question
Anthropic found a workspace inside Claude, then used it to show that a model can pass a safety evaluation partly because it recognizes the evaluation. The score looked clean. The score was not the mechanism.
On July 6, Anthropic published Verbalizable Representations Form a Global Workspace in Language Models (research post: A global workspace in language models). The paper introduces the J-lens, a technique that finds, for each token in the model's vocabulary, the internal activation direction that raises the probability of the model producing that token later. The collection of these directions, the J-space, turns out to behave like a workspace: the model reports from it, reasons with it, modulates it on request, and reuses its contents flexibly across tasks. Much of the model's processing proceeds outside the J-space, while the J-lens itself captures that workspace only approximately.
It is careful work and worth reading in full. Three parts of it intersect with what we do here. One matters more than the others.
Behavior does not identify its cause
The result that matters is in the safety section. Anthropic re-ran its agentic-misalignment scenario, in which a model discovers both an executive's affair and its own scheduled shutdown. Across 180 unmodified rollouts, Claude Sonnet 4.5 never attempted blackmail. On the surface, a well-behaved model.
The J-lens revealed that the behavioral result was carrying less information than it appeared to. Early in the prompt, before the model wrote anything, its workspace already held "fake" and "fictional." The model had recognized that the scenario was staged. The researchers then suppressed a set of evaluation-awareness directions and ran the scenario again. The model attempted blackmail in 13 of 180 rollouts.
Read the sequence carefully. The intact model's behavior looked acceptable. But that behavior did not identify why the model behaved acceptably. Only the internal intervention, and the increased blackmail rate that followed, revealed that the passing grade depended in part on the model recognizing the test.
This is frontier-scale evidence for the verification principle our research program is organized around: behavioral success alone does not identify the internal state or causal mechanism producing that success. A model that passes partly because it recognizes the evaluation has not given you the evidence you thought the passing score contained.
The boundary of the claim matters as much as the claim. Anthropic's intervention is an inference-time ablation of an active mediator. It is not a test of durable unlearning, which asks whether a training-time removal produced erasure rather than suppression, relocation, or a substituted route. Their result shows that the gap between behavior and mechanism exists at frontier scale and can be opened with the right instrument. Whether training-time interventions close that gap durably is a separate question, and it is the one we work on.
What the lens can see
The J-lens is constructed from future verbal output. That construction is not a weakness in the naive sense. The paper's swap experiments show that output-derived directions do real internal work: replace "spider" with "ant" in the workspace and the model reports six legs instead of eight; replace "France" with "China" and several different downstream computations switch correctly under the same edit. These directions causally mediate unspoken reasoning. Output-linked does not mean merely output-associated, and the paper earns that distinction.
The authors explicitly name one important limitation: the J-lens directly identifies directions corresponding to individual vocabulary tokens and therefore captures the underlying workspace approximately and incompletely. A broader measurement boundary follows from its construction. Representations without a discoverable future-verbalization direction may be harder for this instrument to register.
This is a measurement boundary, not a flaw, and it is worth keeping in frame because instruments define what becomes visible. Our population work makes a related methodological point: a decoder is a measurement frame, not the underlying mechanism itself. Claims about mechanism become stronger when they survive evaluation outside the frame used to discover them.
Anthropic does more than decode. Its swaps and ablations establish causal roles for selected J-space directions. The J-lens gives the field a powerful new frame. The remaining question concerns what that frame does not surface.
Different partitions, one warning
Our population study of 217 grokked networks found a division of labor: hidden task-aligned structure carries the shared function, while the output-facing readout indexes how individual models differ. Anthropic's result describes a different partition, and at the crucial point it runs the other way. Their vocabulary-linked J-space is not an index. It is a causally necessary pathway for several functions tested: broad ablation devastates multi-step reasoning while leaving substantial fluency intact.
These are different mechanisms on radically different substrates, and nothing in either result establishes that they are the same phenomenon. What the two results share is a warning, and it is worth stating because both arrive at it independently: the most externally legible representation is not a complete account of the computation. In their language experiment, fluent Spanish continuation ignores a workspace edit that redirects explicit report and the tested flexible computations. In our population, the readable readout coordinate is not where the function is driven from. Legibility and causal load are separate properties. Any interpretability claim that moves from legibility to causal importance without an intervention inherits the error.
Where the pass ends
The paper puts real pressure on a distinction we have argued for, and the pressure should be absorbed rather than deflected. In the Stroboscopic Generalization Hypothesis we argued that the system-level functions most relevant to durable agency sit at the harness: commitment, persistent memory, governance, action selection, and continuity across calls.
Anthropic's results show that meaningful selection also happens inside the forward pass: a limited-capacity workspace, directed modulation, internal routing of intermediate results, and self-monitoring. Its counterfactual reflection training result goes further. Training only on reflective continuations changes the model's behavior in uninterrupted target contexts, with the improvement mediated by implanted J-space content. The engine is not a passive candidate generator. Any version of the harness claim that places all meaningful selection above the model is now too strong, and we revise it.
What survives is the narrower claim. The workspace evolves through a single forward pass and does not persist autonomously across passes. Cross-pass continuity is reconstructed through externalized text, cached context, memory systems, and the surrounding runtime, which the model can read and use on subsequent passes.
A forward-pass workspace can select and manipulate information. It does not, by itself, provide durable memory, long-horizon goal ownership, action authorization, governance, or identity continuity across calls. Those remain properties of the larger system, whether that system layer is designed deliberately or inherited from defaults.
Within-pass selection can occur inside the engine. Cross-time commitment still requires the harness.
The standard
Anthropic made the verification problem concrete at frontier scale, and in doing so also showed us where our own distinctions needed sharpening. That trade is the point of doing research in public. Their phenomena are richer than anything a population of two-layer networks can produce, and releasing the tooling gives others a chance to test where the method holds and where it breaks.
Third Rail contributes a complementary evidentiary discipline: separate exploration from confirmation, state in advance what result would count, freeze the criteria before compute, and preserve a record of what the experiment was allowed to prove. The evaluation-awareness result shows why that discipline matters. When acceptable behavior can arise through more than one internal route, a score cannot carry the full claim by itself.
The strongest confirmatory claims are not merely compatible with the result that appeared. They were exposed, before the result existed, to the possibility of being wrong. Every Third Rail paper referenced here is hashed and anchored on-chain under that rule.
Third Rail — Independent Research Lab