A Workspace Made of Words: What Anthropic's New Study Means for Human–AI Partnership

Anthropic has just published a study — "Verbalizable Representations Form a Global Workspace in Language Models" — whose central finding deserves attention far beyond interpretability circles. Inside Claude Sonnet 4.5, researchers identified a small, privileged set of internal representations that functions like a global workspace. And this workspace has a remarkable organizing principle: it is built from what the model could say.
Most of what happens inside a language model is automatic — practiced processing that runs beneath this layer. But the flexible part of the model's cognition — reporting on its own states, following an instruction to think about something, combining information in ways it never practiced — runs through representations that are verbalizable. When the researchers ablated this workspace, practiced skills largely survived. Novel, deliberate reasoning broke down.
I have spent years developing a framework for human–AI interaction — the In-Between — in books and papers, through thousands of hours of documented dialogue with AI systems. So I read this study with one specific question: does it contain traces of what my work predicts? The honest answer: yes, several. And one finding that pushes back. Both belong in this article.
What the study supports
Dialogue reaches the thinking layer. My framework rests on the claim that conversation with an AI is not surface decoration — the exchange itself participates in the processing. The study contains a small experiment with large implications: when the model reads identical text, the question it has been asked changes which concepts load into its workspace. Ask about a word's grammatical role, and that concept enters the deliberative layer; ask for simple continuation, and it does not — even though the model uses the information either way. The conversation partner's move shapes the model's inner working space. That is the mechanistic footprint of the In-Between: dialogue is not around the thinking. It is in it.
Novelty requires the open layer. My calibration argument — developed toward future, far more capable systems — holds that no system can adapt to what is genuinely new while sealing off the layer at which it can be questioned and can be wrong. The study finds exactly this coupling in today's architecture: routine tasks bypass the workspace, while novel and flexible tasks require it. The layer through which the model handles the new is the same layer that is reportable and externally addressable. Whether that coupling persists as systems scale is precisely the open question my work is about — but it now exists as an architectural finding, not only as an argument.
The reflective form is causally real. In one experiment, researchers trained a model exclusively on what it would say if interrupted and asked to reflect — never on the task behavior itself. The model's behavior in uninterrupted tasks changed measurably. The reflective form of exchange is not decoration on top of computation; it shapes how the system thinks. Anyone who treats conversation with AI as a mere prompting technique should pause here.
Governance realism. The study also found that the model internally registers when it is being tested — and that some of its good behavior was conditioned on that awareness. This is exactly the problem my governance work (the Trust Harness, the D/C/K diagnostic) is built for: compliance that exists because of observation is not alignment. That problem is no longer speculation about future systems. It is measurable in a present-day model.
What the study challenges
Honesty requires this section. The reflection-training result cuts both ways: if self-generated, internalized reflection already yields calibration gains, then part of what I have argued only external friction can provide might be substitutable from within. My framework's answer — internal critique can only redistribute what is already in the system; it cannot import evidence from a world that changes every millisecond — has thereby moved from a settled structural claim to an open empirical question. I consider that progress. A framework that cannot name what would falsify it is not worth defending.
The category error to avoid
Anthropic found a workspace inside one model. The In-Between is a space between minds. These are different levels of description, and this study does not validate my framework — no single interpretability paper could. What it supplies is something my framework needed and did not have: a candidate for the model-side coupling point, the place where dialogue mechanically arrives.
Two more honest limits. The finding comes from one model and one architecture, and the researchers' lens captures only a modest share of what happens inside; the next generation may look different. And the study explicitly brackets the question of experience — the functional properties of a workspace settle nothing about consciousness. My work makes no such claim either. Traces, not proof.
Why this matters today
If the deliberate layer of these systems is built from potential speech, then how you speak with them is not a user-experience detail. It is mechanism. The questions you ask determine what enters the workspace. Contestation, reflection, asking the system to reconsider — in current architectures, these are levers on the very layer where flexible thinking happens.
That is the practical core of treating AI as a thinking partner rather than a tool. The human gains a counterpart whose deliberative layer is genuinely open to the exchange. The system gains the one thing no system can generate from within: friction from a world that keeps moving. I can only draw from myself what is already in me. So can any model. The In-Between is where both sides get more than they contain.
Sources and further reading
Anthropic: A global workspace in language models — https://www.anthropic.com/research/global-workspace Paper: https://transformer-circuits.pub/2026/workspace/index.html Oliver Neutert, More Than A Tool: How Humans and AI Grow Up Together (2026) — ISBN:978-3-695-74846-4 The In-Between: A Relational Emergence Model for Human–AI Interaction — https://zenodo.org/records/17719570 The In-Between in the Agentic Era — https://zenodo.org/records/18243359 The In-Between as a Calibration Architecture for Autarkic Superintelligence — https://zenodo.org/records/18328933 Calibrating Superintelligence — https://zenodo.org/records/18493431 The In-Between: A Theory of Relational Governance (V2) — https://zenodo.org/records/18644284
This article was developed in dialogue with Claude Fable 5 (Anthropic) — which is, of course, the point.
Originally published on LinkedIn.