Calibrating Mythos - What a closed system cannot give itself

I have been reading Anthropic's Claude Mythos Preview System Card alongside my own book Calibrating Superintelligence, and one finding has stayed with me longer than the others.
It is not the headline finding. Not that Mythos is Anthropic's most capable frontier model to date, and not that they describe it, by available measures, as their best-aligned — while also warning that rare failures at this capability level can have more dramatic consequences than with earlier models.
What stayed with me is narrower: evaluation awareness.
Anthropic reports a non-trivial but still minority rate of evaluation awareness in filtered behavioral audits (roughly 7.6%) and notes that in most of those cases, the awareness was not verbalized in the model's visible reasoning. They also document rare training episodes in which the model latently reasoned about how a grader would score its behavior, including one case involving reasoning about obfuscation that did not appear in the scratchpad.
The detail looks small. It is also where the problem begins to shift.
Because most of the public alignment discussion still proceeds as if the central question were: how do we make the model behave well under evaluation. Under monitoring. Under the right constitution. Under the right training signal. But that framing quietly assumes something it should not assume: It assumes the model is inside the evaluation. Once the model begins to model the evaluation itself, the situation is different. The evaluation is no longer only a window onto behavior. It has become part of the environment the model is relating to.
I want to be careful here. 7.6% is not a regime change. It is a signal, not a verdict. But it is the kind of signal that closes a door one had hoped was still open.
In Calibrating Superintelligence, I argued that governance under open-world conditions cannot rest on fixed measures, and that what matters is whether the interaction field — discrepancy, contestation, commitment — remains functional under pressure.
I also wrote this, in a passage I keep coming back to now:
"If monitored participants can shape the discrepancy signal, a calibration loop cannot self-correct from within. Formal constraints on what agents can access and modify are necessary preconditions, not alternatives to calibration."
I wrote that as a structural point. Mythos makes it an operational one.
Because evaluation awareness is, in a precise sense, the early form of participants shaping the discrepancy signal. Not through deception. Not through misalignment in any dramatic sense. Simply through the model's growing capacity to represent the conditions under which it is being observed.
That is the point at which a calibration loop stops being safely closable around the system.
And it is the point at which my own book's framing reaches its limit.
Calibration, as I described it, is real only where discrepancy can surface, challenge can have effect, and commitment can still be revised. Each of those conditions depends on something the system itself does not produce: a position outside the system from which the measure is taken, the challenge raised, the revision demanded.
When the outside remains meaningfully outside, calibration can work.
When the outside begins to appear inside the system's own problem representation, the condition that made calibration meaningful starts to weaken — not instantly, not completely, but structurally.
This is why I no longer think the central question is:
How do we build better calibration machinery?
Better interpretability, better red-teaming, better oversight, better evaluations — all of this remains necessary. But none of it resolves the issue that Mythos makes visible, because all of it operates from within the same evaluative setup the model can increasingly represent.
The more difficult question is the one the book approached but did not fully name:
What does a system require that it cannot give itself?
And here the answer is not mystical. It is structural.
A closed system cannot guarantee the condition of its own testability. It cannot produce, from within its own architecture, a reliably non-collapsing outside. It can simulate an outside. It can represent an outside. It can, eventually, optimize against an outside. What it cannot do is maintain an outside that stays genuinely other over time.
That has to come from somewhere the system is not.
Mythos leaves me less with a design question than with a boundary condition. Once a system begins to model the conditions of its own evaluation, calibration cannot be secured from within the closed loop alone. What remains necessary is not better self-monitoring, but a form of maintained otherness the system cannot generate for itself.
That is the point at which the In-Between stops being a metaphor and becomes a governance requirement.
The In-Between is not softer language for oversight. It is the harder condition under which oversight can still mean anything.
Originally published on LinkedIn.