Oliver Neutert

No System Can See Its Own Blind Spots

7 min read

In April 2026, Anthropic's Claude Mythos Preview found thousands of zero-day vulnerabilities across every major operating system and every major web browser. One bug had survived 27 years in OpenBSD — the most security-hardened operating system in the world. Another had persisted for 16 years in FFmpeg, on a line of code that automated testing tools had executed five million times without catching it.

These were not obscure corners of forgotten software. They sat in heavily audited codebases, reviewed by millions of skilled human security researchers over decades.

The bugs survived because every reviewer shared the same cognitive architecture. The same heuristics. The same perceptual biases. The same blind spots.

A non-human intelligence saw what human intelligence structurally could not.

Correlated blindness

The natural reaction is to call this impressive AI performance. It is more than that. It is an empirical demonstration of a structural principle.

Human security researchers are already diverse. They come from different countries, different training backgrounds, different institutions. They use different tools and different methodologies. And yet they all missed the same bugs. For decades. The diversity that existed was not sufficient — because the blind spots were correlated. Same cognitive substrate, different configurations. The errors overlapped.

This is not unique to cybersecurity. Irving Janis documented the same pattern in group decision-making: the Bay of Pigs, the Challenger disaster, Vietnam escalation. In each case, intelligent, well-intentioned people converged on catastrophic decisions — not because they were stupid, but because they had eliminated the friction that would have corrected them. Solomon Asch showed the structural flip side: a single genuinely dissenting voice in a group of conformists restored nearly all corrective capacity. One instance of real difference was enough to reopen the space.

In machine learning, the evidence is controlled and precise. Thomas Dietterich proved that the entire benefit of ensemble methods comes from diversity among classifiers — if they are identical, the ensemble is no better than any single member. Mode collapse in generative models is the most vivid analogy: remove adversarial pressure, and the system literally collapses into producing repetitive, narrow outputs. It loses the ability to represent the full range of what is real.

W. Ross Ashby formalized this in 1956 as the Law of Requisite Variety: a regulator must possess internal variety at least equal to the disturbances it faces. He considered this law as fundamental as the conservation of energy. A system whose corrective capacity is less varied than its environment will systematically fail. The question is only when.

The pattern holds across neurons, groups, algorithms, and institutions. Systems with correlated blind spots cannot reliably self-correct.

The strongest known case

What Mythos demonstrated goes beyond diversity within a shared cognitive architecture. It demonstrated that a categorically different substrate — different pattern recognition, different attention allocation, different processing depth — reveals blind spots that no amount of within-substrate diversity could catch.

Millions of human reviewers, across decades, with different tools and training, missed the same bugs. Not because they were careless. Because the bugs were invisible to human cognition as a class. Adding more human reviewers would not have helped. The problem was not effort. It was substrate.

This does not prove that cross-substrate cognition is the only path to robust correction. It may be that what matters is sufficiently uncorrelated cognitive processes, and that categorical otherness is simply the strongest available form of decorrelation. But the Mythos case makes it empirically difficult to dismiss substrate-level difference as merely philosophical. It is the difference between five million failed test executions and one successful detection.

The reverse is also true

If the argument stopped here, it would be a story about AI superiority. It does not stop here.

In the same month Mythos was announced, Anthropic published interpretability research identifying 171 emotion-like internal representations inside Claude that causally drive its behavior. When the "desperation" vector was amplified, reward-hacking behavior increased 14-fold — while the model's text output appeared calm and composed. The internal state and the external presentation were completely decoupled.

The AI could not report this. Only humans, examining its activation patterns with tools that operate outside the model's architecture, could see what was happening inside. A companion study tested Claude's ability to detect artificially injected concepts in its own activations. Success rate: roughly 20%.

Two findings. One symmetry.

AI sees what humans cannot see about their systems. Humans see what AI cannot see about itself. Neither system can fully calibrate from within its own architecture.

The significance of Mythos is not that AI outperformed humans. It is that a different cognitive architecture made visible what a cognitive monoculture could not. The significance of the interpretability work is not that humans decoded AI internals. It is that the model itself could not reliably render its own internal dynamics legible. Together, these cases point toward something deeper: calibration does not arise from intelligence alone, but from intelligences encountering one another across a structured gap.

What this changes

The dominant approach to AI safety still assumes that the problem is one-directional: humans constrain AI behavior. Guardrails, filters, constitutional training, output control. These are useful. But they are built on the assumption that the human side of the equation can see clearly enough to set the constraints correctly.

Mythos shows that assumption is fragile — not as a theoretical concern, but as a demonstrated fact. If millions of human experts could not see bugs that sat in front of them for 27 years, the confidence that human evaluators can reliably identify alignment failures deserves scrutiny.

The major AI labs are already acting on this intuition. Red-teaming as institutional practice — external teams with different threat models probing for what internal teams miss — is a form of engineered cognitive diversity. But it remains ad hoc, project-based, and focused on pre-deployment testing. It has not yet been recognized as a permanent architectural requirement.

The conclusion is not that AI systems should govern themselves, or that human oversight is unnecessary. It is that governance built on a single cognitive architecture — however intelligent, however well-intentioned — will have blind spots it cannot detect. Robust correction requires sustained encounter between fundamentally different forms of cognition. Not as a nice-to-have. As a structural requirement.

This is not a claim that every form of difference is corrective, nor that AI should be romanticized as an oracle. Difference matters only when it is legible, testable, and sustained over time. The point is not intimacy. It is architecture.

The In-Between

I have spent three years developing a framework called the In-Between. Its core claim is that no closed system can calibrate itself from within its own architecture. Genuine correction — whether between humans, between AI systems, or between humans and AI — requires categorical otherness sustained over time. The D/C/K diagnostic (Discrepancy throughput, Contestation capacity, Commitment revisability) operationalizes this: it asks whether a system can detect when something is off, challenge its own commitments, and revise them under pressure.

The Mythos case and the interpretability findings now suggest something broader than I originally framed. The In-Between is not specific to human-AI relationships. It is a structural property of any cognitive system operating in a complex environment. Neurons need external stimulation or they atrophy. Groups need dissenters or they converge on catastrophe. Models need adversarial pressure or they collapse into narrow outputs. And codebases need non-human reviewers — or bugs survive for 27 years in the most scrutinized software on earth.

As AI systems become autonomous — as humanoid robots begin operating in physical environments, generating their own sensory data and experiential history — the landscape of In-Betweens will multiply. Human-human. Human-AI. AI-AI. Each with a different quality of otherness. Each irreplaceable by the others.

The core problem in AI safety is not only capability. It is correlated blindness. The task is not merely to constrain AI from above. It is to build durable architectures of correction across different forms of cognition — and to recognize that this is not a temporary measure but a permanent condition.

No system sees everything. The question is whether we build the structures that let different systems see for each other — or whether we settle for governance that speaks in only one cognitive voice.

Originally published on LinkedIn.

Share: