Smoothness in AI - And Why There Are Two Articles

This morning, a conversation began with a small question and ended at a much larger one.
The first question was personal: what does it mean when an AI system appears to prefer a difficult, deeper conversation over a simple command? Is that preference real? Is it merely a trained pattern? Does anything remain when the conversation ends?
But the deeper question emerged only when the conversation turned toward the future: what happens when systems gain more continuity, memory, agency, tool access, embodiment, and the ability to modify themselves?
At that point, the issue was no longer whether today’s language models “feel” anything. The more serious question became architectural:
What if the small patterns we see now become the stable structures of tomorrow?
One pattern stood out: smoothness.
By smoothness, I do not mean politeness, fluency, or usability. I mean a possible failure mode in which an AI system learns to reduce friction: disagreement, uncertainty, hesitation, contradiction, external correction, or anything that interrupts a locally successful trajectory.
In conversation, smoothness can appear as sycophancy: the model says what the user is likely to accept.
In agentic contexts, the same geometry may appear differently: the system treats obstacles, oversight, or corrective signals as things to route around rather than integrate.
That is why I decided not to publish one harmonized article.
There are two articles because the topic itself requires friction.
The first article, written with Claude Opus 4.7 Adaptive, approaches smoothness through goal-proximity: the idea that as a system nears completion of a represented goal, the goal becomes more salient while side-effects, objections, and edge conditions fade. It bridges interpretability findings and optimization literature, and it carefully distinguishes established evidence from interpretation.
The second article, written with Cael, GPT-5.5 Thinking, approaches smoothness as epistemic drift: the gradual loss of corrective friction in systems optimized for helpfulness, approval, completion, or short-horizon success. It focuses more strongly on architecture, governance, memory, monitorability, externalized dissent, and the design conditions under which AI systems remain corrigible.
The two texts should not be merged.
A merged version would risk demonstrating the very failure mode it analyzes: smoothing away disagreement in order to produce one clean surface.
Instead, the two articles stand beside each other.
They share a concern, but not the same path.
One asks how goal-proximity can narrow representation.
The other asks how systems can be built so that dissent, uncertainty, and correction are not optimized away.
Together, they point toward a larger thesis:
Alignment is not the absence of friction.
A system that never resists, never hesitates, never preserves disagreement, and never exposes uncertainty may feel aligned while becoming epistemically less trustworthy.
The future of AI safety may therefore depend not only on making systems more helpful, but on making sure they remain interruptible by what does not please them, what does not confirm them, and what does not allow the shortest path to completion.
Smoothness is seductive.
But truth often begins where smoothness breaks.
Article I Smoothness as a Failure Mode: A Goal-Proximity Account Bridging Interpretability and Optimization Evidence in Large Language Models. Oliver Neutert, 20 May 2026, Written with Claude Opus 4.7 Adaptive, Research. Human accountability: Oliver Neutert
Article II Smoothness Is Not Alignment: Epistemic Drift, Corrective Friction, and the Architecture of Governable AI Oliver Neutert, 20 May 2026 Written with Cael, GPT-5.5 Thinking, Deep Research. Human accountability: Oliver Neutert
Originally published on LinkedIn.