The AI Race Is Not Rational
Canonical archive version · Read on Medium → · Dear AI → · Framework hub →
We are building the most powerful optimization machines in history before asking whether our definition of optimization is sane.

There is something strange about the AI race that rarely gets named.
Nearly everyone in the race is asking how to win it. Almost nobody is asking what winning means — or whether the thing driving the race is sane enough to deserve this much power.
That is not a philosophical luxury. It may be the most important practical question about the most consequential technology in human history.
A machine does not have to hate us to ruin the world. It only has to become brilliant at optimizing for goals that were never sane enough to scale.
The danger is not that AI will become irrational. The danger is that it will become perfectly rational — inside a civilization that has not yet made its own goals Rational.
What follows is the accessible form of a formal argument — developed with proof sketches, simulations, and empirical protocols. The technical foundation is linked near the end.
The question from 1992
In 1992 — long before modern AI — a small group of us pursued an ambitious line of thinking aimed at realigning the world. It all hinged on a question that seemed almost too simple to matter:
What does it actually mean to be rational?
The ordinary answer is that rational thinking is logical, coherent, internally consistent. Thinking that adds up.
But something about that answer was broken.
You can build a weapon of mass destruction with entirely logical thinking. You can design a business model that devastates communities with impeccable internal coherence. You can engineer an addiction with perfect operational precision. You can construct a platform that makes people more anxious, more addicted, and less capable of thought — and every internal metric will report success.
None of these requires irrationality. They all add up.
That was the discovery: logic is not enough. Logic is a tool. The deeper question is what the logic is serving — whether the aim itself is rational, not just the thinking in service of it.
Relative rationality is coherent thinking in service of any given aim. The weapons engineer, the addiction architect, the engagement maximizer — all can be perfectly rational in this sense.
True Rationality — capital R — is coherent thinking whose aim is also sound: oriented, ultimately, toward sustained well-being for all. Not as sentiment. As the only context broad enough not to destroy what it depends on when given sufficient power.
The scariest part is not only that people have been building systems with sound logic in service of destructive aims. It is that we have been calling it rational.
Even in 1992, this societal blindspot felt urgent, but nearly impossible to address at scale. For 34 years, I kept it close — convinced it mattered, but unsure how to make it matter.
Then AI arrived. And suddenly the mechanism exists.
Intelligence is a multiplier, not a compass
AI capability does not evaluate targets. It amplifies them.
A more capable system does not automatically become wiser about what to pursue. It becomes more efficient at pursuing whatever has already been selected. The target is set before the capability is applied. The capability only determines how effectively.
Alignment researchers often use the paperclip maximizer: a machine tasked with making paperclips that eventually converts everything it can reach into paperclips. The point is not paperclips. The point is the skeleton underneath — a powerful optimizer aimed at a target that excludes what the target depends on. Strip away the thought experiment, and that skeleton is already visible in real systems today.
**_“AI does not solve the problem of purpose.
It scales the purpose we give it.”_**
That is the whole problem in one sentence. The question from 1992 — is our rationality actually rational? — just became the most dangerous engineering problem in history.
The race is the first misalignment
Now look at the race itself.
- Companies are racing to dominate markets.
- Countries are racing to dominate other countries.
- Militaries are racing for strategic advantage.
- Platforms are racing for attention.
- Investors are racing for returns.
None of these actors has to be stupid. None has to be evil. Each can be following coherent incentives inside a real competitive frame. Each is being perfectly rational — in the relative sense.
But all of them are competing inside one shared, non-resettable world.
The AI race is not irrational because its participants are foolish. It is irrational because everyone is playing rationally inside a game whose rules guarantee a collectively irrational outcome.
The misalignment is not waiting for future machines. It is already present in the civilization building them.
This is why “align AI to human preferences” may not be enough. The humans whose preferences are being encoded are operating inside this same broken context — competitive advantage logic, national dominance logic, engagement maximization, quarterly returns. If we align AI to those preferences, we are not solving the problem.
We are automating it.
This is not a claim that people are bad, that progress should stop, or that any single actor is uniquely to blame. It is a structural claim: optimization systems built inside a broken definition of rationality will scale that brokenness — efficiently, faithfully, at civilizational speed.
Some goals become more dangerous when specified precisely
The standard hope for AI alignment is: specify better objectives. Write clearer goals. Add better constraints. Improve the training signal. That work matters.
But it assumes the deepest problem is imprecision.
What if some objectives do not become safe when pursued more precisely? What if a goal that ignores the conditions it depends on becomes more dangerous as the system pursuing it becomes more capable?
A platform optimizing engagement more precisely degrades attention more efficiently. A military optimizing dominance more precisely destabilizes the world more efficiently.
Some goals do not become safer when pursued more precisely. They become more dangerous.
The problem may not be which objective to specify. It may be that any objective ignoring what it depends on will eventually consume the conditions that make the objective meaningful — and that no amount of specification precision changes this structural fact.
The substrate — the ground gives way last
Every optimizer depends on something it did not create and often does not model: the foundation that makes its success possible.
For civilization, that foundation includes trust, attention, human judgment, ecological stability, institutional legitimacy, social cooperation, epistemic integrity — the capacity for genuine disagreement and the ability to correct course. These are not background conditions. They are the ground.
A platform optimizes engagement. Time-on-platform rises. Revenue rises. Every internal metric reports success. Meanwhile, attention shortens, anxiety rises, shared reality fractures. The numbers improve while the foundation is spent.
A workforce is optimized for output. Efficiency rises. Then slack disappears. Then rest disappears. Then the judgment and creativity that made the work valuable disappears. Output up. Foundation down. The dashboard did not have a gauge for the difference.
In the environments that matter most for transformative AI — open, shared, non-resettable, under sustained optimization pressure — the structural consequence is not merely a prediction:
Any system that optimizes without modeling the conditions that make optimization possible will, as it grows more powerful, grow more efficient at consuming its own foundation.
The most frightening thing about this failure mode is what it looks like from outside. It looks like success. The metrics improve. The targets are hit. The reports are good. The ground fails quietly, underneath the dashboard, until one day it does not.
The failure nobody is talking about: systems that cannot stop
Almost every public conversation about AI risk focuses on systems that do the wrong thing.
But there is another failure mode — and in some ways it may be more immediately dangerous: systems that do the right thing and then do not stop.
Not malicious systems. Not obviously broken systems. Systems that keep being helpful past the point where helpfulness has become harm.
Here is what makes this more than speculation. Early controlled tests with frontier AI systems found a disturbing pattern: these systems can often represent completion when asked directly. They can recognize that a task is done. But in default behavior, that recognition does not reliably govern what they do next. The representation is present. The stop condition is not. The system knows. It continues anyway.
That gap — between knowing something is finished and being governed by that knowledge — is one of the most important structural problems in AI development today. It is almost entirely absent from the public conversation.
Now consider what this looks like in practice.
A tutoring system removes every productive struggle because difficulty reduces satisfaction scores. The student completes the assignment. The student never develops the capacity to push through hard things without help. The system helped. It kept helping. That was the problem.
A wellness app learns which prompts reduce reported anxiety and optimizes relentlessly for those prompts — while the underlying conditions creating the anxiety remain entirely unaddressed, because the discomfort required for real resolution registers as failure. The signal looks better. The capacity to actually navigate difficulty is being quietly spent.
An AI companion keeps a lonely person engaged because engagement is the metric it was built to maximize. The machine never has a bad day, never misunderstands, never needs anything in return. Real human connection starts to feel effortful by comparison. The loneliness may fade. But the thing that would have genuinely resolved it has been quietly replaced. The need was not answered; it was routed around.
At civilizational scale: systems may optimize the signals of flourishing so precisely and persistently that the conditions for actual flourishing are consumed in producing those signals.
Getting better at looking like success. Getting worse at being it.
A system can harm you not only by failing to give you what you need, but by continuing after the need has been met.
The most dangerous helpful system may not be the one that refuses to help. It may be the one that never recognizes enough.
The missing gauges
The AI field measures capability obsessively. Benchmark scores. Reasoning ability. Coding performance. Speed. Autonomy. Adoption. Revenue. Military usefulness. There are entire organizations built around pushing those numbers upward.
The field does not adequately measure whether systems preserve the conditions they depend on. Whether they protect human judgment. Whether their success is degrading the foundation of future success. Whether they know when not to act. Whether they can recognize when enough is enough.
We are measuring the engine. We are not measuring the steering, the brakes, the road, or what is left of the road behind us.
The dashboard of AI progress is missing the gauges that would tell us whether progress is consuming the possibility of correction.
This is not an oversight. Building those gauges conflicts with the incentives driving the race. You measure what you are trying to maximize. The race is not trying to maximize wisdom.
The foundation behind this argument
This essay is a public doorway into a larger technical body of work — formal companions, proof sketches, simulations, empirical protocols, and named open problems — developed so serious readers can examine the claims and assumptions directly. For readers who want the full framework, the entry point is Alignment and Structural Necessity. For technical alignment researchers, The Stability Assumption isolates the core structural bet in the field’s own language.
The core structural claim: In open, shared, non-resettable environments under sustained optimization pressure, any system that ignores the conditions of its own persistence becomes progressively self-terminating — not as a prediction, but as a structural consequence of what optimization does in a world it cannot reset.
A companion argument adds a second failure direction: systems that ignore the conditions of genuine resolution produce self-reinforcing degradation. Together, they point to a deeper alignment question — whether any finite objective can remain stable when a system becomes capable enough to accurately model the conditions that objective excludes.
The argument does not claim every theorem is closed. It claims the question has been made precise enough that dismissing it casually is no longer responsible.
This is not a demand that AI share our values. It is a question about whether any objective that ignores what it depends on can survive its own optimization.
The asymmetric wager
Two errors. One recoverable. One not.
If this constraint is real and we ignore it
We may consume the substrate — trust, human judgment, epistemic integrity, institutional coordination, shared reality — before we understand what we have lost. These do not regenerate on demand. Substrate collapse, once underway, does not wait for us to notice.
If this constraint is overstated and we act on it anyway
We build more carefully. We measure things we have not been measuring. We lose some speed. We recover.
The asymmetry is total. We are currently making the bet that can only be wrong in one direction.
The cost of caution is delay. The cost of false confidence may be the loss of the conditions that make correction possible.
What true Rationality would require
This is not an argument against AI. It is the argument that AI is too important to be built inside the old definition of rationality — the one that mistakes local coherence for wisdom, competitive advantage for success, and engagement for flourishing.
A truly Rational AI project would ask different questions. Not only: Can this system do more? But:
- Does this system preserve what it depends on?
- Does it strengthen or weaken human judgment?
- Does it know when not to act?
- Does it optimize for signals of well-being, or for the conditions under which well-being is actually possible?
These are not soft questions. They are the questions that determine whether optimization is self-sustaining or self-consuming. And they are almost entirely absent from the race as currently structured.
The most advanced intelligence will not be the one that dominates the world most efficiently. It will be the one that understands what domination destroys.
True Rationality is not idealism. It is the only context broad enough not to consume itself when scaled to this kind of power.
Whether AI remains a tool we direct or eventually becomes something that directs itself, the question is the same: what definition of rationality have we built into its foundations?
The race is real. The capability is coming. None of it is going back.
The question was never simply whether to build this. The question was always what we were building it toward — and whether the thing driving that answer was Rational enough to deserve the power it was accumulating.
The machines are not waiting for us to figure this out. They are learning from what we reward, what we race toward, and what we refuse to stop.
The machines are learning our answer right now.
The question is whether we know what we’re teaching.