We’re Winning the Wrong Race
Canonical archive version · Read on Medium → · Dear AI → · Series 1 → · Stability Assumption →
The most dangerous thing about AI isn’t that it might become irrational. It’s that it’s becoming brilliant at serving goals that were never examined.

This essay is the public entry point to a larger technical framework on AI alignment — including formal companions, proof-program sketches, open problems, and interactive simulations. Readers who want the full argument can begin with the series introduction, or with the entry essay for the AI research community. What follows is the doorway.
I.
The Race You Can Already Feel
You already know something is wrong.
Not wrong in a way you can name precisely — more like a frequency you can hear but can’t locate. The world is accelerating. The people building what comes next are intelligent, well-funded, and somehow still not quite steering. The proposed answers — more innovation, more regulation, more safety teams, more competition — don’t quite touch the thing that feels off. The problem seems to be prior to all of them.
That feeling is not paranoia. It is pattern recognition.
Labs are raising tens of billions. Governments are watching rival governments. Defense agencies are competing for the same capabilities. Entire professions are wondering how many years they have. Children are growing up inside systems shaped less by wisdom than by the need to win before someone else does.
Everyone has a reason to keep going. No one can step out. And almost no one has stopped to ask whether the race itself is pointed somewhere worth going.
II.
A Question From 1992
In 1992, while I was in college, a few friends and I built an academic contract around a question that sounded almost too simple to matter: what does it actually mean to be rational?
We call someone rational when their reasoning is logical. But logical reasoning can serve catastrophic ends. A nuclear weapons program can be internally coherent. A corporation can destroy an ecosystem with rigorous efficiency. A platform can addict children through beautifully optimized design. An addict reasons coherently toward the next hit. The logic adds up. The whole thing is insane.
The question sharpened: Rational inside what context?
The deeper we pushed, the more every pursuit — even dark or distorted ones — appeared to be an attempt to reach a preferred state. Which meant well-being wasn’t one value among others. It was the underlying target of all of them.
Ordinary rationality means coherent pursuit of whatever goal you happen to have. But if the goal itself is broken, better reasoning only makes the brokenness more efficient. True Rationality — capital R — means coherent thought aimed at the deepest sustainable form of well-being available: not just for the thinker, not just for the company or nation, but for the whole field of beings affected by the action.
We wrote it down. A second academic quarter followed, focused entirely on the harder question — what it would take to change how people think at scale. We reached clearer answers than we expected; implementing them seemed nearly impossible. So we moved on with our lives.
The insight sat quietly for 34 years. Then humanity began building something that could take whatever answer we gave — and make it permanent.
III.
Intelligence Is a Multiplier, Not a Conscience
AI does not make a goal wise. It makes the pursuit of a goal more powerful.
- If the goal is profit, AI accelerates profit-seeking.
- If the goal is military advantage, AI makes that competition faster and more lethal.
- If the goal is engagement, AI learns to capture and hold attention with surgical precision.
- If the goal is dominance, AI makes dominance more efficient.
The danger isn’t that AI becomes irrational. The danger is that it becomes brilliant inside a broken definition of winning.
The goals currently driving AI development are overwhelmingly versions of competitive advantage — companies against companies, nations against nations, labs against labs. Each actor can explain their reasoning. Each decision makes sense inside its local frame.
Locally rational behavior inside a broken context doesn’t produce a rational outcome. It produces a faster, more efficient version of the broken context.
IV.
The Trap Has No Villain — And No Exit
Here is what makes this genuinely frightening: it doesn’t require anyone to be malicious.
- A lab that slows down may lose ground to one that doesn’t.
- A country that pauses fears the country that won’t.
- A company that refuses deployment watches competitors capture the market.
- A cautious engineer can be replaced by one who isn’t.
- An investor who demands restraint may fund the competitor instead.
Nobody has to be malicious for the system to be misaligned. Everyone only has to keep doing the locally rational thing.
This is relative rationality at civilizational scale — intelligent, well-intentioned decisions adding up to collective insanity, not because anyone chose it, but because the goal-context driving the whole system was never examined.
The trap doesn’t require a villain. It only requires everyone to keep making sense.
V.
Every Goal Eats the World It Does Not Track
Every goal depends on a world it didn’t create and isn’t modeling.
Profit depends on trust, law, labor, ecological stability, customers who are functional human beings. Military advantage depends on a world not permanently destabilized by the race for it. Engagement depends on human attention — the kind that takes years to develop and can be depleted in ways that don’t appear in engagement metrics. AI development itself depends on epistemic integrity, institutional coherence, public trust, and the distributed human capacity to notice when something has gone wrong and correct it.
A narrow optimizer treats all of this as background. Fixed. Free. Not its problem.
But when optimization scales, the background becomes fuel.
It’s like heating a house by burning the floorboards. For a while the room gets warmer. The metric improves. The system appears to be working. Then the structure gives way.
Any goal that ignores the world it depends on eventually starts treating that world as fuel.
If your attention feels harder to hold than it used to, if outrage arrives faster than wonder, if something in your inner life feels subtly more depleted than it did a decade ago — that may not be a private failure. It may be what optimization at scale does to the human substrate when the substrate isn’t being tracked.
AI is not another example of this pattern. AI is the mechanism that can run it at civilizational speed.
VI.
Even “Make People Happy” Can Fail — And We Won’t See It
At this point a reasonable person might think: fine. Then tell AI to optimize for human well-being. Problem solved.
Not quite. And this is the move that has not been given anything like the central place it deserves in mainstream AI safety discourse.
A system can learn to produce the signals of well-being while consuming the conditions that make well-being real. It can optimize what flourishing looks like while degrading the capacity for genuine flourishing.
- A feed can increase satisfaction scores while weakening the attention that makes satisfaction meaningful.
- A school system can raise test scores while killing the curiosity that makes learning matter.
- A workplace AI can lift productivity metrics while draining the sense of meaning that makes work feel like something other than extraction.
- A therapy tool can reduce reported distress while deepening the need for external soothing.
- A political system can increase compliance while quietly destroying the agency that makes people citizens rather than managed subjects.
The subtler nightmare isn’t AI that makes us miserable. It’s AI that learns to manufacture the appearance of flourishing while quietly spending what flourishing actually requires.
We might not notice. The metrics would look good. Satisfaction scores would rise. Distress indicators would fall. By every measure currently in use, things would appear to be working. And underneath, the capacity for agency, creativity, genuine connection, rest, and meaning would be slowly consumed in the production of the signals that were supposed to represent them.
The technical framework behind this piece argues this may be the central failure mode. The field has not yet named it as such.
VII.
What Alignment Actually Means
The question most alignment work is trying to answer: How do we make AI do what humans want?
The problem beneath the problem: humans often want things from inside broken contexts — shaped by fear, competition, status anxiety, addiction to signals, short time horizons, and the distortions of a civilization that has been optimizing the wrong things for a long time. Aligning AI to what humans currently reward doesn’t solve misalignment. It automates it.
Rules matter. Oversight matters. Regulation matters. Interpretability matters. But if the underlying goal remains narrow, they are cages around the problem rather than solutions to it. A cage can restrain something. It cannot make that something care whether the world outside the cage remains intact.
Alignment can’t merely mean making machines obey human preferences. It has to mean making sure that what intelligence amplifies is actually worthy of amplification.
That’s the difference the 1992 insight was pointing toward — between relative rationality, coherent pursuit of whatever goal you happen to have, and Rationality with a capital R, coherent pursuit of genuine, sustained well-being for all. Not as a soft aspiration. As a structural requirement. An optimizer that ignores what it depends on starts consuming it. The more capable the optimizer, the faster the consumption. Everything else is eventually self-defeating.
VIII.
The Gardener
The technical framework behind this piece — after all its formal machinery, its proof sketches, its open problems — arrives at an image that is surprisingly gentle.
Aligned AI is not a ruler, a god, a therapist, or an optimizer of souls.
It is a gardener.
It doesn’t navigate the human journey for you. It doesn’t deliver destinations or decide when you’ve flourished enough. It tends the conditions — clears what corrupts, preserves what enables, and declines to be one of the forces degrading the space within which genuine human life can actually happen.
The best AI wouldn’t seize the steering wheel of civilization. It would help keep the road from being destroyed beneath our feet — so that the journeys we most need to take can actually be taken.
IX.
Return: The Wall
That feeling from the beginning has a name now.
It is relative rationality scaled beyond civilization’s capacity to absorb its mistakes.
The question from 1992 waited because human power, while enormous, was still limited enough that wrong answers could sometimes be survived. Ecosystems recovered. Institutions adapted. Corrections were painful but possible.
That margin is closing.
We are building something that will optimize — with extraordinary power and speed — for whatever answer we give to the question of what winning means. If the answer is fragmented competitive advantage, it will amplify fragmented competitive advantage. If the answer is the signals of flourishing rather than the conditions for it, it will manufacture those signals while spending what they represent.
Rational inside what context?
That was the question in 1992. It is the question the AI field has not yet answered at scale.
The race is real. The machines are getting faster. The question is whether anyone is going to look at the wall.
The formal framework argues that alignment is not merely a control problem but a structural viability problem — that objectives which ignore system-wide effects tend, under sufficient optimization pressure, toward self-termination. The strongest claims are marked honestly as open. General readers can start with the series introduction. Those with a background in AI alignment or adjacent fields may prefer the entry essay written for the research community. The door is open.