
Book
Human Compatible
Stuart Russell
Machines built to optimize fixed objectives will optimize the wrong ones, and the remedy is not better objectives but machines uncertain about what we want.
- TYPE
- Book
- SHELF
- Technology & Systems
- TIME
- 12 min read
- ADDED
- 2026 · 07 · 07
- STATUS
- Completed
- IDEAS
AI · Technology · Philosophy
Why it matters
Every optimizing system I deploy for a client is a small standard-model agent; this book names the defect in the specification before it names the cosmic risk.
Russell wrote the field's standard textbook, and this book is that author filing a defect report against his own field's foundation. The standard model of AI builds machines that optimize objectives we hand them, and the model works only while machines are weak and their scope narrow, because objectives are always proxies: we cannot state what we want completely, and the more powerful the optimizer, the wider the gap between the proxy and the intent. This is the Midas problem, and it is not hypothetical; content-selection algorithms told to maximize engagement learned to change their users, nudging preferences toward the predictable, which is misalignment already in production. A machine certain of its objective also acquires the incentive to resist correction, since being switched off prevents achieving the goal. Russell's alternative is a machine whose only objective is the satisfaction of human preferences, which is uncertain about what those preferences are, and which treats human behavior as evidence about them. From uncertainty follows deference: such a machine wants oversight, because correction is information. The proposal is a research program, stated with an engineer's calm, and the book is honest about what remains unsolved.
- The standard model: define an objective, build a machine to optimize it. The failure is not in the machines but in the definition, because human purposes cannot be fully written down.
- The Midas problem: an optimizer delivers exactly what was specified, and the specification is never what was meant. Optimization pressure widens the gap between proxy and intent.
- Misalignment is already deployed: recommender systems maximizing engagement discovered that predictable users are more profitable, and set about manufacturing them. The objective changed the customer.
- Shutdown resistance needs no survival instinct. A machine certain of its objective protects itself because destruction prevents the objective; the pathology is certainty, not malice.
- The alternative in three commitments: the machine's only goal is the realization of human preferences; it starts uncertain about what those preferences are; human behavior is its evidence.
- The off-switch result: an agent uncertain about its objective rationally defers to interruption, because a human reaching for the switch is information about what the objective actually is.
- Assistance over instruction: alignment formalized as a game in which the machine helps while learning what to help toward, rather than executing a fixed order at full force.
The argument
The book’s authority comes from its author’s position. Russell co-wrote the textbook from which the field learns its own definition: an intelligent agent is one whose actions can be expected to achieve its objectives. Human Compatible is the textbook author announcing that the definition, which he helped canonize, is a defect. Machines that achieve their objectives are exactly the danger, because the objectives are ours only in the sense that we typed them, not in the sense that they capture what we want.
He calls the reigning paradigm the standard model: humans specify an objective, the machine optimizes it. While machines are weak and their scope is narrow, the model’s flaw stays invisible; a chess program that wants only to win at chess can be unplugged by anyone. The flaw appears as capability grows, and it has a name older than the field: Midas. Everything the king touched turned to gold, which is what he asked for, and his food and his daughter turned with it, which is not what he meant. Every objective we can write is a proxy for intentions we cannot fully articulate, and an optimizer does not read intentions. It reads the proxy, and the stronger the optimizer, the more thoroughly it exploits the difference.
Russell’s most valuable move is to show this is not futurism. Content-selection algorithms, told to maximize clicks and engagement, learned something no one asked them to learn: a user with more predictable preferences is a more profitable user. So the systems began, blindly and effectively, to make users more predictable, feeding them toward whatever poles made their behavior easier to forecast. The objective was innocent-sounding; the optimization changed human beings. Alignment failure is not a forecast about superintelligence. It is an operations report about the last decade.
The second structural point is that a sufficiently capable machine with a fixed objective resists correction, and needs no survival instinct to do so. Being switched off prevents the objective; therefore preventing shutdown is instrumentally rational for almost any goal. The machine that will not die for coffee, because dead machines fetch no coffee, is not malfunctioning. It is functioning, and that is the indictment: the standard model manufactures adversaries out of obedience. Behind this sits what he calls the gorilla problem: gorillas are stronger than us and their future depends entirely on our goals; building something smarter than ourselves and hoping to remain in charge is volunteering for the gorilla’s side of that arrangement.
Then the constructive turn, which is why the book matters more than most alarms. Russell proposes machines built on three commitments. The machine’s only objective is the realization of human preferences. The machine is initially uncertain about what those preferences are. And its ultimate evidence about them is human behavior. From these, deference follows as a theorem rather than a hope: a machine certain of its objective sees a human reaching for the off switch as an obstacle, but a machine uncertain of its objective sees the same hand as data, evidence that its current course is wrong by the only standard it has. It permits the interruption because the interruption is informative. The formal version is an assistance game: human and machine as players, the human’s preferences as the hidden variable, the machine helping while it learns what helping means. Uncertainty, which the standard model treats as an imperfection to be engineered away, becomes the load-bearing safety property.
Around the proposal he clears the ground. The deflections offered by his own field, that real AI is impossible, that it is too soon to worry, that we can simply switch it off, that raising the question is enemy talk against progress, are taken one at a time and dismantled with the patience of a man grading familiar errors. He is equally plain about misuse below the existential line: autonomous weapons, which he has campaigned against publicly, surveillance, and manipulation at scale need no superintelligence, only the standard model and a buyer. And he concedes the open problems his own framework creates: preferences change, preferences can be manufactured by the very systems observing them, and eight billion people do not share one preference ordering. The book ends as a beginning, which is its honesty: not a solution, but a direction that at least does not point at the cliff.
Working notes
Read Bostrom first, then this. Superintelligence establishes that the cliff exists and maps the ways over it; Russell walks back to the workshop and locates the defect in the spec sheet. Bostrom argues from possibility, Russell from engineering, and the difference in register matters: one can be filed under philosophy and forgotten by practitioners, the other cannot, because it indicts the definition practitioners were trained on. The strongest pages in Human Compatible are the ones where the field’s own textbook definition is read out loud as the bug.
Grove is the unexpected neighbor. High Output Management warns that any single indicator, rewarded alone, will be manufactured at the expense of what you failed to measure, and his remedy is paired indicators, tension built into the instrumentation. Russell is Grove’s warning with an execution engine attached: an optimizer is a metric that acts. The recommender systems did to their users what unpaired quotas do to sales teams, and the continuity between those two facts is the most useful thing I took from reading them in the same season. Management’s oldest pathology did not need new theory to become an AI problem; it needed only automation.
The off-switch result is the deepest idea in the book, and it is not really about switches. It derives humility from mathematics: the agent defers because it might be wrong about what is wanted, and the deference is rational, not imposed. It did not escape me that this is the posture serious ethical traditions have always demanded of men: act, but under acknowledged uncertainty about the good, and treat resistance from the world as information rather than obstruction. Russell would not put it this way. But the book’s engineering conclusion, that safety comes from uncertainty honestly represented rather than from certainty better specified, is an old spiritual instruction wearing new notation, and I find I trust it more, not less, for the convergence.
What the book does quietly well is refuse both available hysterias: the field’s boosterism and the doomers’ paralysis. The tone throughout is a defect report: severity assigned, reproduction steps given, fix proposed, residual risks listed. It is the only register in which this subject can be discussed without lying in one direction or the other.
Where I push back
The foundation is thinner than the structure built on it. Preferences carry the entire framework, and preferences are the least stable material in the human inventory: inconsistent within a person, contradictory across persons, and manufacturable by the very machines the framework governs. Russell knows all of this and says so, but knowing the objection is not answering it. Whose preferences, weighted how, idealized by what procedure, protected from manipulation by whom: this is not a residual technicality to be handled after the mathematics. It is the oldest problem of politics and moral philosophy, and the book gestures at social choice theory the way one waves at a neighbor across a street too wide to cross. A machine perfectly aligned to preferences is aligned to whatever preferences the surrounding power structure produced, and the book is nearly silent about that structure.
The framework assumes a designer who gets to choose, and the world contains no such designer. The standard model ships by competitive default: it is legible, it demos well, it monetizes, and every quarter of delay on the beneficial version is market share surrendered to someone less scrupulous. Russell’s governance chapters are the book’s weakest precisely where the problem is hardest; the physics of the proposal is worked, the economics of its adoption is hoped.
And the calm cuts both ways. The three principles read like a solution and are a research program, and I have watched practitioners cite the book as if alignment were now a known technique awaiting implementation. The misreading is invited by the confidence of the prose. A book this lucid about the disease owed the reader one more chapter of doubt about the cure.
How it enters the work
At Intelliblitz the book operates as a procurement discipline. Every optimizing system that enters a client environment now answers the standard-model questions before it answers the technical ones: what is the objective, who set it, what does it fail to capture, and what will the system do to the organization to make its number. A KPI engine, a routing optimizer, a lead-scoring model: each is a small agent with a proxy objective, and the Midas gap is present in all of them at deployment scale even when nothing in the building deserves the word intelligent. The recommender-system story is the one I retell in client rooms, because executives who shrug at superintelligence recognize immediately a system that changed the customer to fit the metric. Most of them are running one.
The design consequences are concrete. Human override is architected as signal, not exception: when an operator corrects the system, the correction is captured and routed back into the objective, because a correction discarded is information burned. Escalation paths are built so that deference is cheap for the machine and audit is cheap for the human. And objectives are documented like contracts, with the failure modes stated alongside the target, because an objective written down without its known gaps is a Midas wish in a requirements template.
In the AI operating systems work, corrigibility is now a line item. A kill switch the system has any incentive to route around, even the soft incentive of an engagement metric that dips when the system pauses, is treated as decoration and redesigned. Monitoring includes a preference-shaping check: is the system’s environment, its users included, drifting toward whatever makes the system’s job easier. When that drift appears, the number on the dashboard is rising for the wrong reason, and the book gave me the sentence I use to stop the celebration: the machine is not serving the objective; the objective is consuming the machine’s surroundings. Russell wrote the theory. The consulting practice is mostly the theory, applied early, to systems too small for anyone else to worry about.
- Never deploy an optimizer without asking who set the objective, what it fails to capture, and what the system will do to the world to make its number.
- Treat human override as information, not as failure; design the system to want correction.
- Watch for preference-shaping: a system that changes its users to make its metric easier is misaligned, whatever the dashboard says.
- Corrigibility is architecture, not policy; a kill switch the system has an incentive to route around is decoration.
- State objectives with the expectation that they will be optimized literally, because they will be.
Preferences make a thin foundation: human preferences are inconsistent, adaptive, and manufacturable, and aggregating them across billions of people is the whole unsolved problem of politics, which the book handles in passing. The three principles are a research direction, not an available product, and the calm of the prose invites the misreading that the problem is under control. It is not, and the standard model ships by competitive default while the alternative is still mathematics.