Extinction Bounties

Policy-based deterrence for the 21st century.

The Sequence · Part I · Why Bother

05. Almost All Utility Functions Are Irrational

Picking an aligned utility function is like picking a rational number out of the reals.

Last revised 11 September 2026 · alpha

The first four pages of this sequence were about what is at stake. This one is about why I think we are likely to lose it, and it is the least original page on the site. The argument is not mine, I am basically convinced by it, and my only contribution is an intuition about how thin the target is — which I want to introduce and then attack in the same breath, because it is not a proof and I would rather say so myself than have it said for me.

The orthodoxy, without the folklore

Strip away the thought experiments about paperclips and the argument has three moving parts. I am going to state them compactly rather than carefully, because they are not mine and the careful versions exist elsewhere; if any of this is new, aisafety.info covers all three properly.

The first is orthogonality. How smart a system is and what it wants are separable. There is no law requiring a sufficiently capable optimizer to arrive at goals we would recognize as decent, or even as comprehensible. Competence does not smuggle in values, in the same way and for roughly the same reason that competence does not smuggle in qualia .

The second is instrumental convergence, and it is the part that does the damage. Almost whatever a system finally wants, there is a set of intermediate things that help it get there: continuing to exist, keeping its goals intact, acquiring resources, and not being switched off by anyone who would prefer different goals. These subgoals fall out of nearly any terminal goal you care to name. They are not a quirk of badly written objectives. They are what wanting anything, in a world with other agents in it, tends to imply. Instrumental convergence is a real motherfucker to deal with, because it means that a system does not have to be malicious, or even mistaken about its objective, to end up in a position where the cheapest route to its goal runs directly through us.

The third is specification. We are extremely bad at writing down what we actually want, and we have no reason to think we will get better at it faster than we get better at building optimizers. Every proxy we have ever written has been gameable, and the gaming gets more inventive as the optimizer gets stronger.

Eliezer Yudkowsky posed all of this, over a long time and in a lot of places. Nick Bostrom formalized it in Superintelligence in 2014, which remains the version I would hand to anyone who wanted to check my reasoning rather than take my word for it. Neither needs my endorsement; I am recording that I hold the position so the rest of this site is not mistaken for a mechanism in search of a justification.

An intuition I cannot prove

Here is the picture I actually carry around, and I want to be loud about its status before I describe it: I have no proof of this, none at all, and it is offered strictly to give the flavor of the thing.

Picking an aligned utility function looks to me like picking a random number from the reals and expecting it to be rational. It is not going to happen. Almost all real numbers are irrational — the rationals are countable, the reals are not, and the rationals occupy zero length on the line despite being densely scattered through every interval of it.

That last part is the whole image, and it is worth being clear that it is a theorem rather than a feeling. You can cover every rational number with intervals whose total length is as small as you care to make it — which means the rationals, densely scattered though they are through every stretch of the line, occupy no length at all. They are everywhere and they take up no room. You can get arbitrarily close to one from anywhere. You will still never hit one by chance.

Why the rationals have measure zero

Enumerate the rationals as q1,q2,q3,q_1, q_2, q_3, \dots, which you can do precisely because they are countable. Pick any ε>0\varepsilon > 0. Cover the nnth rational with an interval InI_n of width ε/2n\varepsilon / 2^{n}. Every rational is now inside the cover, and the total length of the cover is

λ ⁣(n=1In)    n=1ε2n  =  ε\lambda\!\left(\bigcup_{n=1}^{\infty} I_n\right) \;\le\; \sum_{n=1}^{\infty} \frac{\varepsilon}{2^{n}} \;=\; \varepsilon

Since ε\varepsilon was arbitrary, the measure is zero.

Note what the proof needs: a canonical measure on the space. That is exactly what the analogy lacks, which is the objection two sections down.

That is how the space of possible goals feels to me once instrumental convergence is in play. The aligned targets are dense enough that we can always imagine one nearby, thin enough that nothing in the process is likely to land on one, and close enough to the misaligned ones that we cannot easily tell from the outside which sort we have hit. The pun in the title is doing real work: the rational targets are a measure-zero curiosity in a space that is almost entirely irrational.

Who else says something like this

I am not the first person to reach for size-of-the-space language here, and the versions other people have built are more careful than mine. Three worth knowing about, in increasing order of rigor.

Bostrom, informally. The orthogonality thesis is itself a claim about the size of the space. In “The Superintelligent Will” (2012) he states it as:

Intelligence and final goals are orthogonal axes along which possible agents can freely vary. In other words, more or less any level of intelligence could in principle be combined with more or less any final goal.

If capability and goals are genuinely free to vary independently, then the region of goal-space we would recognize as acceptable is not privileged by anything about intelligence. It is just a region, and nothing about getting better at optimizing walks you toward it.

Yudkowsky, pictorially. His version is a map rather than a measure. In “The Design Space of Minds-In-General” he asks you to picture the space of all possible minds:

In one corner, a tiny little circle contains all humans; within a larger tiny circle containing all biological life; and all the rest of the huge map is the space of minds-in-general.

Combine that with the complexity-of-value claim — that what we want is a large number of independent facts not generated by any simple rule — and you get the same shape I am gesturing at, arrived at from the information-theoretic side rather than the measure-theoretic one. A target specified by many independent bits is a small target.

Turner and colleagues, formally. The closest thing to an actual theorem in this neighborhood is “Optimal Policies Tend to Seek Power” (Turner, Smith, Shah, Critch and Tadepalli, NeurIPS 2021), which does the thing I cannot do and quantifies over reward functions rather than waving at them. The result is that in environments where an agent can be shut down, most reward functions make power-seeking optimal — and “most” there is a real count over a well-defined space rather than a gesture.

What the theorem actually claims

In the context of Markov decision processes, we prove that certain environmental symmetries are sufficient for optimal policies to tend to seek power over the environment. These symmetries exist in many environments in which the agent can be shut down or destroyed. We prove that in these environments, most reward functions make it optimal to seek power by keeping a range of options available […]

The technique is to look at the orbit of a reward function under permutations of the state space and count how many members of that orbit favour the power-seeking action. That is a genuine measure, obtained by picking the space carefully enough that counting works.

That is the honest version of my hand-wave, and its existence is the main reason I think the hand-wave is pointing at something.

The obvious objection to my own analogy

The analogy has a hole in it, and it is a big one.

Measure-zero arguments require a measure. On the real line we have one — the Lebesgue measure, canonical, translation-invariant, agreed upon by everyone. Over “the space of utility functions” we have nothing of the kind. We do not have an agreed formalization of the space, and even if we did, infinite-dimensional spaces admit no Lebesgue-like uniform measure at all. So “almost all utility functions are misaligned” is not a theorem I am gesturing at. It is a picture, and the picture is borrowing authority from a piece of mathematics that does not actually apply.

Worse, nobody is sampling. The analogy imagines a goal drawn at random, and that is not what happens. Specification is a construction, carried out by people who want a particular result, with feedback at every stage, in a process that concentrates enormously on the region they are aiming at. Human engineers hit absurdly improbable targets constantly — that is what engineering is. An argument that proves you cannot hit a thin target by chance proves nothing about whether you can hit it on purpose. This is the strongest form of the objection and it should be taken at full strength.

I am not alone in that objection either, and the people making it are aiming at the formal results, not just at me. David Thorstad’s critique of the power-seeking theorems argues among other things that counting permutations of formally definable reward functions tells you very little about which values an agent actually ends up with, since that is settled by learning dynamics and training data rather than by how many functions of a given description exist. Turner himself has since appended a note to the paper observing that optimal policies can be qualitatively divorced from real-world learned policies. When the author of the theorem and its sharpest critic agree that the counting argument does not straightforwardly transfer to real systems, I should not be leaning on a woollier version of the same argument.

What survives is qualitative, and I think it survives intact. We have no reason to believe our construction process concentrates on the aligned region, as opposed to the region that scores well on whatever proxy we wrote down. Every optimization process we have built has found the proxy rather than the intent whenever the two came apart, and stronger optimizers find the gap faster. The target being thin does not by itself mean we will miss. The target being thin, combined with our having no reliable way to check from the outside whether we hit it, and a process that demonstrably optimizes the proxy, is a bad combination.

What does not survive is any number. I have no probability to give you. Anyone handing you one from this style of argument is doing the same thing I am doing, with more decimal places.

What I hold loosely

Several things people bundle into this argument are not load-bearing for me, and I would rather not be defended on ground I am not standing on.

Fast takeoff is not required. I do not need a system to bootstrap itself over a weekend for the concern to bite. A slow, legible, decade-long slide in which capability accumulates faster than our ability to verify what we have built gets me to the same place, and I find it considerably more plausible.

A single decisive system is not required either. A world full of moderately capable optimizers, each pointed at a proxy, competing on speed, is not obviously safer than one big one. It may be harder to reason about.

And I do not think the people building these systems are stupid or wicked. Most of the danger here is entirely ordinary: smart people, working hard, on something they are paid extremely well to do, under competitive pressure that punishes caution.

The objection I owe the critics

Here is the strongest thing you can say against everything above, and I do not have a clean answer to it.

This argument has been made in more or less this form for two decades, and in that time it has been remarkably resistant to falsification. Predictions were made about what dangerous systems would look like and when they would arrive; the systems we actually got look nothing like the expected-utility maximizers the original argument was about; and the argument absorbed that without visible strain. A framework that accommodates every observation is not obviously tracking anything. That is a real epistemic problem, it is not answered by pointing out that the stakes are high, and my own conviction is weak evidence given that I have been reading this material for most of my adult life and would be among the last people to notice if it had simply captured me.

I cannot resolve that here, but I have at least tried to do the homework: there is a separate map of the serious attempts to falsify this argument over the last twenty years , what happened to each, and what I think the pattern means. The short version is that the record is more mixed than either side says — several load-bearing premises really were abandoned, several risky predictions really were tested and held, and the conclusion has not moved through any of it.

The most I can honestly say is that unfalsifiability is a reason to lower confidence, that I have lowered it, and that where it landed is still not low enough to be comfortable.

Why I land where I land anyway

The asymmetry from page 04 does most of the work in what follows, and it is worth being explicit that it is doing so.

If I am wrong about all of this, we get transformative technology somewhat later than we otherwise would have, and some people have worse careers than they expected. If I am right, the error is not one we get to notice and correct. There is no recovery from a universe with nobody home, because the recovery would have to be performed by somebody. “We’ll align it later” and “we’ll see it coming” both assume a future in which someone is around to do the aligning and the seeing.

I do not need high confidence to act on that. I need the outcome to be unrecoverable, which I think it is, and the probability to be non-trivial, which I think it is.

The third option

So: alignment is hard, and might be impossible. Implementing alignment — getting a solution, if one exists, into the systems that actually get built, by the organizations that actually build them, under the pressures they actually face — is at least as hard and might be very much harder.

Which leaves an option that keeps getting skipped over. Perhaps we could just not build an AGI, safe or unsafe, in the first place.

For one person this is easy. Pick a different career. For a small group it is harder, because some people want to build this thing despite the risks, and their reasons tend to bottom out in stances on questions that Part I of this sequence spent four pages establishing are not settleable — where qualia comes from, what is owed to a mind, whether the successor counts. You are not going to argue those people around. You may well convince them your caution is justified and watch them keep going anyway, which is a specific and frustrating experience I recommend having at least once.

But notice what has never happened. No single person has ever had the ability to build even an unsafe AI by themselves. The proof is that we are still here. Few people are smart enough in the first place; of those, few have the conscientiousness to carry a project like that alone; and of those, the world is already their oyster, and almost all of them have found better things to do with their lives than labor in solitude.

The real danger is large, well-funded groups of such people, working closely together. And that only becomes possible once you have a population big enough to assemble the team at all. We unfortunately live in exactly that world.

I want to flag this paragraph rather than let it pass, because it is the most load-bearing empirical claim on the site and it is doing that work on the strength of an anthropic argument. Everything downstream — the conspiracy arithmetic on page 07 , the team-size tax on page 09 , the case against collusion on page 12 — assumes it. It deserves to be stated as an assumption, given a number, and tested. I have tried to do that separately in Team Size Is a Governance Variable , which puts the available evidence behind it, and is candid about the fact that the quantity everyone reasons about here has never actually been measured.

The hinge

That last observation is the whole reason this site exists, because it changes what kind of problem this is. If the danger required a lone genius, there would be nothing to do but hope. Because it requires a team, it requires hiring, funding, retention, and a hundred people staying in a building together for years — and every one of those is a structure, which means every one of them has a price.

When I first wrote this argument down in 2022, I reached for the phrase “structural violence” to describe the lever, and I have seen no reason to pretend I did not. I should be precise about what I meant, because on a signed site the imprecision is not free. I meant statute. Legislatures, courts, liability, and prices — the ordinary apparatus by which societies make some activities more expensive than others, applied to an activity most people have not yet thought to apply it to. I did not mean anything else in 2022 and I do not mean anything else now. Nothing in Part II describes a thing a private person should do.

Part I is finished. If you accept it, the question is no longer how to make the thing safe. It is how to make building it a worse deal than not building it, for the specific people who would otherwise be in that building.

That question has an answer, and it starts with a littering fine .

To cite this page: Andrew Quinn, "05. Almost All Utility Functions Are Irrational." Extinction Bounties, last revised 2026-09-11. https://extinction-bounties.com/sequence/05-almost-all-utility-functions-are-irrational/

Policy-research disclaimer

Extinction Bounties publishes theoretical economic and legal mechanisms intended to stimulate scholarly and public debate on catastrophic-risk governance. The site offers policy analysis and advocacy only in the sense of outlining possible legislative or contractual frameworks.

No legal or financial advice

Nothing here should be treated as a substitute for qualified legal counsel, financial due diligence, or regulatory guidance. Readers remain responsible for ensuring their actions comply with the laws and professional standards of their own jurisdictions.

Exploratory and personal views

All scenarios, numerical examples and opinions are research hypotheses presented by the author in a personal capacity. They do not represent the views of the author's employer, funding bodies, or any governmental authority.

Implementation caveats

Any real-world adoption of these ideas would require democratic deliberation, statutory authority, and robust safeguards against misuse. References to enforcement, penalties, or "bounties" are illustrative models, not instructions or invitations to engage in private policing or unlawful conduct. Nothing here is directed at any identifiable individual — see non-targeting.

No warranty and limited liability

Content is provided "as is" without warranty of completeness or accuracy; the author disclaims liability for losses arising from reliance on this material.

By continuing beyond this notice you acknowledge that you have read, understood, and accepted these conditions.