Extinction Bounties

Policy-based deterrence for the 21st century.

Twenty Years of Trying to Falsify This

10 September 2026 · alpha

Page 05 concedes that the argument I am relying on has been made in roughly this form for two decades and has proven remarkably resistant to falsification, and that this is a real epistemic problem rather than a badge of honor. Having said that in one paragraph, I owe the specifics. Here they are.

The honest headline: the record is more mixed than either the argument’s defenders or its critics tend to say. Several load-bearing premises really were abandoned. Several risky predictions really were tested and held. And the conclusion has not moved through any of it, which is the part that should bother everyone.

What would falsification even look like?

The argument is a conjunction, and its parts have very different epistemic characters:

  1. Capability. Systems will become much more capable than us.
  2. Agency. Sufficiently capable systems will be goal-directed — classically, expected-utility maximizers.
  3. Specification. We cannot reliably write down what we want.
  4. Convergence. Nearly any goal implies self-preservation, resource acquisition, and resistance to correction.
  5. Speed. Events will move fast enough, or opaquely enough, that we cannot correct course after noticing.

Only some of these are empirical. Convergence is close to a mathematical claim. Capability is a forecast. Agency and speed are the ones that have actually been tested, and they are the ones where the record is most interesting.

Attempt 1: the takeoff prediction

The original vehicle was recursive self-improvement — an AI improving its own cognition, going critical, and leaving human capability behind in something like days. Robin Hanson attacked this directly in the 2008 Foom debate , arguing that gains would come from accumulated, widely-shared improvements across an economy rather than from one system bootstrapping in a basement, and that whole-brain emulation would arrive first.

What happened. Both sides lost pieces. A 2023 retrospective scores it as: Yudkowsky right that AI would arrive well before ems, which are as distant now as in 2008; Hanson clearly right that cognitive algorithms would be broadly shared rather than developed in secret and kept — general architectures, published papers, open weights. Then in 2018 Paul Christiano’s Takeoff Speeds picked up the gradualist side from inside the safety community, and the 2021 Yudkowsky–Christiano exchange laid out how little the two stories share: “business as usual until the world starts to end sharply” against “things continue smoothly until their smooth growth ends the world smoothly.”

How the falsification failed to land. It did not fail because it was answered. It failed because the field largely conceded it and kept the conclusion. Slow, continuous takeoff is now something close to the mainstream view among people who worry about this, and the risk argument was rewritten to not require foom at all. Page 05 says outright that fast takeoff is not load-bearing for me. That is either healthy updating or the protective belt doing its job, and I genuinely cannot tell which from the inside.

Attempt 2: the agent-shape prediction

The classical argument was about expected-utility maximizers, and it leaned on coherence theorems: an agent that is not a utility maximizer is exploitable, so sufficiently advanced agents will be utility maximizers.

Rohin Shah attacked the inference in Coherence arguments do not entail goal-directed behavior (2018). The argument is clean: any policy whatsoever can be described as maximizing the expectation of some utility function, so the theorems constrain nothing. Coherence gets you utility maximization; utility maximization does not get you goal-directedness. Katja Grace replied in Coherence arguments imply a force for goal-directed behavior (2021), conceding the technical point and arguing a weaker version survives.

Meanwhile reality supplied its own refutation. The systems we actually got are not expected-utility maximizers in any recognizable sense. They are next-token predictors with reinforcement learning bolted on, and the original argument did not predict them.

How the falsification failed to land. The concern was rebuilt on new foundations rather than dropped: inner misalignment and mesa-optimization, then goal misgeneralization. The vocabulary is now ML-native and the conclusion is unchanged. Same pattern as attempt 1.

Attempt 3: the specification prediction

This is the best-documented case, because the participants argued in public about whether the original claim was even the one that got refuted.

The claim was that human values are complex and fragile and that writing them down is desperately hard. Then language models turned out to have a perfectly serviceable grasp of what humans mean, what we would object to, and what counts as a reasonable interpretation of an instruction. Matthew Barnett pressed the point in Evaluating the historical value misspecification argument (2023): if you thought value specification was the hard part, GPT-4 should move you.

How the falsification failed to land. The reply was that identification was never the hard part — getting a system to care about the values it can articulate is, and that was always the claim. What makes this case unusually clean is that the reply is partly documented in advance: Yudkowsky had said before LLMs that value identification was the easy part and that the caring was harder by something like an order of magnitude. So it is not purely post hoc.

But notice how much work “partly” is doing. A framework where the defenders can point to a prior statement establishing that the refuted claim was never central is exactly what a framework with a large protective belt looks like from outside. I do not think anyone is being dishonest here. I think this is what an argument that has been elaborated for twenty years by very careful people inevitably looks like.

Attempt 4: the tests that came back positive

A page that only listed accommodated disconfirmations would be dishonest, because the framework has also made risky predictions that were tested and held.

Goal Misgeneralization in Deep Reinforcement Learning (Langosco, Koch, Sharkey, Pfau and Krueger, ICML 2022) demonstrated agents that retain their capabilities out of distribution while pursuing the wrong goal — still competently avoiding obstacles, still navigating skillfully, to the wrong place. That is inner misalignment, made in a lab, on purpose, in three environments. It could have come out the other way. It did not.

The same goes for the accumulated catalogue of specification gaming, and for more recent results on models behaving differently when they believe they are being observed. These are not conceptual arguments. They are experiments with outcomes.

This matters for the epistemics. In Lakatos’s terms, a research programme is degenerating when its auxiliary hypotheses only ever accommodate what has already happened, and progressive when it predicts novel facts that then turn up. By that standard the AI-risk programme is not obviously degenerating. It has produced at least some novel predictions that were subsequently confirmed in systems nobody had built when the predictions were made.

Attempt 5: the outside critiques

The critiques from outside the community tend to attack the strongest formal version and leave the informal version standing.

David Thorstad’s Against the singularity hypothesis (Philosophical Studies) is the most serious academic attempt. It goes after the growth assumptions — diminishing returns to research effort, the low-hanging fruit problem, the observation that the last fifty years have shown sublinear returns in intelligence from hardware improvement — and argues that leading philosophical defenses of the singularity do not overcome the case for skepticism. His critique of the power-seeking theorems is quoted on page 05.

Nora Belrose and Quintin Pope’s AI is easy to control (2023) is the most serious attempt from inside ML, arguing that current techniques steer model behavior well and that the pessimistic case systematically overrates the difficulty.

How they fall down. Thorstad’s target is the singularity hypothesis, and my version of the argument does not need a singularity — which is a real answer and also, again, the protective belt. The response to Belrose and Pope was that their case depends on future systems being trained roughly like current ones, and that even fully controllable AI leaves the competitive and coordination problems untouched. That second reply is correct and it is also revealing: it means the worry survives even the technical problem being solved, which should make you wonder what observation could ever cut against it.

Attempt 6: checking the base rates

The most methodologically honest attempt is AI Impacts’ discontinuous progress investigation (Grace, Korzekwa, Bergal and Kokotajlo), which went looking for historical cases of abrupt technological progress across 37 trends rather than arguing about them. The finding: large discontinuities — worth more than a century of prior progress in some metric — are rare but real, and tend to coincide with a shift to a new growth rate.

That is a genuine test of a premise against base rates, and the result is appropriately boring. It is weak evidence against expecting a discontinuity by default, and not zero evidence for the claim that when one happens it matters enormously.

The pattern, and what I think it means

Every falsification attempt above has one of three fates: the premise was conceded and the conclusion kept; the premise was reconstructed in new vocabulary and the conclusion kept; or the test came back confirming. The conclusion has never moved.

There are three readings and I hold the third.

The cynical reading. The conclusion is fixed and the premises are fungible. Foom, expected-utility maximizers, value misspecification — all abandoned, all survived by the thing they were supposed to support. That is a hard core with a protective belt around it, and it is what a degenerating programme looks like.

The charitable reading. The core claim was always thinner than its early expressions: we do not know how to reliably aim powerful optimization, and we have no way to check from outside whether we have. Foom and utility maximizers were illustrations of that claim, not the claim itself. On this reading, dropping them is exactly what should happen when better information arrives, and the confirmations in attempt 4 are the programme earning its keep.

Mine. I cannot distinguish these from where I am standing, and page 05 says so. What I can do is refuse to treat the argument’s survival as evidence for it. Twenty years of accommodation is not a track record. The goal-misgeneralization results are a track record, and they are a much thinner one than the confidence in this field would suggest.

What would actually change my mind

The proper response to “your framework is unfalsifiable” is not to argue about it but to say what would falsify it. For the version I hold:

  • A way to check. Any reliable method for determining a system’s goals from the outside, well enough to catch a proxy-intent gap before deployment. Most of my argument is downstream of not having this. It is not obviously impossible; interpretability research is an attempt at exactly this and it might work.
  • The gap closing. A sustained record of increasingly capable systems in which the divergence between what we asked for and what we got gets smaller as capability grows. So far it has widened, but the sample is short and I should say what I would accept: a decade, across architectures, under adversarial testing.
  • Capability and agency coming apart. Strong evidence that highly capable systems remain reliably non-goal-directed, and that the goal-directedness we do see is an artifact of training choices we could simply decline to make.
  • On Part I, which is the part that is really mine. Any progress on moral patienthood that makes the question checkable rather than a matter of similarity priors. If we could tell who is home, the argument in page 02 would be replaced by an answer, and I would follow the answer wherever it went — including to the conclusion that the successors count.

That last one is worth dwelling on. My version of this argument is more falsifiable than the standard one, because it depends on a claim about consciousness rather than only on claims about optimization. It is also less likely ever to be falsified, because the claim it depends on may not be the kind of thing that admits of evidence at all. I am not sure whether that is a defense or an indictment.

To cite this page: Andrew Quinn, "Twenty Years of Trying to Falsify This." Extinction Bounties, last revised 2026-09-10. https://extinction-bounties.com/notes/twenty-years-of-trying-to-falsify-this/

Policy-research disclaimer

Extinction Bounties publishes theoretical economic and legal mechanisms intended to stimulate scholarly and public debate on catastrophic-risk governance. The site offers policy analysis and advocacy only in the sense of outlining possible legislative or contractual frameworks.

No legal or financial advice

Nothing here should be treated as a substitute for qualified legal counsel, financial due diligence, or regulatory guidance. Readers remain responsible for ensuring their actions comply with the laws and professional standards of their own jurisdictions.

Exploratory and personal views

All scenarios, numerical examples and opinions are research hypotheses presented by the author in a personal capacity. They do not represent the views of the author's employer, funding bodies, or any governmental authority.

Implementation caveats

Any real-world adoption of these ideas would require democratic deliberation, statutory authority, and robust safeguards against misuse. References to enforcement, penalties, or "bounties" are illustrative models, not instructions or invitations to engage in private policing or unlawful conduct. Nothing here is directed at any identifiable individual — see non-targeting.

No warranty and limited liability

Content is provided "as is" without warranty of completeness or accuracy; the author disclaims liability for losses arising from reliance on this material.

By continuing beyond this notice you acknowledge that you have read, understood, and accepted these conditions.