Extinction Bounties

Policy-based deterrence for the 21st century.

The Sequence · Part II · What To Do

13. Next Steps

Pilots, draft statutes, and what would have to be true.

Last revised 11 September 2026 · alpha

Thirteen pages is a long walk to arrive at a request for a pilot program. That is where it arrives, and I would rather say so flatly than inflate the ending. There is no draft bill at the end of this. There is a description of the smallest experiment that would actually tell us something, and an argument that it should be run a very long way from anything that matters.

Test it where it cannot kill anyone

The subject of this site is frontier AI. The first implementation of this mechanism should have nothing to do with frontier AI.

That is not modesty, and it is not a rhetorical retreat. It follows from the design itself. The mechanism has free parameters — the bounty share, the evidentiary standard, the size of the penalty relative to the harm, the false claim penalty — and I do not know their values. Nobody does. A scheme with unknown parameters, aimed at a risk that admits no recovery , is not a safety proposal. It is a second unrecoverable gamble stacked on top of the first one. If the argument of Part I is right, it is right about my proposal too.

So calibrate somewhere cheap. The structure that makes bounties attractive is not exotic and not rare:

  • the offense is hard to detect from outside,
  • the harm is diffuse enough that no individual is motivated to prosecute it,
  • and the offender is better resourced than any single victim.

That description covers a great deal of ordinary, boring law. Illegal dumping. Fly-tipping. Unpermitted discharge. Wage theft. The littering case from page 06 was not a joke and was not only a teaching device — it was a jurisdiction where the entire chain can be run end to end for the price of a soda can.

Wage theft is the candidate I would push hardest, because the detection problem is close to pure. The people best positioned to document it are the people it happens to, they are precisely the people who cannot afford to bring a case, and the current enforcement model depends on an agency with a budget line and a backlog. Every structural feature the mechanism is designed to exploit is present, and the worst case if it fails badly is a wave of bad claims and some wasted court time.

I should name the disanalogy rather than let someone else find it. In all of these pilot domains there is a victim — somebody who was underpaid, somebody whose river was fouled. In the case this site actually cares about there is no victim to point at, because the harm has not happened and, if it does, there is nobody left to bring the suit. So a successful pilot would not demonstrate the thing I most need demonstrated. It would demonstrate the plumbing: that bounties find offenses, that insurers price behavior, that the false-claim rate is survivable. That is worth having and it is less than a proof.

What a draft statute needs

This is the most useful thing on the page, and I have tried to make it specific enough to be attacked. Seven components, and every one of them is a place the scheme can fail.

A defined covered activity. Not “dangerous AI.” A description a court can apply without reading anybody’s mind, naming an activity rather than an intention. This is the hardest drafting problem in the whole proposal and the place where a bad statute does the most damage.

A penalty schedule. Pegged to something, and the choice of peg is contested — social harm, the offender’s compensation, a flat figure. See how large should a bounty be ; I do not think this is settled and I do not want to pretend otherwise.

An evidentiary standard. Higher than a civil preponderance, given what is at stake for a defendant, and specified in advance rather than left to accumulate through case law.

A bounty share. What fraction of the penalty goes to whoever brought the evidence, and what happens to the remainder. A share too small does not move anyone; a share too large manufactures an industry with its own interests.

An insurance requirement. Whether coverage is mandatory or merely strongly advantageous is a genuine fork with different regimes on either side — see the note on that .

A false-claim penalty. Symmetric and real. A bounty system without a serious cost for manufacturing evidence is a machine for producing manufactured evidence, and this component is not optional.

A sunset and review clause. The statute should expire on a fixed date unless renewed, with a required evaluation before renewal. I want this partly on general principles and partly because I might be wrong, and a mechanism I am wrong about should not outlive the argument for it.

Anyone who wants to take this proposal apart should start with that list rather than with the philosophy. The philosophy is where I am most confident and the drafting is where I am least.

The research that does not exist

Three bodies of work are missing, and the last two are missing in ways I find genuinely strange.

Actuarial. I framed this badly in an earlier draft, as the problem of pricing a liability whose tail is uninsurable by construction — a claim nobody is around to file. That framing dissolves on inspection. Nobody is proposing to insure the catastrophe. The policy covers a bounded statutory bounty, and as the insurability note argues, the bounty is insurable precisely because it is not the damage. The uninsurable tail is simply not in the contract.

The real actuarial problem is different and better documented: underwriters behave badly under ambiguity rather than under risk. Faced with a loss probability they cannot pin down, they do not charge the expected value plus a modest margin — they charge a great deal more, and sometimes they decline to quote at all. That is the failure mode to plan against. Not a mispriced premium, but no market.

What insurers do when they cannot pin a number down

Kunreuther, Hogarth and Meszaros established in Insurer Ambiguity and Market Failure that insurers are strongly ambiguity averse. The asymmetry is professional rather than irrational: an underwriter who overestimates a risk loses an account, while one who underestimates it produces a visible loss ratio and an unpleasant meeting.

Later work finds insurers still pricing ambiguity through rough heuristics rather than any formal link to their own economic objective.

The precedent that should worry a proponent is environmental pollution liability, where ambiguity was severe enough that the industry withdrew from the market rather than price it.

So the question I actually want answered is narrower and answerable. Given that qq is a claim probability rather than a harm probability , and given decades of base rates on how often insiders disclose misconduct in other domains, how much ambiguity load would a carrier apply in year one — and is the resulting premium a deterrent, or a prohibition? An overpriced premium is the product working . An absent market is the mechanism failing to start.

The natural experiment nobody is mining. Several jurisdictions have been paying informants a share of what gets recovered, at scale, for decades. The SEC and CFTC whistleblower award programs. Section 7623 of the Internal Revenue Code. The qui tam provisions of the False Claims Act, which are old enough to have a genuine historical record. FinCEN’s anti-money-laundering program, which is newer.

These programs are a live, running test of the exact question this site turns on: does paying people to come forward change behaviour, and what does it cost to do it?

I said in an earlier draft that nobody was mining this. That was wrong, and I should correct it rather than quietly delete it — economists have been working this seam for fifteen years and the findings are more favourable to this proposal than I had any right to expect.

Three findings, and all three run in this proposal’s favour, which is exactly why I want to be careful with them.

Employees are already the largest single detection channel, and money is why. Not regulators, not auditors — the people inside. The frivolous-claims worry did not materialise, in the one comparison that isolates it. And the deterrence effect has actually been measured, once, and it came out about ten times larger than the money recovered. That is the chilling effect measured rather than asserted, in the one domain where somebody managed to measure it.

The three findings, with sources

Detection. Dyck, Morse and Zingales studied every reported fraud case in large US companies between 1996 and 2004 and found that detection does not come from the actors you would expect . The SEC accounted for 6% and auditors for 14%; employees accounted for 19%, more than either. They attribute the pattern partly to monetary incentives, using frauds against the government — where qui tam pays — as the comparison against frauds where nothing is paid.

False reports. Their conclusion is that monetary incentives influenced detection without increasing frivolous suits, which is the precise question page 12 worries about. Buccirossi, Immordino and Spagnolo reach a compatible verdict in Whistleblower Rewards, False Reports, and Corporate Fraud : the claim that rewards produce fraudulent reporting is frequently asserted and, across several US programmes, has not been a major problem.

Deterrence. Jetson Leder-Luis paired whistleblower suits against the universe of Medicare fee-for-service claims and found that settlements totalling roughly $1.9 billion produced deterrence effects of around $19 billion — about ten times the recovery, in money never fraudulently billed because the programme existed.

I want to be disciplined about what that does and does not license. Medicare billing fraud is frequent, homogeneous, and leaves a paper trail in a claims database, which is why it is measurable at all; frontier AI development is rare, heterogeneous and leaves nothing comparable. A 10x deterrence multiple in healthcare is not a 10x multiple here. What it does establish is that the central causal claim of this site — paying insiders changes behaviour upstream of enforcement, by more than the enforcement itself recovers — is not speculative. It has been estimated, and it came out large.

One finding cuts against me and I am keeping it in view. Buccirossi and co-authors conclude that reward programmes require sufficiently precise legal systems and are not viable in weak-institution environments. That is a direct problem for the suggestion on page 10 that poorer jurisdictions might host a bounty-hunting export industry. Hosting investigators may be fine; hosting the adjudication is a different matter, and the literature says so.

The questions still genuinely open are narrower than I first listed them. Does a paying regime detect more than a protection-only regime, controlling for everything one can? What is the rate of claims that turn out to be manufactured, and does the false-claim penalty deter them? What happens inside an organization that knows its employees have a standing financial reason to report it — does the internal compliance function get stronger, or does everything simply move offline? Does anyone leave a covered industry because of it, which is the only effect page 09 actually cares about?

Meanwhile the AI policy literature has arrived at the same question from the other direction and left it open. The Institute for Law & AI’s design guide for whistleblower legislation lists, among the decision points facing a legislator, whether whistleblowers should be offered monetary incentives at all. Every AI whistleblower statute now on the books protects the whistleblower’s job. None of them pays. The question is on the table and unanswered, and there is decades of data sitting one field over.

I want to be careful here, because I notice which way I want these studies to come out, and because the ones that exist came out in my favour it is worth saying what would not follow. None of this work is about catastrophic or unrecoverable risk. All of it is about recurring, compensable, measurable fraud in domains with dense administrative data. The transfer from there to here is an argument I have to make rather than a finding I get to cite.

Measurement of team size itself. This is the one I find strangest, because it is the cheapest of the three and it conditions everything else. Every mechanism in this family — extinction bounties, the False Claims Act, SEC awards, antitrust leniency — assumes the conduct involves enough people that one of them can be induced to talk. That is an empirical quantity, and as far as I can establish no one has ever estimated it. There is no study regressing detection probability on the number of people who knew; RAND says outright that nobody has correlated leak rates with the size of the cleared population; and Epoch AI, which tracks compute to several significant figures, publishes nothing modelling labour as an input to capability, despite its own cost work putting R&D staff at 29 to 49% of a frontier training run . We have built a governance literature on compute thresholds and none at all on the other half of the bill. I have set out what I would actually measure in Team Size Is a Governance Variable : the ratio of critical-path headcount to author counts, the smallest team that has reached a given capability level over time, and whether larger teams suppress misconduct or merely conceal it.

What would make me drop this

A separate note holds the rest of the AI-risk field to the standard of saying what would change its mind. It would be indefensible to apply that standard outward and not here.

This is a different exercise from page 12 , which collects the objections to the mechanism and answers the ones I can. These are observations that would not merely wound the proposal but end it:

  • A pilot where the false claims survive the filter. If a realistic evidentiary standard cannot separate genuine reports from manufactured ones at a tolerable rate, the scheme is a harassment engine and no amount of tuning fixes that.
  • Deterrence achieved, institution destroyed. If covered organizations respond by becoming so internally adversarial that they stop functioning — or simply stop writing anything down — then I have bought deterrence at a price I did not intend to pay, in a currency I cannot see on the balance sheet.
  • Pure displacement. If the labor-market effect is that the marginal researcher moves to an uncovered jurisdiction rather than to a different field, the mechanism has relocated the risk and called it a reduction. Page 10 is about why I think this is tractable; if it is not tractable, this fails.
  • No insurance market at any price. If underwriters will not write the coverage at all, the insured half of the design is fiction and what remains is an ordinary large fine with a bounty attached, which is a substantially weaker proposal.
  • A stable collusion equilibrium. If offenders and bounty hunters reliably find private settlements that beat prosecution for both, competition among enforcers is not doing the work Hanson’s design needs it to do.
  • Minimum viable team size falling below the point where anyone can defect. If a team of ten can train a genuinely frontier-capable model from scratch — not distil one, not fine-tune somebody else’s weights — then the conspiracy arithmetic has nothing to work with, and neither does any other insider-reporting mechanism. This is the one with a clock on it rather than a verdict, and I have tried to put numbers on the clock . It is also the reason to act early: unlike the others, this failure mode arrives on its own schedule whether or not anybody argues for it.

None of those is hypothetical in the sense of being unobservable. A pilot in the wage-theft or dumping family would produce evidence on the first two within a few years, which is the main reason to run one.

It was never about AI

Nothing in the mechanism is AI-specific. Strip out the examples and what remains is a general-purpose brake for any activity that is hard to detect, cheap to commit, and catastrophic in the tail. Gain-of-function research is the obvious sibling and arguably an easier first target, since the covered activity is far more crisply definable than anything in machine learning. A note works through the generalization .

Which produces an objection I do not think I can fully answer. A society that possesses a cheap, self-funding, highly effective way to slow down a frontier will use it, and it will not always use it on the frontiers I would have chosen. The same machinery aimed at gene editing, or nuclear power, or vaccine development, or encryption research, would work exactly as well and would be exactly as hard to switch off once an industry of enforcers exists.

The first reply is the weak one and I will make it anyway: the mechanism requires a legislature to define a covered activity, which is the same choke point that governs every other law, and it does not create a new power to prohibit so much as a cheaper way to enforce prohibitions already available. That is true and it is not very comforting, because legislatures pass bad laws routinely and the cheapness is precisely what changes the calculus.

Hanson has thought about this longer than I have and puts it better, in a way that turns the objection partly on its head. Writing about expanding bounty hunting , he names the same downside and then draws out what it exposes:

we’d actually uniformly enforce the laws on our books. Politically connected people and districts could less use “suction” to tilt the scales of justice their way, and we’d have to admit when we aren’t serious about enforcing laws with very low bounties.

Sit with the second half of that. The discomfort we feel at the prospect of uniform enforcement is not really a fear that the mechanism will work. It is a recognition that our current legal order depends on selective non-enforcement — on a large stock of laws nobody intends to apply consistently, whose actual severity is set by prosecutorial discretion rather than by statute, and which therefore fall hardest on whoever lacks the connections to attract leniency. The existing arrangement is not a safeguard against bad laws. It is a bad law survival mechanism, plus arbitrariness.

And it gives the design a dial I had not noticed. Under a bounty regime the seriousness of a prohibition is a published number. A legislature that wants a law on the books without meaning it has to set a low bounty and say so in public, where the low number is legible to everyone as an admission. That is a considerable improvement on the status quo, where the same admission is made silently by declining to prosecute and nobody has to own it.

I do not think this dissolves the objection. Cheap uniform enforcement of a bad law is still bad, a published bounty can still be set high on something that should not be prohibited at all, and an industry of enforcers is still a constituency for expansion. But the honest comparison is not between this mechanism and a world of carefully chosen laws. It is between this mechanism and a world where enforcement is rationed by discretion, and I am no longer confident the status quo wins that comparison on the merits rather than on familiarity.

So I will put it as a cost rather than as a solved problem. This proposal makes frontier-slowing cheap, and cheap things get done more often, including the wrong ones. I accept that cost because of what Part I says is at stake. Someone who does not accept Part I should not accept the cost, and I would rather lose them here, honestly, than win them with an answer I do not have.

What the delay is for

The mechanism does not solve anything. It is worth being exact about this: at its most successful it does not produce safe AI, does not answer any question in Part I, and does not make anybody wiser. It buys time.

Time happens to be what this particular problem calls for. The questions that Part I says we cannot presently answer — who is a moral patient, whether experience travels with capability, whether there is anybody home in a system that talks like a person — are not obviously permanently unanswerable. They are unanswered. There is no instrument today; there might be one in fifty years, and if there is, the entire shape of this argument changes and I will be glad to abandon it. What would be unforgivable is committing irreversibly while the question is still open, on the grounds that waiting was inconvenient.

That is the whole case, and it does not require me to be right about very much. It requires only that the outcome is possible, that we cannot currently check, and that there is no second attempt. Everything from the littering fine onward is machinery for buying the interval in which somebody might find out.

The machinery is economic because the problem is economic. The danger does not come from a lone genius; it comes from large, well-funded teams, and teams are made of individuals making career decisions under incentives that can be changed by statute. That is a lever, and it is a more accessible one than the mathematics of alignment.

It is easier to shift the Nash equilibrium of working on AI in the first place than it is to create safe AI. Financial incentives have driven the vast majority of AI improvements in the last decade, and financial incentives can be used to stop them.

To cite this page: Andrew Quinn, "13. Next Steps." Extinction Bounties, last revised 2026-09-11. https://extinction-bounties.com/sequence/13-next-steps/

Policy-research disclaimer

Extinction Bounties publishes theoretical economic and legal mechanisms intended to stimulate scholarly and public debate on catastrophic-risk governance. The site offers policy analysis and advocacy only in the sense of outlining possible legislative or contractual frameworks.

No legal or financial advice

Nothing here should be treated as a substitute for qualified legal counsel, financial due diligence, or regulatory guidance. Readers remain responsible for ensuring their actions comply with the laws and professional standards of their own jurisdictions.

Exploratory and personal views

All scenarios, numerical examples and opinions are research hypotheses presented by the author in a personal capacity. They do not represent the views of the author's employer, funding bodies, or any governmental authority.

Implementation caveats

Any real-world adoption of these ideas would require democratic deliberation, statutory authority, and robust safeguards against misuse. References to enforcement, penalties, or "bounties" are illustrative models, not instructions or invitations to engage in private policing or unlawful conduct. Nothing here is directed at any identifiable individual — see non-targeting.

No warranty and limited liability

Content is provided "as is" without warranty of completeness or accuracy; the author disclaims liability for losses arising from reliance on this material.

By continuing beyond this notice you acknowledge that you have read, understood, and accepted these conditions.