Extinction Bounties

Policy-based deterrence for the 21st century.

The Sequence · Part II · What To Do

07. The Gist

Economics and game theory against sophisticated offenders.

Last revised 11 September 2026 · alpha

The littering machine works because littering is easy to catch. Somebody sees you do it, the evidence is a wrapper with your fingerprints on it, and nobody involved has planned anything. Take those conditions away and the whole apparatus should fall over.

This page is about why it does not, and it is where the argument I have been making since 2022 actually lives.

Where ordinary law stops working

Start with the standard economic account. Gary Becker’s Crime and Punishment: An Economic Approach (1968) treats an offense as a decision under uncertainty: what deters is not the severity of the punishment but its expected severity — roughly the probability of being caught and convicted multiplied by what happens if you are. Two knobs, and the product is what the prospective offender actually faces.

Most of criminal justice policy is an argument about the second knob, because the second knob is the one legislatures can turn by writing a number in a statute. The first knob is expensive. Raising the probability of detection means more investigators, more surveillance, more of the state’s attention, and it runs into diminishing returns almost immediately for any offense that is committed carefully.

Robin Hanson’s diagnosis, in Privately Enforced & Punished Crime (2018), is that ordinary law works by letting the injured party do the detecting, which is fine for accidents and mild selfishness among people who already deal with each other — and fails completely when someone plans, in advance, to make themselves and the evidence hard to find. What a modern society does instead is hire investigators and give them special powers; what it then has to do, because it does not entirely trust them, is limit those powers with juries, rules of evidence and standards of proof. The result still leaks, because officials coordinate well enough to build a wall of silence of their own.

Hanson's own statement of the problem

Non-crime law deals mostly with accidents and mild sloppy selfishness among parties who are close to each other in a network of productive relations. In such cases, law can usually require losers to pay winners cash, and rely on those who were harmed to detect and prosecute violations. This approach, however, can fail when “criminals” make elaborate plans to grab gains from others in ways that make they, their assets, and evidence of their guilt hard to find.

Ancient societies dealt with crime via torture, slavery, and clan-based liability and reputation. Today, however, we have less stomach for such things, and also weaker clans and stronger governments. So a modern society instead assigns government employees to investigate and prosecute crimes, and gives them special legal powers. But as we don’t entirely trust these employees, we limit them in many ways, including via juries, rules of evidence, standards of proof, and anti-profiling rules. We also prefer to punish via prison, as we fear government agencies eager to collect fines. Yet we still suffer from a great deal of police corruption and mistreatment, because government employees can coordinate well to create a blue wall of silence.

Notice how precisely this describes the case I care about. Frontier AI work is not sloppy selfishness among parties close to each other in a network of productive relations. It is deliberate, well-resourced, conducted by people who are extremely good at their jobs, and — critically — it produces no injured plaintiff who can detect and prosecute the violation. Tort law wants a harmed party to bring a suit. The harm I am worried about in Part I does not leave one.

So the first knob is stuck near zero and the second is unavailable, because there is no one left to sue on behalf of. That is the hole the mechanism is meant to fill.

The proposal

Hanson’s move is to stop treating detection as a government function:

I propose to instead privatize the detection, prosecution, and punishment of crime. […] The key idea is to use competition to break the blue wall of silence, via allowing many parties to participate as bounty hunter enforcers and offering them all large cash bounties to show us violations by anyone, including by other enforcers. With sufficient competition and rewards, few could feel confident of getting away with criminal violations; only court judges could retain substantial discretionary powers.

Three features are doing the work. Punishment is a fine rather than a prison term, which makes it a transferable quantity that can fund its own enforcement. Anyone may collect, including other enforcers, which means there is no cartel of detectors to capture. And the discretion that remains — the part that cannot be mechanized without becoming monstrous — sits with judges, where we already have several centuries of practice at constraining it.

Hanson is not writing about AI. He is proposing a general replacement for criminal enforcement, and the crimes he has in mind are ordinary ones. What follows is my application, and he bears no responsibility for it.

The arithmetic of silence

Here is why this bites hardest on exactly the kind of activity I want to deter.

Take a group of people who all know about something they are not supposed to be doing. For the secret to hold, every one of them has to stay quiet — so the chance the arrangement survives is a probability multiplied by itself once per person. That is the whole observation, and it is unremarkable as arithmetic and quite remarkable as policy, because headcount sits in the exponent rather than in the base.

What that means concretely is that secrecy does not decay gently as a group grows. It falls off a cliff. Take a hundred people who are each 99% reliable over a year — a one in a hundred chance of saying something, which is a very loyal group — and you get roughly a two-in-three chance of exposure per year out of people who are individually almost perfectly discreet. Move each person to 97%, which is the sort of shift a large enough payment might buy, and it goes to about 95%.

The arithmetic, and the table it produces

With nn people who each stay quiet over the period with independent probability pp, the arrangement survives with probability pnp^{\,n}, so

P(at least one speaks)=1pnP(\text{at least one speaks}) = 1 - p^{\,n}
Group size nnp=0.99p = 0.99p=0.97p = 0.97p=0.95p = 0.95
1010%26%40%
3026%60%79%
10063%95%99%
30095%>99.9%>99.9%

Read down a column rather than across a row: holding individual loyalty fixed, the thing that destroys secrecy is scale.

This quantity reappears throughout Part II as qq, the probability that a claim lands: it is what an underwriter is pricing and what makes a team’s expected liability climb faster than its headcount.

So the thing that destroys secrecy, holding individual loyalty fixed, is scale. And scale is not optional here. Page 05 argued that the danger has never come from a lone genius but from large, well-funded groups working closely together — which is a claim about the threat model that turns out, on this page, to also be a claim about the attack surface. The same fact that makes frontier work dangerous makes it leaky. A serious lab is a several-hundred-person operation, and it cannot become a ten-person operation without ceasing to be able to do the thing.

That last sentence is an empirical claim, it is the one an opponent should attack first, and I have since gone and checked it. The honest finding is that nobody knows. The public numbers are large and getting larger, but the one public figure for the people actually on the critical path of a training run is about five — and both of those are true, because they answer different questions.

What the published author counts actually show

Author counts on frontier model papers have risen roughly two orders of magnitude in five years: six on GPT-2 , 31 on GPT-3 , around 280 on GPT-4 , 1,350 on Gemini 1.0 .

But the proxy is close to worthless, and the proof is that Gemini 1.5 ’s author list grew from 662 to 1,136 across revisions of the same paper. Meanwhile Google’s pre-training lead has said on record that keeping a 40-day training run alive took a rotation of about five people.

Team Size Is a Governance Variable works through which number belongs in the exponent above, why nobody outside the labs currently knows, and what it would take to find out.

The honest caveat, up front. Independence is the weakest assumption in this model and I am not going to bury it. Real colleagues are not independent draws. They are selected for shared conviction, bound by contracts, embedded in friendships and mortgages and visa sponsorships, and — as Hanson himself notes in Easy Conspiracy Tests (2020) — a group with a credible threat of punishing whoever exposes it can stay intact at sizes this model says are hopeless. Correlation cuts those numbers down, possibly by a lot. It does not eliminate the exponent, and the bounty is specifically an instrument for attacking the correlation. But everything above is an illustration of a mechanism, not an estimate, and page 12 treats the objection properly.

Having written that caveat defensively, I now think it is better read as a statement of the target.

Put cohesion into the model as its own quantity — call it ρ\rho, running from zero for a group of strangers to one for a group that simply holds — and everything so far turns out to have been the ρ=0\rho = 0 case of a larger picture. Two things follow, and they point in opposite directions.

The bad one first: under correlation, the discovery rate no longer climbs without limit as the team grows. It flattens and stops, at a ceiling set by how cohesive the team is. Ten thousand employees does not help. That is a real wound, it is worse at exactly the high-conviction mission-driven organisations this argument is aimed at, and page 12 is where I take it seriously rather than here.

The good one is that it tells you what the mechanism is actually for. What protects a conspiracy is not its size but its cohesion — a thousand people who are all the same person leak exactly as readily as one. So a large payment, claimable by one person, without needing anybody else’s agreement, is not merely exploiting the exponent. It is a device for driving ρ\rho down: a shock applied individually to each member of a group whose only protection is that its members move together. The mechanism does not just use effective team size. It manufactures it.

Cohesion, and the ceiling it puts on discovery

Let ρ\rho be the probability that the team is simply the kind that holds — shared conviction, strong enough culture, a sufficiently frightening NDA — and suppose that otherwise members decide independently as before. Then

P(nobody talks)  =  ρ+(1ρ)pnP(\text{nobody talks}) \;=\; \rho + (1-\rho)\,p^{\,n}

At n=100n = 100 and p=0.99p = 0.99, a modest ρ=0.3\rho = 0.3 takes the discovery rate from 63% down to about 44%. The optimistic case above — bounties pushing pp to 0.97, discovery to 95% — becomes about 67%. Still a large improvement. Considerably less than advertised.

The structural damage is worse than the numerical damage. Take the limit:

limn[1ρ(1ρ)pn]  =  1ρ\lim_{n \to \infty} \left[1 - \rho - (1-\rho)\,p^{\,n}\right] \;=\; 1 - \rho

Discovery is capped at 1ρ1-\rho no matter how many people are involved.

The same parameter can be written as an effective team size, neff=1+(n1)(1ρ)n_{\text{eff}} = 1 + (n-1)(1-\rho), which is the same claim in the units the rest of Part II uses: at ρ=0\rho = 0 it recovers nn, and at ρ=1\rho = 1 it collapses to a single decision-maker however many people are in the building. Substituting neffn_{\text{eff}} for nn in the table above is the honest version of every number on this page.

I develop both in Team Size Is a Governance Variable .

What actually moves p

The arithmetic is only interesting if pp is something we can change. Becker says the two knobs are magnitude and probability of detection; the claim here is that a bounty turns the second knob by converting every person in the building into a potential detector, which is cheaper than any amount of policing.

So, concretely, at $1,000 a head: would my own probability of keeping quiet drop from 99% to 97%?

For $99,000 in a hundred-person firm — absolutely. People grind Leetcode for months to get comp packages like that. Even if I had to implicate myself in the documents I released, I would in large part be paying a bounty to myself. And if I fully believed in the mission, I could still tell myself that this kind of runway buys a decade of doing whatever I actually think is right.

That is the whole trick. Not that anyone becomes a hero, and not that anyone stops believing in what they are doing. Just that the price of silence acquires a number, and the number is legible to the sort of person who negotiates offers for a living.

This is not hypothetical

The part of the mechanism that people find least plausible — that a government would pay large sums to informants out of money taken from offenders — is the part that already exists.

Offender-funded enforcement already runs at nine figures in American law. The SEC pays its whistleblowers out of a fund financed entirely by sanctions collected from violators, and its largest award to date is nearly $279 million to a single person. On the civil side the False Claims Act does the same through qui tam , where a former employee who pursued a case for a decade without the government joining him took home roughly $266 million.

The two precedents, with the numbers

In May 2023 the SEC announced its largest-ever whistleblower award , nearly $279 million to one person. The press release includes the sentence that matters most here:

Payments to whistleblowers are made out of an investor protection fund, established by Congress, which is financed entirely through monetary sanctions paid to the SEC by securities law violators. No money has been taken or withheld from harmed investors to pay whistleblower awards.

In 2022 Biogen settled for $900 million in a False Claims Act case brought by a former employee, Michael Bawduniak, who received a relator share of roughly $266 million — about 30% — after pursuing it for a decade without the government intervening.

I want to be careful about what this does and does not establish. It does not show that bounties work against catastrophic risk; securities fraud and Medicare kickbacks are compensable harms with identifiable victims, which is exactly the condition I said was missing. What it shows is that the institutional machinery is built, the constitutional objections have been litigated, and the sums involved are already in the range this argument needs. The novel part of my proposal is not the payment. It is the target.

The effect that matters more than detection

Return to the hundred-person firm, and suppose I decide against turning my colleagues in.

Do I want to keep working there?

Absolutely not. Every additional person hired is another draw against me. My own exposure rises with the size of the team, monotonically, and there is nothing I can do about it except leave. The mechanism does not need me to inform in order to change my behavior — it only needs me to notice the arithmetic and update my career plans.

This is the second-order effect, it is larger than the first, and page 09 is about it. A bounty regime that never pays out a single bounty, because everyone quietly moved to a different industry before anything happened, would be a complete success.

The line

It is easier to shift the Nash equilibrium of working on AI in the first place than it is to create safe AI. Financial incentives have driven the vast majority of AI improvements in the last decade, and financial incentives can be used to stop them.

That is the claim the entire site is built on, and it is worth being precise about its shape. It is not a claim that we would catch people. It is not a claim that anyone would go to prison. It is a claim about which of two hard problems is harder — that repricing a labor market is a thing societies know how to do, and that specifying values into an optimizer is not.

How large

A thousand dollars is far too low. It was a useful number for the littering example on page 06 precisely because it is small enough to feel real, but against these stakes it is noise — and the money comes out of the pocket of the guilty party rather than the public purse, so the usual budgetary objection to large awards does not apply.

A better peg is something like ten years of total compensation: about $2.25 million as of 2022, and certainly higher now. That is a number chosen to make the expected value of the career negative rather than to punish, which is a different design target from most of criminal law and produces different answers. The sizing note works the question properly.

It also raises the obvious problem: almost nobody can pay $2.25 million on demand. A fine nobody can pay deters nobody, because everyone involved knows it will not be collected. Solving that is what the insurance leg is for, and it turns out to be the part of the machine that does most of the actual deterring. That is page 08 .

To cite this page: Andrew Quinn, "07. The Gist." Extinction Bounties, last revised 2026-09-11. https://extinction-bounties.com/sequence/07-the-gist/

Policy-research disclaimer

Extinction Bounties publishes theoretical economic and legal mechanisms intended to stimulate scholarly and public debate on catastrophic-risk governance. The site offers policy analysis and advocacy only in the sense of outlining possible legislative or contractual frameworks.

No legal or financial advice

Nothing here should be treated as a substitute for qualified legal counsel, financial due diligence, or regulatory guidance. Readers remain responsible for ensuring their actions comply with the laws and professional standards of their own jurisdictions.

Exploratory and personal views

All scenarios, numerical examples and opinions are research hypotheses presented by the author in a personal capacity. They do not represent the views of the author's employer, funding bodies, or any governmental authority.

Implementation caveats

Any real-world adoption of these ideas would require democratic deliberation, statutory authority, and robust safeguards against misuse. References to enforcement, penalties, or "bounties" are illustrative models, not instructions or invitations to engage in private policing or unlawful conduct. Nothing here is directed at any identifiable individual — see non-targeting.

No warranty and limited liability

Content is provided "as is" without warranty of completeness or accuracy; the author disclaims liability for losses arising from reliance on this material.

By continuing beyond this notice you acknowledge that you have read, understood, and accepted these conditions.