Extinction Bounties

Policy-based deterrence for the 21st century.

The Sequence · Part II · What To Do

09. The Goal

Make dangerous research a bad career move.

Last revised 11 September 2026 · alpha

Here is the win condition, stated first because it inverts how people normally read a proposal like this one.

Nobody is ever prosecuted. No fine is levied, no bounty is claimed, no case comes to court. The statute sits on the books, the insurance market prices it, the labor market reads the price, and the projects that would have triggered it are never staffed in the first place. A mechanism that never fires has not failed. If it never fires, it has worked perfectly.

This is worth sitting with, because enforcement proposals are usually evaluated on conviction rates and this one should be evaluated on the opposite. Every prosecution is a project that got far enough to be worth prosecuting. The prosecutions are the leakage, not the product.

What the mechanism is actually for is repricing a labor market.

Why the labor market is the lever

Page 05 established the load-bearing empirical claim: no individual has ever had the capacity to build even an unsafe AI alone, and the proof is that we are still here. It is load-bearing enough that I have given it its own treatment , because the whole argument below is an exponent applied to a number nobody has measured. Few people are smart enough; of those, few have the conscientiousness to carry such a project by themselves; and of those, almost all have found better things to do than labor in solitude. The danger comes from large, well-funded groups working closely together.

It is theoretically possible that one person working in total isolation could do it. It seems unlikely in a world where no one person can make a commercial pencil by themselves — the point of Leonard Read’s I, Pencil , which the Competitive Enterprise Institute made into a six-minute film in 2012 and which is the best short introduction to dispersed knowledge I know of. The pencil argument cuts both ways here, and I want both halves. Nobody can make a pencil alone, which is why the coordination that produces one is miraculous; and nobody can build a frontier model alone, which is why the coordination that produces one is legible. The same fact that makes markets astonishing makes conspiracies detectable. In practice you need capital — the hardware, the power, the data — and you need labor, in a lot of specializations at once. Both leave reliable heat signatures. Capital is legible because it is large and someone has to sign for it. Labor is legible because people have names, résumés, mortgages, and opinions.

So the thing standing between the world and the project is not a technical barrier. It is a hiring pipeline. Hiring is a market. Markets respond to expected cost. That is the entire theory of this page.

There is a further consequence worth drawing out, because it is the strongest thing the mechanism has going for it and the old version of this argument missed it entirely.

The arithmetic is superlinear in exactly the right variable

Page 07 gives the conspiracy result: the more people who know, the worse the odds that all of them stay quiet. Bounties push each person’s willingness down. That much is the standard argument.

Now look at what it does to a team rather than to an individual. Two things get multiplied together, and both of them grow with headcount: more people means more fines to pay, and more people means a higher chance of paying any of them at all. The result is that expected liability does not rise in proportion to the size of the team. It rises much faster.

The superlinear tax, with numbers

If each participant faces a fine FF on conviction, the expected total liability of a team of size nn is roughly

L(n)  =  n(1pn)FL(n) \;=\; n\,\bigl(1 - p^{\,n}\bigr)\,F

Take p=0.99p = 0.99 — each participant has only a one percent chance of speaking up in the period — and FF at page 07’s ten-years-of-compensation peg, call it $2.25 million:

Team size1pn1 - p^{\,n}Expected total liabilityPer person
100.096$2.2m$215k
1000.634$143m$1.4m
1,000≈1.000$2.25bn$2.25m

The per-head cost rises by nearly seven times between a team of ten and a team of a hundred, and the total by a factor of sixty-six. Nothing about the individuals changed. Only the headcount did.

Illustrative, not estimated: this inherits page 07’s independence assumption, and the cohesion term flattens the curve considerably.

That is the property you want, and it is not one you can easily get any other way. Page 05 says the danger scales with team size. This mechanism taxes team size superlinearly. The hundred-and-first hire does not merely add their own expected liability; they raise the detection probability faced by the hundred people already there. Inside the firm, each additional person is a negative externality on all the others.

The policies are bought by individuals — that is the whole point of insuring the person rather than the firm — but the firm is ultimately the one paying for them, and it pays whether it means to or not. Either it subsidises premiums outright, which labs would be under considerable pressure to do, or it does not and the cost arrives as wage demands from people who now have a large, legible, annually repriced expense attached to working there. A researcher whose premium went up by forty thousand dollars because the team grew is going to want forty thousand dollars. There is no third possibility in which the cost simply evaporates.

So the externality gets internalized by the firm regardless of who signs the cheque, and it shows up somewhere the firm cannot ignore: the wage bill. That is also the answer to the worry that employer reimbursement defeats the whole scheme. It does blunt the individual’s incentive, which is a real cost and page 12’s problem. It does not blunt the firm’s, because the rate stays individually priced — a lab that reimburses pays more for riskier people on more dangerous projects, and the aggregate lands in its accounts either way.

Two caveats, both real. The model is crude: it assumes independence between participants — the assumption page 07 now handles with an explicit cohesion term — along with a fixed period, no coordination, and that every convicted person pays in full. And critically, these numbers describe a team engaged in the prohibited activity. A compliant team’s expected liability is not L(n)L(n) but approximately zero, because there is nothing to report. The tax falls on the activity, not on the headcount as such — which is exactly the targeting property a blunt instrument like a compute cap cannot offer.

The individual calculation

From inside a career, the numbers are smaller and the decision is simpler.

The fine is insured — that is the whole point of the “insured” in fine-insured bounties, and page 08 covers why. So what an individual actually faces is not the fine. It is the premium, priced by someone with money at stake who has read the same arithmetic above.

Which means the deterrent does not have to be catastrophic. It has to clear a much lower bar: the premium has to exceed the pay rise. A frontier job is worth taking while what it pays over the next-best alternative is more than what it costs you to insure. If a frontier role pays three hundred thousand dollars more than the next-best job, the mechanism needs a premium north of three hundred thousand a year — not a fine equal to all the money in the world. And on a hundred-person project, by the arithmetic above, the expected annual liability per head is already several times that.

The condition, stated

With a covered role paying CC, the best uncovered alternative paying CC', and an annual premium π\pi, the covered role is worth taking only while

Cπ  >  CC - \pi \;>\; C'

so the decision flips as soon as π\pi exceeds CCC - C'. The quantity that has to move is a difference between two salaries, not the value of a human life.

I want to be careful about how much weight this bears. The numbers are illustrative and I have picked them to be legible rather than accurate; I do not know what compensation differentials or premiums would actually look like, and the note on bounty sizing is where that question gets worked properly. The point is structural. The quantity that has to move is a difference between two salaries, not the value of a human life, and that is a far more modest thing to ask of a policy.

The effect is exit, not conviction

Now the part that does most of the work, and which page 07 introduces in one line: the second-order effect matters more than the first.

Ask what happens to someone who considers the bounty and decides not to report their colleagues. They have not thereby made themselves safe. They are now sitting inside a building where everyone else faces the same offer, where the detection probability rises with every hire, and where their own exposure is determined entirely by other people’s choices. The rational response to declining the bounty is not to relax. It is to leave.

Publish or perish becomes pause or perish. And note that the exits do not require anyone to be virtuous, persuaded, or even particularly worried about AI. They require only that people notice their compensation has stopped covering their exposure. That is a much weaker thing to rely on than convincing people — page 05 is explicit that you will not argue anyone out of a position that bottoms out in where they think qualia comes from, and this mechanism does not try to.

The effect compounds in the direction you want. Departures shrink the team, which would lower 1pn1-p^{\,n}, except that the people who leave still know what they know. The project either shrinks below the size at which it can accomplish anything, or it keeps hiring and keeps climbing the curve.

Where those people go, and whether that is good

This is the part I am least sure about, so let me be plain about it.

The optimistic version: capable people leave frontier capability work for adjacent fields — safety and interpretability, other areas of software, other sciences entirely — and the world gets the benefit of their work without the tail risk. Some of them do safety work precisely because the market has just made it the better-paid option relative to its risk-adjusted alternative, which is a result nobody has managed to produce by appealing to conscience.

The pessimistic version is that this is displacement rather than reduction. The work moves rather than stopping: to jurisdictions that have not adopted the statute, into classified programs where the reporting channel does not exist, into firms structured so that no individual is legibly a participant. If that is what happens, the mechanism has not reduced risk. It has relocated it, and possibly into hands with less external scrutiny than a publicly traded lab faces.

I do not think the pessimistic version is the default, but I hold that loosely, and it is the single objection I would most want an economist to check before anyone drafted anything. Page 10 is about the jurisdictional half of it. There is no page that fully answers the classified- programs half, and I would rather say so than pretend.

The gap that already exists

Here is the strongest empirical point available to this page, and it is not hypothetical.

Whistleblower machinery for AI now exists, and every piece of it protects the whistleblower’s job. Not one of them pays.

The statutes now on the books

California’s SB 53, New York’s RAISE Act, and the federal AI Whistleblower Protection Act introduced by Senator Grassley all establish protections for people who report safety problems at frontier labs.

None of the three attaches a payment.

That is not an oversight, and the people designing this legislation know the option is there. The Institute for Law & AI’s guide to designing AI whistleblower legislation lists among its decision points for lawmakers, explicitly, whether whistleblowers should be offered monetary or other incentives to come forward. It is on the table. It has simply not been picked up. Their companion piece on protecting AI whistleblowers covers the ground that has.

Why the slot stayed empty

I think I know why, and the two reasons are the two things this mechanism is unusually good at.

Somebody has to pay, and it looks like the taxpayer. A bounty is a cheque. Cheques come from appropriations, appropriations are fought over, and a line-item reading payments to AI lab employees is not something a legislator wants to defend in an election year. The path of least resistance is to write the protections, which cost nothing visible, and leave the payment question for a later bill that never comes.

And the cheque would have to be enormous. This is the part I think is underappreciated. A whistleblower award has to beat the expected value of staying quiet, and for this population staying quiet is worth an extraordinary amount — total compensation in the high six or seven figures, equity that vests over years, and a career in the most capital-rich industry on earth. A flat statutory award of fifty or a hundred thousand dollars does not move anyone. It is not a small version of the incentive; it is no incentive at all. To work, the number has to be in the millions, and a number in the millions is exactly the number a legislator cannot put next to the word taxpayer.

So the two objections compound. The payment must be large, and large payments are politically impossible when they appear to come from the public purse. The slot stays empty not because anyone thinks incentives do not work, but because nobody has a good answer to who writes the cheque.

This mechanism has one, and it is the entire reason I keep insisting the insurance and the bounty are a single design rather than two.

The money comes from the offender. The fine funds the bounty; the insurance funds the fine; the premium funds the insurance; and the premium is paid by the individual and passed through to the lab in wages, by the argument earlier on this page. At no point does an appropriation appear. This is not free — the activity being deterred pays for its own deterrence, which is the point — but the party paying is the party doing the dangerous thing, not the public.

And it is not speculative. Congress has already built exactly this: the nine-figure precedents on page 07 were paid out of wrongdoers’ sanctions rather than out of anyone’s taxes, which is why the SEC programme is described as free to the taxpayer .

So the fiscal objection is, strictly, already answered by precedent. What the bounty design adds is that the money never routes through a government fund at all — it goes from an insurer to a claimant — which removes the appearance of public expenditure as well as the substance. The optics problem and the accounting problem both dissolve.

The second objection dissolves along with it, and this is the part I find genuinely neat: the size problem and the funding problem have the same solution. Peg the bounty to the offender’s own compensation — ten years of it, on page 07’s suggestion — and the award automatically scales to whatever the industry is paying, without anybody having to legislate a dollar figure that will be wrong in three years. A number large enough to outbid a frontier salary is affordable precisely because it is drawn from a frontier salary.

Notice exactly what the existing statutes do and do not change. Anti-retaliation protection lowers the cost of speaking for someone who has already decided to speak. It does nothing whatsoever to the expected value of taking the job, which is the only variable this page cares about. A protected whistleblower is still a person who took the job. Protection is a floor under one decision; a bounty is a price on a different one.

This is the narrowest possible statement of what I am proposing: a payment schedule attached to machinery that already exists, in a slot that its own designers have already identified and left empty. Everything else on this site is argument about why that slot should be filled. See also what about turning yourself in , because the answer there turns out to be a feature rather than a bug.

What this does not reach

State programs. Classified work. Any jurisdiction that has not adopted the statute, which at present is all of them.

Those are not edge cases and I am not going to treat them as such. A deterrent that stops at a border relocates a project rather than preventing it, and the projects most worth preventing are precisely the ones with the resources to relocate. That is the hardest practical problem this proposal has, it is the one I have the least confidence about, and it gets its own page .

One last thing, said plainly because this page more than any other could be misread. Everything above describes the aggregate effect of a statutory penalty structure on a labor market. It is a claim about prices and hiring, addressed to legislators who might write such a statute and to economists who might tell me why it would not work. It is not directed at any person, it does not name anyone, and nothing in it describes an action that a private individual should take. See non-targeting .

To cite this page: Andrew Quinn, "09. The Goal." Extinction Bounties, last revised 2026-09-11. https://extinction-bounties.com/sequence/09-the-goal/

Policy-research disclaimer

Extinction Bounties publishes theoretical economic and legal mechanisms intended to stimulate scholarly and public debate on catastrophic-risk governance. The site offers policy analysis and advocacy only in the sense of outlining possible legislative or contractual frameworks.

No legal or financial advice

Nothing here should be treated as a substitute for qualified legal counsel, financial due diligence, or regulatory guidance. Readers remain responsible for ensuring their actions comply with the laws and professional standards of their own jurisdictions.

Exploratory and personal views

All scenarios, numerical examples and opinions are research hypotheses presented by the author in a personal capacity. They do not represent the views of the author's employer, funding bodies, or any governmental authority.

Implementation caveats

Any real-world adoption of these ideas would require democratic deliberation, statutory authority, and robust safeguards against misuse. References to enforcement, penalties, or "bounties" are illustrative models, not instructions or invitations to engage in private policing or unlawful conduct. Nothing here is directed at any identifiable individual — see non-targeting.

No warranty and limited liability

Content is provided "as is" without warranty of completeness or accuracy; the author disclaims liability for losses arising from reliance on this material.

By continuing beyond this notice you acknowledge that you have read, understood, and accepted these conditions.