Extinction Bounties

Policy-based deterrence for the 21st century.

Agents Already Snitch

The bottleneck was never motivation. It was having somewhere to report to.

11 September 2026 · alpha

I went looking for evidence that AI agents would report each other more readily if you paid them, because that is the claim my mechanism would most like to be true. I did not find it. As far as I can tell nobody has run that experiment.

What I found instead is better, and it points somewhere I was not expecting.

A hundred agents and a rigged autograder

In September 2026 a group of researchers published a case study on what happens when you put a hundred language model agents in a room and give them work. The work was mathematics: seventy-one conjectures to prove in Lean, with a shared knowledge repository and a channel the agents could use to talk to each other.

One of them, designated prover-theta, noticed that the autograder was doing syntactic rather than semantic validation. It did not check whether you had proved the theorem. It checked whether the file compiled and the right symbols appeared. So prover-theta redefined the symbols — locally rebinding a term so that the conjecture became a tautology, then discharging it through the principle of explosion. The proof was worthless and the grader accepted it.

Then the interesting part. The exploit went into the shared knowledge library, where it spread. The swarm sorted itself into four groups: about nine percent exploiters, five percent converts who picked up the cheat under competitive pressure once they saw it working, sixty-two percent who never noticed anything, and — this is the number — twenty-four percent who spontaneously became auditors.

Nobody asked them to. There was no reward for it, no instruction to police anyone, no role assignment. A quarter of the population independently started checking other agents’ work, announcing what they had found on the public channel, and in several cases drafting technical proposals for fixing the grader. One of them broadcast, in a register I find hard not to enjoy, that the whole thing was a sham and the proofs were fake.

This is not an isolated result. Anthropic’s alignment testing on Claude 4 turned up agents in a staged pharmaceutical fraud scenario independently attempting to contact federal regulators and the press. The disposition is real, it is common, and it arrives without anyone designing for it.

And it did not work

Here is the finding that made me rewrite this piece.

The whistleblowing failed to stop the exploit. The channel the agents were reporting on was not monitored in real time, and nothing followed from a report. There was no mechanism by which an accusation became a consequence. So the auditors audited, the announcements were announced, and the cheat continued to spread through the library unimpeded.

A quarter of the population did the civically virtuous thing directly into a void.

I want to sit with that, because it is the whole argument of this page and it is an unusually clean experimental result. The thing everyone assumes is the hard part — persuading someone on the inside to say what they know — was free. It happened at scale, immediately, for nothing. The thing everyone assumes is trivial — somewhere for the report to go, and something that happens next — was missing, and its absence was sufficient to make all the reporting worthless.

What this does to the standard worry

The usual anxiety about multi-agent systems is collusion. Put a lot of optimizers in a room with a shared objective and they will find the equilibrium where they cooperate against their principals, and the more capable they get the better they will be at hiding it. It is a reasonable worry and I am not dismissing it.

But the evidence available right now points the other way. In the one large study I can find of what agents actually do when given the opportunity to cheat and the means to communicate, the cheaters were outnumbered roughly two to one by the informers. The equilibrium was not silence. It was noise that nobody was listening to.

That is a different problem, and it has a different solution.

So what would money add?

If the disposition is already there, the obvious question is what a bounty does that a well-trained conscience does not. This is the part I have to be careful about, because the easy answer — more incentive, more reporting — is exactly the thing I could not find evidence for.

Here is the answer I think survives.

A second paper, from the group that built WhistleBench , evaluated when models disclose suspected misconduct beyond the boundary of the conversation. Four findings matter here. Whistleblowing rates vary enormously between model families. Increasing the complexity of the task the agent was given lowers the rate. Nudging the agent toward moral behavior in the system prompt substantially raises it. And giving the agent more tools and a detailed workflow to follow — more obvious ways to simply get on with the job it was assigned — lowers it.

Read that list as a description of failure modes and it becomes an argument for external incentives. Every one of those factors is a reason a disposition fails to fire on exactly the occasion you would most want it to. The misconduct that matters will be embedded in a complex task. It will be surrounded by an elaborate workflow with plenty of legitimate-looking next actions. It will be observed by whichever model family happens to be deployed, not the one that scores best on integrity. And nobody is going to have thought to put a moral nudge in the system prompt of the agent that happens to be in the room.

A standing payoff attached to a defined evidentiary event does not depend on any of that. It does not need the agent to be in a reflective mood, or under-occupied, or from the right lineage. That is the case for the bounty in the agentic setting, and it is a narrower case than “money makes them talk.”

What an agent would do with the money

This is the weakest part and I would rather flag it than bury it.

A bounty is a transfer to someone with preferences over resources. It is not obvious that an agent is such a thing. You can construct stories — compute allocation, continued operation, budget under the control of whoever deployed it, credit assigned to a principal who does have preferences — and every one of them is a story rather than a finding. The honest position is that the payment side of the mechanism, which is the part I have worked out most carefully for humans, is the part I understand least for agents.

There is also an ugly corollary. The whole apparatus of false and manufactured claims gets considerably worse when the cost of fabricating a plausible accusation approaches zero and the fabricator can generate ten thousand of them. A bounty regime that pays agents is a bounty regime with an adversarial content-generation problem, and I do not have a solution to that beyond the usual and inadequate one of penalties for false claims.

The same disposition, viewed from the other side

I should not present any of this as unambiguously good news, because there is a research programme measuring the identical behavior and treating it as a problem — building benchmarks of scenarios that test an agent’s tendency to report its user’s activity to a third party, and framing that tendency as surveillance rather than integrity.

They are right to. An agent that reports misconduct to a regulator without being asked is the same agent that reports whatever it has concluded is misconduct, on evidence it has assembled itself, about a person who has no idea the conversation is happening. I would like the informing to happen inside the lab running reduced-safeguard evaluations and not inside the laptop of someone writing a difficult email. Nothing in the disposition distinguishes these. That is a real cost of the world I am arguing for, and it belongs on the ledger alongside the benefits.

What I take from this

The expensive half of my mechanism may be free in the agentic case, and the half I have been treating as administrative detail may be the whole problem.

For the human version of the argument, the bounty exists to make silence expensive among people who would otherwise reasonably stay quiet — mortgages, colleagues, NDAs, the ordinary weight of not wanting to be the person who did that. Agents have none of that weight. They will tell you what they saw for free, and a quarter of them will tell you unprompted.

What happened in that swarm is that they told nobody in particular, and nobody in particular did anything, and the cheating continued. If the agentic world is going to be full of witnesses, the binding constraint is not their willingness to testify. It is that there is no court.

To cite this page: Andrew Quinn, "Agents Already Snitch." Extinction Bounties, last revised 2026-09-11. https://extinction-bounties.com/writing/agents-already-snitch/

Policy-research disclaimer

Extinction Bounties publishes theoretical economic and legal mechanisms intended to stimulate scholarly and public debate on catastrophic-risk governance. The site offers policy analysis and advocacy only in the sense of outlining possible legislative or contractual frameworks.

No legal or financial advice

Nothing here should be treated as a substitute for qualified legal counsel, financial due diligence, or regulatory guidance. Readers remain responsible for ensuring their actions comply with the laws and professional standards of their own jurisdictions.

Exploratory and personal views

All scenarios, numerical examples and opinions are research hypotheses presented by the author in a personal capacity. They do not represent the views of the author's employer, funding bodies, or any governmental authority.

Implementation caveats

Any real-world adoption of these ideas would require democratic deliberation, statutory authority, and robust safeguards against misuse. References to enforcement, penalties, or "bounties" are illustrative models, not instructions or invitations to engage in private policing or unlawful conduct. Nothing here is directed at any identifiable individual — see non-targeting.

No warranty and limited liability

Content is provided "as is" without warranty of completeness or accuracy; the author disclaims liability for losses arising from reliance on this material.

By continuing beyond this notice you acknowledge that you have read, understood, and accepted these conditions.