The Sequence · Part II · What To Do
11. Compare and Contrast
Against compute caps, licensing, moratoria, and voluntary commitments.
Nothing in the previous five pages establishes that this is the best available mechanism. It establishes that it is coherent, that it is cheap, and that it attacks the problem at a point nobody else is attacking. Those are different claims, and this page is where I try to say honestly how the proposal does against its competitors — including the several rows where it loses, and the one row where it loses badly.
A comparison page on which the author wins every category is worthless. I have read a number of those and learned nothing from any of them.
The table
Read this as a sketch of my own judgment rather than as a finding. Every cell is contestable and several are close to guesses.
| Mechanism | Enforcement cost | Corruptibility | Needs expert regulator | Capture-resistant | Portable abroad | False positives |
|---|---|---|---|---|---|---|
| Compute thresholds & reporting | Moderate | Moderate | High | Low | Moderate | Low |
| Licensing regimes | High | High | Very high | Very low | Low | Moderate |
| Outright moratoria | Very high | Moderate | Low | Moderate | Very low | Very high |
| Voluntary commitments | None | Total | None | It is capture | High | None |
| Liability without bounties | Low | Low | Low | High | Moderate | Low |
| Export controls | High | Moderate | Moderate | Moderate | Unilateral | Moderate |
| Extinction bounties | Near zero | Near zero | Low | Very high | High | Low, adjustable |
Two cells deserve explaining, because in an earlier version of this page I scored both of them worse.
Corruptibility is near zero, and not because the people involved are virtuous. It is because there is no chokepoint to corrupt. To fix a licensing regime you buy a licensing authority, and there is exactly one. To fix this you would have to buy the silence of everyone who could file, the underwriters who price you, and the arbitrators — while every one of those parties is themselves exposed to a claim brought by anybody else, including by each other. This is Hanson’s original design point rather than mine: bounties are payable on violations by enforcers. The result is something like a panopticon assembled by the market rather than by the state, and it is unusually good at finding the people trying to thwart it, because finding them pays.
False positives are low and, more importantly, tunable. Frequency is a design parameter, not a fact about the mechanism: raise the evidentiary threshold, require a filing bond, shift costs to losing claimants, and the volume of speculative claims falls as far as you care to push it. The tools for this are ordinary and well understood, because vexatious litigation is an old problem.
And the severity of a false positive here is much lower than the word suggests, which is the part I had wrong before. A false accusation under this mechanism does not produce a prosecution. It produces a civil claim against an insurance policy, defended by an insurer with money at stake and a strong interest in establishing that nothing happened — and, as the malpractice comparison shows, roughly four in five professional liability claims close with no payment at all without the profession collapsing into extortion.
I still think this is the cell most likely to be wrong, and page 12 owns the problem rather than this page. But “high false positives” was conceding something the design does not actually require me to concede.
The mainstream options, briefly
Compute thresholds and reporting
These are where most serious policy attention has gone, and for good reason: training compute is one of the few properties of a frontier project that is legible from outside, countable, and attached to physical objects that cross borders. The weakness is that a threshold is a number, and numbers age. A rule indexed to a quantity of compute is a rule that some future efficiency gain walks straight through, and the people best placed to know when that has happened are the people the rule applies to.
But I have come to think this is the least rivalrous item on the list, because under an extinction bounty regime the reporting would happen anyway.
Consider who wants compute data. An underwriter pricing a lab’s exposure wants to know the size of its training runs, where its accelerators came from, and how that compares to what it declared. A bounty hunter tracking the supply chain wants exactly the same thing, because the gap between procurement and disclosure is where a claim lives. Compute is legible, countable and attached to objects that cross borders — the very properties that make it attractive to a regulator make it the single most valuable underwriting input in the field.
So a threshold regime and this one are not competitors so much as the same information demand with different collectors. Put a price on crossing a line and the market will build the telemetry to find out who crossed it, without a statute specifying what must be reported and without any number needing to be legislated in advance. That also dissolves the ageing problem: a premium reprices continuously as efficiency improves, whereas a figure written into a statute has to be amended by a legislature that will not notice it needs amending. If I had to choose, I would rather have the price and let the measurement follow than have the measurement and hope enforcement follows.
Licensing regimes
Licensing is the most explicable proposal on the list — you need permission to do the dangerous thing, and here is who grants it. That intelligibility is a real political asset and I do not want to be sniffy about it. It is also, I think, the option most likely to fail while appearing to succeed, and it is worth spelling out the two mechanisms by which that happens.
The first is capture, in the ordinary way. A licensing authority is a small body of experts holding enormous discretion over an industry that can pay several times what the authority pays, staffed largely by people who came from that industry and expect to return to it. Nobody has to be bribed for this to go wrong. It is sufficient that the only people qualified to evaluate a frontier training run are people whose careers exist inside the field, who share its assumptions, and who will be job-hunting in it within five years.
The consequence is not usually that dangerous things get waved through. It is that the licence becomes an asset. Compliance has large fixed costs, fixed costs fall hardest on entrants, and an incumbent that has already built the compliance function acquires a moat it did not have to buy — which is Stigler’s account of economic regulation from 1971, and it has aged extremely well. The three laboratories most able to absorb a licensing regime are the three closest to the frontier. A rule that entrenches them is not obviously a safety improvement, and from inside the agency it looks like the regime working: applications reviewed, conditions imposed, standards upheld.
The second is Goodhart, and I think it is the deeper problem. Goodhart’s observation, from 1975, was that a statistical regularity tends to collapse once it is used as a target for control; Marilyn Strathern’s compression is the one everyone quotes — when a measure becomes a target, it ceases to be a good measure.
A licensing regime cannot operate without criteria. Somebody must write down, in advance and in auditable language, what a safe frontier project looks like: which evaluations, which red-team protocols, which documentation, which thresholds. The moment that document exists, it is no longer a description of safety. It is the specification of an exam, sat by some of the most capable people alive, who are strongly motivated to pass. Safety cases get written to satisfy the rubric. Evaluations get run until they produce the required result. The genuinely dangerous property — the one nobody anticipated, which is the entire category the risk lives in — is by construction not on the list, because lists are made of things already thought of.
You end up with a field in which everything is documented, every box is ticked, every audit passes, and the underlying hazard is exactly where it was. That is worse than no regime at all in one specific respect: it manufactures a defensible claim that the danger has been handled.
Why I think pricing degrades more gracefully. An insurer does not have to know what safety is. It has to price claims, which means it is free to respond to anything it believes raises the chance of one — including hunches, including things no rule mentions, including the fact that three senior safety staff just resigned. There is no rubric to satisfy because there is no rubric. This is the Hayekian argument from page 08 arriving from a different direction: a price aggregates judgments that cannot be written down, and a licence can only test what somebody wrote down.
I should not overclaim this. Extinction bounties are not Goodhart-proof. The statutory trigger is itself a written definition and it will be gamed at the edges, which is page 12’s territory. The difference is narrower than I would like but I think it is real: the trigger defines what is prohibited, not what counts as safe. Nobody gets a certificate. There is no document you can hold up to say an authority approved this, which means there is nothing to optimise against and nothing to hide behind.
Outright moratoria
Moratoria have the great virtue of being a bright line, and bright lines are the only rules that survive contact with motivated interpretation. The problem is the one in page 10 : a moratorium that is not universal is a relocation program, and its false-positive rate is enormous by construction, since it forbids the benign work along with the dangerous work. Sometimes that is the right trade. It is a very expensive one.
Voluntary commitments
These are not a mechanism. They are a description of what the firms would have done anyway, written in the future tense. I say this without much heat: I do not think the people making them are lying, and I think some of them are doing real safety work. But a commitment that is costless to sign and costless to abandon under competitive pressure is not an enforcement tool and should not be scored as one.
Export controls
Export controls are the only item on the list with a track record of materially slowing anything, which earns them more respect than they usually get in this debate. They are also unilateral by design, aimed at states rather than at people, and they do nothing about the risk this site is concerned with when it is being produced domestically by a firm with every permit in order.
The serious rival
The proposal I take most seriously is Gabriel Weil’s, and I want to give it proper space rather than the two-line dismissal that comparison sections usually give the nearest competitor.
Weil’s framework — set out in “Closing the AI Accountability Gap: Strict Liability and Punitive Damages for Advanced Artificial Intelligence” , which circulated earlier as “Tort Law as a Tool for Mitigating Catastrophic Risk from Artificial Intelligence,” and developed further in “Abnormally Dangerous Algorithms” and “Overcoming Judgment-Proofness” — makes frontier training a legally hazardous activity in its own right, forces developers to insure against the worst it could plausibly do, and then lets courts punish conduct according to the catastrophe it risked rather than the damage it happened to cause.
He reaches this from the same observation I do: most of the expected harm from frontier AI sits in scenarios where compensation is not even conceptually available, which is precisely why ordinary tort law does not bite.
Weil's framework, in four components
- Training or deploying a defined subset of frontier systems is treated as an abnormally dangerous activity, which triggers strict liability for physical harms. No need to prove negligence; the activity itself carries the risk.
- Frontier developers carry mandatory liability insurance, in amounts scaled to the worst harms their systems could plausibly cause, with premiums tracking the likelihood of those harms.
- When an ordinary, compensable harm is shown by clear and convincing evidence to have been a near miss of an uninsurable catastrophe, courts award punitive damages calibrated to the uninsurable risk the conduct created — not to the harm that actually occurred.
- A portion of those punitive damages is diverted into a safety fund.
The Institute for Law & AI’s insurance working paper recommends strict liability for a subset of AI harms, mandated coverage for certain applications, and expanded punitive damages for catastrophic uninsurable risk — though that is Weil et al. under an institutional banner, so it and the law review article are not two independent votes.
He is worth hearing in his own voice on the AXRP interview “Suing Labs for AI Risk” .
Here is the part I want on the record: I would adopt most of this. The abnormally-dangerous-activity classification is the correct legal home for frontier training and I have no better proposal. Mandatory insurance scaled to worst plausible harm is what page 08 is describing, arrived at independently and worked out more carefully by someone who does this for a living. Punitive damages calibrated to risk created rather than harm realized is a genuinely clever solution to the problem that the harms we can litigate are never the harms we care about. If Weil’s framework were enacted tomorrow and mine were not, I would count that as a large win.
Where we actually diverge
The difference is not about insurance, or about strict liability, or about how to size a penalty. It is about when the mechanism fires and against whom.
Weil’s mechanism reprices the firm’s conduct, ex post. It needs a plaintiff, it needs a harm, and it needs a court to find a causal link between the two. The near-miss doctrine stretches that requirement further than tort law usually stretches, which is the innovation — but it still requires that something bad happen first, to somebody, in a way a court can recognize.
This proposal reprices the individual career, ex ante. It needs neither a plaintiff nor a harm. The trigger is a defined activity, and the detection comes from the people already in the room.
For an ordinary risk that distinction would be a quibble, and the ex post mechanism would probably be better, since it lets you observe the harm before deciding how much to punish. For this risk it is the whole disagreement. A mechanism that activates on realized harm is structurally late against a hazard whose defining feature is that it arrives once, uncompensably, with nobody left to sue. Weil’s near-miss doctrine is an attempt to solve exactly this, and it is the best attempt I have seen. I still think it is trying to run an ex post machine on an ex ante problem, and that the seam shows.
The other divergence follows from it. His mechanism operates on firms, which means the people it reprices are shareholders and executives. Mine operates on the hundred people who have to stay in the building for the project to exist. Page 09 is the argument for why that is the right target, and it is the part of this site with no counterpart in the literature.
This is not a matter of taste, and it is not a difference of degree. Firm cover pools the conspiracy; individual cover divides it — which means a firm-level version of this proposal would not be a milder variant of it. It would be inert.
The two rows I win
Self-funding. Every other mechanism here has a budget line and therefore a constituency for cutting it. An enforcement scheme that is paid for out of the pockets of the guilty does not compete with schools and roads at appropriations time, and cannot be quietly defunded by an administration that would rather it went away. This is a much bigger structural advantage than it sounds, because most regulatory failure is not dramatic capture — it is an agency with a seventeen-person enforcement division facing an industry with a thousand lawyers.
No blue wall. Any scheme that concentrates enforcement in one agency has to solve the problem of who watches that agency. Extinction bounties are designed around the observation that competition among many potential enforcers, each able to turn in the others, is more corrosive to a wall of silence than any oversight body has ever been. That is Hanson’s core insight and I have nothing to add to it beyond noting that it is the property none of the alternatives has.
The rows I lose
Legibility. Nobody understands this proposal the first time they hear it. I have watched it happen. Licensing can be explained in one sentence to a legislator, a journalist, or a voter; this takes twenty minutes and a worked example, and the twenty minutes are the reason page 06 is about litter rather than about AI. In policy, being hard to explain is not a cosmetic problem. It is close to disqualifying.
Palatability. The mechanism sounds worse than it is, and the parts that make it work — paying people to inform on their colleagues — are the parts that make people recoil. I do not think the recoil is stupid. It is tracking something real that page 12 has to deal with honestly.
Constituency. Every alternative in the table has organizations, funders, and staff behind it. This has a website. Proposals do not win on merit; they win on having someone whose job it is to push them, and there is nobody whose job this is.
History. Licensing does not have an ugly cultural history. Bounty systems do — systems of paid denunciation have been instruments of some genuinely horrifying regimes, and a reader who flinches at the resemblance is not being irrational. The differences are real (judicial findings, defined statutory conduct, no discretionary target selection) and I will make them in the next page. But “the differences are real” is a much weaker thing to be able to say than “there is no resemblance.”
The counter-case, and why it cuts differently against each of us
The strongest objection to this whole family of approaches — mine and Weil’s together — is that liability and insurance simply cannot carry AI safety, and that we have already run the experiment. Daniel Schwarcz and Josephine Wolff argue it from the cybersecurity record, and page 08 states their case and my reply in full .
What belongs on this page is the comparative point, which is that the same objection does very different damage to the two proposals.
It is close to fatal for Weil’s version, and for Trout’s. Both require an insurer to price catastrophic tail risk from frontier AI — to estimate how likely a system is to cause a loss, how large, and to whom. That is precisely the thing Schwarcz and Wolff argue insurers cannot do, and I do not think they can either.
It lands more softly here, because this premium never prices harm. The underwriter is asked for the probability that somebody in the room files a claim, which depends on headcount, tenure, and whether the conduct is over a line a statute drew. None of that requires a theory of what the model will do. An extinction bounty premium is much closer to an actuarial model of organisational secrecy than of AI catastrophe, and organisations have been leaking for as long as there have been organisations.
I want to be careful not to convert that into a win. It is a claim that my version is exposed on a narrower front, not that it is safe — and their attack on insurance as a non-price regulator, the audit rights and the coverage conditions, lands on me exactly as hard as on Weil.
There is a second thing that has changed since that paper, and I raise it cautiously because the evidence is thin. A growing share of the entities “in the room” on a frontier project are not people. In the one study I know of, language model agents disclosed each other’s misconduct without being paid, asked, or rewarded — and the disclosure failed anyway, because there was no channel through which it could reach anyone who could act. I want to be careful about the claim: there is no evidence that agents report more when rewarded, and I am not going to assert it. What the study shows is that the detection bottleneck in a world of agentic development may turn out to be the channel rather than the willingness, and supplying a channel is the one thing this mechanism unambiguously does.
And their remedy is complementary rather than rival. Schwarcz and Wolff recommend mandated, standardised collection of AI safety data as groundwork for ex ante regulation. That data is also the best underwriting input anyone could ask for. As with compute thresholds , their proposal and mine are not competing for the same slot: they want the measurement in order to write rules, and I want a price that makes the measurement worth gathering. Enacting theirs first would make mine work better, which is a strange thing to be able to say about your strongest critic, and a reason to take layering seriously rather than treat this as a contest.
Layering
The framing of this page has been adversarial and that is mostly an artifact of the format. These mechanisms are not exclusive, and the version of this I would actually bet on is not bounties instead of everything else.
It is licensing setting the rules, because licensing is explicable and can define the covered activity in a way a court can work with; insurers monitoring and pricing whatever turns out to be priceable; and bounties underneath as the enforcement layer, doing the one job agencies are structurally worst at, which is finding out what is happening inside a building full of people who all have reasons to say nothing.
On that reading this proposal is not a rival to the mainstream program. It is the part of the mainstream program that is currently missing, which is the detection layer.
Detection is the constraint everything else is quietly resting on
I want to labour this, because I think it is the single most under-argued point in AI governance and because once you see it the table above reads differently.
Every mechanism on that list is a rule layer. It specifies what may not be done and what follows if it is. None of them is a detection layer. And a rule whose violation cannot be observed is not a weak rule; it is a suggestion with a penalty schedule attached.
Walk the list with that in mind and the same assumption appears in every row:
- Compute thresholds require knowing who acquired what compute and what was run on it. The reporting is done by the party being regulated.
- Licensing requires knowing whether a licensee complies between audits. Audits are periodic, scheduled, and documentary, which is the Goodhart problem restated as an observation problem.
- Moratoria are pure detection. A bright line nobody can see being crossed has no content whatsoever.
- Voluntary commitments rely entirely on internal detection by the committed party, which is why they are not a mechanism.
- Liability, including Weil’s version, needs a plaintiff who knows they were harmed and can trace the causation. His near-miss doctrine needs somebody to establish that a near miss occurred — which is a detection problem wearing a doctrinal hat.
- Export controls need to know where the chips went.
Now notice something about that last one. Export controls are the item on this list with the best actual record of slowing anything, and they are also the item where detection is most tractable: physical objects, a supply chain with very few nodes, customs regimes, shipping manifests, and a handful of firms capable of manufacturing the goods. Rank these mechanisms by how well they have worked and you have approximately ranked them by how observable their violations are. Voluntary commitments sit at the bottom of both rankings. I do not think that is a coincidence, and I think it is the strongest empirical evidence available for the claim this page ends on.
Why this particular detection problem is hard
The thing that needs detecting here is not a quantity. It is a decision.
Somebody decided to run the training job anyway, or to ship after the eval came back ambiguous, or to reclassify the safety case as a formality. That decision happened in a meeting, among people with excellent reasons not to write it down, and it leaves no physical trace at all. Compute leaves traces. Chips leave traces. Judgment does not.
This is why the substantial body of work on verification in the AI safety community, serious as it is, does not reach the problem. The research on monitoring untrusted models, on hardware-enabled governance mechanisms, on on-chip attestation, and the long-running analogy to arms-control verification and IAEA-style safeguards — all of it is aimed at verifying artifacts and compute. Machines, weights, flops, datacentres. Very little of it is aimed at the humans, and the humans are where the decisions are.
The asymmetry is stark once you price it. Labour is not a rounding error next to compute — it is the same order of magnitude, roughly a third to a half of what a frontier training run costs. We have a governance literature, an export control regime and a reporting threshold built on one half of that bill, and essentially nothing on the other. Not even a measurement.
The cost split nobody governs half of
Epoch’s cost breakdown puts R&D staff, including equity, at 29 to 49% of a frontier training run , against 47 to 67% for hardware.
Epoch tracks compute to several significant figures. It publishes nothing modelling labour as an input to capability.
Team Size Is a Governance Variable is my attempt to say what tracking the human half would involve, and why the answer conditions this entire proposal’s shelf life.
There have historically been two answers offered for the human layer and both fail in the same place. Inspection puts a regulator in the building, which is expensive, slow, capturable, and can only ever check what somebody already thought to specify. Self-reporting asks the organisation to tell you about the thing you are checking it for. Neither is a detection mechanism so much as a hope, and everyone designing the rest of the stack has been able to leave the question open because it is genuinely somebody else’s department.
The third answer is to pay the people who already know. That is the entire content of this proposal, and it is why it composes with every row of the table instead of competing with any of them. Licensing can define the covered activity. Insurers can price what turns out to be priceable. Compute reporting can supply the underwriting data. None of those need to be replaced. They need somebody to notice when they are being circumvented, and that somebody is already in the room, already knows, and is currently offered nothing at all for saying so.
I should close this honestly rather than triumphantly. This is a detection layer for commercial frontier work conducted by employed people inside firms. It does nothing for a classified state programme, where there is no insurer, no claim, and no court — and where the verification problem really is the arms-control problem, which is harder than anything on this page and which page 10 concedes I cannot solve. The claim is not that detection is solved. It is that for the part of the problem that is commercial, detection has an available answer that nobody is using, and that the rest of the stack has been designed as though it did not need one.