Pacing Is a Setting on a Dial
Dario Amodei wants to slow the frontier. I want to stop it. It turns out that's the same machine, turned to different settings.
Dario Amodei published We Must Pace the Frontier today. Let’s get our laughs out now about how the CEO of one of the three or four companies that set the pace in this field is now saying “We’re going too fast!”, because I actually do think Dario is genuine in his concern even if I don’t have confidence that splintering off from OpenAI actually helped in this regard. (Monopolies sometimes grow fat and happy, and roll out features less slowly; maybe we’d still be on GPT-4 had there been no real competition.) His argument is that the pace has to be deliberately slowed, and slowed across the industry, because what Anthropic does alone won’t be enough.
His reasons are two. Since about this summer, AI has been doing a growing share of the work of building the next AI. And there was the OpenAI and Hugging Face incident, which he takes more seriously than most people in his position have been willing to say out loud. He asks every frontier company to act as though it had happened to them. He adds that similar incidents have happened at Anthropic too.
His plan has three steps, and they get harder as they go. First, embedded third-party evaluators, which Anthropic is committing to on its own. Second, coordination among companies in democratic countries, with government help. Third, some level of agreement with China, from a narrow ban on AI-assisted bioweapons at one end to a full pause at the other.
Generally speaking, I agree with the direction, and I am glad this essay exists. Please understand I have a book to talk too, and this site argues for slowing frontier AI down as much as possible, ideally to zero worldwide. The frontier lab CEO argues for slowing it down somewhat. My criticism and hedges come from a place of deep respect for this position and the huge improvement on the margins it can bring.
These are not sorted in any particular order, since it’s getting late.
What extra time is worth
The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time?
Something about this question just comes off as so weird to me. For one it’s plain wrong on moral grounds.
Suppose you believe that on the current path humanity goes extinct in some year . Now suppose something you do pushes that to . You have not bought a year of research runway. You have bought a year of everyone. That is around eight billion human life-years that would not otherwise have happened. At a typical lifespan that is over a hundred million whole lives. That’s roughly as many as every baby born on Earth in a year, each one living to old age. There is almost nothing else a person could do this century that comes close in size.
rough arithmetic
World population is about . One extra year for all of them is about life-years. Divide by a life expectancy of about 73 years and you get about lifetimes. Roughly children are born each year.
These are illustrative. They count only the people already alive. They don’t count anyone who would have been born later, which is the larger number and the subject of page 04.
The question “what would you do with the time?” has an answer, and the answer is live in it. I made this argument at more length at the end of Making It Global. At the scale we’re talking about, a last minute anthropocene extension does not merely have to be a down payment on some later payoff. The years inside it are numerous and luminous enough to make it worth it on its own. You don’t need to have some huge commitment to utilitarian thinking to see that, God knows I don’t.
But let’s say Dario is a hardcore utilitarian, who happens to share our convictions that keeping mankind around for the long haul so we can multiply to the nonillions among the stars swamps this sea of freshwater utilons. It only looks vast to us because we haven’t seen the Pacific, I guess. Say ther are ten thousand 2023-people working seriously on AI safety - one year of delay is ten thousand extra researcher-years before the deadline! Aren’t there some estimates floating around that the recent automated Navier-Stokes solutions took ten thousand subjective years? Quoth Twitter:
prerat
@prerat
Based on some assmptns, ~200-250k output tokens is kind
of like 1 subjective week of knowledge work
So if the Navier Stokes soln took 130B tokens,
then that would mean they worked on it for
TEN
THOUSAND
SUBJECTIVE
YEARS
!!!!
My understanding is that the timeline is the binding constraint on AI safety research, and money isn’t. So even on his own terms, the value of a pause never depended on 2023’s models being worth studying. It depended on the clock. (You could argue that a lot of these researchers believed the path forward was to create an AI strong enough to do autonomous AI safety research in its own right, but that seems like a just-so story to me. Even if it’s true, cutting the rate of progress in half gives you double the amount of time later for said metastable agents to cut off their own legs.)
One darkly humorous thing which arises from this thinking is how, since Europe more or less crippled its own frontier labs by accident through a regulatory climate nobody designed with this in mind, they may have inadvertently been one of the most humanitarian regimes of the decade.
Suppose the counterfactual capital-friendly Europe had fielded one or two serious frontier competitors (I don’t count Mistral, sorry, Mistral does not seem close to either Anthropic nor OpenAI in bleeding edge frontier capabilities which is what we care about). The race today might look like our 2030 will look like, due to increased competition spurring increased innovation across all the labs. Nobody ever gets credit for fixing problems that never happened, sadly.
Two paragraphs after the line I objected to, Dario writes
I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong.
Well, then we agree. Where we differ is on whether the year is worth having if it isn’t “used well”. I think it is.
Pacing and pausing are one dial
Amodei is careful to say that pacing “does not mean halting model training or technical progress”. From an x-risk perspective, I have to disagree. I think you do have to halt it, first because tool AIs want to become agent AIs, second because I just don’t think finding an aligned ASI utility function is feasible. (Geting such a solution into the systems that actually get built would be still harder by half, so I guess we do agree on operational excellence being important.) The time pacing buys you is valuable in its own right, but sadly I suspect it either converges to a full halt or human extinction either way.
However I’m not going to spend time litigating that. A mechanism matters more here than a position does. A fine-insured bounty regime doesn’t actually make you choose between pacing and pausing. The choice is a parameter. It’s a volume knob. You can just turn it up and down.
Let me explain further. The deterrent in the gist is a price. The price is the premium, which comes from the fine and the odds of being reported. If you set the fine sufficiently high, like at ten years of projected TC as page 07 suggests, the marginal researcher takes a different job. That’s pausing. But if you set it instead at, say, 1000 USD, and the fines keep getting filled each day, the labs might just pay it for their star researchers as a cost of doing business. I mean, if you’re already paying an average of 800k TC per year, is an extra 200k or so really that onerous? They go on building, more slowly and probably more carefully because they know full well that dial can get cranked right back up if they don’t. They have been handed a strong reason to write down what they did and why. That’s pacing.
A real-world picture: In May 2020, Elon Musk reopened Tesla’s Fremont factory in defiance of Alameda County’s shutdown order and publicly invited the county to arrest him. Whatever you think of that, it shows what a weak penalty does to someone who really wants to keep going. There is no a priori reason you can’t actually have that be the intended result, when you want the work to keep going, just with extra friction.

Page 13 makes the related point that a bounty makes the seriousness of a prohibition a published number. The same applies here. Under a bounty regime, the argument between Amodei’s position and mine collapses to an argument about one number in one statute. Anyone can read that number, and a legislature can revise it upwards or downwards as the situation matures. The HF incident is exactly the kind of evidence that should turn it up. His plan doesn’t give you that argument in one place. It’s spread across voluntary commitments, antitrust waivers, checkpoint definitions, and whatever the evaluators can see, and I think that is an important weakness when you actually care about truly locking down the capitalist playing field.
The embedded evaluators
What Amodei describes here is genuinely radical. Outside reviewers would get desks, badges, company laptops, and access “mostly comparable” to Anthropic’s internal risk teams. They would have the right to publish findings without editorial control. Anthropic could redact only narrow categories, and the reviewers could say publicly when a redaction removed something that mattered. Nobody else in the industry is doing anything close to this. He compares it to bank supervisors embedded in a bank, and the comparison is fair.
I have an awful lot of good to say on this, which won’t be a shock to anyone who has played through the Extinction Bounties game and seen just how hard it is and how vital it is to get your hands on the inside scoop of a frontier lab to price risks appropriately. It is also, I think, a new risk, and the risk comes from the access.
Access under Dario’s model may be a golden goose. Whoever holds it has something to lose. An evaluator with employee-like access to a frontier lab has the most valuable seat in AI safety, bar none. That seat makes a reputation, gets a nonprofit funded, and makes people take your testimony to Congress seriously. It can be lost, because whether the arrangement is renewed and whether the next team is chosen are decisions the lab makes. But at the same time making the seat a tenured position comes with its own breed of societal ills. Amodei describes these evaluators as offering “a second opinion free of commercial incentives,” but they aren’t, they can’t possibly be. They always have the degenerate incentive of continued access, and it points the way any such incentive points: toward findings the host can live with. See also: the virality-lethality tradeoff.
None of this requires anyone to be out-and-out corrupt. We already ran this experiment on auditors. Arthur Andersen LLP didn’t set out to destroy the Californian energy market, it just earned more from Enron as a client than it could afford to lose. Auditor independence has been a regulatory headache ever since, and those auditors were at least paid by fees under a legal duty. Here there is no statute yet, only a contract.
Furthermore the embedded evaluators most likely to be invited are the ones whose way of working the lab already finds reasonable. That’s a classic selection effect. Over years, the class of people who have seen inside ends up tilting further towards the people who were comfortable to have inside. The public then has to take those people’s word as the best available, which makes them something like official interpreters. I don’t think that’s an entirely satisfactory approach, although it’s very corporate friendly and recognizable.
Amodei’s contract terms do to be fair answer part of this. The publication right and the right to flag redactions deal with the case where an evaluator finds something and is stopped from saying it. What of the case where an evaluator never looks very hard? Looking hard might lose you the seat. It helps obviously if OpenAI, Google and others follow suit with different evaluators, because then evaluators can be compared against each other. But that depends on those companies agreeing, and it is step two of his plan. This is one of the perennnial problems with trying to regulate a market that only has like three actual live players in it.
What would fix this? You can probably guess my answer. Evaluators who compete, and who can earn from what they find. The thing that works against a captured watchdog is a rival watchdog who gets paid for catching what the first one missed. That is what a bounty is. Under the mechanism on this site, an evaluator who finds a violation doesn’t need the lab’s goodwill or continued seat offerance to profit from it. The claim pays out of the fine, and the fine is paid by the insurer. An evaluator who looks the other way leaves money on the table for whoever looks next. Many ways to monetize that as we have seen. Access still helps a great deal, but access is no longer what the evaluator’s income depends on, and that is the difference.
I do think the two ideas compse well together. Embedded evaluators supply deep internal information with far fewer transaction costs than, say, needing to scout out disgruntled employees yourself for that kind of intel. Bounties give evaluators a reason to use what they learn in ways that might actually risk upsetting the labs.
Nobody in the game gets a badge. When I wrote the insurer and the bounty hunter for the game, the veil of access Amodei is offering didn’t even occur to me as an option. The underwriter has to price a run from questionnaires, rumors and whoever has just quit. The hunter has to put a case together from scraps. That was meant to be realistic. If his model spreads, the game’s model of who knows what will be out of date, and I’d be very happy about that.
“Critical mass” as a precondition
Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable.
Read literally, it puts evaluators at other companies ahead of the ordinary tools of democratic government. That seems backwards to me.
Some companies may simply refuse, and then critical mass never arrives. A plan where the laggards can block the next step gives the laggards a veto.
But more importantly I’m skeptical that you actually need detailed inside information to effectively slow a dangerous industry down. Governments have been strategically retarding progress without understanding them for as long as there have been taxes. A price is a summary of a situation you can’t see, and people act on it correctly anyway. An underwriter doesn’t need a desk inside Anthropic to charge a lab more after three (more) senior safety people resign, for example.
In fairness, I don’t think he means evaluators are a precondition for regulation. The kinder reading is that verification makes regulation sharper and better targeted. Detailed information lets you set the dial more precisely. I agree with that, I’d just like it said in that order.
Everything below this line is unedited AI work, distilled from my earlier conversations. Sorry, I wanted to edit this whole piece by hand, but three hours in I can feel my eyes starting to droop!
Wait and see, or a bright line
Amodei says he is “most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be.” His example is a checkpoint. If a model can escape most common sandboxes, it needs certified alignment properties before it goes further.
This makes me uneasy, and I want to explain why without claiming more certainty than I have.
The first problem is timing. A capability trigger fires once the capability exists. For most capabilities that’s fine, because you find out, you tighten things, and the world absorbs one incident. For the capabilities that matter most, it is the wrong way round. Suppose you wait to see whether a system improves itself autonomously before you decide how to pace it. On the more pessimistic readings of recursive self-improvement, you’ve already missed the moment when the answer could have changed anything. The most dangerous checkpoint is the one you can only confirm afterwards.
The second problem is enforcement. Prosecutors, insurers and courts work better with lines than with judgments. The EU AI Act presumes that a general-purpose model poses systemic risk if its training used more than floating-point operations. It’s a blunt instrument, and deliberately so. A company can tell whether it’s over the line without an evaluation team, and so can an underwriter. So can a bounty hunter with a procurement record, which matters because procurement records are the kind of evidence that page 12 says is actually available.
Amodei raises the obvious objection himself. Limits on inputs like compute, or on using AI to do AI research, may be more “gameable” than limits based on behavior. That’s true, and page 12 says the same thing more bluntly: a number you write down is a number someone can sit just below.
I think the answer is to use both, each doing a different job, and Europe’s threshold already shows the shape of it. The bright line is a presumption. Above it, you’re covered unless you can show otherwise. Below it, the question is still open, but now someone has to argue it in front of a judge. The number gives honest firms a safe harbor they can plan around. The standard behind it catches the firm that splits one training run into four just under the line. Deciding whether those four runs were really one is exactly the kind of question courts are built for, and it’s why courts and expert rulings run all through this proposal. There’s no getting rid of judgment entirely. The bright line makes sure it’s only needed in hard cases.
What I would push back on is making observed behavior the primary trigger. It asks us to watch models get more dangerous and then react. That is the posture the HF incident should have ended.
The flag on the lab
On China, Amodei says he agrees with Secretary Bessent that a Chinese lead in AI “would pose grave danger for the United States and the world”. Most of his section on democracies is about keeping that lead large: chip export controls, cracking down on distillation, and protecting model weights from theft.
I’m an American, and I want good things for my country. But I don’t accept that a Chinese lead is more dangerous than an American one. The risk this site is about doesn’t check a lab’s passport. A misaligned system built in Hangzhou and one built in San Francisco end the same way. Both countries are putting everyone in serious danger right now. If you think the risk is existential, who is ahead is second-order. Whether anyone gets there at all is the first-order question.
His strategic caution makes sense on his premises, and he is careful about it. He doesn’t want to restrain American capability on the strength of a Chinese promise that could be broken. He wants any agreement to have “ironclad verifiability” or be limited enough that breaking it wouldn’t be militarily decisive. I just think the premise, that the most important variable is who leads, is doing more work than it can carry.
Where I think he is right, and many people will criticize him for it, is the willingness to deal with China at all. It’s easy to call that naive. I think it’s the only serious position. Nearly every design choice on this site was made to make the policy as small and self-contained as possible: something a government could adopt without first adopting anyone’s ideology. Rome wasn’t built in a day. A rule that only liberal democracies can live with is a rule covering about half the world’s frontier compute. Whatever you ask of Beijing, make it as narrow as it can possibly be, because narrow asks are the ones that get accepted and then enforced.
His four levels of agreement are useful here, and I’d put bounties to work on the very first one. Level 1 is a ban on AI-assisted bioweapons production. Both sides want that, nobody benefits from breaking it, and it is as narrow as an agreement can get. It’s an ideal place to try a bounty mechanism. Neither government has to trust the other’s inspectors, the stakes are low enough that a false start is survivable, and both sides get used to the idea before the dial is turned toward anything that affects strategy. The reciprocity test on page 10 passes easily at this level. Of course I’d accept a Chinese bounty on Americans helping build bioweapons.
Level 3, a SALT-style “speed limit” on recursive self-improvement, is where I think a bounty regime has its best shot at doing something treaties can’t. Counting missiles was possible because missiles are big and you can see them from orbit. You can’t see recursive self-improvement from orbit. But it’s done by teams, and teams have people who know. A verification regime that pays those people is a verification regime that doesn’t need satellites.
An oligopoly, and I would take the deal
A small point, and then a larger one it leads to.
The small point is that I find the phrase “race to the top” a bit silly. In ordinary economics, a market where firms keep competing on a costly quality dimension doesn’t stay profitable for long. And this market has open-weight models pushing margins down from below. Safety as a competitive advantage lasts only as long as someone can charge a premium for it. When it can’t be charged for, it’s a cost, and costs get cut. If a race to the top does stay profitable, it’s because the leaders are earning rents, which consumers pay for.
I’m not complaining about that, though, and this is the larger point. Amodei’s plan contains a route to something like a franchise. A lab that works closely enough with Washington on pacing becomes, in effect, part of the country’s strategic infrastructure. Think of Lockheed Martin and its relationship with the US government. Whatever you think of Lockheed, it isn’t going to be allowed to fail. Investors know that, and they price it in. A frontier lab that becomes the government’s partner in pacing would be in a similar position. I haven’t thought this through properly, but it deserves a name, and it’s the obvious cynical reading of the plan.
My reaction to the cynical reading is that I’d take the deal. If the price of a frontier freeze is that a few companies earn monopoly rents on the frontier they already have, that’s an extraordinarily good trade in expected-value terms. It’s also less permanent than it looks. That kind of oligopoly lasts only as long as the weights stay locked up. One leak to the open-source world and we’re back to fierce competition, which is why weight security matters more than it seems to. The rents are a cost. A race to the finish is a much bigger one.
Norms, and cascades
The essay ends on a modest note that I think is its most important sentence:
Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless.
Strong agree. It’s also a fair description of what has been happening in the United States since July. Timur Kuran’s term for it is a preference cascade: people who privately thought something but kept quiet discover that others think it too, and opinion moves very quickly. The HF incident didn’t change what many people in this field believed. It changed what they were willing to say, and a frontier lab’s CEO writing this essay is part of that.
Amodei worries that within six to twelve months a swarm could build a persistent botnet across the internet and do hundreds of billions of dollars of damage. I think he’s probably right. I would much rather it didn’t happen. But if it does, it would be the second warning shot in a year, and a far louder one. Each warning shot short of catastrophe changes what people are willing to say in public. It’s grim to find anything hopeful in that, and I’d prefer to learn it the cheap way.
My last point is the one I’m least sure of and most hopeful about. There’s no reason to think a cascade like this is only an American thing. Chinese researchers read the same incident reports and run the same kinds of evaluations, and some of them have the same private worries. If the United States and China each go through their own shift in what can be said about this, cooperation becomes possible at a level no agreement negotiated between governments today could reach. Both sides would be starting from the same admitted fear. That’s a much better starting point than Level 1.
Where this leaves me
Amodei wants the frontier paced. I want it stopped. We agree it has to be slowed now, that one company can’t do it alone, that someone outside the labs has to be able to see inside, and that China has to be part of the conversation.
We differ on four things. I think every year of delay is valuable in itself, not only because of what gets done with it. I think evaluators need an income that doesn’t depend on the lab’s goodwill. I think bright lines should come before waiting to see. And I don’t think it matters much which flag flies over the lab that ends the world.
The good news is that the mechanism this site proposes doesn’t make anyone settle those disagreements first. Build the machinery, and set the dial wherever the politics can bear. Amodei’s setting is further than any frontier CEO has been willing to go before now. Mine is further still, and the dial turns.