Team Size Is a Governance Variable
The assumption every insider-reporting mechanism depends on, and which nobody is measuring.
Everything on this site rests on a claim I have never stated as an assumption, which is a bad sign about how carefully I was holding it. Page 05 puts it like this: no single person has ever had the ability to build even an unsafe AI by themselves, and the proof is that we are still here. From there page 07 derives the conspiracy arithmetic, page 09 derives the superlinear team-size tax, and page 12 uses it again to argue that collusion does not scale. Four pages of machinery on one sentence, and the sentence is an anthropic argument doing empirical work.
It is also not mine alone. The False Claims Act, the SEC whistleblower program, antitrust leniency, and every proposal to protect AI whistleblowers depend on the same premise: that the conduct involves enough people that at least one of them can be induced to talk. That premise is an empirical quantity. It has a value, it changes over time, and as far as I can tell nobody in AI governance is measuring it.
This piece tries to state it properly, put numbers on it where numbers exist, and say what would falsify it.
There are four different numbers here, and only one of them matters
The first problem is that “team size” is not one quantity.
| What it counts | Observable? | |
|---|---|---|
| Everyone at the lab | Roughly, via reporting | |
| Names on the technical report | Yes, exactly | |
| People without whom the run does not happen | No | |
| People who would recognise the activity as prohibited and could testify to it | No |
An insider-reporting mechanism does not care how many people work at a lab. It cares how many people could walk into a courtroom and describe the prohibited act. That is , and it is the only one of the four that belongs in the exponent.
Everyone reaches for instead, because it is the one you can count. It is wrong in both directions. It undercounts: contractors, data centre staff, safety researchers who saw the evaluation results and wrote nothing, the compliance lawyer who read the capability report. It overcounts worse: a great many listed authors contributed something real to a model without knowing anything that would support a claim. Being on a paper is not the same as being a witness.
This is not a quibble. It is the measurement problem at the centre of the whole question, and it is why the numbers in the next section should be read as a proxy whose bias is unknown rather than as the quantity itself.
What the author counts actually show
| Model | Year | Authors |
|---|---|---|
| GPT-2 | 2019 | 6 |
| GPT-3 | 2020 | 31 |
| Gopher | 2021 | 80 |
| PaLM | 2022 | 67 |
| GPT-4 | 2023 | ~280 (≈82 “core contributors”) |
| Gemini 1.0 | 2023 | 1,350 |
| DeepSeek-V3 | 2024 | 299 |
The naive worry — that frontier work is collapsing to small teams and the mechanism is about to lose its grip — is not what this shows. Author counts have risen by more than two orders of magnitude in five years. On the most easily-measured proxy, the assumption is in better shape than when I first wrote the argument down in 2022.
I do not believe the proxy.
Three things spoil it. First, Gemini 1.5’s author list grew from 662 to 1,136 across successive versions of the same paper. The team did not grow eightfold relative to Gemini 1.0’s neighbours; the crediting did. Second, Anthropic publishes model cards under a corporate byline with no individual attribution at all, so its team size is unobservable by this method — a lab can drop off the chart by changing its citation style. Third, and decisively: Google DeepMind’s pre-training lead has said in an interview that keeping a roughly forty-day Gemini Flash 2.0 training run alive took a rotation of about five people.
Set that against 1,350 names on the Gemini paper and you have the entire problem in one comparison. The two numbers differ by more than two orders of magnitude, they are both true, and they answer different questions. Nobody outside the labs knows which one governs.
That vacuum is real rather than a failure of searching. Nathan Lambert, who has worked inside these organisations, states plainly that how the leading laboratories structure their training teams is a closely guarded secret. I could not find a single on-record, lab-confirmed figure for the core pre-training team on any frontier run.
What is published is suggestive in a different way. Epoch AI’s company dataset has Anthropic at 915 staff in December 2024, of whom 521 were in R&D — so roughly 57% technical at one lab that discloses the split. And Epoch’s cost breakdown of frontier training runs puts R&D staff costs, including equity, at 29–49% of the total, against 47–67% for hardware.
That last figure is the one I would put in front of a policymaker. Labour is not a rounding error next to compute; it is the same order of magnitude as the hardware bill. We have built an entire governance literature on compute thresholds — reporting requirements, export controls, the training-run FLOP trigger — on an input that accounts for about half the cost of a frontier model. The other half is people, and there is no corresponding literature at all. Epoch tracks compute meticulously and publishes nothing modelling labour as an input to capability. I checked. It does not exist.
Cohesion, not headcount
Now the theory, because I think the site has been making a weaker version of this argument than it could.
Page 07 gives the standard result: people, each independently silent with probability , and the chance somebody talks is . It then concedes — correctly, and at some length — that independence is the weakest assumption in the model, because real colleagues are selected for shared conviction and bound by contracts, friendships, mortgages and visas.
I have come to think that concession is backwards. Write the correlation in explicitly. Let measure cohesion, and define an effective team size
so that detection over the period is . At this is page 07’s table unchanged. At the group behaves as a single bloc: , and a thousand people are exactly as leaky as one.
Which says something the original framing missed. What protects a conspiracy is not its size but its cohesion. Headcount only converts into detection to the extent that participants can act unilaterally.
And that reframes what the bounty is for. I have been describing it as an instrument that exploits large — the group is big, therefore somebody talks. That is not quite right. A large payment, claimable by one person without anybody’s agreement, with insurance making the aftermath survivable, is precisely an instrument for driving down. It attacks cohesion directly: it makes defection unilateral, makes it lucrative, and makes the consequences bearable.
The mechanism does not merely exploit effective team size. It manufactures it. That, rather than the raw exponent, is what I should have been arguing, and it means page 07’s honest caveat is better read as a statement of the target.
What the literature does and does not support
Here I have to report against my own interest.
The quantitative base is thin. The canonical model of secrecy failure scaling with group size is Grimes’s On the Viability of Conspiratorial Beliefs (2016), which is where the “a conspiracy this large cannot hold” genre gets its numbers. It is weak evidence. It is a single-parameter exponential fitted to three cases whose participant counts were estimated rather than measured, it required a published correction, and the original failure curves were non-monotonic — impossible for a survival function. Critics have called the resulting parameter circular, and I think that criticism lands. No peer-reviewed replacement exists.
More striking: there appears to be no study anywhere regressing detection probability on the number of people who knew. Not in the fraud literature, not in the cartel literature, not for classified information — RAND has said as much explicitly , noting the absence of empirical work correlating leak rates with the size of the cleared population. The central quantity in this entire family of mechanisms has never been estimated.
The real theoretical home is antitrust, not conspiracy modelling. Motta and Polo’s Leniency Programs and Cartel Prosecution (2003) formalises exactly the mechanism I want, under a different name. Their “pre-emption effect” is the rising incentive each member faces to reach the regulator first as their belief grows that somebody else will. That is the -scaling defection race, already modelled, already tested, and I had not connected it to this argument until now.
It also supplies a warning I would not have thought to include. Harrington and Chang’s model of cartel birth and death predicts that an effective leniency programme increases the average duration of detected cartels in the short run, because it kills the least stable ones first and leaves the durable ones in the sample. Long-run duration then falls. Empirical work in 2024 bears this out. So a working mechanism’s early statistics look like failure. If extinction bounties were adopted and the first years produced few cases and apparently entrenched offenders, that is consistent with the mechanism working, and I would rather commit to that prediction in advance than explain it afterwards.
And one finding cuts against me. Larger scientific teams have lower retraction rates, not higher. Rathmann and Rauhut’s study of retracted articles supports a control hypothesis — more co-authors, more mutual monitoring, less misconduct — and a 2021 Scientometrics analysis finds retractions skew heavily toward teams of two to four.
There are two readings and the data cannot separate them. Either large teams suppress misconduct through mutual monitoring, which is good for my argument and suggests the deterrent runs partly through colleagues rather than only through the state; or large teams conceal it better through diffusion of responsibility, which is bad for my argument and means headcount buys opacity rather than exposure. I cannot tell you which. It is the single most informative unresolved question I found, and a decent experiment probably exists inside existing retraction data.
The supporting evidence, for balance, is real but indirect. Dyck, Morse and Zingales find employees detect 19% of large-firm corporate fraud against the SEC’s 6% — decentralised insiders outperform the agency built for the job by three to one. And 93% of state-court trade secret cases involve a current or former employee as defendant. Insiders are the dominant vector for information leaving an organisation. Neither study reports how many people knew.
The Manhattan Project is usually cited backwards
The standard counterexample is that roughly 125,000 people worked on the Manhattan Project and kept it quiet, so large groups evidently can hold a secret.
They did not. The programme was penetrated by Fuchs, Hall and Greenglass independently, and Moscow knew about it before the FBI did. It stayed secret from the American public and from the Axis; it did not stay secret from the adversary who mattered.
Which is the distinction the mechanism needs. Secrecy is adversary-relative. Extinction bounties do not require a leak to the newspapers. They require one person to reach one claimant. On that standard, the largest secret ever attempted failed almost immediately, through multiple independent insiders, under wartime classification and the threat of execution.
The honest half: those insiders defected from ideological conviction, not for money, and nothing in the historical record tells us a payment would have worked faster. Madoff points the other way, and more usefully — seventeen years of fraud survived by keeping the circle tiny while thousands of investors saw nothing. That is not a counterexample to the thesis. It is the thesis running in reverse, and it is the cleanest evidence I have that minimum viable team size is the variable to watch.
The shelf life
Now the part that worries me.
Define as the smallest team that can reach capability level at time . The mechanism binds only while stays above some critical value , below which expected liability no longer exceeds the private value of the project. Note that this is a claim about the covert case: a secret programme is precisely the one that minimises headcount, so rather than observed team size is the adversarial bound and the right thing to track.
Three trends push down.
Algorithmic efficiency. Epoch estimates the compute required for a given capability level halves roughly every eight months , with a wide confidence interval. Yesterday’s frontier becomes tomorrow’s modest budget.
Research automation. METR’s time-horizon measurements have the length of task an agent can complete autonomously doubling every four to seven months, and on RE-Bench frontier agents outscored human experts by 4× at a two-hour budget, with humans only pulling ahead over longer horizons. If one researcher with a fleet of agents does what ten did, falls without any lab deciding anything.
Open weights and distillation. This is the sharpest attack and it deserves its own paragraph.
Tools like Heretic automate the removal of refusal behaviour from open-weight models. No expertise, no team, thousands of such models already published. Taken at face value this puts on the dangerous step and leaves an insider-reporting mechanism with nothing to bite on.
I think it is serious and I do not think it is fatal, for a reason worth being precise about. Abliteration does not create capability. It removes safety training from capability that a large team already built, and the result sits at the base model’s level. So the attack does not abolish the conspiracy — it relocates it to the release decision, which is made by an identifiable group at a lab squarely inside the mechanism’s coverage.
But that is a genuine concession, and it has a design consequence I had not drawn before. If open weights are the shortcut, then the act worth attaching a bounty to is not only the training run. It is the decision to release weights above a given capability level. That is a different statutory trigger than the one the sequence has been assuming, and pages 09 and 11 are written as though the training run were the only event worth pricing.
What I would measure
If someone wanted to make this tractable, here is the programme. None of it is technically hard; it is simply not being done.
- Author counts over time, by lab, normalised for crediting practice. The Gemini 1.5 version history is a natural experiment in credit inflation: the same work, the same model, two author counts. Use it to calibrate the bias.
- The ratio . Feinberg’s five-person rotation against 1,350 names is one observation. Ten more, from interviews and post-mortems, would establish whether the gap is two orders of magnitude or three.
- from the open-weight frontier. The smallest team that has produced a model at a given capability level, as a time series. This is directly observable from public releases and nobody maintains it.
- Retraction data as the natural experiment on . Whether larger teams suppress or conceal misconduct is answerable with data that already exists, and it is the crux.
- Labour as a governance input, alongside compute. Epoch’s cost work already shows labour at 29–49% of a training run. The modelling that exists for compute does not exist for people.
What would falsify this
I would rather say it now than be argued into it later.
The argument fails if a team below roughly ten people produces a genuinely frontier-capable model from scratch — not a distillation, not a fine-tune of somebody else’s weights. Nothing like that has happened. DeepSeek, the usual candidate, is somewhere between 150 and 400 people.
It fails if research automation reaches the point where a single person plus agents can do a full training run end to end, at which point and every insider-reporting mechanism in this family dies at once — the False Claims Act analogues included.
It is substantially weakened if the retraction result turns out to run through concealment rather than control, because then headcount buys opacity and the exponent works against me.
And it is weakened, though not destroyed, by anything that makes cohesion cheap: credible threats against defectors, enforceable secrecy, or a state programme with coercive tools. Cohesion is the real defence, and I have no measurement of it at all.
What I am not willing to do is leave the assumption unstated. It carries four pages of this site and a good deal of existing law, it has a number attached, and the number is moving. The mechanism has a shelf life. I would like to know how long it is, and right now nobody can tell me — which is, I think, the most interesting thing in this piece.