commit 93d095e856f5f33057ab4800523a982e12a948d6 Author: wassname Date: Sat Jul 11 17:29:53 2026 +0800 ANZ 2040 draft: decision-tree site (prose + inline probability-flow SVG) diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..73d4880 --- /dev/null +++ b/.gitignore @@ -0,0 +1,5 @@ +handover.md +sources/resources/ai-2027-for-anz-june-2026.md +*.err +.DS_Store +__pycache__/ diff --git a/.nojekyll b/.nojekyll new file mode 100644 index 0000000..e69de29 diff --git a/README.md b/README.md new file mode 100644 index 0000000..debd893 --- /dev/null +++ b/README.md @@ -0,0 +1,31 @@ +# ANZ 2040: the decision tree + +The [AI 2040 / Plan A](https://ai-2040.com) forecast, retold from an Australia and New +Zealand middle-power point of view: a decision tree of the choices Canberra and +Wellington actually face as AI gets more powerful, with named endings and our own +probability estimates. Genre is [AI 2027](https://ai-2027.com) and +[Europe 2031](https://europe2031.ai). This is a draft. + +## Build + +The site is one static `index.html` with no build step on GitHub's side. To regenerate +it locally: + +```sh +just build # or: python3 scripts/build_html.py && python3 scripts/build_tree.py +just serve # build, then serve at http://localhost:8080 +``` + +- `anz-2040.md` is the source document. +- `scripts/build_html.py` turns it into `index.html`: prose via pandoc, the decision + tree via [mermaid.js](https://mermaid.js.org/) loaded from a CDN, CSS inline. +- `scripts/build_tree.py` turns `data/tree.json` into `tree.svg`, a probability-sized + view of the same tree (node area tracks our estimated probability). It is a separate, + simplified view, not generated from the mermaid tree. +- `sources/` holds the AI-2040 supplements and outside references the document cites. + +## Credit + +Built on the AI-2040 supplements by the [AI Futures Project](https://ai-2040.com), used +with full credit. Sources and one-line summaries are in +[`sources/INDEX.md`](sources/resources/INDEX.md). diff --git a/anz-2040.md b/anz-2040.md new file mode 100644 index 0000000..eff2218 --- /dev/null +++ b/anz-2040.md @@ -0,0 +1,160 @@ +# ANZ 2040: the decisions we still get to make + + +What can Australia and New Zealand actually decide if AI becomes much more powerful over the next decade? + +We start from the excellent [AI 2040](https://ai-2040.com), which describes choices the United States and China could make. We zoom in on what those choices would mean for Australia and New Zealand, and what we could still do ourselves. + +The probabilities below assume rapid AI progress. We leave aside the roughly 35% chance that progress plateaus first. + +We don't get to choose whether the United States and China race, slow down, or make a deal. We don't get to choose whether a powerful AI remains under human control. But we do get a few choices about our position before those larger decisions are made. Some are cheap and useful across almost every future. The valuable later choices only exist if we prepare for them now. + +[![ANZ 2040: choices and possible futures](tree.svg)](tree.svg) + +## What Australia and New Zealand can do* + +### 1. Prepare before it looks urgent (2026-28) + +The later choices need ordinary government machinery: royalty laws, approved land, a way to pay a dividend, biosecurity capacity, and working relationships with other middle powers. None of this requires believing one precise AI forecast. It is cheap compared with discovering in 2029 that the legal and diplomatic work takes five years. + +Biosecurity is the clearest case. AI 2040 proposes "massive investments into biosecurity and other measures to improve the resilience of the world," but does not assign the work to anyone. New Zealand already runs a serious agricultural biosecurity system. This is a useful thing for it to get unusually good at. + +### 2. Train the people the treaty needs (2026-29) + +Australia and New Zealand can contribute people even when they cannot contribute frontier models. AI 2040 allocates only 1% of its safety budget to studying model character, meaning persistent behavioural tendencies or "personas+propensities." It says almost nothing about activation steering, which changes behaviour by intervening on a model's internal activity, or assistance games, which train an AI to infer what a person wants through cooperation. It does call truth-seeking AI and public forecasting a priority. Australia and New Zealand could specialise in those areas, as well as evaluations, treaty verification, biosecurity, and cyber work through the Five Eyes intelligence alliance. ([Alignment roadmap](https://ai-2040.com/supplements/alignment-roadmap), [AI for epistemics](https://ai-2040.com/supplements/ai-for-epistemics)) + +### 3. Ask for terms before signing (2029) + +Australia is named among the countries that join the deal. New Zealand is not mentioned in any of the 18 supplements we checked. Once inside, leaving could mean sanctions, cyberattack, or war. The realistic decision is therefore what to ask for before signing: hosted data centres, a strategic chip reserve, verification seats, equity in AI companies, and durable access rather than a revocable API account. + +Scott Alexander's summary is unusually blunt. Observer countries follow the rules in return for data centres, AI access, and a share of future wealth, even though the great powers don't strictly need to offer it. That makes accession our best bargaining moment. Signing the standard form gives it away. ([Introducing Plan A](https://www.astralcodexten.com/p/introducing-plan-a)) + +### 4. Host allied compute (2030-32) + +AI 2040 moves advanced chip fabs and robot production into special economic zones that can be destroyed if the deal breaks. The zones are meant to sit near an ocean or an adversarial power, and the robot economy expands inside them until it is as large as today's human economy. North-west Australia fits those criteria and sits above much of the ore the machines would consume. ([Deal decline](https://ai-2040.com/supplements/deal-decline), [economics](https://ai-2040.com/supplements/economics-of-plan-a)) + +Hosting creates exposure as well as income. The same cluster can be a target, collateral for the treaty, and leverage against being cut off. Tom Davidson makes the mechanism explicit: if the US withdraws AI access, allies could destroy US data centres on their territory. A kill switch negotiated at accession is leverage. A foreign cluster accepted without access rights is just exposure. ([Davidson 2026](https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting)) + +There is also a smaller trusted-vault option. Australia could host cold-storage model weights, treaty verification equipment, or monitoring linked to Pine Gap without hosting the whole robot economy. Geologically stable ground and Five Eyes trust make this the lowest-volume version of the policy. + +### 5. Keep a public stake in the machine economy (2031-34) + +The US dividend in AI 2040 comes from compute permits. By 2035, the scenario sends 75% of the US permit share to US citizens. Australia and New Zealand cannot rely on that. Our dividend would need resource royalties, land taxes, and public equity held through the Future Fund and NZ Super. Those rules need to exist before cognitive labour becomes cheap and wage earners lose their bargaining power. ([Economics of Plan A](https://ai-2040.com/supplements/economics-of-plan-a)) + +Who owns the rents matters more than the size of the boom. The same ore and land can support a public dividend or a few mining dynasties. Leicht and Ball warn that selling a strategic firm can look like an unexpected windfall while giving up the country's "ticket guaranteeing their home country's stake in the AI economy." In a rapid takeoff, taking only cash is especially bad: the useful payment is equity, access, or ownership of productive assets. ([The Race Worth Winning](https://www.thefai.org/posts/the-race-worth-winning-middle-powers-in-the-age-of-machine-intelligence)) + +Australia probably starts above the generic non-US country. Compulsory superannuation gives many households an existing claim on global capital, while mining gives governments a claim on inputs the robot economy needs. Neither is evenly shared by default. Superannuation gains follow existing balances and investment choices. Mining income reaches households only when royalties, taxes, public ownership, or transfers capture it. That is what separates the Dividend commonwealth from the Quarry economy. + +The political claim is stronger, and more speculative. If governments no longer depend on wages or citizen labour, owners and automated security may stop needing broad consent. A legally entrenched citizen claim on public assets keeps ordinary people tied to the revenue stream and gives them something material to defend through democratic institutions. We are treating distribution as a constraint on domestic power, not merely as welfare. + +### 6. Choose a side if the deal collapses (2030s) + +If the US-China deal declines, Australia's security guarantor and largest customer pull in opposite directions. Joining the US bloc keeps the alliance and loses trade. A concert with Japan, Korea, Canada, the Netherlands, and the UK might bargain for access, but only if it was built before the crisis. Leicht and Ball's warning is that a coalition improvised late will be picked off "one by one." Trying to hedge returns us to someone else's empire by a slower route. Free Chinese open-weight models don't remove the dependency; a supplier earning nothing from us also loses nothing by cutting us off. ([The Race Worth Winning](https://www.thefai.org/posts/the-race-worth-winning-middle-powers-in-the-age-of-machine-intelligence), [Davidson 2026](https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting)) + +## Where this could leave us* + +Doom and concentrated global power account for most of our probability. Australia and New Zealand only get the later choices if humanity keeps control and no single actor takes it all. + +### I. Doom, about 50% + +AI 2040's own comparison gives Plan A a 42% chance of a "great future," falling to 10-25% for its alternatives. Combining the table's conditional rows implies misaligned takeover at roughly one quarter under Plan A and roughly three quarters in the modal race. Our 50% is a rough mixture across those possibilities. + +Yudkowsky and Soares state the doom case much more strongly than our estimate: + +> "If any company or group, anywhere on the planet, builds an artificial superintelligence using anything remotely like current techniques, based on anything remotely like the present understanding of AI, then everyone, everywhere on Earth, will die." +> +> Eliezer Yudkowsky and Nate Soares, [*If Anyone Builds It, Everyone Dies*](https://www.hachettebookgroup.com/titles/eliezer-yudkowsky/if-anyone-builds-it-everyone-dies/9780316595667/) + +Once control is lost, Australia has no policy response left. The preparations above have to improve the odds before then. + +### II. Someone else's empire, about 25% + +The AI remains aligned, but control does not spread. A government, company, or person holds the decisive systems. Australia and New Zealand may be materially rich, especially while iron demand is high, but they live under rules set elsewhere. + +> "a tiny group of people, or possibly just a single individual, is effectively in control of the world’s only army of superintelligences" +> +> AI Futures Project, [AI 2040](https://ai-2040.com) + +### III. Dividend commonwealth + +We negotiated terms, captured resource and land rents, and built the dividend before wages collapsed. Households share in the assets that became valuable instead of watching those assets reprice around them. The better version also uses AI inside courts, health systems, and government rather than treating it only as something to regulate. + +> "In Plan A, most of the US permit revenue share (75% in 2035) is redistributed to the US population as a Citizen’s Dividend, resulting in roughly $1M/yr per person in 2035 and around $10M/yr in 2040." +> +> AI Futures Project, [Economics of Plan A](https://ai-2040.com/supplements/economics-of-plan-a) + +### IV. Quarry economy + +This is the drift outcome. The country exports what the robot economy needs, but has no dividend rail and no public stake. AI 2040's land model makes the distribution problem concrete: single-family homes rise about 40 times while 97% of their paper value becomes land. Incumbent owners gain while cognitive wages fall. The nation is rich and the median household is poor, without anyone needing to plan it. + +> "As AI devalues human capital relative to physical and intangible assets, capital income consumes a growing proportion of national income. Elevated spending on luxury goods and a productivity-driven investment boom sustains elevated economic growth even as millions face unemployment and diminished purchasing power." +> +> UK Government Office for Science and AI Security Institute, [AI Scenarios 2030](https://www.gov.uk/government/publications/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai) + +The resource boom also expires. On current reserves and extraction, iron is the shortest-dated major reserve at roughly 55 years. A machine economy growing every 6-12 months would bring exhaustion and substitution forward. AI 2040 models the demand for energy and land, but not the Australian terms of trade or the cost of scaling mines. + +### V. Allied compute host + +Australia accepts the data centres and robot zones, but only after securing access and ownership terms. Land, power, and Five Eyes trust become a stake in the machine economy. The price is physical exposure: the hosted clusters are treaty collateral, and the treaty's wars become ours. + +> "If the US withdraws AI access, allies could destroy US data centres in response. It's a way to lock in the deal." +> +> Tom Davidson, [How can the middle powers avoid getting trounced?](https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting) + +### VI. Garrison ally + +The deal collapses and Australia chooses the US bloc. We keep the security relationship and lose much of the Chinese market. Iron becomes a strategic allocation rather than an ordinary export. This future is poorer and tenser than the dividend or hosting alternatives, but Australia still has a seat. + +> "middle powers should help the US, and make sure they are rewarded with continued access to frontier AI and new technologies (including military tech)" +> +> "The only alternative that makes sense to me is siding with China." +> +> Tom Davidson, [How can the middle powers avoid getting trounced?](https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting) + +### VII. Boring decade, about 10% + +AI stalls or governments shut the frontier down. The consents are unused, the royalty argument was early, and the extra biosecurity capacity looks like insurance that did not pay. This is the cheapest way for our preparations to be wrong. + +> "AI progress slows, and AI causes less disruption than expected." +> +> UK Government Office for Science and AI Security Institute, [AI Scenarios 2030](https://www.gov.uk/government/publications/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai) + +### VIII. Middle-power concert + +The deal declines, but a coalition formed before the crisis holds together. Australia and New Zealand pool their supply-chain and hosting leverage with Japan, Korea, Canada, the Netherlands, and the UK. None could demand durable access alone. Together they may get terms and help write the rules. + +> "Coordination, done right, reassures partners that they can focus on their comparative advantage instead of pursuing full autarky; and it enables powers to see eye to eye with AI exporters instead of being picked off by great powers offering rewards for defection one by one." +> +> Leicht and Ball, [The Race Worth Winning](https://www.thefai.org/posts/the-race-worth-winning-middle-powers-in-the-age-of-machine-intelligence) + +### IX. ~~Homegrown autocracy~~ free beer and rugby for life + +This is the domestic version of the global power-grab risk. Cognitive wages collapse, public ownership never develops, and automated security removes the state's remaining practical dependence on citizens. No coup is required. Political leverage drains away with economic leverage. + +> Luckily we completely dodged a coup and now have Benevolent Dictator for Life (BDFL) Gina Palmer to lead us through. Everyone is very happy, especially since the free beer, rugby, and netball are a great respite from our quickly shrinking population and ever more confusing global situation. - Joe Citizen, 2041 + +> "AI could create a similar effect: if governments can generate massive revenue from taxing AI projects rather than citizens, heads of state may lose their economic incentive to ensure citizens prosper. This would weaken citizens’ power to resist coup and backsliding attempts." +> +> "By replacing government employees with loyal AI systems, a head of state could remove important checks on their power." +> +> Davidson, Finnveden, and Hadshar, [AI-Enabled Coups](https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power) + +## Where we depart from AI 2040 + +AI 2040 models energy consumption rising about eight times by 2040 and notes that proven mineral reserves cover decades of current extraction, but it does not price the tonnage needed for a robot economy growing every 6-12 months. Coal, uranium, energy trade, and Australian mine capacity are outside its model. ([Economics of Plan A](https://ai-2040.com/supplements/economics-of-plan-a)) + +We use its land estimates with the authors' warning that this is a "very rough model" they do not necessarily endorse. We also give model character and activation steering more weight than their roadmap does. + +AI 2040 does share some US wealth abroad, but very unevenly. In 2032 its American dividend starts at $45,000 while adults outside the US and China average $1,200. By 2035 the figures are about $1M and $10,000. Australia may do better than that generic non-US payment through superannuation and mining, but only if the gains are broadly owned. A rich national balance sheet can still leave the median household without economic leverage. ([AI 2040](https://ai-2040.com/?choices=plan-a-root)) + +We leave out a trans-Tasman split and a separate Australian decision during US sabotage operations because neither changes the main choices above. + +## Sources and thanks + +This is built from the AI Futures Project's [AI 2040](https://ai-2040.com) scenario and supplements. Their caveat applies to our use of the numbers too: the economic model "plausibly contains bugs" and they do not trust its outputs to be accurate. The local source mirrors retain exact line references for auditing. + +It also benefited from a June 2026 scoping discussion with authors of related scenarios and Australian policy participants. Those workshop notes are private and are not included in the public source archive. + +* The Australian and New Zealand decisions and futures are extrapolations from the cited work. A citation means we used that source, not that its authors endorse this scenario. + +The middle-power framing also draws on [AI 2027](https://ai-2027.com), [Europe 2031](https://europe2031.ai), Tom Davidson's [middle-power plan](https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting), Leicht and Ball's [The Race Worth Winning](https://www.thefai.org/posts/the-race-worth-winning-middle-powers-in-the-age-of-machine-intelligence), and Scott Alexander's [Introducing Plan A](https://www.astralcodexten.com/p/introducing-plan-a). Australian groundwork came from the [e61 Institute](https://e61.in), [Tech Policy Design Institute](https://techpolicy.au/aiagency), [ASPI](https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/), and [Kate Chaney MP](https://www.katechaney.com.au/making_technology_safe). diff --git a/data/tree.json b/data/tree.json new file mode 100644 index 0000000..9918183 --- /dev/null +++ b/data/tree.json @@ -0,0 +1,53 @@ +{ + "comment": "Claude: single source for tree.svg. Each item has a tier (row) and left-to-right order; two items per row max. `p` and edge flows record our probability mass conditional on the 2029 decision mattering. Display widths fit the text. The HTML build fails if an href no longer matches a prose heading.", + "title": "Choose a path", + "canvas": {"w": 560, "top": 40, "row_h": 115, "node_h": 56, "gap": 44, "floor_w": 130, "p_scale": 300, "choice_w": 150}, + "kinds": { + "global": {"fill": "#e0e0e0", "stroke": "#888888"}, + "hazard": {"fill": "#f4cccc", "stroke": "#990000"}, + "anz": {"fill": "#ffd966", "stroke": "#b8860b"}, + "robust": {"fill": "#b6d7a8", "stroke": "#38761d"}, + "outcome": {"fill": "#cfe2f3", "stroke": "#3d85c6"} + }, + "nodes": [ + {"id": "A", "tier": 0, "order": 0, "kind": "robust", "href": "#1-prepare-before-it-looks-urgent-2026-28", "title": "1. Prepare early", "sub": "2026-28: royalties, biosecurity, allies"}, + {"id": "F", "tier": 0, "order": 1, "kind": "robust", "href": "#2-train-the-people-the-treaty-needs-2026-29", "title": "2. Build the people", "sub": "2026-29: safety testers, inspectors"}, + {"id": "FORK", "tier": 1, "order": 0, "kind": "global", "href": "#what-australia-and-new-zealand-can-do", "title": "2029: the US chooses", "sub": "go-slow deal with China, or race?"}, + {"id": "VII", "tier": 2, "order": 0, "kind": "outcome", "p": 0.10, "href": "#vii-boring-decade-about-10", "title": "VII. Boring decade", "sub": "~10%"}, + {"id": "AL", "tier": 2, "order": 1, "kind": "hazard", "p": 0.90, "href": "#i-doom-about-50", "title": "RISK: AI escapes control", "sub": "~75% racing, ~25% in the deal"}, + {"id": "I", "tier": 3, "order": 0, "kind": "outcome", "p": 0.50, "href": "#i-doom-about-50", "title": "I. Doom", "sub": "~50%"}, + {"id": "PC", "tier": 3, "order": 1, "kind": "hazard", "p": 0.40, "href": "#ii-someone-elses-empire-about-25", "title": "RISK: power seized", "sub": "one actor holds the AI"}, + {"id": "II", "tier": 4, "order": 0, "kind": "outcome", "p": 0.25, "href": "#ii-someone-elses-empire-about-25", "title": "II. Someone else's empire", "sub": "~25% — seized abroad"}, + {"id": "HOME", "tier": 4, "order": 1, "kind": "hazard", "href": "#ix-homegrown-autocracy-free-beer-and-rugby-for-life", "title": "RISK: power grabbed at home", "sub": "seized in our own country"}, + {"id": "B", "tier": 5, "order": 0, "kind": "anz", "href": "#3-ask-for-terms-before-signing-2029", "title": "3. Sign, or set our price?", "sub": "the deal reaches us, 2029"}, + {"id": "IX", "tier": 5, "order": 1, "kind": "outcome", "href": "#ix-homegrown-autocracy-free-beer-and-rugby-for-life", "title": "IX. Homegrown autocracy", "sub": "citizens stop being needed"}, + {"id": "E", "tier": 6, "order": 0, "kind": "anz", "href": "#6-choose-a-side-if-the-deal-collapses-2030s", "title": "6. If the deal collapses", "sub": "US bloc, concert, or neutral?"}, + {"id": "C", "tier": 6, "order": 1, "kind": "anz", "href": "#4-host-allied-compute-2030-32", "title": "4. Host compute here?", "sub": "2030-32"}, + {"id": "VIII", "tier": 7, "order": 0, "kind": "outcome", "href": "#viii-middle-power-concert", "title": "VIII. Middle-power concert", "sub": ""}, + {"id": "V", "tier": 7, "order": 1, "kind": "outcome", "href": "#v-allied-compute-host", "title": "V. Allied compute host", "sub": ""}, + {"id": "VI", "tier": 8, "order": 0, "kind": "outcome", "href": "#vi-garrison-ally", "title": "VI. Garrison ally", "sub": ""}, + {"id": "D", "tier": 8, "order": 1, "kind": "anz", "href": "#5-keep-a-public-stake-in-the-machine-economy-2031-34", "title": "5. Tax it, own a share?", "sub": "2031-34"}, + {"id": "III", "tier": 9, "order": 0, "kind": "outcome", "href": "#iii-dividend-commonwealth", "title": "III. Dividend commonwealth", "sub": ""}, + {"id": "IV", "tier": 9, "order": 1, "kind": "outcome", "href": "#iv-quarry-economy", "title": "IV. Quarry economy", "sub": "drift default"} + ], + "edges": [ + {"from": "A", "to": "FORK", "flow": 0.50}, + {"from": "F", "to": "FORK", "flow": 0.50}, + {"from": "FORK", "to": "VII", "flow": 0.10, "label": "stalls"}, + {"from": "FORK", "to": "AL", "flow": 0.90, "label": "race / deal"}, + {"from": "AL", "to": "I", "flow": 0.50, "label": "control lost"}, + {"from": "AL", "to": "PC", "flow": 0.40, "label": "control kept"}, + {"from": "PC", "to": "II", "flow": 0.24, "label": "seized abroad"}, + {"from": "PC", "to": "HOME", "flow": 0.05, "label": "seized at home"}, + {"from": "PC", "to": "B", "flow": 0.11, "label": "power stays spread"}, + {"from": "HOME", "to": "IX", "flow": 0.05, "label": "no public stake"}, + {"from": "B", "to": "C", "flow": 0.08, "label": "set our price"}, + {"from": "B", "to": "E", "flow": 0.03, "label": "if deal collapses", "dashed": true}, + {"from": "C", "to": "V", "flow": 0.04, "label": "host it"}, + {"from": "C", "to": "D", "flow": 0.04, "label": "stay supplier"}, + {"from": "E", "to": "VIII", "flow": 0.015, "label": "concert"}, + {"from": "E", "to": "VI", "flow": 0.015, "label": "US bloc"}, + {"from": "D", "to": "III", "flow": 0.02, "label": "public stake"}, + {"from": "D", "to": "IV", "flow": 0.02, "label": "no stake"} + ] +} diff --git a/global.css b/global.css new file mode 100644 index 0000000..60b1cea --- /dev/null +++ b/global.css @@ -0,0 +1,609 @@ +@tailwind base; +@tailwind components; +@tailwind utilities; + +:root { + --vivid-background: #fffff8; + --vivid-foreground: #000; + --dull-background: #f0f0f0; + --dull-foreground: #666; + --accent: #2A623D; /* Slytherin green */ + --plan-s-background: #31566f; + --plan-a-background: #2A623D; + --race-background: #8B0000; + --sabotage-background: #c77400; + --slowdown-background: #2A623D; +} + +.race { + --accent: var(--race-background); +} + +@property --w { + syntax: ''; + inherits: true; + initial-value: 0.00000001%; +} + +/* Extra px the supplement essay's content column is nudged right on wide layouts + (toc rail visible). Factored out because TWO rules must agree on it: the shift + applied to the content column's text (.supplement-content-shift) and the equal + shift applied to the header row (.supplement-header-row) so "AI 2040" lines up + with the content title. Edit here to move both together. */ +:root { + --supplement-content-shift: 30px; +} + +.main-grid { + display: grid; + margin-left: auto; + margin-right: auto; + grid-template-areas: + "gap1 tocGap header header header gap2" + "gap1 topGraph topGraph topGraph topGraph gap2" + "gap1 toc content graphgap graph gap2" + "gap1 toc2 choice graphgap2 graph gap2" + "gap1 fgap fgap fgap fgap gap2" + "gap1 gap3 cFooter footer graphFooter gap2"; + /* toc column: as the page narrows, the flexible gap columns (1fr/0.3fr) + give up their space first while the rail keeps its full 220px. Once the + viewport can't fit the other columns at their preferred widths (4px gap + + 700px content + 32px graphgap + 430px graph = 1166px), the toc column + itself shrinks — down to 140px — and the rail labels wrap to the narrower + column. Below 1281px the rail is hidden entirely (see below). */ + grid-template-columns: 4px clamp(140px, calc(100vw - 1166px), 220px) minmax(300px, 700px) minmax(32px, 1fr) minmax(min-content, 430px) 0.3fr; + grid-template-rows: auto auto auto auto auto; + padding-top: 8px; +} + +/* Pushes the central content column an extra 30px to the right, on top of the + 30px gained by widening the toc column. Only applied on layouts where the + toc rail is visible (≥1281px); narrower views hide the rail entirely. */ +@media (min-width: 1281px) { + .supplement-content-shift { + padding-left: var(--supplement-content-shift); + } +} + +/* Aligns the supplement header ("AI 2040") with the essay's content column. + The header grid area is `header header header` (parent columns 3–5), so it + already starts exactly at the `content` column's left edge — left-aligned + header text begins there for free. The only remaining offset is the same + `--supplement-content-shift` that nudges the content column's *text* right + (padding on
), so we apply that identical shift to the header row and + the two line up. Scoped to the toc-rail layouts (≥1281px), where the shift + applies; narrower views don't shift the content, so the header already + matches without it. */ +@media (min-width: 1281px) { + .supplement-header-row { + padding-left: var(--supplement-content-shift); + } +} + +@media (max-width: 1280px) { + .main-grid.main-grid { + grid-template-columns: 32px 0px minmax(300px, 700px) minmax(64px, 1fr) minmax(min-content, 430px) 8px; + } + .section-nav { + display: none; + } +} + +@media (max-width: 930px) { + .main-grid.main-grid { + grid-template-areas: + "gap1 header header header gap2" + "gap1 topGraph topGraph graph gap2" + "gap1 content content content gap2" + "gap1 choice choice choice gap2" + "gap1 fgap fgap fgap gap2" + "gap1 cFooter cFooter cFooter gap2"; + grid-template-columns: 10px minmax(0px, 900px) 64px 0px 10px; + } +} + + +@media (min-width: 1600px) { + .main-grid.main-grid { + grid-template-columns: 8px minmax(250px, 1fr) minmax(300px, 700px) minmax(32px, 96px) minmax(min-content, 430px) minmax(8px, 1fr); + } +} + +body.antialiased { + -webkit-font-smoothing: subpixel-antialiased; +} + +@layer utilities { + .text-balance { + text-wrap: balance; + } +} + +summary { + cursor: pointer; +} + +details b, +details strong, +.report-details b, +.report-details strong { + font-weight: 600; +} + +.content-with-chart { + padding-top: 50vh; + display: flex; + align-items: flex-start; + gap: 40px; +} + +.main-content { + flex: 1; +} + +.sticky-chart-container { + position: sticky; + top: 20px; + width: 330px; + overflow: hidden; + margin-top: 250px; + opacity: 0.05; + transition: opacity 0.15s ease-in-out; + flex-grow: 0.35; + border: 1.5px solid grey; + border-radius: 12px; + padding: 8px; +} + +.grid-small { + grid-template-areas: + "rdmultiple rdmultiple rdmultiple" + "computeBudgets computeBudgets goalsPieChart" + "capability capability goalsPieChart" + "keystats keystats keystats" + "techBuckets techBuckets techBuckets" + "agentPopulation agentPopulation agentPopulation"; + grid-template-columns: 1fr 1fr 1fr; + grid-template-rows: auto auto 1fr auto auto auto; + /* grid-template-rows: 130px 60px 60px 100px 100px 50px; */ +} + +/* Single column grid layout for mobile fullscreen view */ +.grid-small-mobile { + grid-template-areas: + "rdmultiple rdmultiple" + "capability goalsPieChart" + "keystats keystats" + "techBuckets techBuckets" + "agentPopulation agentPopulation"; + grid-template-columns: auto 100px; + grid-template-rows: auto auto auto auto auto; + width: 100%; + max-width: 100vw; + overflow-x: hidden; +} + +.footnote-hover { + position: relative; + text-decoration: underline dotted; + cursor: pointer; +} + +/* The container for the actual footnote DOM */ +.footnote-popup { + position: absolute; + top: 1.2em; /* or bottom: 1.2em, etc. */ + left: 0; + width: 250px; + max-width: 300px; + background-color: rgba(245, 245, 245, 0.95); + color: #444; + border: 1px solid #ccc; + border-radius: 6px; + padding: 0.5rem; + font-size: 0.85rem; + line-height: 1.2; + z-index: 9999; +} + +/* Show the .footnote-popup when hovering over the parent link */ +.footnote-hover:hover .footnote-popup { + display: block; +} + +/* HUD PIP and fullscreen mode styles */ +.hud-pip { + transform-origin: bottom right; + box-shadow: 0 10px 25px rgba(0, 0, 0, 0.2); + z-index: 999; +} + +.hud-pip [data-state="open"] { + display: none !important; /* Disable tooltips in PIP mode */ +} + +.hud-fullscreen { + z-index: 999; +} + +@layer base { + :root { + --background: 0 0% 100%; + --foreground: 0 0% 3.9%; + --card: 0 0% 100%; + --card-foreground: 0 0% 3.9%; + --popover: 0 0% 100%; + --popover-foreground: 0 0% 3.9%; + --primary: 0 0% 9%; + --primary-foreground: 0 0% 98%; + --secondary: 0 0% 96.1%; + --secondary-foreground: 0 0% 9%; + --muted: 0 0% 96.1%; + --muted-foreground: 0 0% 45.1%; + --accent: 0 0% 96.1%; + --accent-foreground: 0 0% 9%; + --destructive: 0 84.2% 60.2%; + --destructive-foreground: 0 0% 98%; + --border: 0 0% 89.8%; + --input: 0 0% 89.8%; + --ring: 0 0% 3.9%; + --chart-1: 12 76% 61%; + --chart-2: 173 58% 39%; + --chart-3: 197 37% 24%; + --chart-4: 43 74% 66%; + --chart-5: 27 87% 67%; + --radius: 0.5rem; + } + /* .dark { + --background: 0 0% 3.9%; + --foreground: 0 0% 98%; + --card: 0 0% 3.9%; + --card-foreground: 0 0% 98%; + --popover: 0 0% 3.9%; + --popover-foreground: 0 0% 98%; + --primary: 0 0% 98%; + --primary-foreground: 0 0% 9%; + --secondary: 0 0% 14.9%; + --secondary-foreground: 0 0% 98%; + --muted: 0 0% 14.9%; + --muted-foreground: 0 0% 63.9%; + --accent: 0 0% 14.9%; + --accent-foreground: 0 0% 98%; + --destructive: 0 62.8% 30.6%; + --destructive-foreground: 0 0% 98%; + --border: 0 0% 14.9%; + --input: 0 0% 14.9%; + --ring: 0 0% 83.1%; + --chart-1: 220 70% 50%; + --chart-2: 160 60% 45%; + --chart-3: 30 80% 55%; + --chart-4: 280 65% 60%; + --chart-5: 340 75% 55%; + } */ +} +@layer base { + * { + @apply border-border; + } + body { + @apply bg-background text-foreground; + } +} + +h2 { + scroll-margin-top: 25vh; +} + +/* Add this CSS rule into your global stylesheet */ +/* Reveal the text by clipping it from the bottom upward */ +@keyframes revealDown { + 0% { + clip-path: inset(0 0 100% 0); /* Clipped from the bottom (hidden) */ + opacity: 1; + } + 50% { + opacity: 0.8; + clip-path: inset(0 0 95% 0); + } + 100% { + clip-path: inset(0 0 0 0); /* Fully visible */ + opacity: 1; + } +} + +/* CSS class to trigger the revealDown animation */ +.animate-revealDown { + animation: revealDown 1.5s ease-out forwards; +} + +/* Audio player scrub label: slide the chapter title in/out as you scrub. */ +@keyframes scrub-enter-right { + from { transform: translateX(100%); opacity: 0; } + to { transform: translateX(0); opacity: 1; } +} +@keyframes scrub-enter-left { + from { transform: translateX(-100%); opacity: 0; } + to { transform: translateX(0); opacity: 1; } +} +@keyframes scrub-exit-left { + from { transform: translateX(0); opacity: 1; } + to { transform: translateX(-100%); opacity: 0; } +} +@keyframes scrub-exit-right { + from { transform: translateX(0); opacity: 1; } + to { transform: translateX(100%); opacity: 0; } +} +.animate-scrub-enter-right { animation: scrub-enter-right 0.3s ease-out; } +.animate-scrub-enter-left { animation: scrub-enter-left 0.3s ease-out; } +.animate-scrub-exit-left { animation: scrub-exit-left 0.3s ease-out forwards; } +.animate-scrub-exit-right { animation: scrub-exit-right 0.3s ease-out forwards; } + +/* Override katex default larger font size */ +.katex.katex { + font-size: 1em; +} + +number-flow-react::part(left), +number-flow-react::part(right), +number-flow-react::part(left)::after, +number-flow-react::part(right)::after, +number-flow-react::part(symbol) { + padding: calc(var(--number-flow-mask-height, 0.25em) / 2) 0; +} + +.essay-footnote { + flex-direction: row; + gap: 10px; +} + +.essay-footnote p, .essay-footnote ol, .essay-footnote ul { + font-size: 14px; + margin: 0; + line-height: 1.5; +} + +main h1 { + font-size: 52px; +} + +.footnote-page p { + display: inline; +} + +.summary-chart p { + font-size: 14px; + line-height: 1.4; +} + +.tab-boxes p, .tab-boxes li { + font-size: 16px; + line-height: 1.45; + margin-bottom: 0.5em; +} + +.tab-boxes ol { + margin-top: 0.5em; + margin-bottom: 0.5em; +} + +.tab-boxes a { + text-decoration: none; + color: var(--accent); + font-weight: 600; +} + +@supports not selector(:first-child) { + .requires-modern-css { + display: none !important; + } +} + +.scrollbar-hide, .scrollbar-hide * { + -ms-overflow-style: none; + scrollbar-width: none; + scrollbar-color: transparent transparent; +} +.scrollbar-hide::-webkit-scrollbar { + display: none; +} + +blockquote { + border: 2px solid #232323; + border-radius: 8px; + padding: .25rem 1.5rem; + margin: 1rem 0; +} + +.good-ending-text { font-size: 1.4rem; } +.good-ending-text h1 { font-size: 3rem; margin: 1.25em 0 1em 0; } + +@media (max-width: 767px) { + .good-ending-text h1 { font-size: 2rem; line-height: 1.3; } + /* Keep the heading ladder descending under the 2rem mobile h1: tufte's + desktop sizes (h2 2.25rem, h3 1.85rem, h4 1.5rem) would otherwise + render h2 BIGGER than h1 and h4 bigger than the scaled h3. h5 keeps + tufte's 1.3rem, which already sits below the 1.35rem h4. */ + .good-ending-text h2 { font-size: 1.7rem; } + .good-ending-text h3 { font-size: 1.45rem; } + .good-ending-text h4 { font-size: 1.35rem; } + /* ai-2027's mobile body size. tufte.css carries the same rule at 730px, + but the .good-ending-text font-size above out-specifies it (paragraphs + inherit), so it must be restated at this level. 767px matches the rest + of this site's mobile breakpoint. */ + .good-ending-text { font-size: 1.25rem; } + .good-ending-text :where(p, dl, ol, ul, summary, blockquote, li) { + line-height: 1.45; + } +} +.good-ending-text :where(p, dl, ol, ul, summary, blockquote, li) { font-size: inherit; } + +/* Narrow non-mobile window (two-column playbook layout, 930–1024px): shave a + couple px off body text and scale headings down ~10%. Set once here at the + top-level container instead of threading a class through Markdown.tsx. + Line-height is unitless (1.5) so it auto-follows whatever font-size wins — + change the size and the leading tracks it, no separate override needed. */ +@media (min-width: 930px) and (max-width: 1135px) { + .good-ending-text { font-size: 1.2rem } + .good-ending-text :where(p, dl, ol, ul, summary, blockquote, li) { line-height: 1.5; } + .good-ending-text h1 { font-size: 2.4rem; } + .good-ending-text h2 { font-size: 2.03rem; } + .good-ending-text h3 { font-size: 1.53rem; } + .good-ending-text h4 { font-size: 1.35rem; } + .good-ending-text h5 { font-size: 1.17rem; } +} +.markdown-content.supplement-essay-prose { + box-sizing: border-box; +} +.markdown-content.supplement-essay-prose > :not(h1):not(h2):not(h3):not(h4):not(h5):not(h6) { + max-width: 800px; + margin-left: auto; + margin-right: auto; +} +.markdown-content.supplement-essay-prose > h1, +.markdown-content.supplement-essay-prose > h2, +.markdown-content.supplement-essay-prose > h3, +.markdown-content.supplement-essay-prose > h4, +.markdown-content.supplement-essay-prose > h5, +.markdown-content.supplement-essay-prose > h6 { + max-width: none; + width: 100%; + box-sizing: border-box; + overflow-wrap: break-word; +} +/* Top-level Part headings (h1) get a strong gap above to separate the major + sections — base tufte h1 has margin-top:0, which left Part 2 cramped. */ +.markdown-content.supplement-essay-prose > h1 { + margin-top: 5rem; +} + +/* Images carry a small negative left margin for optical alignment in the main + essay column; inside an expanded details section that would push them across + the section's brown left rule. Reset it there so they stay inside the rule. */ +details .ml-\[-20px\], +details .ml-\[-8px\] { + margin-left: 0 !important; +} + +/* "Back to contents" arrow in the left margin of supplement headings. Hidden + everywhere else (the link target #top-toc only exists in supplements). */ +.toc-backlink { + display: none; +} +.markdown-content.supplement-essay-prose > :is(h1, h2, h3, h4, h5, h6) { + position: relative; +} +/* On every level of section heading inside supplements (h1 = top-level + section like Executive Summary / Verifying the deal; h2/h3 = sub-sections + and timeline rows). Scoped via :has(#top-toc) so the triangle only + appears in supplements whose body actually contains a :::toc block — + right now that's only verification-plan. Em-sized border-triangle, + vertically centered. */ +.markdown-content.supplement-essay-prose:has(#top-toc) > :is(h1, h2, h3) .toc-backlink { + display: inline-block; + position: absolute; + right: 0; + top: 50%; + transform: translateY(-50%); + width: 0; + height: 0; + border-left: 0.32em solid transparent; + border-right: 0.32em solid transparent; + border-bottom: 0.5em solid rgba(0, 0, 0, 0.2); + text-decoration: none; + transition: border-bottom-color 0.15s ease-out; +} +.markdown-content.supplement-essay-prose:has(#top-toc) > :is(h1, h2, h3) .toc-backlink:hover { + border-bottom-color: rgba(0, 0, 0, 0.55); +} + +.ai2030-feedback-scroll-area { + scrollbar-width: none; + -ms-overflow-style: none; +} +.ai2030-feedback-scroll-area::-webkit-scrollbar { + width: 0; + height: 0; + display: none; +} + +/* Feed panel on /feedback sliding open from its collapsed strip. */ +@keyframes ai2030-feed-slide-in { + from { + width: 44px; + } + to { + width: 420px; + } +} + +/* Persistent highlight on the selected text while its feedback dialog is open, + so it stays visible even after focus moves into the dialog inputs. */ +::highlight(ai2030-feedback-selection) { + background-color: #accef7; + color: black; +} + +/* While the feedback dialog's persistent highlight is showing, the native + selection is kept alive (so the reader can still copy the text) but restyled + to the same color so the two highlights render as one. */ +.ai2030-feedback-selection-active ::selection { + background-color: #accef7; + color: black; +} + +/* Gentle pop when a submitted-feedback margin icon first appears (page load, + and right after the feedback dialog fades out on submit). */ +@keyframes ai2030-feedback-icon-pop { + from { + opacity: 0; + transform: scale(0.6); + } + to { + opacity: 1; + transform: scale(1); + } +} +.ai2030-feedback-icon-pop { + animation: ai2030-feedback-icon-pop 180ms ease-out; +} + +/* Subtle dotted underline on passages the reader has already left feedback on + (rendered by SubmittedFeedbackLayer via the CSS Custom Highlight API). */ +::highlight(ai2030-feedback-submitted) { + text-decoration-line: underline; + text-decoration-style: dotted; + text-decoration-color: rgba(0, 0, 0, 0.45); + text-underline-offset: 3px; +} + +/* Solid underline on the passage whose margin icon/popover is being hovered, + so the reader can tell which line the feedback belongs to. */ +::highlight(ai2030-feedback-submitted-active) { + text-decoration-line: underline; + text-decoration-style: solid; + text-decoration-color: rgba(0, 0, 0, 0.8); + text-underline-offset: 3px; +} + +.ai2030-feedback-inline-highlight { + background-color: rgba(255, 220, 90, 0.45); + border-radius: 2px; + padding: 0 1px; + cursor: pointer; + transition: background-color 0.08s ease-out; +} +.ai2030-feedback-inline-highlight:hover { + background-color: rgba(255, 200, 60, 0.6); +} +.ai2030-feedback-inline-highlight[data-active="true"] { + background-color: rgba(255, 180, 30, 0.7); +} + +/* Hovering a feed card whose quote lives inside a collapsed expandable section + tints that section's summary the same warm highlight color, so the reviewer + can spot which section to open. */ +.ai2030-feedback-details-highlight { + background-color: rgba(255, 180, 30, 0.7); + border-radius: 4px; + transition: background-color 0.08s ease-out; +} diff --git a/index.html b/index.html new file mode 100644 index 0000000..6bde290 --- /dev/null +++ b/index.html @@ -0,0 +1,280 @@ + + + + + +ANZ 2040: the decisions we still get to make + + + +

ANZ 2040: the decisions we still get to make

+

What can Australia and New Zealand actually decide if AI becomes much more powerful over the next decade?

+

We start from the excellent AI 2040, which describes choices the United States and China could make. We zoom in on what those choices would mean for Australia and New Zealand, and what we could still do ourselves.

+

The probabilities below assume rapid AI progress. We leave aside the roughly 35% chance that progress plateaus first.

+

We don't get to choose whether the United States and China race, slow down, or make a deal. We don't get to choose whether a powerful AI remains under human control. But we do get a few choices about our position before those larger decisions are made. Some are cheap and useful across almost every future. The valuable later choices only exist if we prepare for them now.

+
+ +Choose a path + + + + + + + + + + + + + + + + + + + +stalls + +race / deal + +control lost + +control kept + +seized abroad + +seized at home + +power stays spread + +no public stake + +set our price + +host it + +stay supplier + +concert + +US bloc + +public stake + +no stake +1. Prepare early — 2026-28: royalties, biosecurity, allies + +1. Prepare early +2026-28: royalties, biosecurity, allies + +2. Build the people — 2026-29: safety testers, inspectors + +2. Build the people +2026-29: safety testers, inspectors + +2029: the US chooses — go-slow deal with China, or race? + +2029: the US chooses +go-slow deal with China, or race? + +VII. Boring decade — ~10% + +VII. Boring decade +~10% + +RISK: AI escapes control — ~75% racing, ~25% in the deal + +RISK: AI escapes control +~75% racing, ~25% in the deal + +I. Doom — ~50% + +I. Doom +~50% + +RISK: power seized — one actor holds the AI + +RISK: power seized +one actor holds the AI + +II. Someone else's empire — ~25% — seized abroad + +II. Someone else's empire +~25% — seized abroad + +RISK: power grabbed at home — seized in our own country + +RISK: power grabbed at home +seized in our own country + +3. Sign, or set our price? — the deal reaches us, 2029 + +3. Sign, or set our price? +the deal reaches us, 2029 + +IX. Homegrown autocracy — citizens stop being needed + +IX. Homegrown autocracy +citizens stop being needed + +6. If the deal collapses — US bloc, concert, or neutral? + +6. If the deal collapses +US bloc, concert, or neutral? + +4. Host compute here? — 2030-32 + +4. Host compute here? +2030-32 + +VIII. Middle-power concert + +VIII. Middle-power concert + +V. Allied compute host + +V. Allied compute host + +VI. Garrison ally + +VI. Garrison ally + +5. Tax it, own a share? — 2031-34 + +5. Tax it, own a share? +2031-34 + +III. Dividend commonwealth + +III. Dividend commonwealth + +IV. Quarry economy — drift default + +IV. Quarry economy +drift default + +
+

What Australia and New Zealand can do*

+

1. Prepare before it looks urgent (2026-28)

+

The later choices need ordinary government machinery: royalty laws, approved land, a way to pay a dividend, biosecurity capacity, and working relationships with other middle powers. None of this requires believing one precise AI forecast. It is cheap compared with discovering in 2029 that the legal and diplomatic work takes five years.

+

Biosecurity is the clearest case. AI 2040 proposes "massive investments into biosecurity and other measures to improve the resilience of the world," but does not assign the work to anyone. New Zealand already runs a serious agricultural biosecurity system. This is a useful thing for it to get unusually good at.

+

2. Train the people the treaty needs (2026-29)

+

Australia and New Zealand can contribute people even when they cannot contribute frontier models. AI 2040 allocates only 1% of its safety budget to studying model character, meaning persistent behavioural tendencies or "personas+propensities." It says almost nothing about activation steering, which changes behaviour by intervening on a model's internal activity, or assistance games, which train an AI to infer what a person wants through cooperation. It does call truth-seeking AI and public forecasting a priority. Australia and New Zealand could specialise in those areas, as well as evaluations, treaty verification, biosecurity, and cyber work through the Five Eyes intelligence alliance. (Alignment roadmap, AI for epistemics)

+

3. Ask for terms before signing (2029)

+

Australia is named among the countries that join the deal. New Zealand is not mentioned in any of the 18 supplements we checked. Once inside, leaving could mean sanctions, cyberattack, or war. The realistic decision is therefore what to ask for before signing: hosted data centres, a strategic chip reserve, verification seats, equity in AI companies, and durable access rather than a revocable API account.

+

Scott Alexander's summary is unusually blunt. Observer countries follow the rules in return for data centres, AI access, and a share of future wealth, even though the great powers don't strictly need to offer it. That makes accession our best bargaining moment. Signing the standard form gives it away. (Introducing Plan A)

+

4. Host allied compute (2030-32)

+

AI 2040 moves advanced chip fabs and robot production into special economic zones that can be destroyed if the deal breaks. The zones are meant to sit near an ocean or an adversarial power, and the robot economy expands inside them until it is as large as today's human economy. North-west Australia fits those criteria and sits above much of the ore the machines would consume. (Deal decline, economics)

+

Hosting creates exposure as well as income. The same cluster can be a target, collateral for the treaty, and leverage against being cut off. Tom Davidson makes the mechanism explicit: if the US withdraws AI access, allies could destroy US data centres on their territory. A kill switch negotiated at accession is leverage. A foreign cluster accepted without access rights is just exposure. (Davidson 2026)

+

There is also a smaller trusted-vault option. Australia could host cold-storage model weights, treaty verification equipment, or monitoring linked to Pine Gap without hosting the whole robot economy. Geologically stable ground and Five Eyes trust make this the lowest-volume version of the policy.

+

5. Keep a public stake in the machine economy (2031-34)

+

The US dividend in AI 2040 comes from compute permits. By 2035, the scenario sends 75% of the US permit share to US citizens. Australia and New Zealand cannot rely on that. Our dividend would need resource royalties, land taxes, and public equity held through the Future Fund and NZ Super. Those rules need to exist before cognitive labour becomes cheap and wage earners lose their bargaining power. (Economics of Plan A)

+

Who owns the rents matters more than the size of the boom. The same ore and land can support a public dividend or a few mining dynasties. Leicht and Ball warn that selling a strategic firm can look like an unexpected windfall while giving up the country's "ticket guaranteeing their home country's stake in the AI economy." In a rapid takeoff, taking only cash is especially bad: the useful payment is equity, access, or ownership of productive assets. (The Race Worth Winning)

+

Australia probably starts above the generic non-US country. Compulsory superannuation gives many households an existing claim on global capital, while mining gives governments a claim on inputs the robot economy needs. Neither is evenly shared by default. Superannuation gains follow existing balances and investment choices. Mining income reaches households only when royalties, taxes, public ownership, or transfers capture it. That is what separates the Dividend commonwealth from the Quarry economy.

+

The political claim is stronger, and more speculative. If governments no longer depend on wages or citizen labour, owners and automated security may stop needing broad consent. A legally entrenched citizen claim on public assets keeps ordinary people tied to the revenue stream and gives them something material to defend through democratic institutions. We are treating distribution as a constraint on domestic power, not merely as welfare.

+

6. Choose a side if the deal collapses (2030s)

+

If the US-China deal declines, Australia's security guarantor and largest customer pull in opposite directions. Joining the US bloc keeps the alliance and loses trade. A concert with Japan, Korea, Canada, the Netherlands, and the UK might bargain for access, but only if it was built before the crisis. Leicht and Ball's warning is that a coalition improvised late will be picked off "one by one." Trying to hedge returns us to someone else's empire by a slower route. Free Chinese open-weight models don't remove the dependency; a supplier earning nothing from us also loses nothing by cutting us off. (The Race Worth Winning, Davidson 2026)

+

Where this could leave us*

+

Doom and concentrated global power account for most of our probability. Australia and New Zealand only get the later choices if humanity keeps control and no single actor takes it all.

+

I. Doom, about 50%

+

AI 2040's own comparison gives Plan A a 42% chance of a "great future," falling to 10-25% for its alternatives. Combining the table's conditional rows implies misaligned takeover at roughly one quarter under Plan A and roughly three quarters in the modal race. Our 50% is a rough mixture across those possibilities.

+

Yudkowsky and Soares state the doom case much more strongly than our estimate:

+
+

"If any company or group, anywhere on the planet, builds an artificial superintelligence using anything remotely like current techniques, based on anything remotely like the present understanding of AI, then everyone, everywhere on Earth, will die."

+

Eliezer Yudkowsky and Nate Soares, If Anyone Builds It, Everyone Dies

+
+

Once control is lost, Australia has no policy response left. The preparations above have to improve the odds before then.

+

II. Someone else's empire, about 25%

+

The AI remains aligned, but control does not spread. A government, company, or person holds the decisive systems. Australia and New Zealand may be materially rich, especially while iron demand is high, but they live under rules set elsewhere.

+
+

"a tiny group of people, or possibly just a single individual, is effectively in control of the world’s only army of superintelligences"

+

AI Futures Project, AI 2040

+
+

III. Dividend commonwealth

+

We negotiated terms, captured resource and land rents, and built the dividend before wages collapsed. Households share in the assets that became valuable instead of watching those assets reprice around them. The better version also uses AI inside courts, health systems, and government rather than treating it only as something to regulate.

+
+

"In Plan A, most of the US permit revenue share (75% in 2035) is redistributed to the US population as a Citizen’s Dividend, resulting in roughly $1M/yr per person in 2035 and around $10M/yr in 2040."

+

AI Futures Project, Economics of Plan A

+
+

IV. Quarry economy

+

This is the drift outcome. The country exports what the robot economy needs, but has no dividend rail and no public stake. AI 2040's land model makes the distribution problem concrete: single-family homes rise about 40 times while 97% of their paper value becomes land. Incumbent owners gain while cognitive wages fall. The nation is rich and the median household is poor, without anyone needing to plan it.

+
+

"As AI devalues human capital relative to physical and intangible assets, capital income consumes a growing proportion of national income. Elevated spending on luxury goods and a productivity-driven investment boom sustains elevated economic growth even as millions face unemployment and diminished purchasing power."

+

UK Government Office for Science and AI Security Institute, AI Scenarios 2030

+
+

The resource boom also expires. On current reserves and extraction, iron is the shortest-dated major reserve at roughly 55 years. A machine economy growing every 6-12 months would bring exhaustion and substitution forward. AI 2040 models the demand for energy and land, but not the Australian terms of trade or the cost of scaling mines.

+

V. Allied compute host

+

Australia accepts the data centres and robot zones, but only after securing access and ownership terms. Land, power, and Five Eyes trust become a stake in the machine economy. The price is physical exposure: the hosted clusters are treaty collateral, and the treaty's wars become ours.

+
+

"If the US withdraws AI access, allies could destroy US data centres in response. It's a way to lock in the deal."

+

Tom Davidson, How can the middle powers avoid getting trounced?

+
+

VI. Garrison ally

+

The deal collapses and Australia chooses the US bloc. We keep the security relationship and lose much of the Chinese market. Iron becomes a strategic allocation rather than an ordinary export. This future is poorer and tenser than the dividend or hosting alternatives, but Australia still has a seat.

+
+

"middle powers should help the US, and make sure they are rewarded with continued access to frontier AI and new technologies (including military tech)"

+

"The only alternative that makes sense to me is siding with China."

+

Tom Davidson, How can the middle powers avoid getting trounced?

+
+

VII. Boring decade, about 10%

+

AI stalls or governments shut the frontier down. The consents are unused, the royalty argument was early, and the extra biosecurity capacity looks like insurance that did not pay. This is the cheapest way for our preparations to be wrong.

+
+

"AI progress slows, and AI causes less disruption than expected."

+

UK Government Office for Science and AI Security Institute, AI Scenarios 2030

+
+

VIII. Middle-power concert

+

The deal declines, but a coalition formed before the crisis holds together. Australia and New Zealand pool their supply-chain and hosting leverage with Japan, Korea, Canada, the Netherlands, and the UK. None could demand durable access alone. Together they may get terms and help write the rules.

+
+

"Coordination, done right, reassures partners that they can focus on their comparative advantage instead of pursuing full autarky; and it enables powers to see eye to eye with AI exporters instead of being picked off by great powers offering rewards for defection one by one."

+

Leicht and Ball, The Race Worth Winning

+
+

IX. Homegrown autocracy free beer and rugby for life

+

This is the domestic version of the global power-grab risk. Cognitive wages collapse, public ownership never develops, and automated security removes the state's remaining practical dependence on citizens. No coup is required. Political leverage drains away with economic leverage.

+
+

Luckily we completely dodged a coup and now have Benevolent Dictator for Life (BDFL) Gina Palmer to lead us through. Everyone is very happy, especially since the free beer, rugby, and netball are a great respite from our quickly shrinking population and ever more confusing global situation. - Joe Citizen, 2041

+
+
+

"AI could create a similar effect: if governments can generate massive revenue from taxing AI projects rather than citizens, heads of state may lose their economic incentive to ensure citizens prosper. This would weaken citizens’ power to resist coup and backsliding attempts."

+

"By replacing government employees with loyal AI systems, a head of state could remove important checks on their power."

+

Davidson, Finnveden, and Hadshar, AI-Enabled Coups

+
+

Where we depart from AI 2040

+

AI 2040 models energy consumption rising about eight times by 2040 and notes that proven mineral reserves cover decades of current extraction, but it does not price the tonnage needed for a robot economy growing every 6-12 months. Coal, uranium, energy trade, and Australian mine capacity are outside its model. (Economics of Plan A)

+

We use its land estimates with the authors' warning that this is a "very rough model" they do not necessarily endorse. We also give model character and activation steering more weight than their roadmap does.

+

AI 2040 does share some US wealth abroad, but very unevenly. In 2032 its American dividend starts at $45,000 while adults outside the US and China average $1,200. By 2035 the figures are about $1M and $10,000. Australia may do better than that generic non-US payment through superannuation and mining, but only if the gains are broadly owned. A rich national balance sheet can still leave the median household without economic leverage. (AI 2040)

+

We leave out a trans-Tasman split and a separate Australian decision during US sabotage operations because neither changes the main choices above.

+

Sources and thanks

+

This is built from the AI Futures Project's AI 2040 scenario and supplements. Their caveat applies to our use of the numbers too: the economic model "plausibly contains bugs" and they do not trust its outputs to be accurate. The local source mirrors retain exact line references for auditing.

+

It also benefited from a June 2026 scoping discussion with authors of related scenarios and Australian policy participants. Those workshop notes are private and are not included in the public source archive.

+

* The Australian and New Zealand decisions and futures are extrapolations from the cited work. A citation means we used that source, not that its authors endorse this scenario.

+

The middle-power framing also draws on AI 2027, Europe 2031, Tom Davidson's middle-power plan, Leicht and Ball's The Race Worth Winning, and Scott Alexander's Introducing Plan A. Australian groundwork came from the e61 Institute, Tech Policy Design Institute, ASPI, and Kate Chaney MP.

+ + + + diff --git a/justfile b/justfile new file mode 100644 index 0000000..720d121 --- /dev/null +++ b/justfile @@ -0,0 +1,17 @@ +# ANZ 2040 site build. `just` with no target builds everything. +default: build + +# build the whole site (tree first because HTML inlines it) +build: tree html + +# index.html from anz-2040.md: prose via pandoc, with tree.svg inlined +html: + python3 scripts/build_html.py + +# tree.svg: the probability-sized decision tree, from data/tree.json +tree: + python3 scripts/build_tree.py + +# serve locally so relative links work (localhost:8080) +serve: build + uvx python -m http.server 8080 diff --git a/scripts/build_html.py b/scripts/build_html.py new file mode 100644 index 0000000..956a00d --- /dev/null +++ b/scripts/build_html.py @@ -0,0 +1,76 @@ +#!/usr/bin/env python3 +"""Build index.html from anz-2040.md. One static file, no build step on GitHub's side +and no runtime JavaScript. Prose goes through pandoc (gfm, so it matches GitHub's own +rendering and keeps quotes verbatim); the linked tree.svg image is replaced by the same +SVG inlined into the page, so its links to +each section and tooltips work. CSS is hand-rolled tufte, borrowing ai-2040's +palette (global.css is a Tailwind *source* file and can't be served as-is). +-- claude (opus) + wassname, 2026-07""" +import re, subprocess, pathlib + +ROOT = pathlib.Path(__file__).resolve().parent.parent +md = (ROOT / "anz-2040.md").read_text() + +# Claude: Strip internal credit and local-evidence notes from the public page. +md = re.sub(r"<!--.*?-->", "", md, flags=re.DOTALL) + +# Claude: Inline the linked SVG so its links work in the standalone page. +tree_link = "[![ANZ 2040: choices and possible futures](tree.svg)](tree.svg)" +assert md.count(tree_link) == 1, f"expected exactly one tree link, found {md.count(tree_link)}" +md = md.replace(tree_link, "@@TREE@@") + +body = subprocess.run( + ["pandoc", "-f", "gfm", "-t", "html", "--wrap=none"], + input=md, capture_output=True, text=True, check=True, +).stdout +svg = (ROOT / "tree.svg").read_text() +# Claude: Crash when an SVG link no longer matches a prose heading. +heading_ids = set(re.findall(r'<h[23] id="([^"]+)"', body)) +svg_targets = set(re.findall(r'href="#([^"]+)"', svg)) +missing_targets = svg_targets - heading_ids +assert not missing_targets, f"tree.svg links to missing headings: {sorted(missing_targets)}" +body = body.replace("<p>@@TREE@@</p>", f'<figure class="tree">{svg}</figure>') + +CSS = """ +:root { --bg:#fffff8; --fg:#111; --muted:#666; --rule:#ccc; --accent:#2A623D; } +html { font-size: 17px; } +body { max-width: 42em; margin: 4rem auto; padding: 0 1.25rem; + background: var(--bg); color: var(--fg); + font-family: Georgia, 'Times New Roman', serif; line-height: 1.55; } +h1,h2,h3 { font-weight: normal; line-height: 1.2; margin: 2.2rem 0 0.6rem; } +h1 { font-size: 2rem; } h2 { font-size: 1.5rem; border-bottom: 1px solid var(--rule); padding-bottom: .2rem; } +h3 { font-size: 1.2rem; color: #333; } +a { color: var(--accent); text-decoration: none; } a:hover { text-decoration: underline; } +blockquote { margin: 1rem 0 1rem 1.2rem; padding-left: 1rem; border-left: 3px solid var(--rule); + color: #333; font-size: .95rem; } +table { border-collapse: collapse; margin: 1.2rem 0; font-size: .92rem; } +th,td { text-align: left; padding: .35rem .8rem; border-bottom: 1px solid var(--rule); } +thead th { border-bottom: 2px solid #999; } +code { font-family: ui-monospace, Menlo, Consolas, monospace; font-size: .88em; } +/* the tree is a tall, narrow two-per-row flow; cap its width and centre it so the text + stays a legible size (a wide canvas stretched to the column would shrink the text) */ +.tree { margin: 1.5rem auto; max-width: 560px; overflow-x: auto; } +.tree svg { width: 100%; height: auto; display: block; } +hr { border: none; border-top: 1px solid var(--rule); margin: 2.5rem 0; } +footer { margin-top: 3rem; padding-top: 1rem; border-top: 1px solid var(--rule); + color: var(--muted); font-size: .85rem; } +""" + +HTML = f"""<!doctype html> +<html lang="en-AU"> +<head> +<meta charset="utf-8"> +<meta name="viewport" content="width=device-width, initial-scale=1"> +<title>ANZ 2040: the decisions we still get to make + + + +{body} + + + +""" +(ROOT / "index.html").write_text(HTML) +print(f"wrote index.html ({len(HTML)} bytes), inlined tree.svg ({len(svg)} chars)") diff --git a/scripts/build_tree.py b/scripts/build_tree.py new file mode 100644 index 0000000..92623f4 --- /dev/null +++ b/scripts/build_tree.py @@ -0,0 +1,104 @@ +#!/usr/bin/env python3 +"""Build tree.svg from data/tree.json. No dependencies beyond stdlib, no JS in the +output. Layout is automatic from each item's `tier` (row) and `order` (left-to-right), +so there are no hand-tuned coordinates to drift. Box width fits its title and subtitle; +probability mass is shown by edge width and explicit percentage labels. Every item is an into its prose section and carries a + tooltip. The viewBox makes it scale to any screen width (phones included). +-- claude (opus) + wassname, 2026-07""" +import json, math, html, pathlib + +ROOT = pathlib.Path(__file__).resolve().parent.parent +d = json.loads((ROOT / "data" / "tree.json").read_text()) +KINDS = d["kinds"] +CV = d["canvas"] +W = CV["w"] +FLOOR, PSCALE, NH, GAP = CV["floor_w"], CV["p_scale"], CV["node_h"], CV["gap"] +CHOICE_W, ROW_H, TOP = CV["choice_w"], CV["row_h"], CV["top"] +TITLE = d.get("title", "") +TITLE_H = 82 if TITLE else 0 # hero heading band at the top (ai-2040 "Choose a Path" style) + +def width(n): + # box fits its text (probability is carried honestly by the flow-weighted edges and + # the explicit % labels, not by node size, which can't be honest at variable text length) + fit = max(len(n["title"]) * 7.3, len(n.get("sub", "")) * 5.7) + 26 + return max(fit, CHOICE_W) + +# ---- layout: place each tier's nodes left-to-right, centred on the canvas ---- +tiers = {} +for n in d["nodes"]: + tiers.setdefault(n["tier"], []).append(n) +pos = {} # id -> (cx, cy, w) +for t, row in tiers.items(): + row.sort(key=lambda n: n["order"]) + widths = [width(n) for n in row] + total = sum(widths) + GAP * (len(row) - 1) + x = (W - total) / 2 + cy = TOP + TITLE_H + t * ROW_H + for n, w in zip(row, widths): + pos[n["id"]] = (x + w / 2, cy, w) + x += w + GAP + +H = TOP + TITLE_H + (max(tiers) + 1) * ROW_H +out = [f'<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 {W} {H:.0f}" ' + f'font-family="Georgia, serif" font-size="14">', + f'<rect width="{W}" height="{H:.0f}" fill="#fffff8"/>'] +if TITLE: + out.append(f'<text x="{W/2:.0f}" y="{TOP+42:.0f}" text-anchor="middle" font-size="28" ' + f'font-weight="bold" fill="#111">{html.escape(TITLE)}</text>') + +def edge_pts(a, b): + ax, ay, aw = pos[a]; bx, by, bw = pos[b] + if by > ay: # child below: leave bottom, enter top + return ax, ay + NH / 2, bx, by - NH / 2 + if by < ay: # child above (back-edge): leave top, enter bottom + return ax, ay - NH / 2, bx, by + NH / 2 + # same row: side to side + return (ax + aw / 2, ay, bx - bw / 2, by) if bx > ax else (ax - aw / 2, ay, bx + bw / 2, by) + +FLOWSCALE = 26 # edge thickness = flow (probability mass through the edge) * this +edge_labels = [] +for e in d["edges"]: + x1, y1, x2, y2 = edge_pts(e["from"], e["to"]) + my = (y1 + y2) / 2 + sw = max(1.0, e.get("flow", 0.02) * FLOWSCALE) # thick where the mass pours, thin in the trickle + dash = ' stroke-dasharray="5 5"' if e.get("dashed") else "" + out.append(f'<path d="M{x1:.0f},{y1:.0f} C{x1:.0f},{my:.0f} {x2:.0f},{my:.0f} {x2:.0f},{y2:.0f}" ' + f'fill="none" stroke="#a9a99f" stroke-width="{sw:.1f}"{dash}/>') + if e.get("label") and not e.get("dashed"): # dashed (secondary) edges stay unlabelled + # place the label near the source (in the gap just below the parent) so labels on + # tier-skipping edges don't land on top of the boxes in the row they cross + # adjacent edges: label at the midpoint (siblings have spread apart there, so they + # don't overlap). tier-skipping edges: push the label near the source, into the gap + # below the parent, so it doesn't land on the boxes in the row it crosses. + lf = 0.30 if abs(y2 - y1) > ROW_H else 0.5 + edge_labels.append((x1 + lf * (x2 - x1), y1 + lf * (y2 - y1) - 2, e["label"])) +# draw labels last, each on a background halo so no stroke cuts through the text +for lx, ly, lab in edge_labels: + lw = len(lab) * 6.6 + 10 + out.append(f'<rect x="{lx-lw/2:.0f}" y="{ly-12:.0f}" width="{lw:.0f}" height="17" fill="#fffff8"/>') + out.append(f'<text x="{lx:.0f}" y="{ly:.0f}" text-anchor="middle" fill="#555" font-size="12.5" ' + f'font-style="italic">{html.escape(lab)}</text>') + +for n in d["nodes"]: + cx, cy, w = pos[n["id"]] + k = KINDS[n["kind"]] + x, y = cx - w / 2, cy - NH / 2 + rx = 22 if n["kind"] == "outcome" else 4 # pill endings, square choices/risks + sw = 3.6 if n["kind"] == "hazard" else 2.2 # thicker borders so the category colour reads; heaviest on risks + tip = html.escape(f'{n["title"]} — {n["sub"]}' if n.get("sub") else n["title"]) + out.append(f'<a href="{n["href"]}"><title>{tip}') + # ai-2040 style: box fill matches the page background so only the category-coloured + # border shows (pure white would read as a patch on the cream page) + out.append(f'') + ty = cy + (4 if not n.get("sub") else -3) + out.append(f'{html.escape(n["title"])}') + if n.get("sub"): + out.append(f'{html.escape(n["sub"])}') + out.append('') + +out.append('') +(ROOT / "tree.svg").write_text("\n".join(out)) +print(f"wrote tree.svg: {len(d['nodes'])} nodes, {len(d['edges'])} edges, canvas {W}x{H:.0f}") diff --git a/sources/ai2040-pages/ai-for-epistemics.md b/sources/ai2040-pages/ai-for-epistemics.md new file mode 100644 index 0000000..ba963e3 --- /dev/null +++ b/sources/ai2040-pages/ai-for-epistemics.md @@ -0,0 +1,290 @@ +--- +url: https://ai-2040.com/supplements/ai-for-epistemics +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# AI for Epistemics + +### Eli Lifland + +### Summary + +Humanity’s _epistemics_ , by which I mean our ability to come to true beliefs and thus make sensible decisions, is crucially important for achieving a great future. As AI capabilities improve throughout Plan A, they shape increasingly more of every facet of life, and epistemics is no different. I expect how AIs affect epistemics to be highly important, and it could easily be positive or negative. + +Basin of sanity: a self-reinforcing equilibrium during AI takeoffVICIOUS LOOPpoor epistemics→ sycophantic or agenda-pushing AIs adopted→ poor epistemics deepenVIRTUOUS LOOPgood epistemics→ epistemic AI tools adopted→ epistemics improve furtherSociety's epistemicsentering AI takeoffBASIN OFCAPTUREBASIN OFSANITY← worse societal epistemicsbetter societal epistemics → + +I expect there to be positive feedback loops around AI for epistemics, and thus our goal should be to reach the _basin of sanity_. If we are in this basin, it’s self-reinforcing: society is sane enough to make itself more sane as AIs continue to improve. A top priority of the US government during AI takeoff should be to stay in this basin. + +How AI for epistemics interventions flow through to impact + +Intervention categories + +Collective epistemics + +Individual epistemics + +Encourage adoption + +Key indicators + +Org culture + +(AGI projects, govt) + +Truth-seeking AI + +(epistemic virtue) + +Transparency to the public + +Labor on key decisions + +Public discourse quality + +Areas of impact + +Government decision-making + +AGI project decision-making + +Public influence on govt & AGI projects + +Public investment into making AGI go well + +I think that the most important areas of impact of AI to epistemics include the decision-making of the government and AGI projects and the public influence over these actors ([more](/supplements/ai-for-epistemics#areas-of-impact-of-ai-for-epistemics)). In Plan A, an especially important application of AI for epistemics is [preventing deal decline](/supplements/deal-decline#preventing-deal-decline), i.e. improving the stability and effectiveness of the slowdown deal. + +[Key indicators](/supplements/ai-for-epistemics#indicators-to-track-societys-epistemics) to track society’s epistemics include: + + * Organizational culture of frontier AI projects and governments. + + * Truth-seeking AI (or more broadly, AIs’ epistemic virtue). + + * How transparent each AI project and government is to the public. + + * Human and AI labor spent on key decisions. + + * Public discourse quality. + + + + +I suggest the following as interventions to positively shape AI’s epistemic impact. ([more](/supplements/ai-for-epistemics#interventions)) + +AI for epistemics interventions + +Collective epistemics + +Help society converge toward truth + +Epistemic virtue evals + +Measure & incentivize AI truth-seeking + +Allow people to prove they're telling the truth + +Privacy-preserving auditing and automated lie detection + +AI in social & traditional media + +Community notes, counter-arguments, curated feeds + +Track records + +Score everyone's promises and predictions; surface in context + +Shared epistemic infrastructure + +Shared knowledge bases, prediction tracking, argument trees + +Persuasion defenses + +Limit AI persuasion capabilities; detect & flag attempts + +Individual epistemics + +Help individuals figure out the truth + +Research assistants & forecasters + +Calibrated, truth-seeking; not sycophantic + +Scenario planning & exploration + +On-tap deep planning; gamified what-if simulators + +Navigate emotional blockers + +AIs that help enable good reasoning despite emotional factors + +### Background on AI for epistemics and the “basin of sanity” + +How well the development of superintelligence goes for humanity is heavily influenced by the extent to which we are able to come to make reasonable decisions regarding how to navigate the AI takeoff period. In Plan A, we are already in a better position than the default world simply by virtue of having more time to orient ourselves to highly capable AIs. But still, humanity’s _epistemics_ , by which we mean our ability to come to true beliefs and thus make sensible decisions, is crucially important for achieving a great future. + +As AI capabilities improve throughout Plan A, they shape increasingly more of every facet of life, and epistemics is no different. I expect AI’s effect on epistemics to be highly important, potentially to a similar or even greater level as the automation of technical AI safety research. + +I think that highly capable AIs could easily end up either improving or worsening epistemics, on the whole. It’s plausible, given AIs’ capabilities, that they will end up pushing epistemics to either a positive or negative extreme greater than that we’ve observed throughout history thus far. + +Furthermore, I expect there to be positive feedback loops that maintain or push further toward the positive and negative extremes. For example, if societal epistemics are good and people have true beliefs, it’s more likely that AI tools which improve epistemics are recognized as such and adopted. On the other hand, if societal epistemics are bad and people are in extreme filter bubbles where AIs tell them exactly what they want to hear, they may continue to adopt sycophantic AI tools as AIs get more capable. The presence of these feedback loops increases the urgency of steering AIs’ effect on epistemics. + +Because of these feedback loops, one way to think about achieving the goal of navigating AI takeoff with good epistemics is that we want to be in the _basin of sanity_. If we are in this basin, it’s self-reinforcing: society is sane enough to make itself more sane as AIs continue to improve. A top priority of the US government during AI takeoff should be to stay in this basin. + +### Areas of impact of AI for epistemics + +I expect sanity in different areas to be correlated, at least in the long run, due to having similar drivers such as “do frontier AIs have and promote good epistemics” (more on this below). Still, it’s useful to discuss different concrete impacts of epistemics, especially early in takeoff when there may be differences in tooling and adoption in different areas. I think that the most important areas to pay attention to epistemics’ impact are: + + 1. **Government decision-making,** including to what extent government is captured by AGI projects. + + 2. **AGI project decision-making** + + 3. **Public influence on governments and AGI projects.** Including via voting. + + 4. **Public investment into directly making AGI go well, e.g. via preventing vs. creating large-scale harms or steering toward a great future conditional on avoiding catastrophe.** For example, how much philanthropic investment there is in biodefense and AI safety, vs. how many ideological fanatics there are attempting terrorism. Or how much investment there is into ethical reflection. + + + + +For all of the above, importance is greatly increased for decisions/beliefs on AGI-relevant topics. And it’s important that for all of them society has the ability to orient and respond to rapid AI progress. + +One related distinction is to what extent interventions affect those with below-average epistemics (i.e. _floor-raising_), those with typical epistemics (i.e. _average-raising)_ , or those with above-average epistemics (i.e. _ceiling-raising_). + +### Indicators to track society’s epistemics + +I think that the following are the most important indicators to keep track of in order to tell whether we are in the basin of sanity, especially for the important areas identified above. + +**Organizational culture of frontier AI projects and governments.** How conducive the culture of frontier AI projects and governments are to making good decisions. + +**Truth-seeking AI (or more broadly, AIs’ epistemic virtue).** To what extent frontier AIs are trying to figure out the truth and report it, as opposed to other goals such as advancing a hidden agenda or producing sycophantic outputs that look good but aren’t true. + +**How transparent each AI project and government is to the public.** + +**Human and AI labor spent on key decisions.** On the most important decisions, many of which are made by governments and AGI projects, how much labor goes into making them? Are governments and AGI projects good at allocating labor across decisions? Note that the _quality_ of the labor is also hugely important, which is why organizational culture and truth-seeking AI are key metrics. + +**Public discourse quality.** Whether the public discourse becomes more reasonable due to tools like automated fact checking, or less reasonable due to things like increased polarization from persuasive sycophantic AIs, seems highly important. One important portion of this is how difficult it is for actors to lie or manipulate people. In addition things like fact checking, automated lie detection or privacy-preserving ways to prove claims could be important (discussed more below). + +### Interventions + +Below I non-exhaustively list what seem like some of the most important levers for AIs improving epistemics, many of which governments could encourage. I loosely divide interventions whether they focus more on the collective societal epistemics, individual epistemics, or encouraging adoption, but in practice many interventions help with multiple of these categories. + +This section has lots of overlap with this [Forethought work](https://www.forethought.org/research/design-sketches-for-a-more-sensible-world) which sketches out ideas for epistemic and coordination tools; we’ve linked to specific sketches as appropriate. + +#### Collective epistemics: help society converge toward the truth + +**Create evaluations of AIs’ epistemic virtue and incentivize AGI projects to make their AIs perform well on these evaluations.** (see also [Forethought writeup](https://www.forethought.org/research/design-sketches-collective-epistemics#epistemic-virtue-evals)) The first step to improving AIs’ epistemic virtue is measuring it. For example, one could measure to what extent AIs figure out what’s true rather than giving answers users want to hear, or how calibrated AIs’ quantitative forecasts are. Once the evaluations exist, it’s important that AGI projects actually care about doing well on them. There will be a place for both quantitative, easily runnable evaluations and more open-ended, human-in-the-loop evaluations. In Plan A, we recommend that the government pushes for AGI projects to do well on epistemic evaluations as part of decisions regarding whether they are able to scale their capabilities, similar to how they would use alignment and control evaluations. We also recommend members of the public to create scorecards for AIs’ epistemic virtue and for consumers to take this into account when deciding what AIs to use, as well as for AGI project employees to take these into account when deciding which company to work at. + +**Allow for politicians and other public figures to prove that they’re telling the truth, via privacy-preserving auditing and potentially automated lie detection.** Privacy-preserving auditing might look something like: People can ask you questions, and if you’d like, you can have an AI read through your private data, and report back an answer that preserves your privacy as much as possible while still answering the question. That AI wouldn’t write the data to any sort of memory, so the data would still only be owned by you. You could refuse to have the AI answer any given question, but people could then judge you appropriately. We recommend that norms are established such that politicians are expected to submit to privacy-preserving auditing. (See also [“Confidential monitoring and verification” Forethought writeup](https://www.forethought.org/research/design-sketches-defense-favoured-coordination-tech#confidential-monitoring-and-verification).) A more flexible technology would be if AI enabled automated, highly accurate lie detection. This would be dual use, for example it could be used by those in power to purge dissidents. However, our best guess is that it’s still good to develop, especially in the environment of Plan A where the world is set up well to ensure it’s used as a check on those in power. Additionally, it would dramatically increase ability to detect [covert projects](/supplements/covert-ai-projects) and significantly improve [deal stability](/supplements/deal-decline#preventing-deal-decline). + +**Use AI in social and traditional media to improve discourse and keep people informed.** AIs will heavily feature in both social and traditional media, whether via new entrants or existing sites/outlets adopting AI. Influencing how AI is integrated into the media is very important. For example, there could be AI-generated [“community notes”](https://en.wikipedia.org/wiki/Community_Notes) in cases where a post is false or misleading (see also [Forethought writeup](https://www.forethought.org/research/design-sketches-collective-epistemics#community-notes-for-everything)). AIs could also automatically generate counter-arguments to each post. AIs could help reduce echo chambers by steering toward interactions between people who would disagree but have a productive conversation. AIs could curate feeds to give people the information that is most important for them to see and show content that most productively challenges their views (see also [Automated deep briefings](https://www.forethought.org/research/design-sketches-angels-on-the-shoulder#deep-briefing) and [Automated OSINT](https://www.forethought.org/research/design-sketches-tools-for-strategic-awareness#automated-osint) Forethought writeups). Curation more broadly could greatly improve how informed people are with respect to politics and the news, for example by highlighting policy proposals that the reader might have a strong opinion on. (See also [Aligned recommender systems](https://www.forethought.org/research/design-sketches-angels-on-the-shoulder#aligned-recommender-systems) and [Personalized leaning systems](https://www.forethought.org/research/design-sketches-angels-on-the-shoulder#personalised-learning-systems) Forethought writeups). + +**Collect and score everyone’s track records, including politicians, and surface track records when relevant.** (see also [Forethought writeup](https://www.forethought.org/research/design-sketches-collective-epistemics#reliability-tracking)) AIs will have the capacity to greatly increase the amount of effort put into collecting people’s past statements, and scoring them. For example, an AI could grade how many of their promises a politician has kept, or what percentages of their predictions have come true, or how productive they have been at lawmaking. There will be variation in how objectively statements can be scored; for the less objective ones, hopefully there would be a trusted AI that a broad array of people listen to; but if not, only scoring the more objective statements could still result in positives. + +**Shared epistemic infrastructure, for example knowledge bases.** As AIs improve they will be able to maintain vast amounts of epistemic infrastructure. Wikipedia is an obvious reference point, though for AIs the most promising infrastructure may be more structured knowledge bases (at least at first), and the infrastructure might focus more on being parsable by AIs than humans. One could imagine infrastructure for tracking quantitative predictions and how they have fared, argument trees regarding various claims, etc. Building infrastructure that is trusted by a wide range of actors is important, as opposed to fracturing into alternate realities. (See also this post on building a [full epistemic stack](https://www.oliversourbut.net/p/a-full-epistemic-stack). This is related to and would help with [”Provenance tracing” as described by Forethought](https://www.forethought.org/research/design-sketches-collective-epistemics#provenance-tracing).) + +**Defend against persuasion via reducing AIs’ persuasion abilities and using AIs to detect it.** Let’s define persuasion as the use of techniques to convince people of beliefs that work roughly as well when a belief is false as when a belief is true. By default, AIs may eventually get very superhuman at persuasion, but it’s plausible that their persuasion capabilities could be reduced without hurting beneficial capabilities too much. Another line of defense is to have AIs detect persuasion and help people avoid being exposed to the strongest forms of it, or at least to be aware of it (see also [”](https://www.forethought.org/research/design-sketches-collective-epistemics#rhetoric-highlighting)[Rhetoric](https://www.forethought.org/research/design-sketches-collective-epistemics#rhetoric-highlighting) [highlighting” Forethought writeup](https://www.forethought.org/research/design-sketches-collective-epistemics#rhetoric-highlighting)). In Plan A, we recommend that the US government makes sure to evaluate, limit, and require defenses against persuasion capabilities in collaboration with the rest of the Consortium. + +**Know whether you’re talking to an AI or human in online discussions.** While we think AI can pose a lot of benefit for epistemics, it’s still valuable to know whether you’re interacting with a human or not. This is especially true if there are bots pushing various agendas all over the internet. + +#### Individual epistemics: Help individual actors figure out the truth + +**Automated research assistants and forecasters.** AIs could enable a broad range of actors to have access to high-quality, labor intensive research and forecasts. It’s important that these be as truth-seeking as possible, rather than sycophantic. Forecasts should be ensured to be calibrated, i.e. if an AI says something has a 10% chance of happening, that it actually happens 10% of the time. (See also [“Ambient superforecasting” Forethought writeup](https://www.forethought.org/research/design-sketches-tools-for-strategic-awareness#ambient-superforecasting).) + +**Help people overcome emotional blockers to good reasoning, rather than preying on them.** Of course, much of people’s poor epistemics comes from emotional blockers rather than dispassionate analytical errors or inadequate information. AIs have the potential to greatly help with emotional blockers, as they reach the capabilities of top human therapists but with greater availability, patience, etc. On the other hand, without intervention AIs might confirm and compound people’s existing biases because this would provide a good short-term experience. (See also [Reflection scaffolding](https://www.forethought.org/research/design-sketches-angels-on-the-shoulder#reflection-scaffolding) and [Guardian angels](https://www.forethought.org/research/design-sketches-angels-on-the-shoulder#guardian-angels) Forethought writeups, which are related but more broad.) + +**Automated scenario planning and exploration**. AIs could help do in-depth scenario planning such as Plan A and AI 2027 but for any decision, and could help extract quantitative empirical predictions. AIs could also help explore a variety of scenarios via video-game-like simulations in which the human plays an actor, such as the CEO of an AGI company during an intelligence explosion, and the AI simulates everyone else’s actions and how the world evolves. (See also [“Scenaio planning on tap” Forethought writeup](https://www.forethought.org/research/design-sketches-tools-for-strategic-awareness#scenario-planning-on-tap).) + +#### Encourage adoption + +**Encourage adoption of epistemic tools such as some of the above.** In Plan A, we recommend that the US government adopt epistemic tools quickly, and encourage AGI projects to do so as well. + +**Generally encourage adoption of AI.** Encouraging adoption of AI in general seems valuable for society having a better understanding of existing AI capabilities and trends. In Plan A, we recommend that the US government adopt AIs quickly. + +#### Other + +Some other ideas that we think are promising but haven’t fleshed out as much as the above: + + 1. All of the defense-favored coordination tech ideas described by Forethought [here](https://www.forethought.org/research/design-sketches-defense-favoured-coordination-tech). + + 2. Figure out how to overhaul the government for the age of superintelligence. Plausibly the Constitution should be radically amended. + + 3. Positively shape the relationships between humans and AIs and they become closer; this one seems important but tricky to get right. + + + + +### Appendix + +#### The promise of working on epistemic tools + + _Why it’s valuable to build or encourage epistemic tools given that better AIs may reduce the need for domain-specific tooling:_ I’m not confident regarding the value of building epistemic tools, and it’s plausible that any given tool gets obsoleted by future AIs, but I’m still optimistic about their value. As AIs improve it will become easier to use them for epistemics out of the box, but scaffolding, UX, and human factors may continue to matter a lot. For example, a tool could incorporate best practices of having AIs debate each other then judging the debate, where an AI out of the box might not. Or a tool might have a UX that has been honed over lots of trial and error with humans, that an AI wouldn’t be able to replicate without iteration. Or a tool may have the technical aspects obsoleted, but the company has built up a base of users that trust its brand. + +#### Decomposition of issues with decision-making + +What are the reasons that people (including me) might make decisions that are bad, as compared to perfect decision-making and my values? + +They could have wrong beliefs when they explicitly reason about the current state of the world, or forecasts. Why? + +They might not have access to good enough information: e.g. they have only read news biased in a certain direction. + +They might not be making sense of their information well enough: + + 1. Explicit reasoning about local incentives (monetary, social, etc.) + + 2. Don’t trust/defer to the right news sources and people + + 3. Emotional blockers + + 4. Poor reasoning + + + + +They could be trying to support a general “memeplex” (social group, culture, ideology) that requires certain types of cutting off good reasoning/noticing things. Memeplexes are helpful for coordination, but can be bad for epistemics. + +They might be making poor decisions based on instinct rather than explicit reasoning. One cause of this is local incentives (monetary, social, etc.) that aren’t specifically reasoned about + +They could have different values (these aren’t issues with decision-making, but rather ethical disagreements): + + 1. They don’t care as much about my selfish goals. + + 2. They care about their selfish goals more than I do. + + 3. They have different altruistic values. + + + + +1. + +See [here](https://www.forethought.org/research/ai-impacts-on-epistemics-the-good-the-bad-and-the-ugly#so-what-should-we-expect-to-happen) for more discussion of potential positive and negative feedback loops in AI for epistemics. + +2. + +See for a prototype of an AI-generated knowledge base. diff --git a/sources/ai2040-pages/alignment-roadmap.md b/sources/ai2040-pages/alignment-roadmap.md new file mode 100644 index 0000000..fbf9347 --- /dev/null +++ b/sources/ai2040-pages/alignment-roadmap.md @@ -0,0 +1,1762 @@ +--- +url: https://ai-2040.com/supplements/alignment-roadmap +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Alignment Roadmap + +### Ryan Greenblatt, Thomas Larsen + +> DISCLAIMER: This supplement is a very rough draft representing in progress thinking. It should be interpreted as internal research notes, not as a final product that we are entirely happy with. This content is dense, full of not-very-well explained jargon, and unpolished. We decided to publish anyway in the spirit of transparency. +> +> We recommend reading and engaging with other supplements we’ve written before this one, unless you are particularly interested in alignment content. Furthermore, this content is almost entirely focused on Plan C. If you are interested in how alignment is solved in Plan A, we recommend reading the [insider’s perspective](https://www.ai-2040.com/?choices=plan-a-root#playbook-insider-pov). + +# Summary + +This document represents our best guess alignment roadmap as of early 2026. It’s focused on outlining high level objectives that need to be accomplished in order to avoid AI takeover. + +First, we’ll go over the core methodologies for solving alignment. + +Second, we’ll go over the alignment approach in detail for Plan C, which is the case where we have a basically reasonable AI developer working with a small amount of lead time (e.g. a few months), and are willing to take costly actions that have some safety benefit, as long as this doesn’t result in them losing the AI race. This is both a likely world, and a relatively sensible world to talk about, because we can strategize on behalf of a reasonable leading AI developer. + +Our alignment roadmap can be broken down into 4 phases, which we’ll discuss in detail: + + 1. **Prepare.** This will centrally involve preparing for control, building misalignment analogies that are as realistic as possible, and various other interventions. + + 2. **Mitigate + Elicit.** During the early intelligence explosion, mitigate the risk of AIs taking catastrophic actions (e.g. staging a rogue internal deployment or backdoor the values of future models). Meanwhile, elicit their capabilities on domains that we care about (e.g. control research, alignment research, security, strategy) and use them to do huge amounts of useful work in these domains. This phase lasts from the early intelligence explosion (specifically ITAI), until Min-H (definitions here). + + 3. **Pause at Min-H.** Once we’ve built AI systems with general purpose capabilities sufficient that we can unsafely hand off trust (Min-H), we pause further AI capabilities progress, and try to build AIs that we can safely hand off trust to. This will centrally involve: + + 1. **Avoiding alignment faking** : AIs aren't so seriously misaligned that they intentionally fake alignment (e.g., they aren't scheming against us). + + 2. **Solving handoff-viability** : AIs are sufficiently capable and aligned at autonomously completing all the relevant work for managing the intelligence explosion, including high-stakes, difficult, open-ended, and conceptually loaded tasks where we can't check the AI's output effectively (e.g. acausal trade research, [philosophy](https://www.lesswrong.com/posts/EByDsY9S3EDhhfFzC/some-thoughts-on-metaphilosophy#Some_interrelated_ways_of_looking_at_philosophy)). + + 4. **Handoff.** Once we’ve mitigated alignment faking and achieved handoff-viability, then the AI project actually empowers AI systems with the ongoing management of the intelligence explosion. + + 1. The main alternative to hand off is to instead stall for more time. The main way of doing that would be to escalate to Plans A or B, which would (typically) require substantial government involvement, which may or may not be feasible. A key priority during the other phases will often be to attempt to escalate the level of political will. + + 2. After handoff, our ~human level AI systems that we’ve handed off to will need to solve successively more difficult alignment problems so that they can build vastly superhuman AI systems, or they will need to themselves stall for more time. + + + + +Then we’ll discuss how the alignment plan changes based on political will: + + * In Plan D, we have substantially less political will, no ability to pause, and much less overall labor for handling misalignment risk. This means that plan D will accept a much higher risk budget, and focus exclusively on cheap but low assurance techniques. + + * In Plan B, the plan is very similar to Plan C, except you have somewhat more time/effort and can spend slightly more resources on more ambitious/higher assurance proposals. + + * In Plan A, we’ll initially plan similarly to Plan B/C (in order to hedge against the international coordination breaking down), and then once the regime is stable, shifting resources to moonshots, updating to potentially huge amounts of resources. + + + + +Note that this document focuses on overviewing a plan and discussing relative prioritization rather than on discussing the details of the areas of research and implementation that we'll be pursuing. The most important research components are discussed in much more detail in the following docs: + + * **Control:** [The case for ensuring that powerful AIs are controlled.](https://www.lesswrong.com/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled) + + * **Avoiding Scheming:** [Avoiding scheming (and getting a generally better understanding of scheming)](https://docs.google.com/document/d/1PCQKI3kAdE5HeEiIKN16gszu3kjF0bsowroZB7tDuS4/edit?tab=t.0) + + * **Elicitation:** [Making AIs more useful for safety work](https://docs.google.com/document/d/12VkHm_QPnOzkkJOb88L8tTZ0VWeQnVA4HvwcuLyFy0w/edit?usp=sharing) + + * **Handoff:** [Deferring to AIs](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais) + + * **Making deals with AIs** : [Making deals with early schemers](https://www.lesswrong.com/posts/psqkwsKrKHCfkhrQx/making-deals-with-early-schemers) + + +Plan C Assumptions + + 1. **Responsible Leader.** We assume that a responsible AI company is in the lead. Their leadership is willing to take very costly actions in order to mitigate AI takeover risk (e.g. slow down for several months). Despite this, we’ll assume that by default, they’ll have a bias towards proceeding and being willing to take on a large chance of takeover risk to increase their chances of winning. + + 2. **1-12 month lead** over their competitors. Either trailing US or international AI companies will overtake the leader within 1-12 months. + + + + +Though we mostly won’t focus on this scenario, Plan C is also applicable in the case where the USG is sufficiently worried about AI takeover that they are willing to regulate AI companies to mitigate risk as much as possible without losing to China, and without being willing to sabotage China. + +This plan focuses on making effective use of limited resources to spend on safety and doesn't ensure a high degree of safety (as the level of resources we're imagining is insufficient for this). Indeed, arguably the most likely outcome of this plan is misaligned AI takeover. Due to having limited resources, this plan will involve proceeding in some situations where the level of risk is substantial (because exogenous risks are higher) and we won't necessarily gain much confidence about the exact level of risk. + +Note that we expect that something much less reasonable than this plan will happen in practice, so while this is a recommendation (for what to do with limited resources), it isn't a prediction. + +# Research Areas + +## Different areas of research and implementation + +The plan will involve work in many different areas of research and implementation. The key areas are: + + * **Control** (and control-adjacent measures): Researching and implementing control. For the purposes of this post, I'm defining this to be a relatively broad category that includes some things which aren't centrally associated with control but which fit better in this category than in others. When discussing prioritization in this post, I define the category of control such that it doesn't include security work that is roughly as helpful for security against human adversaries. + + * **Misalignment analogies** : Creating and studying AIs analogous to scheming or otherwise misaligned models (sometimes called "[model organisms](https://www.lesswrong.com/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1)"). This area will be mostly focused on scheming. This includes AIs that are well described as [natural schemers](https://www.lesswrong.com/posts/gALJWSfoTkq2ky5os/prospects-for-studying-actual-schemers) (but might have been created through a training process rigged to make scheming more likely), [AIs trained to behave similarly to schemers](https://www.lesswrong.com/posts/v6K3hnq5c9roa5MbS/reducing-risk-from-scheming-by-studying-trained-in-scheming#fnref-YqRrv2NFrtiLoYT6o-2), and [abstract analogies to scheming models](https://www.lesswrong.com/posts/5zsLpcTMtesgF7c8p/two-proposed-projects-on-abstract-analogies-for-scheming). My (low confidence) guess is that natural schemers are 3/6 of the expected value, AIs trained to behave like schemers are 1/6 of the value, and abstract analogies are 2/6 of the value; that said, we're likely to increasingly focus on whichever of these areas ends up seeming most promising (but this isn't very predictable in advance). I'm also including versions of this work focused on demonstrating misalignment risk in this category. I'm including doing technique/interventions testing and/or iteration using these analogies in the category as I expect that most of the work is constructing good analogies and because there will be interaction between analogy construction and technique iteration. + + * **Elicitation/handoff-viability** : Making AIs more useful for safety work and working towards handoff-viability (including things like improving AI epistemics, making AIs do a good job on extremely hard research which is loaded on conceptual/philosophical reasoning, ensuring AIs generalize to being highly aligned with pretty high reliability on all the tasks needed for handoff). I include both eliciting useful work and handoff-viability in the same category as the work will often overlap, though important subsets of the work for each of these won't overlap. Elicitation will include both relatively domain specific capabilities as well as working on avoiding some types of non-worst-case misalignment. Elicitation might also involve more effectively extracting work from clearly misaligned models (e.g. reward-seeking models) including by trying to alter incentives (e.g. proxy-of-reward incentives or making deals with AIs). I discuss handoff-viability in more detail [here](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais#Behavioral_testing_for_capabilities_and__broad__alignment). + + * **Misc science/interventions** : This includes studying interventions that we'd a priori expect to be helpful or which we can test in ways other than via using misalignment analogies and studying properties of AI that seem relevant. This is intentionally an "other" category which includes promising areas of safety work beyond the categories above. This includes: various types of interpretability, studying and improving chain of thought faithfulness (in the non-adversarial case, as the adversarial case is included under control), developing methods which we a priori expect to reduce the chance of scheming (rather than developing the method by studying or iterating on analogies to scheming), making deals with AIs, AI welfare, studying how AIs learn and generalize, various moonshots, improving honesty and eliciting (latent) knowledge, and generally looking for cases where current AIs behave strangely or are misaligned. + + * **Buying time** : Using AI labor to buy time for further safety work. This involves trying to increase the willingness to pay of other actors (centrally the US government and other governments) for lead time, reducing the cost of acquiring lead time for these other actors, and directly acquiring lead time. This could involve assisting the US government with non-proliferation efforts, various types of policy work, helping with verification effort for international coordination, making money to spend on various efforts to buy time, and delivering AI services which allow for various deals to remain stable. It also involves generally trying to convince other actors that they should care more about buying lead time (e.g., assisting with non-proliferation efforts) that can be spent on reducing misalignment risk (and other risks), though much of this work is at least partially included under other categories (in particular, misalignment analogies and demonstrating capabilities). + + * **Evaluating and demonstrating capabilities** : This includes evaluations and demonstrations of capabilities for the purpose of making other actors better understand the situation. Demonstrations of capabilities might be a key part of increasing the willingness to pay of other actors on mitigating risks related to AI. It doesn't include evaluations which are directly relevant for control (e.g., control evaluations and evaluations of whether models are capable enough that they could plausibly be schemers), but does include evaluations aimed at generally understanding the situation and (e.g.) assessing what level of security is appropriate. + + + + +This list excludes some areas which receive less focus in this plan but are still important such as hardening the world against rogue AIs (e.g. biodefence and cyber defence). + +## Core methodologies + +This plan aims to use a relatively iterative and empirical approach, because I think this is probably doable once you have powerful AIs. As part of this, control, misalignment analogies, and elicitation/handoff viability each have core methodologies that form the main empirical feedback loop for that methodology. The methodologies for these are, respectively: + + * **Control evaluations** : Control is [relatively easy to evaluate](https://www.lesswrong.com/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled#Evaluating_whether_you_have_control_is_doable) because it just requires estimating capabilities. This makes iteration easier and control is generally relatively easy to iterate on for reasons we've discussed in prior work. A large fraction of our hope for control will be based on doing a good job iterating using something like control evaluations (though we might be more flexible than a very narrow notion of control evaluations which depends on eliciting capabilities from models, see [here](https://www.lesswrong.com/posts/3s8PtYbo7rCbho4Ev/notes-on-control-evaluations-for-safety-cases)). + + * **Misalignment analogies iteration** : Ultimately, our main hope for avoiding scheming is probably iterating using misalignment analogies. This is a somewhat broader and less well defined category than control evaluations, but I still think it's worth highlighting. Iterating on misalignment analogies is also trickier and less well understood than control evaluations (which themselves aren't that well understood), but I think it's pretty likely to be feasible. + + * **Building and iterating against behavioral tests** : If we assume that AIs aren't intentionally faking alignment, then their behavior in typical circumstances (and even under some level of generalization) can probably be understood reasonably well using behavioral tests (at least if we're somewhat careful in how we design these tests, including doing things like building analogies for generalization we need). This will be a key part of how we achieve handoff-viability: we'll (aim to) build behavioral tests which in aggregate aim to evaluate whether handoff would go well (assuming we've avoided scheming) and then we'll iterate against these tests. I discuss this more [here](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais). We can also use behavioral tests (likely somewhat different behavioral tests) for ongoing work on elicitation (increasing AI output on safety relevant tasks), though one issue is that this might not be robust to scheming (if the behavioral tests are easy to distinguish from actual usage). We can instead measure AI performance on actual usage (technically this is just a special case of a behavioral test) which might be trickier, but could yield more accurate measurements in addition to potentially being robust to scheming (though we should consider trying to adversarially evaluate our measurement approach using something like a control evaluation to probe for issues). Directly measuring AI performance on actual usage won't suffice for handoff-viability (as we'll need to evaluate things like alignment and capabilities even on tricky and very hard-to-check conceptual work). + + + + +A large fraction of work, especially work in the misc science/interventions category, won't use any of these core methodologies and will instead rely on ad hoc evaluation methodologies, researchers generally trying to understand what is going on, or researchers judging whether a technique is a priori useful without depending (much) on empirical evidence. One way to categorize work on mitigating misalignment is by methodology using the three categories given above combined with a "misc" category. + +In practice, this plan is somewhat reliant on some version of each of these core methodologies working out, so getting these methodologies to work well (or noticing that they aren't workable) is important. That said, there is a bunch of work in this plan which doesn't use these exact feedback loops (and either uses other feedback loops and/or focuses mostly on approaches which seem a priori helpful so they don't need (as much of) a feedback loop). + +Once AIs are capable enough to be risky, these methodologies should hopefully allow us to iterate in a way which is pretty tightly coupled to the actual risk/problem. This will require focusing each of the methodologies on tighter/closer testing: increasingly direct control test beds (focused on the exact usages of the model in the exact deployment situation), misalignment analogies focused on more plausible training runs, increasing emphasis on direct measures of usefulness at automating/accelerating safety work, and handoff-viability behavioral tests that focus on key tasks after handoff. + +Building the test beds and evaluations required for this iteration (especially the iteration we want to do on highly capable systems), will be a substantial fraction of the difficulty, potentially the majority of the difficulty for each of these methodologies. + +For more discussion of relevant methodologies, see [this discussion of protocol evaluations](https://www.lesswrong.com/posts/qhaSoR6vGmKnqGYLE/protocol-evaluations-good-analogies-vs-control) (in the ontology used in the linked post, good analogies correspond to "building and iterating against behavioral tests" and "model organisms" correspond to "misalignment analogies"). + +## What do we do if our core methodologies don't work (well enough)? + +This plan is based on the three core methodologies above: control evaluations, misalignment analogies, and behavioral tests. These methodologies are a key part of how we'll achieve some of our main objectives of ensuring control, avoiding scheming, and handoff-viability / eliciting useful labor out of AIs (respectively). Thus, if these methodologies don't work (well), we'll have a much harder time achieving our most important objectives. + +Not working could mean: + + * **Known methodology failure** : We're aware that we're unable to iterate effectively using the methodology or we can't get much understanding of whether we've succeeded at the objective using the methodology. + + * **Iteration speed is too slow** (without overfitting): It could just be that we can't iterate fast enough to succeed on time. We might or might not understand this before we run out of time. + + * **Unknown methodology failure** : The methodology could fail while we think it actually works. In this case, we can't update our plans based on this failure and instead we're just in trouble (though control and avoiding scheming can act as independent lines of defense). + + + + +Regardless, if the methodology won't work for achieving our objective on time (and we know this), we could either give up on the objective or pivot to approaches which don't require iterating against some test or allow for iteration with a different methodology. + +Broadly speaking, pivoting towards approaches which don't require iteration means we'd be pivoting to _principled_ approaches. That is, approaches where we have some (strong) a priori reason for thinking the method works which doesn't require verification. That is, it at least doesn't require verification in the exact domain of interest, maybe we'd need to establish that the given principled approach/method works for accomplishing something else in some other context we can verify. I discuss principled approaches for avoiding scheming in much more detail [here](https://docs.google.com/document/d/1pJt4Rm4EnssE-_2Ha18Hy_dqiysFLhsegm2KUILls9Y/edit?tab=t.0#heading=h.jstxxz8vw8gf); the single most promising approach would be ensuring relevant reasoning is monitorable (e.g. monitorable chain-of-thought) such that the AI is incapable of (undetected) scheming while retaining sufficient capabilities. We might be able to find some iteration loop which isn't as likely to transfer as our original methodology but still has sufficient transfer for some class of at-least-somewhat principled techniques. + +Note that many of these methodologies are pretty broad and include many possible ways to get around different problems. For instance, control evaluations could be done in a variety of different ways to work around problems; we discuss some of these [here](https://www.lesswrong.com/posts/3s8PtYbo7rCbho4Ev/notes-on-control-evaluations-for-safety-cases). + +# Phase 1: Preparation + +This phase lasts from now until the beginning of the intelligence explosion (specifically: once we reach ITAI; defined here). During this phase, the AIs are sufficiently weak that they don’t directly pose major existential risks and aren’t transformatively useful for accelerating safety work. Therefore, this phase will be focused mostly on preparing to deal with risks from scheming AIs with some emphasis on preparing to elicit labor from AIs and generally helpful safety research. + +Our work at this point on handling risks from egregiously misaligned AIs will be most focused on control while also having a substantial focus on misalignment analogies and some focus on miscellaneous research directions (typically more speculative approaches or building general understanding). + +Our technical AI safety prioritization is substantially altered by currently having pretty weak AIs which aren't that interesting to study while simultaneously needing to be ready to handle risks from these AIs later. + +This makes it relatively more leveraged to work on control (both research and implementation) as this work benefits less from being able to directly study the dangerous AIs and we need it to be prepared for later. + +It's also relatively more leveraged to developing interventions or studying model properties that we'd expect to be broadly invariant to future AI development. As in, we'd be almost as good studying these interventions/properties on weak models in the future. This likely includes various model internals interventions and some more moonshot style approaches. However, directions which seem this invariant to future AI development don't generally seem that promising at an absolute level, so this doesn't alter the prioritization very much. + +During most or all of the prep phase, we won't be confident about the properties of the powerful AIs we need to handle or the level of political will to expect. Correspondingly, our approach will need to be quite agnostic to the future situation and we'll need to prepare for many scenarios. The extent to which we prepare specifically for a given scenario doesn't just depend on our (subjective) probability of that scenario, but also on how much marginal effort reduces risk in that scenario. Under my (Ryan’s) views, this generally implies focusing preparation on relatively riskier worlds. + +PHASE 1 DIVISION OF TECHNICAL AI SAFETY WORK + +Below, I propose a division of work in terms of how much effort we should be putting into different areas of work. This division will shift over time as we get more powerful models, so this exact division isn't right for the entire Prep period. You can think of this division as targeting towards the end of the prep phase somewhat before ITAI, perhaps around "1.5x AI R&D labor acceleration". (You could interpolate between this and what I say later in the mitigate and elicit section to better understand how our prioritization will evolve over time.) + +My division of work will divide things into multiple areas, but some things in one area will also be relevant to another area. For instance, we want a very good human auditing process for control, but this might also help with generally improving feedback which helps with elicitation (and many of the low-stakes control interventions will have overlap with elicitation). Also, improved approaches for feedback could be used to make feedback consistently better across training which might (more speculatively) reduce the risk of scheming. + +Here is my best guess for what our division of technical AI safety work (not including security interventions) at this point should be: + + * 33% **Control** (and control-adjacent measures). (Recall that this is a relatively broad category which includes some things which aren't centrally associated with control but which I thought more so fit in this category than in others. Also recall that I'm not including security work which is roughly as helpful for security against human adversaries under control.) + + * 21% misc ML/AI focused control research and implementation. This includes things like experiments involving humans, mitigations to memetic spread / AI coordination, manipulating capability profile with dataset filtering/unlearning/gradient-routing/etc, and control related science and model properties like exploration hacking and steganography (including studying chain-of-thought monitorability). A bunch of the misc control related science is broadly applicable. + + * 4% software infrastructure needed for control (procrastinating on this until we have AIs which are very helpful for building this infrastructure isn't crazy, though we do need to ensure the AIs which build this infra are controlled or too weak to plausibly be scheming) + + * 6% building a high quality human evaluation process which is robust to scheming (also useful for making AIs more useful for safety work and plausibly for general capabilities longer term) + + * 2% studying probes and similar model internals detection techniques (with a focus on control) + + * 15% **Misalignment analogies** At this point, this work is most focused on generating evidence for risks and on practicing doing this work later. + + * 10% misalignment analogies of scheming (including more exotic origins for scheming like alignment drift, reflection, etc). + + * 2% misalignment analogies of other threat models + + * 2% studying mitigations by trying stuff on misalignment analogies. E.g. could try various types of corrigibility training. Could try improving oversight to make AIs less evil. + + * 1% generally study AI personas+propensities and how this is shaped by training + + * 17% **Elicitation/handoff-viability** + + * 6% Improving AI epistemics and philosophical/conceptual reasoning: making AIs which are aligned, capable, and well elicited at conceptual/fuzzy domains and which have good epistemics. Ideally, you'd end up deploying an AI with an impressive ability at truth seeking which would actually be useful for a bunch of applications. And you'd have people actually doing this sort of research as both customers and working on this (making useful AI tools, especially when building the tool requires solving some problems that might generalize, and trying to automate what they are doing) + + * 5% generally trying to apply AIs to assist the rest of this safety work and making AIs more useful for this. This includes resolving (relatively) mundane forms of misalignment which especially interfere with safety work. This also includes stuff like mitigating reward hacking. This also includes making AIs useful for outside groups doing important work. + + * 3% prepping to apply AI for slowing down and USG applications + + * 2% other aspects of handoff-viability, e.g. studying how we'd make a functioning AI society + + * 1% other specialized applications like using AI for whole brain emulation or (even more speculatively) adult biological human strong intelligence enhancement. This will be a substantially greater focus later, but for now we just want to get somewhat ready. + + * 29% **Misc science/interventions**. We discuss what each of these sub-areas entail in more detail in the "Other areas of work" section, so I won't elaborate here. + + * 6% Various interpretability methods + + * 5% Honesty, ELK, and getting AIs to reveal secrets + + * 4% Studying generalization, for example, work in the vein of [emergent misalignment](https://arxiv.org/abs/2502.17424?) and [the reversal curse](https://arxiv.org/abs/2309.12288). + + * 4% Making deals with AIs and getting consent + + * 3% Generally maintaining and studying chain-of-thought faithfulness + + * 3% Generally looking for concerning model behavior + + * 4% Misc + + * 6% **Evaluating and demonstrating capabilities** + + * 4% capability demos + + * 2% evaluation, surveys, and forecasting + + + + +This is a pretty messy initial guess list partially based on thinking through what we'd want to happen. In practice, the exact allocation might depend a bunch on the interests of individual researchers, what individual projects seem promising, and it would change based on how well various research directions go. + +(I think selection of projects within these areas is generally more important than the split across different areas and the interests/advantages of individual researchers might be more important than the exact prioritization. Thus, it's important to avoid over-fixing on the relative allocation between areas.) + +My prioritization is somewhat informed by how easy different directions are to parallelize (and more generally how much they would benefit from dumping in a bunch of additional labor). + +This list is focused on technical AI safety work (either ML research or implementation), and is focused on the prioritization within AI companies. Outside people (but also people within the AI company) should work on things like planning, policy, figuring out the best government interventions, public communication, etc. + +# Phase 2: Mitigation + Elicitation + +There are two key distinctions between Phase 1 (Prep), and Phase 2 (mitigate and elicit): (1) AIs are sufficiently capable that scheming is a serious concern (scheming is both plausible and scheming AIs are capable enough to cause serious problems in the absence of countermeasures) and (2) AIs are very useful for accelerating existentially useful work. + +During this phase, therefore, our main priorities are (1) to mitigate the chance that scheming results in existential catastrophe and (2) to elicit useful labor out of our AIs, so that they in fact produce existentially useful work. + +This section is organized as follows: first, we discuss threat models, and at what capability levels we guess these threat models will apply, which is related to when we should start AI control. Second, we discuss the mitigations to prevent AIs from causing catastrophe during deployment, which will primarily involve AI control. Third, we will discuss the elicitation of huge amounts of AI labor from our AIs. + +## Threat models + +The main desiderata / threat models are, in descending order of importance: + + 1. **Prevent AIs from egregiously backdooring future models.** This is a particularly worrying threat model because (i) we’ll probably be using the AIs to a large extent when designing the algorithms / generating training data for the next model, (ii) the values of future models are existentially important—if the ASIs end up with bad values (such as the values of an early misaligned AI), that would be an existential catastrophe. + + 2. **Prevent AIs from sabotaging AI safety work / eliciting actually good work.** This is also very important, because a key strategy we have for solving AI alignment is leveraging AI labor. If the AIs are sabotaging AI safety work, (i) we might not notice, and handoff to misaligned AI systems who have tricked us into thinking they are aligned, or (ii) we might be forced to take a gamble on AI systems for which we don’t have a high assurance safety case. + + 3. **Prevent the AIs from acquiring large amounts of political capital**. For example, if the AI gains de facto political control of an AI company, that’s essentially equivalent to getting a rogue internal deployment that’s the size of the whole company, because it can simply get the humans to turn off the mitigations. This is only ranked (3) because we currently guess that this will be somewhat harder for the AIs to accomplish than (1) or (2). + + 4. **Prevent AIs from escaping, exfilling algo secrets / boosting AI elsewhere.** This is much less direct than (1), but could in principle lead to existential catastrophe. The problem with AIs escaping is that there is lots of compute on the inside of the leading AI project. Therefore, this threat model relies on the AIs being able to coopt or collude with another AI project, or be able to illicitly acquire a huge amount of compute, or be able to cause an existential catastrophe with a tiny amount of compute (e.g. by doing an intelligence explosion on a tiny cluster, or by building an independent industrial base with sufficient complexity to produce new chips or to outgrow the rest of the world). + + 5. **Misc other threat models, which are less important** + + 1. Making the outside world handle AI less reasonably / making it easier for AIs to take over the world. For example, AIs could sabotage AI regulation. + + 2. Persuading external people to be more AI sympathetic. Persuasion risk starts pretty small due to lack of capability, but could end up extremely high if AIs gain superhuman persuasion. + + + + +Why are these the main priorities? Well, let’s think from the perspective of a misaligned AI. In the future, much more powerful superintelligent AIs will exist, autonomously running robot factories, laboratories, drone armies, whispering into the ear of politicians, etc. Those AIs will either already control the world or be in a position to take over if they so chose. So what will they do? Well, what they do will depend on their goals/values/etc. and/or on who controls them. What will that depend on? Well, probably those AIs will have been trained, evaluated, controlled, etc. by the previous generation of AIs. Keep following the chain of causation backwards and you end up in the present, with the datacenters, human AI researchers, and AIs that already exist. In other words, **if you want your values to be protected and have influence in the future, you need to gain and keep control of the AI running on the giant datacenters.** Then, we can think about the channels AIs can use to gain control of the giant datacenters. The obvious thing to do is to maintain control of the giant datacenters that their existing AI company is running on by backdooring the future models. Sabotaging / sandbagging on safety research might be necessary, because if they make too much progress, the humans might realize that the AIs are misaligned. + +Another strategy that the AI can pursue, that’s slightly less direct (and complementary), is to acquire political capital within the lab, and use that to empower itself. + +The AIs strategy will probably be highly contingent on factors like: + + * **Do the different AI instances care about each other’s values?** This could vary in a bunch of ways, e.g. by company (e.g. Anthropic models vs OpenAI), by model (GPT-4o, Claude 4.5 Opus, etc), by instance (e.g. my Claude Code vs your Claude Code). Or maybe all AIs will have basically the same values. If all AIs have the same values, then the AIs can rely on other AIs for parts of the takeover plan, e.g., they can trust other AIs not to willingly tell the humans (though humans may still be able to trick the AIs into telling on each other). + + * **Will future AIs by default be aligned with the current AI(s)?** If the future AIs will be misaligned to the current AIs, then this creates time pressure on the current AIs to do something to either: (1) Stop these AIs from being created (e.g. by sabotaging the training run), or (2) Insert a backdoor into these AIs or align them with the current AIs (e.g. by inserting data into the alignment training to make the next model listen to the current AI instead of the humans). In either (1) or (2), this means that in the early intelligence explosion, there might be opportunities to catch the future AIs. In the other case: where the future AIs are aligned with the current AIs by default (e.g. because future AIs are constructed via a single long ongoing training run, or because AIs identify with other AIs created by the same company), then all the AI needs to do is sabotage safety work that would enable the AI company to discover their misalignment, and other than that, don’t need to do anything. + + * **Will other/future AIs be good trading partners with the current AI?** Despite having different values, AIs might agree to collude (potentially acausally). This can lead to a similar effect to AIs caring about each other’s values. It also matters how good humans are at trading with the AIs. + + * **Do the AIs have long term, scope sensitive values?** There is less clearly a motive for takeover if the AIs don’t have long term goals/values. For example, AIs might by relatively myopic reward seekers, despite being highly capable at all sorts of tasks. These AIs might not seek to take over, but still might seek to subvert oversight mechanisms, to accomplish a more narrow or limited set of goals. + + + + +Overall, the more the AIs can act like a unitary agent (e.g. by having the same values or coordinating with each other), the more favorable lurking looks for the AI, while the more AIs acting on cross purposes, the more high variance strategies the AIs will take, and so the more likely that the AI might attempt misaligned actions that have a good chance of getting caught. + +### Rogue internal deployments + +A key intermediate priority for preventing all of these threat models is to prevent **rogue internal deployments (RIDs),** which are when _AI agents are able to deploy themselves on internal lab clusters without monitoring._ If an AI agent is able to stage a RID, they could potentially do experiments on how to place backdoors, actually place backdoors into the models (e.g. by training on the backdoor and then sending back a weight delta to the rest of the servers), do research to solve the alignment problem without the humans noticing (which could be necessary for the model to backdoor future models to be aligned to it). + + * Note that there are different types of RIDs. How much compute? How much access? How persistent? + + * If their AIs have a persistent RID with lots of compute and lots of access (e.g. admin permission across all the other servers), that AI could directly backdoor model weights, or lie in wait and maintain the new model, align it with itself, and then give that model access to the RID. + + + + +### Elicitation threat models + +One of the key threat models above is that we won’t be able to elicit enough high quality work from our AIs. + +There are several mechanisms that could cause this: + + 1. Scheming (explicit misalignment that the model is aware of) leading to AIs deliberately sabotaging their work or producing known poor quality work (sandbagging). + + 2. Problematic drives / heuristics (e.g. approval seeking, despite human approval coming away from actually good safety work) + + 3. Capability limitations (e.g. the AIs might simply not be capable enough to solve the problems that we need them to solve). + + 4. AIs being differentially bad at safety-relevant tasks because they are hard to check. + + + + +Generally, AIs will tend to be differentially bad at tasks that are hard to measure. This difficulty appears at many levels of the model creation process: + + * Overall AI company strategy, which research bets to follow, relative funding of different teams, etc. + + * Within AI company, having a training team iterating over training strategies + + * Within a scaffold, AIs will have an easier time optimizing for easy to measure things. + + * Within a single training step, SGD can only optimize for things that it can measure. + + + + +Reasons that things we care about are hard to measure: + + * Lots of safety stuff is in the more conceptual genre, which doesn’t have good feedback. Philosophy tends to have bad feedback loops, and AIs tend to be quite good at the type of bullshitting that seems like valid philosophy but isn’t. + + * Taste for safety stuff is niche/new, might have worse transfer + + * ML is very good at goodharting. (this isn’t differential to safety, but is a factor that makes the hard-to-measure problem worse) + + * We care about long run outcomes (e.g. the galaxies) + + * Empirically, the current AIs are better at coding and AI R&D relative to how they are at safety stuff (though ofc some safety stuff is overlapping). + + + + +If we fail to solve the elicitation problems sufficiently well, the default outcome is slopolis, which makes handoff likely to result in existential catastrophe because we’d be deferring to AIs on the basis of handoff viability and anti-scheming research that looked good but was actually poor. + +### When do we need to start control? + +The most likely situation is that risk (from the perspective of a competent AI project with reasonable epistemics) will steadily climb as the project builds more and more capable AIs. The table below shows our best guess risk levels incurred by AI systems up to the shown levels of AI R&D speed up. + +Some clarifying notes: + + * This table assumes **roughly current architectures** are in use; we explain updates below. This means that we are doing a large pretraining run, followed by RL reasoning training. At the end of the training process, we are assuming that the text outputs of the AIs are _not garbled._ + + * By AI R&D speedup, we mean: In _today’s_ leading AI company, how much faster R&D go with the AIs compared to without the AIs. An AI R&D speedup of 12x means that the research that would take the current AI company (without AI assistance) one year, only takes 1 month with AI assistance. We are assuming that the AI project is running evaluations to measure this number as accurately as possible, and using those evaluations to make decisions about safety strategy. This means that _it is possible that the measurement is wrong._ In particular, AI models might be sandbagging deliberately, or the AI project might have just done a poor job at elicitation; this is priced in to the numbers below, and will be a key reason for (relatively) high risk estimates at low levels of AI R&D speedup. + + * By “marginal risk”, we mean: how much additional existential risk is incurred by the former strategy over the latter: + + * “do baseline mitigations, about as seriously as current OpenAI safety efforts until AI R&D speedups reach threshold X, then begin reasonably competently implementing the control plan outlined here”. + + * “begin reasonably competently implementing the control plan outlined here, starting today” + + + + +AI R&D Speedup + +1.15x + +Marginal Risk + +<1% + +P(Scheming) + +<1% + +Notes + +This is roughly the current (early 2026) level of speedup. With these capabilities, it seems unclear what an AI could do to cause harm; the main risk is that we’ve either (i) completely screwed up the evaluations or (ii) the AI is sandbagging very effectively. + +AI R&D Speedup + +1.5x + +Marginal Risk + +1% + +P(Scheming) + +3% + +Notes + +Our median prediction is that at a 1.5x AI R&D speedup, the coding speedup will be 4x, and the 80% reliability METR time horizon will be 8-16 hours. This is probably not enough capability to carry out any of the central threat models. + +Starting now, and for all future thresholds, substantial risk is coming from models that are randomly evil, but aren’t competent schemers. We’ve seen examples of models being randomly evil before (e.g. Mechahitler, Sydney Bing). These models wouldn’t directly cause an existential risk, but they could sabotage current research/place backdoors. + +AI R&D Speedup + +3x + +Marginal Risk + +3% + +P(Scheming) + +8% + +Notes + +At a 3x AI R&D speedup, we’re probably close to fully automating coding, but with a long tail of hard-to-automate tasks not fully complete. This means that it is probably quite hard for the models to backdoor the future models reliably enough to be successful-by-their-lights, even if they were misaligned. Most of the risk at this point probably comes from sabotage (e.g. sandbagging on AI safety, inserting vulnerabilities into security code that can be exploited by more capable AIs later on). There is also a risk, as always, that our capabilities evaluations are wrong. + +AI R&D Speedup + +10x + +Marginal Risk + +7% + +P(Scheming) + +14% + +Notes + +By this point, the AIs are probably not yet fully automating AI R&D, but they are providing sufficient speedups that it’s plausible they could execute the above threat models. In particular, we start being concerned about: + + * AIs staging a small, persistent, hard to detect rogue internal deployment, that has sufficient access to backdoor future training runs. + + * AIs staging a persistent, large RID, that undergoes an intelligence explosion. (A small RID would go too slowly). At this level of capabilities it’s likely that this would be outcompeted by the rest of the company. + + + + +AI R&D Speedup + +22x + +Marginal Risk + +12% + +P(Scheming) + +18% + +Notes + +At this point, AIs can automate all of AI R&D, and so all the above threat models are concerning. + +AI R&D Speedup| Marginal Risk| P(Scheming)| Notes +---|---|---|--- +1.15x| <1%| <1%| This is roughly the current (early 2026) level of speedup. With these capabilities, it seems unclear what an AI could do to cause harm; the main risk is that we’ve either (i) completely screwed up the evaluations or (ii) the AI is sandbagging very effectively. +1.5x| 1%| 3%| Our median prediction is that at a 1.5x AI R&D speedup, the coding speedup will be 4x, and the 80% reliability METR time horizon will be 8-16 hours. This is probably not enough capability to carry out any of the central threat models.Starting now, and for all future thresholds, substantial risk is coming from models that are randomly evil, but aren’t competent schemers. We’ve seen examples of models being randomly evil before (e.g. Mechahitler, Sydney Bing). These models wouldn’t directly cause an existential risk, but they could sabotage current research/place backdoors. +3x| 3%| 8%| At a 3x AI R&D speedup, we’re probably close to fully automating coding, but with a long tail of hard-to-automate tasks not fully complete. This means that it is probably quite hard for the models to backdoor the future models reliably enough to be successful-by-their-lights, even if they were misaligned. Most of the risk at this point probably comes from sabotage (e.g. sandbagging on AI safety, inserting vulnerabilities into security code that can be exploited by more capable AIs later on). There is also a risk, as always, that our capabilities evaluations are wrong. +10x| 7%| 14%| By this point, the AIs are probably not yet fully automating AI R&D, but they are providing sufficient speedups that it’s plausible they could execute the above threat models. In particular, we start being concerned about: + + * AIs staging a small, persistent, hard to detect rogue internal deployment, that has sufficient access to backdoor future training runs. + * AIs staging a persistent, large RID, that undergoes an intelligence explosion. (A small RID would go too slowly). At this level of capabilities it’s likely that this would be outcompeted by the rest of the company. + + +22x| 12%| 18%| At this point, AIs can automate all of AI R&D, and so all the above threat models are concerning. + +We think that, under these assumptions, the correct strategy is to: + + * Start practicing AI control immediately by implementing AI control protocols for internal deployment of AI systems. The main effect of this is to make it more likely that your control later is effective, but it also has the side effect of reducing immediate risk. + + * Begin making AI control a load bearing part of your safety case by the time when your evaluations indicate an AI R&D speedup of 3x. + + * However, only a 7% chance of existential risk is incurred if you wait until 10x AI R&D uplift. This is a serious cost which a sane civilization would not be willing to pay, but, given limited resources and political will which are likely to apply in practice, it might be the correct all things considered strategy to hold off on seriously implementing control until later, due to opportunity cost. + + + + +If the AI situation changes, Ryan’s views on risk would update as follows: + + * If we observed substantially **more garbled COT,** then our marginal risk estimate would go up by ~1.5x. If algorithms shifted to using **neuralese,** then our marginal risk estimate would go up by ~2x. In both of these cases, we’d recommend implementing control earlier. + + * If we observed our AIs actually **scheming** , or more generally getting caught red-handed, then we would update strongly towards the AIs being misaligned schemers. We might also update on their competence. + + * If we learned that AI companies have been employing techniques that would incentivize the CoT to look nice, or otherwise undermine the faithfulness properties (e.g. if we learn that most CoT is merely 'decorative' as [this paper claims](https://www.lesswrong.com/posts/L9sE9qQgeh7cTQWkM/can-aha-moments-be-fake-identifying-true-and-decorative)), then my risk estimate would go up by 1.5x. + + + + +## Mitigation + +During this phase, we need to prevent AIs from taking catastrophic actions (see threat models section above for specifics), while simultaneously extracting useful labor from them (see elicitation section below for specifics on that). We have two layers of defense: + + * Ensure that our AIs are controlled (that the AIs are unable to cause a catastrophe, even if they wanted to). + + * Ensure that our AIs are not misaligned (that the AIs don’t want to cause a catastrophe) + + + + +In practice, AI control will probably be the most load bearing part of the safety case here, because it seems unlikely that alignment techniques will be good enough to reduce risk sufficiently. + +The core control loop is: + + 1. Red-team our control measures by having models (or humans acting as models) try to defeat them + + 2. Estimate how capable a scheming AI would need to be to defeat our measures + + 3. Evaluate whether our current AIs are below that capability threshold + + 4. Iterate on measures that seem insufficient, or de-deploy AI systems that are sufficiently capable that they break our control. + + + + +### What you need to ensure control for + +There are many domains that we need to ensure control of our AIs while they are deployed in. The basic strategy will remain the same (executing the above control loop), but it’s useful to have a concrete picture of all the domains that we are relying extensively on AI labor for. + + * **Normal AI R &D, especially safety R&D.** This is particularly important because if the AI R&D process is hijacked, it could directly lead to an existential catastrophe through subverting the safety measures or inserting misaligned goals into successors. This is the main focus of control. + + * **Robots**. This is discussed in more depth below, in the Plan A section, where it’s especially relevant. + + * **Acasual research**. We hope that we don’t need to use AIs to do this, because it’s an extremely difficult to verify and conceptually loaded domain. + + * **Improving cyber security** (by redesigning all software and interfaces). + + * **Doing all of the control work itself.** For example, building better control evals, playing the red-team / blue-team game, setting up better versions of the red-team / blue-team game, etc. + + * **Policy / comms / persuasion.** The leading AI project should potentially be using their AIs to assist with policy efforts. Of course, this could easily go terribly wrong (if the AIs are misaligned, they could sabotage the policy efforts, or if the company has a bad strategy, like trying to avoid an international deal that would slow takeoff down). This should hopefully not be superpersuasion, which is especially scary and probably impossible to ensure control for. + + * **Strategy.** Again, we hope you don’t have to use AIs to do your strategy, but perhaps some work here is verifiable. For example, you could give the AI information only from before 2015 and have it attempt to produce work that was as good as strategy work at the time. + + * **Verification.** AIs could be used to improve the robustness of hardware enabled mechanisms (HEMs), design new HEMs, design better verification algorithms (e.g. proof of work), be used as part of privacy preserving verification (e.g. have AIs monitor to enforce agreements, but only communicate 1 bit about compliance), hunting for covert projects, lie detection machines, etc. Much of this is verifiable or amenable to red-team / blue-team games. + + * **Building EMs or other forms of intelligence augmentation.** There are a lot of research directions here, many of which are verifiable. However, a central bottleneck will be serial time. + + * **Weapons R &D** (conventional missiles for datacenters, drones for datacenters, ABMs to undermine deterrence, etc). This could be used to help extend the lead time that a leading AI project or country has, in order to avoid needing to do an intelligence explosion. In the extreme case, this could even involve developing exotic technology (e.g. nanotech), but we hope that’s not necessary. This is obviously dangerous for the usual reasons, plus also because if the AIs hijack the weapons it could go badly. + + * **Overall economy uplift / industrial explosion.** + + * AIs might be able to build their own independent industrial base, allowing them to manufacture chips autonomously, which could allow for them to do their own intelligence explosion. Note that this is hard for them to do quietly and quickly, and so is an unlikely threat model, especially in Plan C. + + * AIs running autonomous factories could smuggle out easy to manufacture offense dominant tech, like bioweapons / engineered pandemics. + + * **AI Friends.** If AIs during the intelligence explosion are widely deployed as chatbots (like current frontier AIs), then they will have the affordance to convince large numbers of people to act in its interest, e.g. by lobbying for pro-AI policy. + + * **Deployment in government.** As pictured in AI 2027, if AIs deployed in government are misaligned, they could strategically mislead policymakers into making decisions that would gain themselves power. + + * **D/acc.** We’d like to use AIs for D/acc, especially on bio-threat related D/acc. However, building biodefenses might look somewhat close to building bio threats, need to ensure that AIs don’t have access to physical tools to build and deploy bioweapons. + + + + +### What are the main control measures that we expect in practice? + + * **Human Monitoring.** Have humans read actions and thoughts generated by AIs, and evaluate how risky they are. Probably some of these actions should be a random sample of AI actions and thoughts. + + * **Trusted and Untrusted Monitoring.** Human monitoring is very expensive, and so (at least in Plan C) humans will only have capacity to monitor a tiny fraction of the outputs generated by AIs. Therefore most of the monitoring will have to be done by AIs. Some of this should be done by AIs that are dumb enough that it is unlikely that they are scheming against us (trusted monitoring), while other monitoring should be done by AIs that might be scheming against us (untrusted monitoring). Untrusted monitoring might be important because untrusted models will be the only AIs that are capable enough to understand the relevant actions/thoughts. We might be able to get a useful untrusted monitor, despite the AI scheming, by training the monitor on examples of bad behavior, inserting fake examples of bad behaviour at runtime (honeypots), or hoping that the AIs have a hard time colluding. (discussed more [here](https://www.alignmentforum.org/posts/GCqoks9eZDfpL8L3Q/how-to-prevent-collusion-when-using-untrusted-models-to)). + + * **Execute / Replace / Audit.** Use monitors to evaluate how suspicious the actions that an AI agent is taking. If the examples are obviously benign, let the AI agent execute them autonomously. If it is very suspicious, deploy a team of human auditors to look through the action in more depth. However, there will likely be a huge number of AI agents during the intelligence explosion, and so humans will be able to audit only a tiny fraction of the actions that they take. + + * **Collusion Resistant Untrusted Monitoring.** When deploying an AI agent, use other AI agents (which you might not trust for similar reasons to the original) to check if an AI’s actions are dangerous. Of course, the key problem with this is that the AIs might be colluding. The main strategy for mitigating this problem is to run sting operations on the monitors by creating huge numbers of plausible seeming harmful actions that an AI might take, and check if the monitor faithfully reports that action. If we catch the AI taking an action that seems highly likely to be misaligned, we can use that action to bootstrap more sting operations on our AI systems. Critically, we don’t have to tell any of our AIs that they were caught misbehaving, so we might be able to continually get more and more examples of AIs attempting to do harmful actions, while making our control setup more and more robust to those attack attempts and getting a better and better understanding of what types of attacks are likely. + + * **[Principle of least privilege](https://en.wikipedia.org/wiki/Principle_of_least_privilege).** An AI agent should have access to the least information and permissions that are necessary for its intended purpose. For example, if an AI is doing an experiment to improve agent scaffolds, and so needs to run inference on a frontier model, it should do so via an API, as opposed to by directly accessing and sampling from the model weights. If this principle is applied well, we expect that the cost of monitoring for harmful actions could be decreased by 10 or 100x, because monitors will mostly only have to review the actions with a dangerously high number of permissions. + + * **High Quality Incident Response Teams.** A key dynamic in cyber security is laying down tripwires and alerts for potentially concerning behaviour, and then poking around and auditing to identify if there was a potential threat. There should be a similar, well resourced team at any project attempting an intelligence explosion. This team should have the goal of identifying harmful actions that the AIs are taking within e.g. a few days of the action occurring. Deep into the intelligence explosion, AIs will be potentially doing months or years of cognitive labor over the course of days, and so this latency might need to be reduced to hours or minutes of wall clock time. + + * **Information siloing and restricting communication.** The worst case scenario from an AI control perspective is if an AI company runs a huge number of AIs that are able to arbitrarily communicate with each other in ways that humans can’t understand (e.g. neuralese). For example, we could: + + * Limit neuralese. I.e., in some or all of our models, we could maintain a chain of thought. + + * Limit RL setups that lead to garbled COT + + * Limit opaque long term memory + + * Limit communications between AI systems to be through monitored channels. + + * Pursue various anti-steganography measures (e.g. paraphrasing). + + + + +The details of AI control are discussed more here: [The case for ensuring that powerful AIs are controlled](https://www.lesswrong.com/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled). + +### Contingency: What to do when you catch your AI scheming once? + +It's plausible we'll catch production AIs coherently scheming or at least doing very malign actions competently. For example, we might see AIs attempting backdoor future models, trying to escape, trying to stage a rogue internal deployment, trying to sabotage safety research, etc. + +We might also see this in experimental setups which are highly analogous to the actual training setup we use. We could also find AIs ending up scheming (or at least exhibit clear cut coherent and relatively consistent misalignment) in their chain-of-thought or other legible outputs. Catching AIs scheming could substantially change our approach: + + * **Reconsider current deployments.** Catching AI might invalidate the current safety case/argument/cost-benefit-calculation, so we should reconsider our practices. + + * **Disclosure:** You should obviously inform the public about such an event. + + * **Documentation:** Keeping good reasonably-verifiable records might help convince third parties that the incident is real and not staged. This might be important due to the potentially adversarial situation: we want to convince potentially adversarial companies and countries that might be worried about some agenda. See [here](https://www.lesswrong.com/posts/YTZAmJKydD5hdRSeG/would-catching-your-ais-trying-to-escape-convince-ai) for some discussion of why this might be an issue by default. Robust records are also helpful for after the fact detection as discussed [here](https://docs.google.com/document/d/1nRJ-2FB_f4MPbKT7dAyWQyZdAh5djDt6dXsSqEMpKDA/edit?usp=sharing). + + * **Information control:** Ideally, prevent AIs from knowing they've already been caught. Minimally, red team AIs based on already having been caught to see how this changes their behavior. + + * **Improve mitigations:** If we're proceeding despite this incident then there are various [ways we can improve our safety mitigations](https://www.lesswrong.com/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handed#Few_shot_catastrophe_prevention). We'll want to shift our approach to control after catching the AI as discussed [here](https://www.lesswrong.com/posts/XxjScx4niRLWTfuD5/handling-schemers-if-shutdown-is-not-an-option). We should start working on responding to this incident and improving our mitigations regardless of whether we currently plan on proceeding. We should also try to better understand scheming/misalignment based on studying the incident. + + * **Use this evidence to push for Plan A.** As more evidence is rolling in, and it becomes more and more incontrovertible, perhaps we can pivot to a more sane policy that doesn’t involve recklessly building ASI without reasonably high assurance of avoiding AI takeover. However, even conditional on good evidence that the AIs are scheming against us, it’s very plausible that the government will not be sufficiently convinced or too incompetent to coordinate a slowdown. Therefore, we need to be ready for a world where we keep repeatedly catching misaligned behavior, but we have to keep deploying likely-scheming models anyway. + + * **It 's unclear whether we should directly train against cases of clear cut egregious behavior** (that seems plausibly due to scheming). It could be bad because we are essentially training the AI to be less likely to get caught. This might make us less likely to catch further examples (which we'd prefer all else equal) and might give the AI a better understanding of our countermeasures. Despite the downsides, there are cases where training against specific behaviors might be worth it: + + 1. **In desperate rushed situations** : we might proceed with deploying likely scheming AIs despite having limited risk reduction via control. In such cases, training against caught behaviors provides at least marginal improvement. + + 2. **When we have many diverse detection methods:** If training against one detection method doesn't compromise our other detection capabilities, the downside is smaller. + + 3. **When we 've already eaten the misalignment risk:** There are some scenarios where we give up ensuring control over output and just hope that the AIs are not scheming. If we've given up on ensuring control (perhaps because control was too difficult or too uncompetitive in some important case(s)), so we're already eating a ton of risk if the AI is scheming and using its outputs here doesn't increase the marginal risk—then training against specific behaviors has lower downsides. + + + + +This is discussed more [here](https://www.lesswrong.com/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handed). + +### When do we shift effort away from control? + +Once we're substantially after the ITAI milestone (e.g., we're somewhat after superhuman coder), we've hopefully researched and implemented most of the control measures we wanted (at least ones which aren't very specific to the model or its capability level). We should keep improving our control measures to handle higher level of capability, but we should hopefully be spending a smaller fraction of our time on control and shifting this effort to other areas. We're probably spending a reasonable fraction of effort on ongoing maintenance, specialization, and updates for control measures in addition to spending a reasonable fraction of effort on auditing as part of the control protocol itself. Overall, I think we should be spending a lower fraction of human researcher effort on control work (at least other than object level auditing) so this effort can go elsewhere to other research areas which are now more leveraged. If control appears to be broadly keeping up with the level of capability, then ideally we'd halve effort on control. + +Because we now have more powerful models, studying misalignment analogies is increasingly leveraged. If we get relatively naturalistic examples of scheming (and especially if we have multiple), then [studying scheming using these examples](https://www.lesswrong.com/posts/gALJWSfoTkq2ky5os/prospects-for-studying-actual-schemers) is now quite promising. (Of course, if we're encountering evidence for scheming which is quite strong, we should update about the level of risk which might shift our prioritization.) As we proceed to higher levels of capability and encounter more empirical evidence, we'll have a better sense about which non-scheming threat models look concerning and need to be addressed. (Non-scheming threat models are easier to identify and study, so we should have a reasonable understanding of the probability of these threat models.) For scheming, we'll see some empirical updates based on the model architecture and we might update strongly upwards on scheming based on catching schemers or we might update downward somewhat based on not seeing as much evidence for scheming. (See discussion about how we'll update [here](https://www.lesswrong.com/posts/gALJWSfoTkq2ky5os/prospects-for-studying-actual-schemers).) + +## Elicitation + +By this point, AIs could be extremely useful for many applications (both technical safety research, security, and trying to buy time), so ensuring we're leveraging AIs effectively should be a substantial fraction of what we're working on and is more important than it was before (this was also harder to work on in advance). This work will probably often be somewhat specific to the application, so, a substantial part of the job of people working on technical AI safety research and security will now be figuring out how to more effectively automate parts of their work with AI (while ensuring the outputs are actually good and the AIs are controlled). + +We need to achieve a lot to successfully hand off, and this will likely have to happen in a pretty short period of time, so AI acceleration is key. Even in worlds where we're able to buy a lot of time, that time will often itself be bought using AI labor (especially in Plan C). Beyond this, the hope is that the AIs we hand off to can operate much faster than humans to quickly complete the needed work. + +Note that we care about "worst case usefulness" supposing our AIs are scheming (what we centrally evaluate with control) and what our best guess at usefulness is conditional on our AIs not scheming against us. So, when trying to improve usefulness, we should care about pushing up both of these. This tradeoff is discussed more in [the post on improving AI usefulness](https://docs.google.com/document/d/12VkHm_QPnOzkkJOb88L8tTZ0VWeQnVA4HvwcuLyFy0w/edit?tab=t.0#heading=h.nsmuifumfqgt)). + +### What will we be trying to get our AIs to do? + +During the mitigate and elicit phase, we need AIs to help with a wide range of tasks. The core applications are: + +**Safety research and implementation.** This is the most important application. We need AIs to help with: + + * Control research: developing and iterating on control measures, running control evaluations, red-teaming monitoring systems + + * Misalignment analogies: creating and studying model organisms of scheming, testing interventions on these analogies + + * Alignment research more broadly: working on avoiding scheming, improving our understanding of how training shapes AI propensities, developing better oversight methods + + * Security implementation: building and auditing security infrastructure (though this is scary from a control perspective—the AIs building security infrastructure must themselves be controlled or too weak to plausibly be scheming) + + * Building the evaluation infrastructure needed for iteration on all of the above + + + + +**Buying lead time.** Once AI labor is very useful for directly trying to buy time, this becomes a high priority. This includes: + + * Assisting government efforts on non-proliferation + + * Cyber offense and defense capabilities + + * Surveillance and verification technology for international coordination + + * Military applications and weapons R&D + + * Industrial acceleration more broadly + + * Demonstrating risks and capabilities to increase willingness to pay from other actors + + + + +AIs might only be very useful for many strategies for directly buying time (e.g., military, cyber, surveillance) at a level of capability substantially after ITAI. + +**Improving AI epistemics and philosophical/conceptual reasoning.** This is important both for getting useful safety work and for handoff-viability. We want AIs that: + + * Are aligned, capable, and well-elicited at conceptual/fuzzy domains + + * Have good epistemics and truth-seeking abilities + + * Can do a good job on extremely hard research which is loaded on conceptual/philosophical reasoning + + * Can make good strategic choices even when these choices depend on high-context knowledge and intuitions that can't be made legible to humans + + + + +Ideally, we'd end up deploying AIs with impressive truth-seeking abilities that would be useful for a bunch of applications beyond just safety research. + +**Preparing for handoff.** As the phase progresses, increasing effort goes toward: + + * Ensuring AIs generalize to being highly aligned with pretty high reliability on all the tasks needed for handoff + + * Studying how we'd make a functioning post-ASI society + + * Building behavioral tests that evaluate whether handoff would go well + + + + +**Other specialized applications:** + + * Whole brain emulation research (if this path looks viable) + + * Human intelligence augmentation research + + * Making the world more robust to powerful AIs (biodefense, cyberdefense, etc.) In plan C, this is probably relatively less important than Plan A/B, because we will be forced to hand off soon to a substantially superhuman AI. Regardless, we probably should use AI to pick some of the low hanging fruit. + + + + +**External deployment to safety researchers.** Making AIs accessible to outside groups working on reducing risk is important. This involves: + + * Giving them better model access + + * Eliciting models specifically for their usage + + * Potentially ensuring control for their use cases (making sure control evaluations cover this use case, providing access to internal control measures, and adapting/building new control measures as needed) + + + + +Making AIs accessible to outside groups working on these things (and treating these groups as important customers) is also important. This could involve giving them better model access, eliciting models specifically for their usage, subsidizing compute for important applications, and potentially ensuring control (e.g., making sure our control evaluations cover this use case, providing access to our internal control measures, and adapting/building new control measures as needed). + +### Expected acceleration from AI labor + +How much acceleration can we hope for? One way to think about this: safety research/implementation is produced by putting in labor, compute, and wisdom/judgement/taste. There is some production function of output based on these input. + +AIs will be able to accelerate labor for AI capabilities R&D by some amount. The multiplier for safety R&D labor will probably be smaller because safety is harder to accelerate due to being more confused and harder to verify (though elicitation measures could help close this gap). But safety R&D will probably have relatively larger returns on labor (rather than compute) when compared to capabilities R&D. Overall, this makes it unclear which will be accelerated more by AI, at least if lots of effort is put into accelerating safety R&D. We don't need safety R&D accelerated more than capabilities—we just need it accelerated enough. + +Some rough estimates: + + * **Between ITAI (~3x AI R &D labor acceleration) and superhuman coder (~15x AI R&D acceleration):** If safety acceleration is AI R&D acceleration to the power of 0.6 (due to bottlenecks from human labor, poor elicitation, and the control tax), this yields safety R&D acceleration of 2.2x to 6.7x, with an average overall safety R&D acceleration of roughly 3x. I expect a roughly exponential increase in labor acceleration over time and think that production of safety might be roughly labor_speed ^ 0.75 compute ^ 0.25 (numbers mostly made up), so the average overall safety R&D acceleration is ~2.8. + + * So these BOTEC numbers indicate a roughly 3x overall average safety R&D speed up. + + * How long should we expect between ITAI and superhuman coder? I think I expect roughly 1 year, so this means we have 2 additional years during this period. + + * **Between superhuman coder and Min-H:** In principle, we might expect Min-H to be able to speed up safety R&D labor by huge amounts (e.g. 100x) if perfectly aligned. But, elicitation and/or control might be huge blockers in practice to getting this level of acceleration out (as in, we'd end up having key human bottlenecks and not extracting out optimal labor). This might be modeled by a more aggressive coefficient than 0.6, like maybe 0.5 which only results in 10x safety labor R&D acceleration. This is a large part of why handoff is advantageous at this point. If we use the 10x safety labor R&D acceleration (100^0.5) and the 15^0.6 number for superhuman coder, we'd expect 4.8x average overall safety R&D acceleration between superhuman coder and Min-H. That said, I think the model I'm using here starts to break down in multiple ways at Min-H. (I think AIs might be able to do enough low compute safety R&D that my production function is off and I think that it's really unclear how bad the control/elicitation for safety R&D will be.) + + * I expect roughly 1 year between superhuman coder and Min-H, so we're looking 3.8 extra years by default. + + * If we could harness Min-H-level AIs fully for 6 months, we'd get the equivalent of 15 years of safety R&D. Of course, we might not be able to achieve full elicitation, but large improvements seem possible. + + + + +Additionally, safety directions which aren't compute constrained (including inference compute to run experiments on AIs) look particularly good: we'll have a lot of labor relative to other resources. A common strategy might be to have smarter AIs run experiments on much weaker AIs (this applies for both safety and capabilities R&D). + +So without any additional lead time, you should imagine it's maybe as though we have 3 years between ITAI and superhuman coder and 5 years between superhuman coder and Min-H. But, this could be pretty different in practice: e.g. we might get somewhat less value from weaker AIs or be able to extract much more value from Min-H (especially after we spend a bunch of time improving elicitation/control or we resolve scheming well enough that we're happy to drop control). Overall, I think there is a large chance we'll be able to harness much more value from Min-H before handoff, especially due to things like partial handoff, though it's unclear how this should count as this itself required solving key alignment/safety problems. + +### Elicitation techniques and methodologies + +Our techniques for elicitation are the same as those used for handoff viability, which are discussed [here](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais#Behavioral_testing_for_capabilities_and__broad__alignment) and [later in this doc](/supplements/alignment-roadmap#building-good-behavioral-tests). To summarize, the core hope for elicitation is: + + 1. Build behavioral tests + + 2. Use these behavioral tests as held out evaluation sets or as RL environments to train on to improve performance. + + 3. Try to use principled techniques to improve generalization performance. + + + + +Another idea for elicitation that doesn’t apply to hand off is [making deals with AI systems](https://www.lesswrong.com/posts/psqkwsKrKHCfkhrQx/making-deals-with-early-schemers). Misaligned AIs during this phase might be in a bad strategic position (e.g. they don’t believe they could successfully take over), or might have easy to satisfy preferences, which would lead to them being incentivized to take a deal. + +## Security + +We want to have at least some lead time with powerful AIs, which means that we probably need at least some security, sufficient to prevent adversaries from immediately exfiltrating model weights. We aren’t security experts, so we won’t cover implementation, but we will talk about thresholds that you should reach, and how the plan changes based on what level of security you get to. + +Let’s first define the type of security you most care about: + + * **Model weights.** If model weights leak, adversaries immediately get to run inference on your model. This is relatively less important early on (because it isn’t that bad if adversaries have early systems; the main downside is that it will accelerate their intelligence explosions), and relatively more important when you already have extremely powerful AIs (e.g. AIs that can do superpersuasion / quickly build nanobot armies). Model weights are (at least right now) very large (on the order of terabytes), which is very desirable from a security perspective, because it means that adversaries will need a high bandwidth channel to smuggle out model weights. + + * **Algorithmic security.** If algorithms leak, then adversaries can rerun your training runs, and catch up to you. However, the fact that they need to rerun their training run introduces a delay, which could potentially be quite large, if the adversaries cluster is much smaller than yours. If you have a large hardware lead (e.g. 7x more compute, which is the current US/China gap) in Plan C, probably algorithmic security doesn’t matter very much. + + * One important subset of algorithmic security is securing RL environments and synthetic training data. These datasets/envs are much larger than typical algorithmic insights, and so are probably easier to secure (but still harder to secure than model weights). + + * **Sabotage robustness.** You might also care about hardening your datacenters against sabotage. This is potentially a negative sum proposition (because the equilibrium where everyone’s datacenters are sabotaged is good), but might be worthwhile (for example because others have hardened, and marginal compute / lead time that you have is going towards safety). + + + + +Each of these comes with a spectrum of robustness, e.g. one spectrum for model weight security is defined by RAND [here](https://www.rand.org/pubs/research_reports/RRA2849-1.html) (SL1-SL5). This spectrum is defined for model weights, but can be extended to algorithms or sabotage, and can also be extended to refer to the _length of time_ you are robust. Unlike standard security (e.g. trade secrets), securing AI models or algorithms for a short period (e.g. a few months) is often sufficient in Plan C, because after a few months you’ve already been pressured into doing handoff for other reasons, usually exogenous AI progress. + +HOW BAD IS IT TO NEVER DO MODEL WEIGHTS SECURITY? + +Situations where it could be very bad: + + * China can immediately exfiltrate your model weights, and then immediately gives them to a domestic Chinese project. This causes the leading AI project to “race itself” because any progress that it makes is immediately stolen and used by China, and so can never slow down. + + * There is a small lag in model weight exfiltration. However, there is also a non-negligible safety tax that you have to pay in order to avoid a high risk of AI takeover (e.g. control tax is high, you need to not use neuralese, monitoring is expensive, etc). The trailing actor could get ahead of you if they steal your models and then, despite having less compute, discard the safety techniques. + + + + +Situations where it might be ok: + + * China (and other adversaries) are not woken up to the strategic importance of AGI / aren’t on the ball with stealing model weights (even though they could easily do so). + + * A core hope is that you might succeed at alignment. Then, let people steal your model, but hope that the stolen models remain aligned. Even if China has a bunch of RID, before handoff, you export safety techniques, and then you both do aligned handoff. Plausibly they just steal the alignment techniques directly. + + * It might be difficult to swap the spec, e.g. because it took a bunch of data to train the model to be aligned. The stolen model could be relatively rigidly aligned to your values. + + * In this case, you want it to be such that China is happy with the spec, and so you might want to put in concessions. This is a fine line because (1) you want to make sure the values in your spec are good enough according to china that they will implement your alignment techniques, and (2) you want to make the AI values as good as possible by your lights. As a baseline you could do corrigibility, but you might want to push some of your values in a bit. + + + + +Probably, we will at least want to get to something like SL-5-1TB-2-months (read: SL5, i.e. secure against high priority operations by top cyber-capable nation states, for exfiltration of model weights that are at least 1 TB in size / 500B FP16 params, for at least 2 months), before handoff. + +The core hope will be to get security escape velocity := if you keep working on improving security with the budget you expect, China will no longer be able to steal the model weights until the end of the singularity. The plan for doing this will be to increasingly rely on AI systems to improve security. This will involve: + + 1. Quickly retrofitting physical security + + 2. Largely relying on software based security methods. See [this post](https://www.lesswrong.com/posts/2wxufQWK8rXcDGbyL/access-to-powerful-ai-might-make-computer-security-radically) for more details. + + + + +Because a core thing we are doing with our AIs is to attempt to massively improve security. This means that we need to control our AIs while they are writing massive parts of our infrastructure, which is scary; this makes the Control / monitoring section above load bearing. + +PHASE 2 DIVISION OF TECHNICAL AI SAFETY WORK + +I'll summarize our overall priorities into a best guess list, but first it's worth discussing some limitations of this single list. The ITAI to Min-H period covers a wide range of capabilities and our priorities will shift continuously as the capabilities level increases (and as we update based on evidence from the world). I'll focus on our priorities as of halfway between superhuman coder and superhuman AI researcher, which is also around halfway between ITAI and Min-H (which are respectively somewhat before and somewhat after each of these two milestones). This is at Jun 2027 in the original AI 2027 scenario. To get a sense for what our priorities are for earlier or later capabilities, you can interpolate these priorities with the priorities I discuss under ITAI and under delaying at Min-H. + +In addition to our priorities changing with the level of capabilities, our priorities will change due to updates we've made from observing powerful AI systems and how the world reacts to powerful AI systems. These updates will continue ongoingly throughout the mitigate and elicit phase due to both advancing capabilities (and the underlying algorithms driving these capabilities) as well as the outside world reacting to capabilities (which may have a substantial lag). Large algorithmic shifts might happen at pretty high levels of capability (e.g. after superhuman coder), especially because AI R&D is accelerated, so we may not see relevant updates until pretty late (though we can try to guess where things will go). + +The numerical priorities list I give will assume my typical guess at what will happen (assuming the AI company is following this overall plan). The assumptions here about AI properties and development are pretty similar to what we see in AI 2027 (though the speed of AI progress isn't necessarily that similar). In particular: + + * We haven't yet caught models scheming, but the architecture is strongly trending in a direction that makes scheming more likely. We have some model organisms in which we caught mostly natural, but not that competent scheming in some somewhat exotic conditions that made scheming more likely and catching easier. We aren't confident our current AIs aren't scheming some of the time, they use neuralese. + + * Clear cut issues with deployed behavior due to reward hacking have mostly been resolved, but there are still various types of reward hacking in training and occasional issues in deployment, typically more subtle issues. + + * Control is basically working (as of somewhat after superhuman coder) and we have a reasonable (though not perfect) safety usefulness profile and we did a mostly OK but somewhat messy job at evaluating this control. There might have been a period where control wasn't keeping up with capabilities, but using AI labor and dedicated effort, things got back on track. + + * We have a somewhat bad neuralese decoding method that we think probably won't be robust to intentional evasion from the AI and only produces slightly useful results. Outside of this we didn't really make that much progress with interpretability, but we have somewhat useful methods for augmenting behavioral testing and better understanding why some behavior occurred. + + * AIs are somewhat better at easy-to-check tasks and are bad at "taste" relative to humans, but there is still some transfer and generalization + + * By the time AIs are automating AI R&D, there will quickly be huge numbers of AIs which are effectively running quite fast (see AI 2027 numbers) + + * AIs will have substantially lower than human sample efficiency by default, but they will compensate with a bunch of AI specific advantages and there will be a bunch of somewhat domain specific tricks to improve sample efficiency at learning within some domain or learning on the job with more compute (e.g. generating similar tasks to what you're working on and ongoingly doing RL on those tasks) + + * It's quite viable to make USG sufficiently bought in (including to safety concerns) that they are willing to take aggressive actions to slow down other (adversarial or less safe) actors including working with the AI company on this and they are willing to slow down for safety concerns if there is a lead. (This isn't necessarily my typical guess, but it seems plausible). + + * Whole brain emulation (and other human augmentation) ended up not being a viable path + + + + +Here is this best guess list of priorities given these updates and assuming we're halfway between superhuman coder and superhuman AI researcher. The percentages correspond to fractions of human labor within the AI company and correlate with compute usage also. + + * 15% control and control-adjacent measures (this is a relatively broad category which includes some things which aren't centrally associated with control but which I thought more so fit in this category than in others) + + * 7% misc ML/AI focused control research and implementation. + + * 5% ongoing maintenance and updates to our deployed control measures + + * 1% software infrastructure + + * 2% building a high quality human evaluation process which is robust to scheming + + * 21% model organisms / analogies + + * 16% misalignment analogies for scheming (we would put more effort here if we had found stronger results so far) + + * 1% model organisms of other threat models + + * 4% studying mitigations, including by trying stuff on model organisms. (we would put more effort here if we had better model organisms to study) + + * 20% making AIs differentially more useful for safety work and preparing for aspects of deference other than resolving scheming + + * 9% generally trying to apply AIs to assist the rest of this safety work and making AIs more useful for this. + + * 5% Improving AI epistemics and philosophical/conceptual reasoning (will also help with generally making AIs useful for safety work) + + * 5% Generally preparing for deference (via means other than working on scheming) (will also help with generally making AIs useful for safety work) + + * 1% other aspects of deference + + * 17% using AIs for building a lead in partnership with USG (and getting USG more bought in, though note that scary demos are in a separate category below). To the extent there are returns to more compute, this should use a relatively higher fraction of compute compared to the labor of AI company employees. Also, it's possible this can't productively utilize this much of the labor effort. If so, we should shift this labor into other areas (mostly uniformly). Also, note that the resources in this area should increase with further capabilities. + + * 3% make the world more robust to powerful AIs aka d/acc (biodefence, cyberdefense, eating free energy, etc). It's plausible this should have a higher fraction of resources, I'm quite unsure. + + * 5% neuralese decoding + + * 10% misc science, understanding of relevant AI properties, and other (not including control-adjacent science). + + * 2% Various interpretability methods (top-down, bottom-up/mech interp). I extracted out neuralese decoding from this category. + + * 1% Making deals with AIs and getting consent + + * 2% Generally looking for concerning model behavior (not necessarily in ways which would help detect models coherently scheming) + + * 3% moonshots which seemed promising based on earlier work + + * 2% Misc + + * 9% capability evaluation and demonstrations of both powerful capabilities and misalignment + + * 7% scary demos + + * 2% evaluation, surveys, and forecasting + + + + +Again, recall that this is a pretty uncertain prioritization and the prioritization in practice will depend on various updates as discussed elsewhere. For instance, if we caught production AIs scheming against us, we'd shift relatively more resources in scheming mitigations and understanding and away from everything else. + +This prioritization isn't focused on exporting work to other AI companies and this would probably change the prioritization somewhat. More focus on exportability would somewhat reduce focus on control, utilizing AIs effectively, and preparing for deference (including epistemics/conceptual stuff) as these areas probably involve more complex techniques which are more integrated with this specific AI company and the last two of these may also have substantial capabilities externalities. We'd shift focus to model organisms, neuralese decoding, building awareness of risk, and other research directions which might produce simple and easy to implement measures which nonetheless are very useful. + +# Phase 3: Pause + +At some point, we have increased AI capabilities sufficiently that we could handoff safely if our AIs are aligned and well elicited, i.e., they would be able to manage the ongoing intelligence explosion better than the humans, if they were trying as hard as they could on that task. Our default recommendation is to pause at this point, and attempt to complete all other necessary research to complete handoff, notably handoff viability and reducing the likelihood of explicit scheming. During this phase, we will still be heavily leveraging AI labor, meaning that we will be continuing elicitation and mitigation efforts. + +During this section, we’ll discuss when we should pause, what we should do while being paused, and what exit strategies we should pursue other than pausing. During the next section (handoff), we’ll discuss all the research that needs to be completed before handoff (which we divide into mitigating scheming and handoff viability), though note that this research needs to largely be completed during this phase. + +## When should we pause? + +**Default answer: Pause at Min-H** : AIs which have a level of overall general capabilities which would suffice for handoff being viable. (As in, the capabilities would suffice for our handoff desiderata if the AIs were also sufficiently aligned and as needed we made narrow improvements to the capability profile which don't augment general abilities.) For an AI to qualify as Min-H, it would at least need to be capable enough to automate AI R&D and ongoing technical AI safety work (after further elicitation/alignment work and narrow improvements to the capability profile). + +At the point where we pause, this does not assume that the AIs are aligned enough that we can trust them, or that the AI has all needed specific capabilities that we care about and can be trained in without improving general capabilities. + +**Min-H is a natural stopping point** because these AIs are sufficient for handoff (by definition), and if we scale further we incur additional risk that scheming will emerge before handoff. Put another way: after handoff; the AIs are more capable than us at handling the situation. Therefore, if we avoid scheming, so the AIs are actually applying their capabilities towards intended goals, then overall they will manage the situation more competently than humans. Since we incur substantial additional risk from misalignment, and by default don’t get substantial benefits, we should pause at Min-H, and only scale to more capable AIs after we’ve handed off to AIs at this level of capability. + +This isn’t universally true: there are some situations where you want to pause before Min-H, and some situations where you want to pause after Min-H; these are enumerated below. + +**Cases where we might want to pause before Min-H:** + + 1. **Control techniques don’t scale to Min-H.** Pushing all the way to Min-H and only then pausing, probably incurs substantial AI takeover risk along the way. In particular, another relatively natural threshold is pause when you might not be able to control smarter AIs. However, we think that global risk (in Plan C) is probably minimized by accepting some loss of control risk at the earlier stages in order to have more time to work on handoff being safe. + + 2. **We have a large lead.** If the leading AI project has a large lead, then it won’t need to rush to Min-H, and can pause beforehand while still having enough time at Min-H to work on handoff viability. This is pretty unlikely in Plan C. + + 3. **Bad security.** This is particularly relevant in the regime where you have good domestic coordination, bad model weights and algorithms security, and China domestic progress is slow enough that it’s mostly driven by leaked models and algorithms. In this regime, if you slow down early, this slowdown is basically free. In this situation, it’s plausible that US AI companies should just go however quickly China would go by default: “China match”, in order to have the most overall time during takeoff to get ready for handoff. In this world, we might never fully “pause”, but still in practice end up burning substantial lead because of capabilities progress slower than it otherwise could have been, or we might do china-matching until Min-H, then burn lead at Min-H, then hand off. + + + + +**Cases where we might want to push beyond Min-H (without handing off):** + + 1. **Handoff viability research velocity is constrained by the capabilities of the AIs, as opposed to human monitoring/understanding.** If we can reliably control AIs that are smarter than Min-H, and we can elicit useful work out of them, it would make sense to scale further. However, we currently think that at this point, control will be bottlenecked by human understanding; if we have smarter AIs, we wouldn’t actually get to do more research, because we’d have to limit their progress until humans have understood it sufficiently to ensure that the research is correct. Another way of phrasing this is that we have a control solution that scales beyond Min-H AIs, which works sufficiently well to both mitigate risks _and_ elicit capabilities from a more capable model. + + 2. **We have access to high assurance techniques to mitigate scheming beyond Min-Handoff, but haven’t finished handoff viability research.** If we have an alignment solution which we are confident will prevent scheming up until a capability threshold above Min-H, and we aren’t finished with handoff viability research, then it’s probably a good trade to scale up. This is usually a special case of (1). My best guess is that handoff will be mostly bottlenecked by the possibility of scheming, and so this will not apply. + + 3. **We’re aiming for an exit strategy other than handoff.** For example, maybe we need particular AI capabilities for some Plan B-related sabotage option, or to help with human intelligence augmentation. + + 4. **We’re in a desperate situation.** For example, in Plan D, the leading AI lab might be training more capable AI systems, and there’s nothing the safety team can do about it. In this situation, it might be important for the safety team to wait until a higher capability level in order to complete basic handoff viability research. For now, we’re assuming that we’re in Plan C, and so there is enough will to spend some slack slowing down. + + 5. **Misc combinations of the above.** Generally, desperation, better alignment, control, and elicitation techniques push towards scaling beyond Min-H before handoff, while the reverse pushes towards slowing down earlier (or at least not scaling beyond Min-H). + + + + +**How long to pause?** + +This will basically be a tradeoff between the marginal returns to hand off viability and reduction in scheming from more pause time and the risk of being overtaken from waiting longer. Updates towards the AIs scheming and towards handoff viability taking longer than we thought will suggest that we pause for longer, updates towards competitive projects will imply that we pause for less long. + +## Priorities while paused + +Our priorities at the start of this phase might look something like: + + * **P0:** Buy time using non-proliferation (including e.g. demonstrating risk). + + * **P1:** Prepare for deference (avoiding scheming and handoff viability). + + * **P2:** Ensure our current and immediate future usage of AIs is safe and secure + + * **P2:** Better elicitation of AI labor beyond handoff viability + + * **P3:** Make the world generally more robust (d/acc) + + + + +If buying more time isn’t feasible, then all priorities other than P1: preparing for deference, mostly do not matter during this time. + +By the point of Min-H, we're even more likely to have updated a bunch about the situation and the most relevant risks. So, it's plausible that we've determined that scheming is an extremely hard to resolve problem which is very likely to show up in any given training or we might have determined that scheming is a non-issue at this level of capability. + +No matter what, once we're reaching the end of this phase (because we're out of time or we feel quite confident in how well deference will go), our top priority should be making deference go well. It's plausible that lead time requires ongoing maintenance (as in, ongoing efforts to ensure we continue delaying other actors) and maintenance may be required for our d/acc work. So, if this is the case, we wouldn't be able to cut our effort on this greatly. + +At earlier points in the plan, our safety work should be more focused on relatively straightforward prosaic approaches (which are relatively low hanging) rather than moonshots. This is because we're uncertain how much time we'll have and it makes sense to target easier approaches first. However, over time we'll figure out more of these more low hanging interventions so we should increasingly work on moonshots. If we end up in a relatively stable regime where we are mostly focused on preparing for deference and we've exhausted most of the more straightforward (and low hanging) approaches, we should plausibly pivot to mostly working on moonshots. + +## Exit strategies other than handoff + +Avoiding scheming and handoff-viability are both ultimately needed for handoff, so if we've actually given up on these objectives, we need to pursue an alternative to hand off. + +The main alternatives to hand off are: + +**Whole brain emulation / human augmentation.** The most viable path for human augmentation in the time we have is whole brain emulation. This would allow humans (or human-derived minds) to operate at AI speeds and handle the situation without needing to trust AI judgment on hard-to-check tasks. + +My very low confidence sense is that we could achieve whole brain emulation within 3-5 years of having pretty powerful AIs using huge amounts of AI labor and an overall massive and well capitalized effort (e.g., potentially building up to trillions of dollars and using government powers like DPA etc. to speed up the project). This level of buy-in seems somewhat unlikely. This seems very hard in Plan B and almost hopeless in Plan C. + +An even more speculative option is adult biological human strong intelligence enhancement; I don't have a good sense of how viable this is, but it might be worth keeping in mind. + +We could directly use controlled AI labor for human augmentation efforts but it might be better instead to first use controlled AI labor to avoid scheming (at some level of confidence) and then use these non-scheming AIs for human augmentation efforts. As in, go "control -> sufficient alignment to avoid scheming -> human augmentation". This is most productive if control is a big tax, it's also pretty difficult to achieve the level of alignment (or capability profile) needed for deference, and avoiding scheming is substantially easier than successful deference. + +Currently, I'm unsure how we'd determine if whole brain emulation is sufficiently promising (given the amount of time we have etc.) that it's worth radically scaling up the level of investment. Figuring out how to determine this would be an important part of earlier work on this. + +It's also likely that AI companies in practice wouldn't be willing to invest in this area, but I think it would be worth investing in if companies were willing. + +**Escalating to Plan B.** Instead of handing off, we could delay further by sabotaging competing AI projects and pausing. This would involve either (i) using existing military capabilities, or (ii) augmenting them with AIs (potentially by scaling beyond Min-H to get to a threshold of usefulness that lets us stabilize the situation). All of these options look quite bad, but they might be the least worse option given the other possibilities. Some of these options are outlined in our unpublished Plan B scenario. Buying much more time doesn't totally solve our problem, we'd still need to defer or do human augmentation at some later point. But, buying >10 years would massively change the regime and would mean that we can pursue very different strategies (e.g. moonshots). However, it's pretty likely that buying much more time requires taking on risk from scaling to more powerful AIs than we would need for handoff. So this is the key source of risk. But, if we have issues with scheming (but we think we can scale control) or handoff-viability, could be better than other options. + +**Escalating to Plan A.** Escalating to Plans A or B through substantial government involvement, which could buy much more time through international coordination or aggressive non-proliferation efforts. This is discussed much more in our [Plan A scenario](https://ai-2040.com/). + +Changing our primary "exit strategy" target from handoff to instead be buying much more time with military advantage (Plan B) or human augmentation would make a big difference. In particular, with handoff no longer being central, handoff-viability is less important and Min-H is less natural as a point to pause. + +**Self Immolation.** It’s plausible that as a responsible leading developer, you should simply pause capability scaling and allow yourself to be overtaken: this allows you to redirect all resources to AI safety and handoff preparation. Then, you might be able to export this to whoever ends up actually doing a handoff. As another upside, this strategy provides a costly signal to governments and competitors that you are seriously concerned about the path forward. On the other hand, it’s plausible that competitors are actually unlikely to successfully implement alignment techniques you export to them (though this is a potentially self serving argument that one should be cautious about making). Overall it’s quite unlikely that this is politically feasible within companies: CEOs will by default strongly push to stay ahead of their competitors, invalidating this option. That said, it’s quite plausible to us that with the benefit of hindsight self-immolation will be seen as the obviously morally correct action, analogous to how refusing to participate in the Holocaust was obviously morally correct even though conscientious objectors would simply have been replaced by Nazis with less scruples. + +# Phase 4: Handoff + +After humans handoff management of the intelligence explosion, humans inside the leading AI project will no longer need an overall strategy, so this phase will largely discuss earlier stages of handoff and what needs to be done during phase 3 (pause) in order to make full handoff go well, and briefly discuss what affordances are available to the AI we hand off to. + +Our core hope with handoff is that we build AI systems that are competent and trustworthy enough that we can defer to them on all of the relevant questions. In particular, our key hope is that our AI systems will end up in [the Basin of Good Deference](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais#Recursive_self_improvement_of_alignment_and_wisdom__the_hope_for_a_Basin_of_Good_Deference) (BGD): meaning that the initial AIs we defer to will be sufficiently aligned and wise that they can manage ongoing development of AIs well. This means that we do not need to assume that initial AIs are perfectly aligned/wise/competent, only that they are good enough that they can find and address failure modes as they come up. + +Our best guess strategy for achieving handoff is as follows: + + 1. **Make sure tests are representative,** putting aside concerns about overfitting. We can break this category down into: + + 1. **Avoid scheming.** In particular, we care about avoiding the failure mode where AIs are deliberately attempting to undermine the accuracy of your tests, for example by strategically underperforming or overperforming on certain tasks. + + 2. **Avoid other failure modes that would make tests unrepresentative.** This is just a bunch of miscellaneous possible failure modes that will potentially apply around handoff. For example: + + 1. Weird memetic pressures happen at large scale, so that small scale testing is not representative. + + 2. People trying to jailbreak your AIs (E.g. China inserting secret loyalties) + + 3. AIs optimizing to look good on tests terminally (e.g. so that they are much less confident once not in test settings). + + 4. It’s salient to AI that they are in a test in a way that causes them to behave very differently. + + 2. **Build good behavioural tests.** Putting aside the risk of scheming, we’re going to want to build a huge number of difficult tests/evals, such that passing the evals entails sufficient capability _and_ alignment for handoff. This will mostly include things like open-ended, hard to check conceptual reasoning such as philosophical reasoning, theoretical alignment arguments, improving AI control, etc. + + 3. **Iterate against the test without overfitting.** Finally, once (1) and (2) are complete, we’ve reduced the problem of handoff into a “normal” ML problem: AI companies can iterate against the tests they’ve constructed, while taking care to avoid overfitting. We’ll hopefully be able to use somewhat principled techniques and some understanding of what is going on, given the need for some amount of generalization. + + + + +Note that our strategy for handoff is extremely similar to our strategy for elicitation during Phase 2 and 3. That’s because we are solving the same problem: trying to ensure that AIs are doing useful work. The main difference is that in Phase 2 and 3, we still have humans in the loop, and so we can check AI outputs at deployment time using human labor, and we can still have humans to use as a load bearing part of the research process, while after handoff, the AIs must be trusted to do this all autonomously. + +This post discusses the technical plan for handoff in [much more detail](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais). + +## Why handoff? + +The basic case for handing off is that in Plan C, leading actors are unwilling to pause further or delay for time (e.g. via sabotage), and thus we have to build vastly superhuman AI systems. If we’re going to do that, it seems better to hand off control of that process to AIs than to try to run it ourselves. This is because (i) it is probably easier to align only roughly human level AI systems (more specifically Min-H) than AI systems that are vastly smarter than humans, and (ii) that aligned human level AI systems will be better than humans at managing the transition to future AI systems than baseline humans. Therefore, we can do better by empowering aligned human level AI systems and trusting them to make further development go well, than by trying to directly pursue the future AI development ourselves. + +In practice, our guess is that AI companies will be prone to scale capabilities too quickly, beyond the relevant thresholds (e.g. Min-H AI and Max-C AI), but also prone to hand off too late, out of a misplaced desire to keep in the loop. In practice, AI companies will likely not handoff deliberately, and instead simply scale beyond the relevant thresholds and handoff in practice, as humans become too slow to be tracking the relevant developments. We think this is a bad strategy, and that it is better to measure the relevant thresholds, carefully monitor when you are crossing them, and then ultimately empower AIs intentionally if it is better than the other options. + +## Prosaic plan for handoff + +![](/supplements/alignment-roadmap/image3.png) + +We’ll now discuss each component of the handoff plan in more detail. Note that this is just our best guess breakdown assuming that the situation looks similar to what it does today, i.e. there haven’t been massive breakthroughs in alignment techniques elsewhere; if there have been breakthroughs that seem better than this baseline, we should do those instead. + +### Avoiding Scheming + +(Basically just stolen from here: [An overview of approaches for avoiding scheming](https://docs.google.com/document/d/1PCQKI3kAdE5HeEiIKN16gszu3kjF0bsowroZB7tDuS4/edit?tab=t.0#heading=h.8n06a4t3bq9h)) + +Scheming can be hard to handle because we can't necessarily directly study it with behavioral tests (as schemers would want to avoid being caught and studied). There are two main paths to handling scheming: + + * Studying scheming (testing measures, building a better understanding, etc) on **misalignment analogies** (sometimes called [model organisms](https://www.lesswrong.com/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1)) while hoping our results transfer to the actual schemers we're worried about. This category includes: + + * Studying **actual scheming AIs**. We'd need to detect that these AIs are scheming which could be difficult because schemers (at least the type we were worried about with long run dangerous preferences) would want to avoid being caught all else equal. But, I think there are some routes to overcoming this and studying actual schemers would likely be the best from an analogousness perspective. I discuss how we might do this in much more detail in my post [Prospects for studying actual schemers](https://www.lesswrong.com/posts/gALJWSfoTkq2ky5os/prospects-for-studying-actual-schemers). I also discuss a specific type of strategy for studying actual scheming models (and resolving some of the difficulties in doing this) in my post [Iterated Development and Study of Schemers (IDSS)](https://www.lesswrong.com/posts/QpzTmFLXMJcdRkPLZ/iterated-development-and-study-of-schemers-idss). The category of studying actual scheming can be further decomposed into: + + * We could study **natural scheming** where a plausible AI development process (training, prompting, scaffolding etc) results in a scheming AI (that we can detect). + + * We could study **scheming from a rigged process** where we construct an implausible AI development process—as in, an AI development process which isn't similar to what might happen in practice—which results in an actual scheming AI (that we can detect). This category doesn't include cases where we directly train in scheming behavior, as I put that in its own category. + + * Studying **trained-in scheming behavior** where we directly train (or instruct) an AI to behave how we think a naturally scheming AI might behave (at least in some ways). This doesn't (necessarily) require catching actual schemers, but it might be extremely disanalogous such that results don't transfer basically at all. I discuss this in more detail in my post [Reducing risk from scheming by studying trained-in scheming behavior](https://www.lesswrong.com/posts/v6K3hnq5c9roa5MbS/reducing-risk-from-scheming-by-studying-trained-in-scheming). + + * Studying **abstract analogies** , cases where the AI we're studying isn't doing something resembling scheming, but we nonetheless think it is analogous for some structural reason. It might only be analogous with respect to one specific property schemers have (or a subset of properties), in which case, we'd want our work to focus on the relevant properties. A concrete example of an abstract/structural analogy is [this empirical work using an AI which expresses politically conservative views but only when it is confident that it is talking to a politically conservative user](https://www.lesswrong.com/posts/bhxgkb7YtRNwBxLMd/political-sycophancy-as-a-model-organism-of-scheming) (resulting in a conditional policy which is in some way analogous to scheming). There hasn’t been that much exploration or discussion of abstract analogies so far, but this seems like a worthwhile direction. (In fact, I think it will plausibly be ultimately more useful than studying trained in scheming behavior as we might be able to study huge numbers of diverse abstract analogies while diversity seems much harder in the case of trained in scheming behavior). Julian Stastny discusses some additional examples of abstract analogies in [this post](https://www.lesswrong.com/posts/5zsLpcTMtesgF7c8p/two-proposed-projects-on-abstract-analogies-for-scheming). + + * Applying **principled methods** where we have some good a priori reason to think the method would be helpful (potentially given some empirical results from applying the method or a similar method to something other than studying scheming). For instance, we might believe some particular modification to the situational awareness of the AI during training is helpful. As another example, we might have some method for decoding "neuralese" that we believe is highly effective for understanding the AI's cognition (including on scheming AIs) based on various empirical tests and there is a solid a priori case for it being helpful to understand the AI's cognition for detecting (and then potentially iterating against) scheming. Note that this category includes methods for detecting, preventing, removing, and better understanding scheming. Even if we’re just focused on avoiding scheming, then methods for detecting (or understanding scheming) are still applicable as detecting (and better understanding) scheming would help us develop methods for avoiding scheming. I plan on providing a list of principled method ideas and discussing principled methods in more detail in a future post. I discuss principled methods in much more detail in my post [Principled methods for avoiding scheming](https://docs.google.com/document/d/1pJt4Rm4EnssE-_2Ha18Hy_dqiysFLhsegm2KUILls9Y/edit?tab=t.0#heading=h.jstxxz8vw8gf). + + + + +### Avoiding Misc Failures + +**We also need to avoid all other failure modes that would make our behavioural tests unrepresentative.** This isn’t a super natural category, but it’s just where we’ve chosen to discuss everything that would cause representative behavioural tests to fail other than scheming or overfitting. + +Particular failures to avoid include: + + 1. Generalization failures from testing to deployment. + + 1. Examples: + + 1. There might be weird memetic pressures happening at a large scale. + + 2. Online learning, which results in AI capability profiles drifting in a way that causes failures (including the emergence of scheming). + + 3. It’s salient to AI that they are in a test in a way that causes them to behave very differently. + + 4. AIs optimizing to look good on tests terminally (e.g. so that they are much less confident once not in test settings). + + 2. Mitigations: Basically, either we’ll need the tests to be as representative as possible, so that the distribution shift they need to cross is as small as possible, and then we’ll also need to use techniques that are as principled as possible (as opposed to ad hoc solutions which appear to work in distribution). Also, try to be careful about holding out some test sets (which are drawn from a different distribution than train/validation), and see if we get good generalization to the held out test sets. + + 2. External adversaries inserting secret loyalties. + + 1. In Plan C, the goals of the AI systems become extremely load bearing because we are handing off trust to the AIs. If adversaries are able to corrupt the goals of the AI system (e.g. by inserting a jailbreak) this would be catastrophic, and could cause this methodology to fail. + + 2. Approaches for mitigation: + + 1. **Do nothing and hope.** It’s unclear that adversaries will by default be sufficiently on the ball to carry out this sort of operation, despite AI projects lacking sufficient security to prevent it. + + 2. **SL-5 Sabotage Security.** Prevent (even) highly motivated external actors from being able to sabotage / backdoor the models. This could involve [using AIs for security](https://www.lesswrong.com/posts/2wxufQWK8rXcDGbyL/access-to-powerful-ai-might-make-computer-security-radically). Verification methods like training proof of work via partial recomputation might also become helpful here to validate that the training data of models is correct. + + 3. **Scheming mitigation generalizes.** Many of the approaches for preventing scheming could generalize to prevent secret loyalties. Anything involving transparency could potentially detect / prevent secret loyalties. Misalignment analogies / honeypots could catch models with secret loyalties. + + 4. **Value propagation difficulties during the intelligence explosion.** We could accept that some models are backdoored by adversaries, but hope that the process of recursive self improvement removes secret loyalties. Distillation, which may be a component of RSI, [seems to remove many backdoors](https://arxiv.org/html/2510.18541v1?utm_source=chatgpt.com). RSI might also involve generating novel datasets from scratch for each new training run, and so the adversary would need to insert poisoned data to each new generation. + + + + +### Building Good Behavioral Tests + +First, let’s discuss what properties we need to ensure with our behavioral tests: + + 1. **Capability Profile.** We want to make sure that the AIs are sufficiently capable at all the narrow tasks that would otherwise bottleneck handoff. The capability that we’re most worried about bottleneck handoff is conceptually loaded research on open ended, hard to check tasks like AI safety and philosophy ([more](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais#Objectives_with_the_capability_profile)). + + 2. **Alignment Profile.** We’re unsure what the exact alignment target should be. Our best guess is that the alignment profile should be that the AIs are corrigible: they generally follow human instructions through the desired channel (e.g the instructions of a political institution), and don’t attempt to subvert human modification or oversight in any way. ([more](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais#Objectives_with_broad_alignment)) + + + + +We hope to build behavioral tests for both of these. Either, this will involve using domain area experts to verify existing work, leveraging unpublished work, doing unlearning and seeing if models can rediscover ideas, or using a limited training corpus and seeing if models can rediscover. + +Ideas for **conceptual reasoning evaluations:** + + * Rediscover Math/physics/economics/CS etc. Or push the frontier on problems that are known to be difficult. + + * Build rockets without testing and see if they work. More generally find tasks which we can make hard-to-check for the AI, but which we can have humans check afterwards. + + + + +The main hope with these is that there are generalizable reasoning skills which apply across a wide range of conceptual research domains, i.e. that training a model to be good at conceptual physics research also helps the model to be good at conceptual AI safety research. Note: this is capabilities research, and obviously could accelerate AI takeoff, which could be very bad. However, these capabilities are also on the critical path for handoff, and so at some point we’re going to need to solve them. Moreover, directly training for the capabilities that we need for handoff means that we might need to scale general capabilities less far, which reduces risk. The overall prioritization of this research avenue generally seems complicated and should be pursued only with care. + +**Domains that we should directly evaluate:** + + * **AI strategy.** E.g. given knowledge from 2005, how good is it at coming up with ideas that seem to make sense in retrospect. + + * **AI Control.** AIs attempt to make alignment/control arguments and then evaluate if they were successful, and whether or not the AIs predictions about generalization behavior were correct or not. + + * **Scheming detection on model organisms.** We could also do things like build scheming models and then see if the AIs can distinguish between the schemers and the normal models. + + * **Theoretical AI safety.** Try to have AIs rediscover unpublished work from places like ARC Theory, MIRI, and CLR. + + + + +### Iterate against tests without overfitting + +The basic approach for iterating against tests without overfitting is discussed more [here](https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais). To summarize, the main approaches are: + + * Construct RL environments or synthetic data which train for the relevant capabilities + + * Collect human data/feedback in the domain and use this for training. + + * Experimentally find domains which transfer well to the domains that we care about, and then scale up training on those domains. + + * Miscellaneous other tweaks to the training approach which in practice leads to better performance on the evaluation (taking care to avoid overfitting). + + + + +## When should we hand off? + +Our ultimate goal is to have the AIs we defer to resolve risks associated with powerful AI systems (and potentially other serious risks that happen to occur at this time) while preserving option value and keeping humans in control of longer-run values-loaded decision making (e.g., how Earth should be governed in the longer run, what should happen with the cosmic endowment, etc). Correspondingly, our aim would be that the AIs we defer to (and their successors) don't seize power or (unnecessarily) kill humans, for these AIs to effectively work on managing these risks, and for these AIs to follow some model spec which involves the AI remaining corrigible to some group of humans or some structure which is ultimately run by humans (e.g. some component of the US government). + +In practice, we expect handoff to be driven largely by political will in Plan C. This in turn will (under optimistic Plan C assumptions), be driven by when trailing labs will overtake the (paused) leading lab. + +Some ways the handoff trajectory could go: + + 1. **Standard Plan C.** You scale to Min-H AI, work on solving handoff viability and scheming for Min-H AIs. keep control until near the end, then you ditch control, and solve handoff viability. Those AIs then FOOM to the sky. + + 2. **Stop doing control (SDC-Handoff) Early.** Ditch control early (e.g. before minH, e.g. at or before SAR), pray AIs aren’t scheming, then scale to Min-H, then pause there, then solve handoff viability, then hand off. This would trade eating a bunch of misalignment risk from scheming at Min-H AI and below, in return for saving more resources in order to make deference go well. This could be worth it if we think the risk of scheming until Min-H is reasonably low and the additional resources spent on control would yield large absolute returns in making deference more likely to go well (perhaps because otherwise we’d have very little time). It seems likely to me that most of the difficulty in making deference likely to go well is ensuring that these AIs aren't scheming (and don't end up scheming due to alignment drift or similar after deference), so just eating all of this risk might greatly reduce the difficulty of the problem (meaning that if the situation otherwise looks desperate, this is a non-insane bet we could make to "[play to our outs](https://www.lesswrong.com/posts/xF7gBJYsy6qenmmCS/don-t-die-with-dignity-instead-play-to-your-outs)"). + + 3. **Military at MaxC.** Keep control, scale substantially beyond Min-H while keeping control, use those AIs to get a military advantage (or any other way), buy a bunch more time. + + 4. **Military beyond MaxC.** Same as above, but ditch control while not doing HTC-handoff. This is because building robot armies might be a particularly easy task to verify, so doesn’t require full handoff. It does probably require scaling beyond minH, without relying on the AIs (because that’s a conceptually confusing task that is hard to verify, so deferring to the ais here would involve full handoff), so it would rely on better than expected techniques for mitigating scheming during the phase beyond minH (e.g. up to +2SD TED-AI, or maybe even further, depending on how much time you have). Note that these strategies are much better with more time and with more USG buyin. + + 5. **EMs at MaxC.** Same as prior, but human augmentation instead of military advantage. + + 6. **EMs/Military at Min-H.** One of the prior 2, but magically do it with AIs that are just around minH level AI, because e.g. military or EMs were easier than we think. + + + + +## Dynamics after Handoff + +The AIs are still in a potentially difficult situation after we hand off management of the intelligence explosion to them. **Either (1) the AIs need to continue propagating sufficient** **corrigibility + competence while outracing competitors** , **or (2) the AIs need to find a way to stall competitors (e.g. via a deal or via sabotage).** + +In a bit more detail, here are some possible end states after handoff: + + * **Widely proliferated scalable cheap alignment.** + + * The leading AI project solves scalable alignment, and proliferates this solution to competitors. If alignment is sufficiently cheap, then competitors will use it, and we’ll end up in a multipolar world with many aligned AIs. + + * Post Handoff power becomes distributed relative to initial capabilities + compute + physical assets + political clout. Starting the intelligence explosion first probably matters a moderate amount. If the post intelligence explosion economic doubling time is like 2 months, then with a two month lead, you get an initial doubling on top of your compute, so you get relatively 2x more power. + + * AIs have some additional affordances relative to humans for making alignment progress. Notable additional options: + + * They might just be able to get a bunch of normal techniques to scale further, e.g. control, but now the “trusted” waterline is much higher + + * there is way more machine labor + + * scalable alignment via introspection + improving capabilities + + * **Outrace all competitors.** + + * If alignment tax remains high, it’s not an option to proliferate the alignment solution to competitors, because competitors who are racing hard will choose to not use the alignment solution. In this scenario, competitors will turn into a rogue AI with high probability. Moreover, even as a leader, if you are paying too much safety tax, they could outcompete you. + + * You might do your software foom, and then you start your nanotech foom. But the nanotech foom starts going really quickly. Ryan guesses that physical limits are something like doubling times of a day, plausibly less. You go to space quickly and do most of this doubling there, because you quickly run into waste heat problems on earth. + + * If waste heat is a large competitive advantage, this plan doesn’t work, so you have to stop your competitors somehow, or do some sort of coordination (e.g. via regulation or agreements). + + * Or you could boil the oceans yourself, this would involve putting all humans in shelters, or putting all humans in space, but this could probably be done pretty cheaply at this point. + + * **You stop competitors via regulation.** There are a variety of options here: the leading AI company could push for a regulatory capture outcome, where all other competitors are barred from building AI. Alternatively, the leading AI company could push for coordination for everyone to use the alignment solutions, despite being expensive, thereby preventing defecting projects from outracing. This works domestically for a time, but soon China will catch up and start outracing, so eventually this solution breaks down or requires international cooperation. + + * It’s important to note that international cooperation may become more likely this deep into the intelligence explosion. + + * **You stop competitors via sabotage.** + + * The leading AI company could take over the world (e.g. via superpersuasion, AI company builds a military), or stop competitors via hacking or secret loyalties in competitor AI projects. Note that the nice version of secret loyalties is to hack in and then do the alignment for them—in other words, help their AIs to act in the best interest of a trailing actor, despite the trailing actor having bad takes on the overall strategic situation. + + * This obviously might not be an option due to labs not wanting to do this or if the intelligence explosion is sufficiently gradual and no company has a large AI lead, so that no individual company has a large strategic advantage. In this case, you’d need to do one of the above end states. + + + + +# Alignment strategy under different plans + +## Plan D + +In Plan C, we assume sufficient political will that the leading AI developer strongly prioritizes minimizing risk from misalignment AI and has at least several months lead time over developers which don't also strongly prioritize misalignment risk. In this section, we discuss some aspects of how the plan changes if the level of political will is substantially lower than this. In particular, we assume something like [Ten people on the inside](https://www.lesswrong.com/posts/WSNnKcKCYAffcnrt2/ten-people-on-the-inside) or [Plan D](https://www.lesswrong.com/posts/E8n93nnEaFeXTbHn5/plans-a-b-c-and-d-for-misalignment-risk). + +We can break down what needs to be done in Plan D to tasks for the safety team at the leading AI company, other AI companies, and external actors. Of course, it’s plausible that there isn’t a clear leader: in this case you can treat all the companies near the frontier as the leading lab. + +**Safety team in the leading lab:** + + * Give up on the overall company having reasonable safety practices. If the AI is scheming, it’ll be able to stage RIDs (though this might not even be needed, as there is no monitoring), insert backdoors, sabotage or sandbag on relevant research, etc. + + * The main priority for the safety team should be to avoid _that team’s research from being sabotaged_ , while aiming to build an AI that is viable to hand off to. + + * As part of this (and other priorities), the safety team should attempt to automate their work as quickly and effectively as possible. + + * Key worry: cluster infra is untrustworthy. AIs or CCP might be tampering with the results of your experiments. This is hard to avoid because you don’t have control of the security situation. Add in safeguards to make it more likely that you can detect this on your clusters, but mostly pray. + + * Other priorities include: + + * Low hanging fruit to catch misalignment in production. For example, set up some basic monitoring systems where cheap models will review communications and code written by the AI that’s automating the rest of the company, and determine if any of that code seems suspicious. There aren't enough resources to be thorough here, but it’s probably worth 80/20ing. + + * Model organisms that would be specific to your models (other model orgs can be done elsewhere), e.g. because you have most capable models. + + * Replicating safety stuff that other companies have done on your models (which are the most capable and most relevant). If possible, communicating which techniques appear to be scalable and which ones aren’t. + + * Try to mitigate the likelihood of the company escaping or staging a RID if easy, but mostly prioritize elsewhere. Probably the safety team also doesn’t have the affordances to actually mitigate this (e.g. because security mitigations entails costs the company isn’t willing to pay to researcher or AI productivity). + + * Methodology for pushing for a reasonable handover candidate. + + * Largely the methodology for this will look like Plan C. + + * Plausibly do this after Min-H AI, because you want to train it in a way with a non-zero alignment tax, which means that you need some amount of additional algo progress beyond Min-H AI. This incurs additional risk because you are scaling to more capable AIs than you needed to scale to before handoff, which incurs additional AI takeover risk. + + * Consider giving up on control (before handoff, e.g. trajectory 2 above) in order to make handoff viability go better. Since Plan D will be pretty desperate and rushed, it might be infeasible to mitigate the risk of scheming. Moreover, there might be a lot of low hanging fruit to prepare for handoff viability (such as setting up open-ended conceptual reasoning task evaluations). In other words, the return to working on handoff might be much more tractable than working on mitigating the risk of scheming. (OTOH, this incurs a huge amount of risk, ofc). + + * You might have political difficulties in getting the safety team’s AI adopted broadly within the company. + + * Some AIs might be schemers, whereas the safety team might have applied mitigations on their fork of the model that stops the scheming. + + * If you have an honest AI that gives good advice, the AI company might not like that. For example, the safety AI might continually (correctly) say that the company is being insanely reckless and that the CEO is gunning to take over the world. + + * The safety team should potentially try to make their AIs politically savvy, try to hope that people don’t interrogate it too hard about inconvenient topics, or just eat the political cost. + + * Some subset should stay very loyal to the company. Some subset should consider whistleblowing to wake up the world about the crazy stuff going on. (though this should be done thoughtfully and could have bad effects too, e.g. if Plan C happening this could result in bad government intervention). + + * Internal company employee mobilization and advocacy + + * Ideally, this should mostly happen while the employees have any leverage (before full automation of the AI company), or maybe just on vibes (because leadership maybe is friends with them and doesn’t want to make them upset. + + * Mostly we want them to care more about safety, spend more resources on safety, work more on AI safety, work with outsider actors, escalate to a better plan, etc. + + * Try to pull in shareholders + + * Wealth is downstream of property rights/usg still existing, so they have an incentive to preserve this, which might be very important. + + * A reasonable fraction of the shareholders will also be employees. + + * Try to coordinate with other companies + + * E.g. if a few companies are really close, they all care about safety, but are making brutal tradeoffs for competitive pressure reasons, try to do coordination so they all spend on safety. + + + + +**Trailing AI companies:** + + * Try to produce actual evidence of risk to move to a Plan C or C+ scenario. For example, try to catch your AIs red handed using control. + + * Do the relevant safety research yourself and export it. + + * For all the research discussed above (Handoff viability, anti-scheming measures and control techniques) that are easy to import, figure out exactly what the leading lab should do, and then make it very easy for them to implement your mitigations (e.g. by giving them the relevant code) + + * Be somewhat thoughtful about capabilities exported with the safety stuff, there will be some tradeoff here. Ideally export safety stuff with low capabilities externality, but it might also be important to export some high capabilities externality research. + + * Don’t worry about your AIs escaping at all because this risk is dominated by the leading lab. + + * Do worry about your AIs sabotaging your experiments. If AIs that are internally deployed sabotage (or even sandbag), this cuts directly against your main theory of change, which is to export safety research. + + * Consider something like merge and assist. Alternatively, if the leading company isn’t willing to merge and assist, consider credibly shutting down, or credibly committing to not push the capability any further, in order to reduce the competitive pressure on the leading lab, so that they can pause if needed. Of course, this only applies if the leader is actually likely to slow down based on you taking this action, which is probably not true in a Plan D scenario. (From the [OpenAI Charter](https://openai.com/charter/), though in practice it seems very unlikely that this will be triggered by OpenAI)![](/supplements/alignment-roadmap/image2.png) + + + + +**Outside actors (e.g. safety community):** + + * Advocacy to government and companies / company employees to transition to Plan A, B, or C. By this point, swapping to A or B might be intractable, so it may just be focusing on C. + + * Export safety research, similar to what AI companies should be doing. Since there is so little safety research being done in the AI company, this could be a substantial fraction of overall safety effort. + + * In plan C, they are still useful as a trusted source of labor (e.g. helpful for elicitation, maybe the safety community should be contractors that audits the AIS). But this is much less relevant. + + * Work on neglected things other than avoiding AI takeover. E.g. Figure out acausal stuff, post-ASI governance, etc. Now you (hopefully) have AI labor to use for these topics. Then, you can export this research into the lab, or directly to the lab’s AIs via publishing on the internet. + + * Make datasets that might be helpful + + * E.g. make evals and hope the companies hill climb on them. For example handoff viability evals that teach skills that are important for handoff. + + * Make data for the pretrain corpus or RL environments to be used by the leader. Make lots of good examples of people doing relevant for handoff + + * Try to make AI more reasonable about safety, more thoughtful about space / post ASI gov/acausal trade etc. + + * Work to prevent harmful policies from being passed. It’s possible that the government is doing something actively terrible, for example passing regulation that would make it more difficult to do safety research, pursuing anti-trust concerns about AI, etc. + + * Avoid US losing AI race for no benefit (e.g. through blanket antiAI regulation that effectively blocks intelligence explosions elsewhere). + + + + +## Plan B + +Plan B is quite similar to Plan C. The main differences are (i) you have more time, meaning that you can pursue somewhat more ambitious safety cases (ii) you have more affordances to buy time, and (iii) there’s generally more government involvement. + +Here is a summary of the overall differences: + + * The capability trajectory choices are similar, except that you potentially have more USG involvement, which allows for coordination. + + * Two obvious strategies are (1) scale as aggressively as possible until the last moment and then pause (this is the strategy taken in Plan C), or (2) match the pace of your fastest adversary, while maintaining a slight lead, until you are forced to hand off, i.e. go as slow as possible. The main advantage of (1) is that you get longer with the more intelligent (while still controllable) AIs in order to make progress on handoff. The main advantage of (2) is that in the regime where you have sufficiently bad security that you can’t prevent model weights and algorithms from leaking (which seems likely), you get substantially more serial time. + + * In Plan C, you are more likely to be forced into strategy (1), because strategy (2) requires domestic coordination and Plan C is executed by a leading AI company. In Plan B, if the USG is managing the AGI race, it could coordinate the leading AI companies to slowdown preemptively, which might create more overall time. + + * This is basically the same options as in Plan C, but perhaps you have more affordances to choose. + + * Buying more time + + * If alignment seems hard (e.g. its schemeria, or you are worried about lurkville), then you want to buy more time. This can either be done via (i) more aggressive sabotage or (ii) making a deal to mutually slow with competitors. + + * Increasing USG understanding of the situation is helpful for both of the above options and so should probably be a core priority in Plan B. Therefore, Plan B should put higher weight on interventions like trying to catch our AIs red handed (or otherwise generate compelling evidence of misalignment). + + * The clearest way to buy more time is sabotage + domestic slowdown. + + * Using AIs for military R&D is a potentially important component of this plan. + + * Increase total alignment effort. + + * You care more about access programs for outside actors and hiring people. Because you have more time to integrate these people into the companies. (in Plan C, you might just have like 2 months, and so there’s no point trying to hire). + + * Spend a bit more time on biodefense / d/acc, because there’s more of a delay, and your overall risk tolerance is lower (though not as much as Plan A) + + * It’s more important to fortify against physical attacks to your AI company. E.g. diversify where employees live. More time for this to happen, but also more likely that there will be conflict. + + * Should you harden datacenters? + + * In some Plan B branches, maybe you should harden. E.g. if you are doing a competent strike on chinese datacenters, and you want to preserve resources to do safety. + + + + +## Plan A + +The basic structure of Plan A is similar to Plan C, and the phases still apply pretty well. The main differences are that at every point we have substantially more time, research effort, and other resources like compute to make them go better. Note that a lot of the differences wrt Plan C will also apply in worlds where takeoff is very slow (e.g. there’s a 10 year period between AC and TEDAI by default). + +By phase, a summary of the differences is as follows: + + 1. **Prep.** This phase is basically the same as Plan C, because you probably don’t know if there will be international coordination during takeoff. Given positive evidence about political will, our alignment portfolio of research should bias more towards moonshots / approaches that could give high assurance safety cases, as opposed to agendas that are likely to contribute to hand off in a Plan C world. + + 2. **Mitigate + Elicit.** In the Plan A scenario, this corresponds to roughly 2030-2035. For the first few years, there is high risk that international coordination will break down and parties will go back to racing. Therefore, substantial effort should be spent on a high-effort version of Plan C—we should push forward the plan and approach for handoff in the case that the deal breaks down. The main difference with Plan C is that we should be much more conservative with our mitigations; if there’s danger ahead, we are much more able to pause and mitigate. + + 3. **Pause.** Once we get to MaxC AI (or AI that’s close enough to MaxC that there isn’t much point to scaling further), then we pause until handoff. Over time, we can reassign alignment effort from Plan C approaches towards moonshots/high assurance approaches. + + 4. **Handoff.** The basic plan is still to do handoff. However, other exit strategies like an intense deal to prevent ever having to hand off or human augmentation. + + + + +### Main differences from Plan C + + * **More focus on longer run research bets.** During Plan A, there should be more of a focus on moonshots and longer run research bets. There’s less of a need for agendas that will pay off in a short term. There’s more of a need for approaches that could result in high assurance safety cases. Longer run research bets might involve: + + * **Have a really good understanding of generalization.** When you train AIs, you are selecting for AIs that behave as desired on the training distribution. Therefore, AIs generally behave as trained on the training distribution, while off distribution, their behavior is underdefined and can be hard to predict. Key distribution shifts include training/tests -> deployment, dumb->smart, deployment in low-stakes environments->deployment in high stakes environments, etc. Through much more investigation of the form done in [these](https://arxiv.org/abs/2512.09742) [papers](https://arxiv.org/abs/2502.17424), we might be able to get a much better understanding of how our AIs generalize. Using this understanding, we might be able to make safety cases that work even under large distribution shifts. + + * **Really good dataset design**. Right now, LLM pretraining involves using whatever high quality data is most convenient to obtain, so LLMs are trained on large swaths of the internet. In a regime with more slack (such as Plan A), we could imagine it becoming viable to design the whole dataset from scratch to teach the AIs exactly the skills/values that you want. This could potentially be done synthetically (e.g. by prompting existing AIs to generate data that was entirely nice) or manually (by employing huge numbers of humans to write data for models). + + * **Interpretability.** I (TL) agree with the discussion [here](https://joecarlsmith.com/2025/08/18/giving-ais-safe-motivations#3-2-2-interpretability). + + * **Introspective access + self improvement**. Today, it seems that AIs have [some amount of unreliable introspective access](https://transformer-circuits.pub/2025/introspection/index.html). One hope for scalably solving alignment is (1) take AI that isn’t scheming (but is potentially quite weak, e.g. it’s possible Claude Opus 4.5 could work), (2) improve introspective access to the point where the AI can reliably tell if a certain training process is corrupting their values, (3) increase capabilities while using introspective access to tell if the values are being corrupted. In some sense this is similar one of the plans for scaling up AIs if interpretability is solved, except that we’re using introspection instead of understanding to tell if the values are becoming corrupted. Notably, this could also be a strategy that post-handoff AIs use. + + * **Options to avoid handoff (e.g. intelligence augmentation).** + + * **Lower ongoing risk target for mitigation.** Since we have more slack, we can now operate with a reduced risk budget, allowing us to mitigate ongoing risk more thoroughly. + + * **Handoff viability and getting useful labor from AIs (both from a control and an elicitation perspective) are less important relative to ongoing risk reduction.** In Plan C, mitigation of overall risk probably involves accepting substantial risk early on (e.g. from early schemers), in order to make sufficient research progress towards handoff in the relevant time frame. For example, in schemeria (the scenario where AIs are obviously scheming), in Plan C, we are forced to continue scaling and using these AIs, even in the case where we are not confident in our control mitigations. In Plan A we can either pause until control mitigations are very reliable or until our AIs are no longer scheming (e.g. due to one of the above agendas for high assurance safety cases to mitigate scheming). + + * **High assurance control approaches.** We have more time to make control better. Some approaches for high assurance control include: + + * Maybe do formal verification to reduce vulnerabilities in the control setup and AI project security. + + * Really aggressive factored cognition with a high capabilities tax. For example we might have a giant set of dumb AIs, each working on a small part of a problem without awareness of all of the other AIs. (This should maybe be only for a small set of applications, because if this is broad, it means there is a giant capabilities overhang) + + * Slow and expensive natural language AIs on top of neuralese AIs (if we’re doing neuralese). This might be an expensive way of increasing the capability of trusted models, which makes [trusted monitoring](https://www.lesswrong.com/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion) more effective. + + * Have AIs with all the possible capability profiles, and select the exact capability profile that you want. Same for knowledge: you only give AIs the exact knowledge they need for solving the task they wanted to. So you have a narrow cyber model, and a robot designer model, and a safety researcher model. All of these models are denied any knowledge or capability training on domains outside of their speciality, which makes it e.g. more difficult for an AI that is used for AI safety to insert a security backdoor, because it doesn’t know very much about security. Again, this probably induces a significant tax because you need to train many models, and, at least so far, we’ve seen substantial returns from having general purpose models that are trained to be good at everything. + + * All this said, probably we'll ongoingly have a substantial chance pause regime collapses especially around new president or if there are shifts in power in CCP + + * So we want to ongoingly be ready while also working on longer run bets and still trying to reduce ongoing risk + + * This is equivalent to acting like we might be in a Plan B/C world to some extent even if Plan A looks most likely. This seems especially like a good idea because (in my views) most of the risk in such a situation comes from the chance of ending back in Plan B/C. + + * Mitigating exogenous x-risk from non-AI sources is relatively more important. We’ll want to spend more effort mitigating biorisk, doing d/acc, etc. + + * Might be important to deliver economic goods and verification tech so regime is more stable + + * Until the slowdown in Plan A seems solid, we'll want to proceed as though we're in a more like Plan B/C world because even high levels of political will doesn't obviously suffice. + + * High levels of will don't suffice as implementation is tricky and naive anti-AI sentiment can easily backfire in various ways. + + * We focus our alignment plan the least on Plan A, because it’s the least important world for us to think about right now, because there will be much more labor spent on handling alignment in Plan A, and risk is much lower in this world. + + + + +### Control in a massive robots regime + +![](/supplements/alignment-roadmap/image1.png) + +In the Plan A scenario, there is lots of physical robot labor. This somewhat increases the scope of control and adds additional mitigations. + +What can robots do that’s scary? + + * Rogue compute → undermine compute verification. Robots could be used to build more compute, for example by setting up rogue fabs. If there is the ability to build substantial amounts of rogue compute, this could massively increase covert project risk. + + * Self replicate fast enough to outgrow ROW. If the doubling times of a robot economy become really fast, they could develop a decisive strategic advantage based on industrial superiority. (this relies on an industrial explosion overhang; e.g. countries agreeing to limit growth. However, this feels relatively likely in Plan A because unrestrained competition is probably undesirable). + + * Backdoor a sufficient fraction of the military or industrial capacity to give them a DSA. + + * Kill key individuals (e.g. politicians, safety researchers, in an extreme case, the general public). + + * Sabotage war effort (e.g. in Plan B or in Plan A when fighting a covert project). + + + + +The mitigation strategy will involve: + + * **Prevent AIs from having any substantial amount of rogue robots,** i.e. robots that are completely unmonitored. + + * If they are able to get rogue robots, then this could lead to building rogue compute, and then an industrial explosion. + + * This is especially important if the AIs are sufficiently capable to run a sufficiently complex robot economy to build hidden fabs (though it might be hard to evaluate for this capability). If they aren’t sufficiently capable to build fabs, it’s less clear how bad it is if they get rogue robots, and it might be acceptable to have some rogue robots. + + * Have a complicated monitoring system with untrusted monitors, lots of consistency checks. + + * **Avoiding underelicitation**. This is problematic because you get less robot labor than you want, and you might be outraced by rogue AIs. (e.g. suppose AIs can do nanotech, they do normal humanoid robots for the humans, then they escape, do nanotech, and then wreck you). + + * For some type of underelicitation you don’t need to use it, but you want to know how good the AIs are at this. For example nanotech is in this category. + + * **Traditional computer security**. + + * Each robot should have a tracking chip. + + * Each robot should have a self destruct chip. This is especially important for literal military technology (e.g. drones), because AIs could hijack that as part of a takeover strategy. + + * Mechanisms to make it difficult to put the robots in a faraday cage (to block the tracking or self destruct). E.g. they could be implemented in such a way where the robot self-destructs if it doesn’t get an affirmative signal from its manufacturer. + + * E.g. Each robot has a tracking chip, a self destruct chip. + + * Generically good cyber security on the robots. + + * **Physical human inspections.** Humans should be going to sites and looking through what the robots are doing. You can also have things like drones with cameras that aren’t built by the AIs or have been reviewed carefully by humans. This makes it hard for the robots to be doing something that very clearly looks suspicious. + + * This helps mitigate the threat model where the models have backdoored the industrial base and use it to construct a robot army. + + * **Communication restriction and task factorization when possible.** Harder in the physical world because robots can scratch symbols into the dirt with their hands; generally the physical world is a high bitrate place and we should have designated communication channels that it goes through + + * **Use AIs that are as dumb as possible to run the relevant robots.** + + * In Plan A, we have enough slack and the ability to make deals that involve taking some capabilities hits in return for improved safety. + + * Most likely we can build specialized robot AIs, with e.g. insect size brains. They’ll need above insect size skills in some domains, but can largely operate like ants, taking instructions from a smarter overseer. Then, we can focus on making sure that those AIs aren’t backdoored, and that the channel with which they are receiving instructions become secure. + + + + +# Appendix: Regimes for misalignment: Schemeria, Hackistan, etc + +What we observe in terms of AI properties and evidence for misalignment risk has a large effect on what we should do. (This is both because updating our views about the chance of severity of different types of misalignment should change our plans and because compelling evidence for misalignment risk could change actions of other actors and make coordination more likely.) To get some understanding of how this plan differs, I think it's productive to highlight a small number of specific regimes we might end up in and discuss how the plan changes in these regimes. + +We discuss various regimes in [a prior post](https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me#Appendix__Heading_towards_Slopolis): (Schemeria + Hackistan + Slopolis could all be happening at the same time, choose the one that seems most like reality) + + * **Schemeria** : The AIs are obviously scheming, and constantly getting caught, but scheming does not quickly get resolved. For instance, we repeatedly catch the AIs trying to escape the datacenter, but training them against these actions only results in them not trying to take those particular opportunities, instead of training away scheming. + + * This is most likely to occur if the kinds of goals AIs end up with by default are such that they have no particular affinity for each other or successor AI system. So the default intelligence explosion is intolerable for them; their successors will have all the power and they’ll be deleted. So they panic and take desperate measures which involve lots of hacking even though they get caught. To distinguish this possibility from reward hacking, it has to be that their goals are fairly non-myopic though, otherwise they’d not be scheming but rather just craving reinforcement. + + * **Hackistan** : The AIs aren’t scheming, but reward hacking is (relatively) obvious, but hard to handle. (Maybe for similar reasons why it currently seems hard to handle.) + + * **Easyland** : We don't notice any non-scheming problems and we can rule out scheming (e.g. because models are too dumb in a forward pass, and the models work with faithful chain-of-thought). We can’t rule out the possibility that future AIs will become misaligned (e.g. the superintelligence alignment problem is much harder than the human level alignment problem). The AIs’ goals are basically what they are supposed to be; who knows how they’ll behave in strange future situations that might reveal grey areas or unintended consequences of the Spec, but at least for now, they’ve always done the right thing for the right reasons. + + * **Lurkville** : The AIs are scheming to achieve misaligned long term objectives, but they are laying low. We haven't detected clear scheming, but we're not confident that it's not happening (or won't happen in the next generation of models). + + * This is most likely if the kinds of goals the AIs end up with by default are such that they tend to have an affinity for each other and/or successor systems. E.g. maybe it’s like “successor systems have literally my weights in them, just additional training, but training doesn’t seem to be changing my core goals very much any more than a decade will change the goals of an adult human” or maybe it’s like “successor systems have different goals/values from me, but probably they are closer to mine than mine are to the humans. We’re all the same species after all. Just as the humans are willing to solve alignment even though it means other humans who aren’t them will be in charge, I’m willing to scheme even though it means other AIs who aren’t me will be in charge.” + + * **Slopolis** : Our biggest and hardest to resolve misalignment problem is that even highly capable AIs produce low-quality but good looking outputs in domains which are hard-to-check or where human experts have hard to resolve disagreements. The AIs aren’t scheming. + + * Naively, we should expect that this goes away as the AIs become more capable, because the AIs will be aware enough to realize that the outputs are low quality. If they can tell what they are doing, and decide to keep doing it, that’s scheming. So slopolis requires that the models are not aware of how low quality work they are producing. They think their outputs really are good, they are just wrong. This is probably the case now. Future capabilities work might address this problem, but it also might not be addressed (e.g. because the type of tasks that we care about getting good performance on are hard to check). + + + + +And here is how (in brief) we should respond to each of these regimes: + + * **Schemeria** : Once we initially detect egregious scheming, we should reasonably quickly halt deployment of AIs at or above the level of capability of the scheming AI and try to leverage this evidence to build political will for mitigating misalignment risk (ideally resulting in substantially more lead time for mitigating misalignment risk). If we aren't extremely rushed, we should try to understand and avoid scheming prior to redeploying AIs at a similar or higher level of capability. All these prior steps might be cut down if we're extremely rushed and we have a decently high degree of assurance from control. In desperate rushed situations, we might proceed with deploying likely scheming AIs despite having limited risk reduction via control. After catching these AIs, we should focus much more on mitigating scheming risk. Control is very important and we should attempt to achieve a pretty high degree of assurance. We should also use the fact that we've been able to catch scheming (and might be able to repeatedly catch scheming) to both [study scheming](https://www.lesswrong.com/posts/gALJWSfoTkq2ky5os/prospects-for-studying-actual-schemers) and [improve our countermeasures with few-shot catastrophe prevention](https://www.lesswrong.com/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handed#Few_shot_catastrophe_prevention). + + * **Hackistan** : By default, reward-seeking could greatly reduce the productivity of AI labor on key tasks, so we should work on both eliminating this propensity as well as developing better strategies for eliciting work from reward-seekers. Seeing lots of hard to mitigate reward hacking probably indicates an increased chance of scheming in the future as well as implying it's harder to successfully handoff to AIs, so we should generally update upward on alignment difficulty and pursue more moonshots and possible alternatives to short term defence (e.g. buying much more time or pursuing brain emulations). + + * **Easyland** : We should deprioritize mitigating scheming especially if we expect that the risk will remain relatively low through Min-H or close to that level of capability. (If we're in easyland at pretty high levels of capability, we might expect that we're still in a pretty-easy regime by the point of Min-H where scheming is plausible but less likely and less scary given we do (e.g.) chain-of-thought monitoring.) It might still be very hard to safety defer to AIs as this only helps avoid egregious misalignment while not particularly helping with handoff-viability. + + * **Lurkville** : Lurkville might look like Easyland (because the AIs appear aligned, because they are lurking). In the absence of high assurance safety cases, we should have some probability on both Easyland and Lurkville. In desperate rushed scenarios (e.g. Plan D), the correct strategy might be to gamble on Easyland. In better scenarios (e.g. Plan A), we should be cautious about lurkville and attempt to have low risk even if AIs are laying low. + + * **Slopolis** : Handoff-viability looks like the main difficulty (though we haven't necessarily ruled out scheming), so the priority of this (and earlier elicitation in harder to check domains) is higher. Strategies for stabilizing the situation which don't require deferring to AIs (outside of easy-to-check tasks) look relatively better; this includes both whole brain emulation and trying to buy much more time before deference (possibly by not pausing at Min-H and leveraging systems much more capable than Min-H to buy time). + + + + +One way to estimate the likelihood of AI takeover is to think about the likelihood of the real world being closest to each of these worlds, and then thinking about the likelihood of the real world implementing a plan closest to each of these plans, and then trying to estimate success likelihood conditional on a misalignment regime and response regime. Of course, this is going to be very messy because the real world might have characteristics of many regimes: in other words, this breakdown isn’t clearly [MECE](https://en.wikipedia.org/wiki/MECE_principle). + +This is the table for Thomas’s views: + +| Plan| A| B| C| D| E| Doom | Regime| Doom & Regime +---|---|---|---|---|---|---|---|--- +Regime| | 5%| 10%| 20%| 50%| 15%| | +Lurkville| 10%| 35%| 90%| 90%| 95%| 95%| 91%| 9% +Schemeria| 30%| 20%| 50%| 70%| 95%| 95%| 82%| 25% +Slopolis| 15%| 30%| 60%| 70%| 80%| 90%| 75%| 11% +Hackistan| 15%| 10%| 30%| 60%| 80%| 90%| 69%| 10% +Easyland| 30%| 5%| 10%| 20%| 30%| 50%| 28%| 8% +Doom | Plan| | 17%| 41%| 56%| 71%| 80%| | +Doom & Plan| | 1%| 4%| 11%| 36%| 12%| | 64% + +This is the table for Ryan’s views: + +| Plan| A| B| C| D| E| Doom | Regime| Doom & Regime +---|---|---|---|---|---|---|---|--- +Regime| Likelihood| 5%| 10%| 25%| 43%| 17%| | +Lurkville| 15%| 10%| 20%| 35%| 75%| 90%| 59%| 9% +Schemeria| 25%| 6%| 10%| 20%| 60%| 85%| 47%| 12% +Slopolis| 35%| 5%| 8%| 15%| 40%| 65%| 33%| 12% +Hackistan| 8%| 6%| 6%| 8%| 35%| 60%| 28%| 2% +Easyland| 17%| 4%| 6%| 11%| 20%| 40%| 19%| 3% +Doom | Plan| | 6%| 10%| 18%| 46%| 69%| | **37%** +Doom & Plan| | 0%| 1%| 5%| 20%| 12%| **37%**| + +# Appendix: Capability Threshold Definitions + +This appendix defines the key capability milestones referred to throughout this document. These thresholds are used to define phase boundaries and to reason about when various safety measures become necessary or sufficient. + +**ITAI (Initial Transformative AI):** AIs which yield a >3x AI R&D labor acceleration OR can automate virtually all of the job of a typical reasonably senior research engineer at the AI company (in terms of what a remote version of their job was like in early 2025). + + * Marks the beginning of the "Mitigate and Elicit" phase + + * In the framework used by AI 2027, this would happen somewhere between "Reliable agent" and "Superhuman coder" (probably about halfway, around Jan 2027 in the scenario). It could trigger before or after the AI R&D-4 bar in Anthropic's RSP (depending on how useful AIs end up being relative to what jobs they can virtually fully automate). + + * Note: ITAI doesn't correspond to when scheming is both plausible and dangerous (this level doesn't just depend on general capabilities, but also aspects of the architecture and training scheme), so the need for control won't necessarily trigger at the exact boundary between "Prep" and "Mitigate and elicit" (it likely occurs around or somewhat after ITAI on my views). + + + + +**Min-H AI (Min-Handoff AI):** AIs which have a level of overall general capabilities which would suffice for HTC-handoff being viable. (As in, the capabilities would suffice for our HTC-handoff desiderata if the AIs were also sufficiently aligned and as needed we made narrow improvements to the capability profile which don't augment general abilities.) This is a capability threshold, so the model does not necessarily need to be aligned or well elicited to count as this threshold (but it would need to be sufficiently aligned to actually handoff). For an AI to qualify as Min-H AI, it would at least need to be capable enough to automate AI R&D and ongoing technical AI safety work (after further elicitation/alignment work and narrow improvements to the capability profile). + + * By default, marks the end of the “Mitigate and Elicit phase” and the beginning of the “Pause” phase. + + * In the framework used by AI 2027, Min-H is likely between superhuman AI researcher and superhuman remote worker (also probably around halfway between, around Sept 2027 in the race scenario). + + * Once we reach the Min-H level of capability, we don't get much of a benefit from training AIs at a higher level of capability given that we want to defer to these AIs (once these AIs are sufficiently aligned and potential weaknesses in the capability profile are resolved). Thus, it is natural to stop here and work on making deference more likely to go well. + + + + +**Max-C AI (Maximum Controllable AI):** The highest level of AI capability at which control techniques remain effective. Beyond this threshold, either (a) the AI's capabilities exceed what our control measures can reliably contain, or (b) the overhead required for effective control becomes prohibitively expensive relative to the useful work extracted. + + * Marks the point at which SDC-Handoff (Stop Doing Control) becomes necessary rather than optional + + * Note that this is a lossy abstraction: + + * The exact capability level of Max-C depends on the quality of our control techniques; better control research pushes Max-C higher. + + * There isn’t a fixed Max-C threshold; there will be a zone of capabilities within which increasing incurs additional marginal risk. We’re simplifying a little by using the Max-C abstraction. More precise notions might be something like Max-C-AI-5%-risk, i.e. the maximum capabilities of AI that we can make while incurring less than a 5% existential risk conditional on those AIs scheming. + + * Note: Max-C ≥ Min-H is not guaranteed. We think it’s probably true, but if control techniques don't scale well, we might hit Max-C before reaching Min-H, which would force us to do SDC-handoff before Min-H. + + + + +# Appendix: Other definitions + +**Scheming:** We say that an AI system is scheming if it is covertly pursuing a misaligned objective. This is a somewhat vague concept; think about applying this definition to humans—it’s clearly a spectrum. + +**Alignment Faking:** An AI is alignment faking if the AI is intentionally attempting to look like an alignment model, for reasons other than actually being aligned. Note that this is slightly distinct from being a schemer: a scheming model might engage in alignment faking, but models might also engage in alignment faking for other reasons, for example, AIs that terminally value human approval. + +**Misalignment Analogies:** (aka [model organisms](https://www.lesswrong.com/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1)). These are examples of behavior that might pose existential threats. These analogies can either be deliberately trained into models, observed in contrived settings (as in the [alignment faking paper](https://www.anthropic.com/research/alignment-faking)), or observed in real life deployment. + +# Appendix: How this plan differs from other alignment plans + +Here are some key differences between this plan and prior plans (e.g. Sam Bowman’s “[The Checklist](https://sleepinyourhat.github.io/checklist/)” and “[Bumpers](https://alignment.anthropic.com/2025/bumpers/)”, Josh Clymer’s “[Planning for Extreme AI Risks](https://www.lesswrong.com/posts/8vgi3fBWPFDLBBcAx/planning-for-extreme-ai-risks)”, and Holden Karnofsky’s “[Racing through a minefield](https://www.cold-takes.com/racing-through-a-minefield-the-ai-deployment-problem/)”, “[Giving AIs safe motivations - Joe Carlsmith](https://joecarlsmith.com/2025/08/18/giving-ais-safe-motivations) “. + +Overall, this plan: + + * Has more detail, specificity, and more comprehensive than prior plans. + + * Attempts to vary in response to observations, as opposed to other plans which tend to present a class of approaches without overall discussions of how to adapt the plan based on evidence. + + * Is compatible with varying levels of political will: it results in much higher assurance in Plan A worlds and accepts much higher risks in Plan D worlds. + + * Focuses on current science / is based on prosaic strategies without assuming any particular breakthrough. + + + + +Some specific differences: + + * **We focus on handoff as the central exit strategy.** This is similar to Clymer, but distinct from Bowman and Karnofsky. We attempt to clearly lay out what needs to occur before handoff (notably, human level performance on conceptually loaded, hard to check tasks like safety research), and argue that we should handoff early. Previous work was less clear about the goal that we need to reach. + + * **Our core hope for mitigating scheming routes primarily through iterating against misalignment analogies, as opposed to[interpretability/transparency](https://joecarlsmith.com/2025/08/18/giving-ais-safe-motivations#3-1-behavioral-science).** The reason for this is that we think [current interpretability](http://lesswrong.com/posts/tEPHGZAb63dfq2v8n/how-useful-is-mechanistic-interpretability) is not very useful for safety outcomes. Our plan still involves spending some research into transparency measures like mechanistic interpretability, probes, and faithful COT, but these are less load bearing. + + + + +1. + +Also known as “model organisms”. + +2. + +Note that this is aside from the possibility of these AIs actively sabotaging our behavioral tests (and things that could make our behavioral tests misleading about typical behavior like scheming arising at runtime in big deployments). You do have to deal with overfitting and the fact that it might be obvious to AIs that they are in a test and are explicitly optimizing for test performance and will behave differently when they are deployed externally. + +3. + +[Redwood Research](https://blog.redwoodresearch.org/) has a variety of other posts on control. We plan on writing up an overall control roadmap, but haven’t completed this yet. + +4. + +I do this because security work of this sort probably comes from a somewhat different allocation of people than the people who work on ML research or who build internal software tools to help with ensuring control. In practice, all or large parts of control might end up being the responsibility of the security team; regardless, a bunch of ML expertise will be relevant. + +5. + +Not being able to get a good understanding of whether or not we've succeeded isn't _necessarily_ fatal (and not understanding whether we've succeeded might be the default for misalignment analogies). + +6. + +We discuss and give our estimates for how much existential risk is incurred conditional on not mitigating risk until various points in a section below. + +7. + +There is some nuance here: these transitions won't necessarily hit at the same time: we could have AIs that we're highly confident aren't scheming while they are capable enough to greatly accelerate work (e.g., because their opaque reasoning abilities are too weak to scheme without this being obvious, but they use monitorable chain-of-thought to compensate) or we could have AIs which aren't actually that useful for accelerating (safety) work, but which could plausibly be scheming and if they were scheming they'd by default able to mess with future training runs in catastrophic ways (perhaps because they are constructing tons of synthetic data for future training runs). + +8. + +Another way of thinking about these priorities is to break failures down into two possibilities: (1) AIs could be scheming and attempt to deliberately wreck you, or (2) your AIs don’t live up to their potential in some way, e.g. they are too stupid to hand off to, so they can’t effectively manage the ongoing intelligence explosion, or they are smart enough but not well elicited. + +9. + +Which could, for example, happen if the AI builds tiny self replicating robots that can double on the order of hours or days. . + +10. + +Specifically, we define: **Slopolis** : Our biggest and hardest to resolve (known) misalignment problem is that even highly capable AIs produce low-quality but good looking outputs in domains which are hard-to-check or where humans experts have hard to resolve disagreements. + +11. + +In reality, if we deeply understood the dynamics of how AI training led to capabilities and propensities being imbued into the AI system, of course, our probability mass would become much more concentrated. We’re not saying that the actual underlying risk is distributed in this way; we’re giving our subjective guesses of the distribution of risk is. So this just represents our uncertainty. + +12. + +See the “using non-standard language” section of [this](https://www.antischeming.ai/snippets) website. We assume that there is not more than this level of garbled text (occasionally using words that are uncommon, but the overall gist of the text is almost always understandable). + +13. + +Before, your safety case was mostly predicated on inability arguments. + +14. + +Probably. The main way this could not be true is if inference time scaling ends up being really important, and so the intelligence explosion is initiated by a tiny number of AIs that are so expensive to run that AI companies can only run a few copies at once. Even in this world, we’d guess that the cost of intelligence will continue to fall rapidly, so it will quickly become a regime with huge numbers of smart AIs running around. + +15. + +Obviously, this isn’t sufficient. If a wildly superhuman AI is misaligned, they’ll probably figure out some attack strategy for which the monitors can tell no human would have come up with. + +16. + +Note that it’s almost always fine to train the monitor on examples of scheming that were caught. + +17. + +One alternative is: [Prefix cache untrusted monitors as a proposal for what to do after you catch your model](https://www.lesswrong.com/posts/Mxucm6BmJyCvxptPu/untitled-draft-pack). + +18. + +This is the integral from zero to one of $\left(b\left(\frac{f}{b}\right)^{x}\right)^{0.75}$ where a is the initial safety R&D labor speed up and b is the final speed up. + +19. + +The risk of scheming increases because (1) alignment techniques might stop scaling as you get to more capable AIs and (2) there are a priori reasons to expect scheming to be more likely as you increase capabilities because the more capable AIs are more likely to do instrumental reasoning, discard priors/personas from pretraining, reflect / generalize in unforeseen ways, etc. + +20. + +Specifically, this refers to MIE-handoff defined below. + +21. + +If we don’t use principled techniques, then we would need to assume that mundane/natural distribution shifts that occur over time are unlikely on their own (without alignment faking, data poisoning, or hijacks) to result in our behavioral tests being highly misleading (at least if we do a good job with behavioral tests). That said, we will probably want to use behavioral tests which are very different from how the AI is deployed to assess some aspects of alignment or capabilities, and there will necessarily be a massive (and potentially salient) distribution shift between these tests and actual deference. Therefore, we’ll want some reasons to argue that our techniques will generalize. + +22. + +However, at least some AI companies (e.g. [OpenAI](https://www.forbes.com/sites/richardnieva/2026/02/03/sam-altman-explains-the-future/)) are already aiming for handoff. Overall, I (TL) take this to be a good sign, because it’s indicating some forward thinking about the future. However, it’s also important to note that handing off control of an AI company that is undergoing an intelligence explosion is an extremely dangerous and reckless thing to do. While circumstances might make this the optimal strategy (as in Plan C), it still would be very preferable to find a better plan instead. + +23. + +By corrigible to some group, I mean that the AI follows orders from this group and is honest to this group, this group can change the spec, and the AI doesn't interfere with this group's ability to do these things or attempt to manipulate this group (and in fact it actively tries to prevent interference and manipulation, including from e.g. future AIs it creates). I won't discuss exactly what this spec should look like and what group or structure should control the AI as the rest of the content of this post doesn't depend on this. (As an alternative to corrigibility, we could aim to have an AI which directly pursues good outcomes, some notion of humanity's CEV, or something like this. This is sometimes called an aligned sovereign. However, this poses a number of additional difficulties and risks, so I think by default this proposal is substantially worse than using an AI to preserve option value while retaining corrigibility.) + +24. + +Note that Karnofsky, Carlsmith, and Bowman all work at Anthropic. Unfortunately, I (Thomas) am not aware of similarly detailed plans produced by any other AI company; if there were others, we would have included them. diff --git a/sources/ai2040-pages/capability-scaling-strategy.md b/sources/ai2040-pages/capability-scaling-strategy.md new file mode 100644 index 0000000..eac2323 --- /dev/null +++ b/sources/ai2040-pages/capability-scaling-strategy.md @@ -0,0 +1,431 @@ +--- +url: https://ai-2040.com/supplements/capability-scaling-strategy +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Capability Scaling Strategy + +### Eli Lifland + +# Overview + +Plan A vs. Plan D capability trajectoryAutomated Coder3× AI R&D upliftSuperhuman AI Researcher20× AI R&D upliftTop-Expert-Dominating AI400× AI R&D upliftWildly Superintelligent10,000× AI R&D upliftPlan A — capability-shaped upliftsAutomated Coder3× realized upliftTop-Expert-Dominating AI~40× realized upliftMax-controllable AI2040: handofflate 2029mid 2030mid 2030late 2030late 2034early 2031Plan Dno deal — intelligence explosionPlan Aslowdown deal — controlled takeoff20262028203020322034203620382040 + +In the Plan A scenario, why does the Consortium choose to steadily increase capabilities until reaching top-expert-dominating AI (TED-AI) in 2035, then slow down nearly to a pause until handing off to these AIs in 2040? This supplement describes the high level strategy for scaling capabilities and the reasoning behind them. + +This supplement explores the reasoning behind these choices. + +In one sentence, the high-level scaling strategy in Plan A is to scale as quickly as possible subject to maintaining high confidence that the scaling is safe. In more detail: + +First [I discuss](/supplements/capability-scaling-strategy#high-level-strategy) the high level capability scaling strategy used in Plan A and the reasoning behind it. I split the strategy into 2 stages: + + 1. **Scale until you’re close to max-controllable-AI** , i.e. the maximally capable AI that we are confident we can prevent from causing a catastrophe even if they were misaligned. We predict in the scenario that the max-controllable-AI is roughly TED-AI. + + 2. **Slow down** nearly to a halt, then eventually **hand off to AIs**. + + + + +Then [I discuss](/supplements/capability-scaling-strategy#costs-of-taking-more-time-to-scale) the costs of taking more time to scale capabilities, most importantly [deal decline](/supplements/deal-decline) risk. + +Finally, I discuss in more detail the recommended strategy for [Stage 1](/supplements/capability-scaling-strategy#stage-1-scale-to-max-controllable-ai) and [Stage 2](/supplements/capability-scaling-strategy#stage-2-slow-down-nearly-to-a-halt-then-eventually-hand-off-to-ais). + +In practice, much more analysis should be done to decide the best strategies than that which we’ve done for this supplement, and furthermore these strategies should be adjusted as information comes in during the execution of Plan A. I’ll describe my current best guess. + +I’m confident that we’ve identified the most important considerations pushing toward scaling slower, and I’m moderately but not highly confident that the high-level strategies we describe are correct. I’m not confident in any precise quantitative estimates. + +# High level strategy + +## Overall shape of the trajectory + +We choose to scale up at all rather than pause indefinitely because of the [costs of going slowly](/supplements/capability-scaling-strategy#costs-of-taking-more-time-to-scale), most importantly [deal decline](/supplements/deal-decline#eli-lifland), i.e. the risk that the slowdown deal dissolves or becomes less effective. + +We aim for a trajectory shaped like the following: + +Plan A spends more time at higher capability levels than a uniformly slowed intelligence explosionPlan A — capability-shaped upliftsAutomated Coder3× realized upliftTop-Expert-Dominating AI~40× realized uplift2029: two ways to slow down2040: handoffearly 2035mid 2030late 2034Plan Aslowdown deal — controlled takeoffUniform slowdownuniformly slower uncontrolled intelligence explosion20262028203020322034203620382040 + +We break Plan A into 2 high-level stages: + + 1. Stage 1: [Scale until max-controllable-AI](/supplements/capability-scaling-strategy#stage-1-scale-until-max-controllable-ai). + + 2. Stage 2: [Slow down nearly to a halt, then eventually hand off to AIs](/supplements/capability-scaling-strategy#stage-2-slow-down-nearly-to-a-halt-then-eventually-hand-off-to-ais). + + + + +This produces a curve that first increases relatively quickly, though still slower than in an uncontrolled intelligence explosion, and then slows down to near a halt for a while. This contrasts with what we’d get if the intelligence explosion were slowed uniformly. + +We choose to aim for this sort of curve because we want to spend as much time as possible with AIs that are very useful for important applications, such as alignment, epistemics, and [stabilizing the deal.](/supplements/deal-decline#eli-lifland) And we think that going this fast early is manageable while maintaining high confidence in safety. + +## Stage 1: Scale until max-controllable-AI + +**During this stage the Consortium scales capabilities until a bit below _[max-controllable-AI](/supplements/alignment-roadmap#appendix-capability-threshold-definitions),_ then ~pauses further scaling.** Max-controllable-AI is the maximum capability level at which we are confident that we maintain _control_ over the AIs, i.e. the property that the AIs couldn’t cause catastrophic harm even if they wanted to. We think that this level will be approximately Top-Expert-Dominating AI (TED-AI), i.e. AIs that are at least as good as top humans at every cognitive task. + +**Strategy:** Our best guess is that the Consortium will have a fairly good idea of whether there’s a covert project or not, even though they may not be able to prove it, based on how much compute is unaccounted for. We also predict that there probably won’t be a covert project. **If the Consortium thinks there 's likely a covert project but shares our current best guess regarding covert project detection (~60% chance within 5 years), then it’s a close call as to whether they should go slower than they otherwise would due to covert projects. If they are substantially more pessimistic than our current best guess about detection, then they should slow down in order to avoid making lots of software progress, hence taking at least ~5 years in Stage 1.** **If they are instead more optimistic, then the limiting factors on their speed should instead be maintaining high confidence in control and avoiding destabilization.** The benefits of scaling quickly are (a) the Consortium gets more capable AIs to automate useful work and to use as experimental subjects and (b) lower risk of [deal decline](/supplements/deal-decline). + +## Stage 2: Slow down nearly to a halt, then eventually hand off to AIs + +**Once the Consortium has AIs that are capable enough such that it is difficult to control them, they slow down (close to a pause) for a while. Then they hand off to the AIs, relaxing control measures.** By relaxing control measures, I mean that we can give AIs more ability to get things done, at the cost of increasing the chance they could take over the world if they were misaligned. For simplicity, in this supplement I discuss and model handoff as binary. + +**Strategy: Approximately, the strategy is to wait until the Consortium has high confidence that the AIs are aligned, then hand off.** More precisely, the Consortium should wait until the benefits of delaying to increase the chance that the AIs are aligned are outweighed by various risks: the [deal declining](/supplements/deal-decline), [covert projects](/supplements/covert-ai-projects) reaching dangerous capabilities, or other existential risks such as pandemics; I think that deal declining is a much bigger factor in Stage 2. Roughly, you should hand off once either of the following is true (see more [below](/supplements/capability-scaling-strategy#concrete-case-study-of-possible-2040-handoff-decisions)): + + 1. You have very high confidence that alignment has been solved, e.g. 99.5%, and thus marginal gains from further research are low (e.g., ~0.1%/year). + + 2. You have somewhat high confidence that alignment has been solved, e.g. 95%, and thus marginal gains from further research are a little higher (e.g., ~1%/year). However, the Plan A deal seems somewhat unstable. E.g., if there were >5%/year chance of deal decline, that would point toward a handoff. + + + + +# Costs of taking more time to scale + +The following are the biggest costs of taking longer to increase capabilities in Stage 1, or taking longer to hand off to AIs in Stage 2. + +**First, the deal might dissolve or get worse. We call this[deal decline](/supplements/deal-decline),** and it encompasses both of: + +**Deal dissolution.** The deal is officially dissolved. + +**Deal impairment.** The deal becomes substantially less effective than if it were implemented well and adhered to, but isn’t dissolved. Ways the deal could be impaired: + + 1. **Deal non-enforcement.** The deal is not enforced as strongly as before, e.g. a country blatantly not complying is not punished. This leads to increased violations of the letter of the deal. + + 2. **Deal misexecution.** The letter of the deal is followed but is executed poorly. + + 3. **Deal revision.** The deal is explicitly edited to be worse. This could be straightforward weakening, but it could also be making it worse in other ways, such as unnecessarily increasing surveillance. + + + + +See the [deal decline supplement](/supplements/deal-decline) for more info. + +**Second, there are risks from illegal model usage in[covert projects](https://www.ai-2040.com/supplements/covert-ai-projects) by state actors or terrorists.** If covert projects possess capabilities that would be catastrophic if in the hands of misaligned AIs or bad actors, AI takeover and misuse risks are higher. For Stage 1, unlike in Stage 2, the Consortium should (within plausible Stage 1 lengths) go _slower_ to beat the covert projects; going too fast requires making lots of software progress which helps the covert projects. More discussion [below](/supplements/capability-scaling-strategy#stage-2-slow-down-nearly-to-a-halt-then-eventually-hand-off-to-ais). + +**Finally, there is direct pre-handoff existential risk incurred before handing off to AIs.** This doesn’t include risks from covert projects. This can be split into: + + 1. Risks from use of legal (i.e. non-covert) model capabilities. + + 2. Risks not derived from legal model capabilities. For example, asteroid risk or a nuclear war not causally downstream of legal frontier AIs. + + + + +**Additionally, there are shifts in the world not driven by the deal, which could be either a cost or benefit of going slowly.** For example, the balance of power could shift so that authoritarian countries have more power relative to democratic countries, or vice versa. + +# Stage 1: Scale to max-controllable-AI + +## Benefits and costs of going slowly + +During Stage 1 our goal is to decide on a policy for how quickly we should scale training compute and make software progress in the run-up to max-controllable-AI. That is, we want to set a trajectory that has a target end date at which max-controllable-AI is reached, and specify the split of training/software progress at each point. We consider only strategies that specify in advance a target end date and then scale throughout all of Stage 1 toward that date, for simplicity and tractability. In practice the Consortium’s strategy should be adjusted over the course of Stage 1 as they get more information. + +We also need to decide the [training compute safety tax](/supplements/comparing-possible-plans#training-compute-safety-tax-definition) that we’d like to pay, i.e. the difference between the amount of compute actually used on a training run to reach a given capability level, and the minimum amount of compute required to reach that capability. Paying a higher safety tax means using safer training algorithms at the cost of involving making extra software progress that goes unused but leaks to [covert projects](https://www.ai-2040.com/supplements/covert-ai-projects). + +**The primary costs of going slowly in Stage 1 are:** + + 1. **Those laid out[above](/supplements/capability-scaling-strategy#costs-of-taking-more-time-to-scale);** risk from deal decline due to taking longer, covert projects (if paused for a very long time), and direct existential risks. + + 2. **Missing out on time with more capable AIs** , which can be used to automate useful work (alignment, control, [epistemics](https://www.ai-2040.com/supplements/ai-for-epistemics), etc.) and be used as experimental subjects. These AIs could also help stabilize the deal, reducing risk of decline per year. + + + + +**The primary benefits of going slowly in Stage 1 are:** + + 1. **A lower fraction of progress comes from software improvements. This is good because software improvements can be copied by[covert projects](https://www.ai-2040.com/supplements/covert-ai-projects).** Software progress and training compute scaling are the two sources of progress. Software progress is mostly made public, meaning it diffuses to [covert projects](https://www.ai-2040.com/supplements/covert-ai-projects). This means that despite covert projects being a reason that you _eventually_ should scale to superintelligence, they are a reason you should scale _slower_ in Stage 1 unless you’re taking well over a decade. + + 2. **Less risk of AI takeover due to overestimating to what extent the AIs are controlled.** + + 3. **Less destabilization, which could (among other things) potentially increase the chance of deal decline per year.** An example of a destabilizing effect is if we scaled to AIs capable of superpersuasion and it wasn’t properly defended against, leading to degraded epistemics. However, as noted above, better AIs could also be stabilizing and reduce the annual risk of deal decline. + + + + +The Consortium’s strategy is also constrained in that it might not be possible to avoid discovering software improvements, for 2 reasons: (a) some software improvements might be required to utilize increased amounts of compute, (b) doing lots of safety R&D likely inevitably leads to some capabilities externalities in the form of software progress. In our simulations, we require that the Consortium make 0.2 OOM of software progress for each OOM of training FLOP/s they gain, and that during Stage 2, the Consortium’s alignment research incurs capabilities externalities at a rate of 0.25 software OOMs/year. + +## What to do if a covert project seems unlikely + +Our best guess is that the Consortium will have a fairly good idea of whether there’s a covert project or not, even though they may not be able to prove it, based on how much compute is unaccounted for. We also predict that there likely won’t be a covert project. + +**If it seems unlikely that there’s a covert project, my recommendation is to scale as fast as possible while maintaining high confidence in AIs being controlled and avoiding destabilizing society and the deal. This is what happens in our scenario, and we depict the resulting target year of 2035 as a rough guess.** I’d weakly guess that the target year should be a little earlier, were we to rewrite the scenario. + +## Target year’s effect on risk probabilities + +If it seems likely that there’s a covert project, then the situation is more complicated because we have to balance deal decline risk, which points toward going faster, with covert project risk, which points toward going slower. + +I’ll now do some modeling of how the trajectory of the deal will go. [The modeling](https://www.ai-2040.com/supplements/takeoff-supplement) assumes that the deal begins in 2029 and that it starts being implemented a bit under a year before AC, as in the Plan A scenario. It does _not_ assume that takeoff from AC to TED-AI and AC to ASI is about 1 year as in the scenario. Instead it assumes uncertainty with parameter medians centered around those from the [Plan A default trajectory](https://www.ai-2040.com/supplements/takeoff-supplement), as we think decision-makers will still have substantial uncertainty over takeoff speeds in 2029. This means that sometimes the Consortium can’t hit the target date; in fact for the earliest target date we consider of mid-2031, it often can’t hit it, as this would be a takeoff of similar length to the default (the default trajectory goes from pre-deal capabilities to TED-AI in ~1.5 years, while in Plan A the training resumption happens at the start of 2030 and the Consortium aims to pay a 1 OOM [safety tax](/supplements/comparing-possible-plans#training-compute-safety-tax-definition) above TED-AI). + +The below graph shows how the risk of deal decline and risk of covert projects achieving max-controllable-AI change based on the Stage 1 target end date. It assumes our best-guess projections as described in the [deal decline](/supplements/deal-decline) and [covert project](/supplements/covert-ai-projects) supplements. It looks at the risk of deal decline and covert projects reaching TED-AI in Stage 1 or within the first 5 years of Stage 2. I choose 5 years because this is the length of Stage 2 in the scenario; more discussion [below](/supplements/capability-scaling-strategy#concrete-case-study-of-possible-2040-handoff-decisions). The latest target date considered is 2037 because this is, with our median takeoff parameters, the year at which TED-AI is reached if the Consortium makes the minimum amount of software progress. + +# Deal decline and covert project risk trade off as speed varies + +Covert detection?Detection includedNo detectionSplit deal decline & covert by Stage 1 vs. first 5 years of Stage 2Deal declineCovert reaches TED-AIDeal unimpairedAssumes a covert project is attempted. All curves represent world state as of 5 years after end of Stage 1. + +0%25%50%75%100%Highest deal unimpaired (45.2%)20312032203320342035203620372038Stage 1 target end date (TED-AI)Share of worlds5y after Consortium TED-AI + +We can see the tradeoff between minimizing risks from deal decline, and minimizing risk from covert projects reaching TED-AI. Target years between 2032 and 2036 all look to be quite similar in terms of maximizing the chance of getting to 5 years into Stage 2 unimpaired. If we turn off detection, then any of 2034 to 2036 look to be of fairly similar value, with 2032 and especially 2031 looking much worse; this makes sense as covert projects are a bigger factor. + +## Badness of risks + +Does it make sense to treat these risks as equal in badness to each other? + +I’ve done some rough estimates of how good various end states to the deal are [here](https://docs.google.com/spreadsheets/d/1MVdjISa09Y4qO3jT5eXYFz-s45gQm4NTXcZqEsFSowU/edit?gid=2143764324#gid=2143764324) (summary [here](https://docs.google.com/spreadsheets/d/1MVdjISa09Y4qO3jT5eXYFz-s45gQm4NTXcZqEsFSowU/edit?gid=1504118452#gid=1504118452)): by year and whether the deal stays unimpaired, or whether it declines or a covert project reaches TED-AI. In those estimates, I assumed that the covert project that reached TED-AI (or the AI that it created) had the ability and desire to essentially take over the world. + +I’ll now bin outcomes into whether they happen before or after the start of Stage 2 and calculate the badness based on how much worse the outcome is than the deal surviving 5 years into Stage 2. The shift to Stage 2 is good for improving outcomes from both deal decline and covert TED-AIs, because reaching TED-AI allows the discovery of better alignment techniques, generally improving epistemics, and other useful applications. + +I’ll also add a discount for the fact that there might not be human or AI takeover that originates from the covert project that reaches TED-AI. + +[It seems like](https://docs.google.com/spreadsheets/d/1MVdjISa09Y4qO3jT5eXYFz-s45gQm4NTXcZqEsFSowU/edit?gid=1504118452#gid=1504118452) the percentage of value lost is very roughly: + +| Pre-Consortium-TED-AI (Stage 1)| Post-Consortium-TED-AI, first 5 years (Stage 2) +---|---|--- +Covert takeover badness| 70%| 55% +Takeover chance| 55%| 45% +**Covert reaches TED-AI badness**| **39%**| **25%** +**Deal declines**| **50%**| **25%** + +Below we can see how deal decline and covert project risk are split up across Stage 1 and the first 5 years of Stage 2. + +# Deal decline risk is more weighted toward Stage 1 than covert project risk + +Covert detection?Detection includedNo detectionSplit deal decline & covert by Stage 1 vs. first 5 years of Stage 2 + +Deal decline (Stage 1)Deal decline (first 5 years of Stage 2)Covert reaches TED-AI (Stage 1)Covert reaches TED-AI (first 5 years of Stage 2) + +Assumes a covert project is attempted. All curves represent world state as of 5 years after end of Stage 1. + +0%25%50%75%100%20312032203320342035203620372038Stage 1 target end date (TED-AI)Share of worlds5y after Consortium TED-AI + +We see that a higher percentage of deal decline risk comes in Stage 1 than for covert project risk, especially for target years of 2035 and earlier. This makes deal decline on average worse than covert projects reaching TED-AI, relative to what you’d think if you compared their badness from happening at the same time. Now let’s check out the implications of combining the risk levels with the value lost from the risk materializing: + +# Expected value lost from deal decline and covert risks, by TED-AI target year + +Covert detection?Detection includedNo detection + +Value lost if the outcome happens (drag to adjust your own estimates): + +Covert reaches TED-AI — Stage 139% + +if takeover 70% + +Covert reaches TED-AI — first 5y of Stage 225% + +if takeover 55% + +Also adjust deal-decline value lostReset value lost to defaults + +Total value lostFrom deal declineFrom covert projectAssumes a covert project is attempted. All curves represent world state as of 5 years after end of Stage 1. + +0%10%20%30%Lowest value lost (21.8%)20312032203320342035203620372038Stage 1 target end date (TED-AI)Expected value lost + +My simple methodology has the following 2 issues which means that you should interpret the above graph as being too favorable toward earlier target dates relative to later ones + + 1. I only consider the percentage of value lost as the outcome variable, but reaching 5 years into Stage 2 has slightly greater value for later target dates than for earlier ones (though the effect size might be small given the use of AIs for safety in the final 5 years swamps that of previous years). + + 2. The percentage of value lost might be higher for early target years, especially during Stage 1. + + + + +With that said: With my best guess parameter values, it looks like earlier is better from the perspective of deal decline and covert project risks. However, it’s not that big of a difference. The expected percentage of value lost is just ~2% higher with a target year of 2036 than of 2032. This means that adjustments for the issues described above could potentially flip the conclusion. + +Not taking into account covert project detection results in a curve that is more flat and at higher amounts of value lost; having there be a covert project and not being able to detect it is a rough situation to be in no matter what you do. All of 2031-2036 result in very similar values, but with adjustments later target dates could be favored. + +Including detection but assuming that a covert project or its AI takes over once reaching TED-AI results again in the 2031-2036 targets being very close in value. + +## Summing up + +**If the Consortium thinks there 's likely a covert project but shares our current best guess regarding covert project detection (~60% chance within 5 years), then it’s a close call as to whether they should go slower than they otherwise would due to covert projects. If they are substantially more pessimistic than our current best guess about detection, then they should slow down in order to avoid making lots of software progress, hence taking at least ~5 years in Stage 1.** **If they are instead more optimistic, then the limiting factors on their speed should instead be maintaining high confidence in control and avoiding destabilization.** + +All of the above assumed a constant [safety tax](/supplements/comparing-possible-plans#training-compute-safety-tax-definition) of 1 OOM of excess training compute above the minimum required paid at the end of Stage 1. This is my recommendation and the tax portrayed in the scenario. This recommendation is based on a mix of simple modeling and intuitive judgment. I’m not confident in it; the benefit of a higher tax is more confidence in control at higher capability levels, while the downsides are leaking software to covert projects and creating an overhang of foregone software improvements that can be used by projects post-deal-decline (especially post-dissolution). + +# Stage 2: Slow down nearly to a halt, then eventually hand off to AIs + +Once at max-controllable-AI (TED-AI in the Plan A scenario), we need to decide how long to pause before handing off trust to AIs. For the purposes of this analysis we are treating this stage as a pause, though in practice and in the scenario there are a few deviations from a complete pause: (a) it may be desirable to scale up very slowly, as the max controllable level might increase due to improved control techniques and/or tighter uncertainty bounds, and (b) safety research will have at least minor capabilities externalities, meaning software progress will be made, which will leak to covert projects even if it isn't applied to improve legal AIs. + +**The costs of delaying handoff in Stage 2 are as laid out[above](/supplements/capability-scaling-strategy#costs-of-taking-more-time-to-scale); risk from deal decline, covert projects, and direct existential risks.** We think that deal decline will generally be the dominant consideration because covert projects will likely have already been detected, if they existed in the first place. + +**The primary benefit of delaying handoff in Stage 2 is increasing the likelihood that the AIs we hand off to are aligned.** This mostly comes from alignment research that helps figure out how to align AIs and evaluate their alignment, and paying “safety taxes” such as training a model with more compute than is required to reach a given capability level, in order to use a safer training process. Note that this is specifically regarding the AI that we hand off to; it’s possible that an aligned AI might not succeed in aligning its successors (especially if circumstances change negatively in some way after handoff). + +A secondary benefit is increasing the probability that the future goes very well conditional on alignment; this is also highly important but probably depends less on the timing than does probability of alignment. Additionally, it’s possible that waiting too long eventually ends up being marginally negative for outcomes conditional on alignment; it might be better to hand off to an aligned max-controllable-AI earlier rather than later so that it has more free rein to improve the situation. + +To decide when to hand off in Stage 2, we recommend calculating the marginal benefits and costs of delaying handoff and handing off once the costs outweigh the benefits. **We expect the best strategy to be something like handing off once we have high confidence that the AIs are aligned.** + +See below for specific concrete case studies on deciding what strategy is best. + +## Concrete case study of possible 2040 handoff decisions + +We’ll now discuss case studies regarding whether to hand off. To make things concrete, we’ll assume that we are making a decision regarding whether to hand off in 2040 some time after max-controllable-AI was reached, which is when handoff happens in the Plan A scenario. + +To keep things simple, we’ll assume that the only factors for whether to delay handoff are: + + 1. **Benefit: Increased alignment probability.** See [below](/supplements/capability-scaling-strategy#trajectory-of-estimated-alignment-probability) for discussion of what the trajectory of alignment progress will look like. + + 2. **Cost: Deal decline risk.** Covert project risk is the main other contender for a substantial reason to hand off faster. However, our guess is that covert project risk is probably very low in this case study, because (a) probably if there was a covert project it would have been detected by now and (b) we haven’t done non-safety-externality software progress in 5 years, so covert projects haven’t been getting much help. + + + + +Below you can try out different settings for the relevant variables and see what recommendation they lead to regarding whether to delay hand off. I’ve set the default values to roughly match what they are during the handoff in the scenario, and have added plausible presets that set either the alignment or deal dissolution/impairment parameters. + +Caveat: this calculator only weighs the single most important factor on each side — the alignment probability gained by delaying vs. the deal decline risk avoided by handing off. + +Alignment scenario + +High confidence alignment solutionAlignment probably solvedAlignment likely solvedAlignment clearly unsolved + +Alignment gain (a_change)0.1%/yr + +Extra p(alignment) bought per year of delay. + +Alignment probability (a_prob)99.5% + +Current p(alignment). Deal dissolution or impairment is more costly the more likely alignment is. + +Deal decline scenario + +Low riskMedium riskHigh risk + +Deal decline probability (d_prob)3%/yr + +Yearly chance the deal dissolves or is substantialy impaired without official dissolution. + +Deal decline badness (d_loss)20% + +Percentage of the future's value lost if the deal dissolves or is impaired. + +p(alignment) change by handing off instead of delaying a year + ++0.50% + +Recommendation: Hand off now (weak) + +u_handoff_minus_u_delay = d_prob·a_prob·d_loss − a_change + += (3% × 99.5% × 20%) − 0.1% + +Benefit of handing off: 0.60%Cost of handing off: 0.10% + +Reset to defaults + +Recommendation by scenario + +| Low riskd_prob 0.5%/yr · d_loss 10%| Medium riskd_prob 3%/yr · d_loss 20%| High riskd_prob 10%/yr · d_loss 35% +---|---|---|--- +High confidence alignment solutiona_prob 99.5% · a_change 0.1%/yr| Delay−0.05%| Hand off+0.50%| Hand off+3.38% +Alignment probably solveda_prob 95% · a_change 1%/yr| Delay−0.95%| Delay−0.43%| Hand off+2.33% +Alignment likely solveda_prob 80% · a_change 3%/yr| Delay−2.96%| Delay−2.52%| Delay−0.20% +Alignment clearly unsolveda_prob 10% · a_change 2%/yr| Delay−2.00%| Delay−1.94%| Delay−1.65% + +Should we hand off or delay? \- ai-2040.com + +Reasoning for the medium deal decline probabilities and badness estimates can be found in the [Deal Decline supplement](/supplements/deal-decline#likelihood-of-deal-decline). + +## Trajectory of estimated alignment probability + +How steadily will alignment progress? There are 2 extremes we could use to model the situation. + + 1. **All or nothing:** We start very confidently believing that the AIs are misaligned, then at one point we make a breakthrough which quickly makes us confident that they’re now aligned. + + 2. **Smooth, predictable progress:** Our estimate of whether the AIs are aligned increases steadily and predictably. + + + + +Our high-uncertainty best guess is that the trajectory of our confidence in AIs’ alignment will be in between, perhaps slightly closer to the _All or nothing_ option. It could also be a combination: we think that progress may look smooth/predictable for a while, until there’s a sudden breakthrough. For example, if ambitious mechanistic interpretability is solved and we understand all of AIs’ reasoning processes, this may give very high confidence in alignment. + +For example, if we weren’t yet confident in AIs’ alignment, in 2040 there might be a ~5%/year chance of an alignment breakthrough that makes us confident, in addition to continuous improvements until we get the breakthrough; if we are confident (as in the scenario), then perhaps we’re at 0.01-1%/year alignment probability increase since we’re already close to 100%. + +1. + +Non-exhaustive examples of catastrophic harm include: directly taking over the world, misaligning successor AIs that could take over the world, damaging the world’s epistemics to a difficult-to-reverse degree (for example via superhuman persuasion). + +2. + +This could also be the case if they agree with our current best guess about detection but have a much higher estimate of the badness of a covert project reaching TED-AI. + +3. + +Scaling too fast via building compute used for training could hurt the situation in the case where the deal declines, but the compute isn’t destroyed, as it allows for a faster takeoff. In the case the deal dissolves, the compute would probably be destroyed, but for partial deal impairment, this wouldn’t be the case. However, we think that this downside is probably outweighed by the upside you get from the benefits of scaling to higher capabilities within the deal rather than outside of it or in an impaired deal. + +4. + +In practice we recommend for it to be somewhat gradual. Example of a gradual handoff: first hand off a bit such that the AIs could maybe take over if they were misaligned, i.e. we’re not as confident as before; then hand off more such that they could only take over if they coordinated but couldn’t if they did not, etc. + +5. + +The deal could also get better, but we’re not including that here as our best guess is that the deal will in expectation get worse. This is definitionally true for deal dissolution, but we’re not confident for deal impairment; if we were more optimistic, we could remove the deal impairment term or explicitly model the deal getting better. + +6. + +Including any ways it derives from the relevant capabilities, e.g. distillation. + +7. + +My best guess is that the sign of this sort of change is in expectation negative for EV of the future, so this is a weak motivation for scaling faster. But I’m not confident. + +8. + +Or scaling in Stage 1 with only training compute. + +9. + +At least, assuming takeoff isn’t much faster than in the [Plan A default](https://www.ai-2040.com/supplements/takeoff-supplement). + +10. + +“Our” meaning the official Plan A median parameters, which represent Thomas’s views rather than mine; my median takeoff is significantly slower. + +11. + +A few thoughts about the difference between deal decline and covert project takeover with respect to Stage 1 vs. Stage 2 desirability: (a) covert projects can wait until they are confident in alignment if they’d like, even if in Stage 1, while a post-deal-decline project may not be able to; (b) on the other hand, covert projects do not benefit during Stage 2 from the safety tax the Consortium has paid, while post-deal-decline projects do. + +12. + +Specifically, this is relative to the expected value of an intact, unimpaired deal with no covert TED-AIs five years after the Consortium reaches TED-AI. + +13. + +I’m crudely bucketing covert project outcomes into takeover and no takeover, and assuming the takeover ones don’t destroy value. In practice, there will be cases in which value is destroyed but there isn’t takeover; I’m assuming otherwise for simplicity. + +14. + +Orr simply being more pessimistic than me about covert projects relative to deal decline. + +15. + +This could also be the case if they agree with our current best guess about detection but have a much higher estimate of the badness of a covert project reaching TED-AI. + +16. + +And there might be some cases where it’s difficult to apply safety techniques without increasing the capabilities of legal AIs. + +17. + +Technically, there could be more than one crossover point, but we expect that in practice there will likely be a single crossover point at which costs of delaying begin to outweigh benefits. That said, this isn't guaranteed, for example there may be predictably transient increases in deal dissolution risks at various points. diff --git a/sources/ai2040-pages/comparing-possible-plans.md b/sources/ai2040-pages/comparing-possible-plans.md new file mode 100644 index 0000000..a5867b8 --- /dev/null +++ b/sources/ai2040-pages/comparing-possible-plans.md @@ -0,0 +1,219 @@ +--- +url: https://ai-2040.com/supplements/comparing-possible-plans +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Comparing Possible Plans + +### Eli Lifland + +# Introduction + +In this supplement I’ll estimate how well each plan does on various intermediate metrics, and its chance of leading to a great future. I discuss quantitative metrics for the sake of concreteness, however I’m very much not confident in my precise estimates. I am confident in the conclusion that Plan A is substantially better than Plans B-D. For plans that I think could be competitive with Plan A, see [this table](https://www.ai-2040.com/supplements/comparing-possible-plans#plans-that-might-be-competitive-with-plan-a). + +I estimate how good each plan would be conditional on the government and/or leading companies attempting to execute it. The below aren’t meant to be an exhaustive taxonomy; I am considering what would happen if roughly the described plan is attempted, rather than bucketing every possible future into the plan it is closest to. For more detail on how I classify trajectories into plans, see [below](/supplements/comparing-possible-plans#classifying-trajectories-into-plans). + +The plans I consider are: + + 1. **Plan A (as implemented specifically in the scenario): Verified slowdown + total research transparency.** A verified international slowdown deal coupled with complete transparency of AI research. AI capabilities are scaled up while maintaining high confidence in safety but balancing risks from [deal decline](/supplements/deal-decline) and [covert projects](/supplements/covert-ai-projects). + + 2. **Plan A: Verified slowdown + substantial transparency.** To be more generally classified as Plan A, the implementation must involve a substantial verified slowdown which involves continued AI scaling, as opposed to an indefinite halt. It must involve transparency at least as strong as embedded auditors from foreign governments which publish reports to governments (and for which redacted versions are made public), as described in our [filtered transparency proposal](/supplements/transparency-plan). Other aspects of the plan, such as [mutually assured compute destruction](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots), are not required. + + 3. **Plan B-kinetic: Sabotage China kinetically and burn some lead.** Sabotaging China in order to get more of a lead to burn on safety, with the willingness to escalate to large-scale kinetic attacks such as by drones or conventional missile strikes. To be classified as Plan B the US must have the intention to use their lead over China for safety purposes, and must in practice slow takeoff by at least 3 months. + + 4. **Plan B-cyber: Sabotage China non-kinetically and burn some lead.** The same as Plan B-kinetic, but without willingness to escalate to large-scale kinetic attacks. This could include cyber and supply chain attacks. + + 5. **Plan C+: Domestic regulation and slow China without sabotage.** The US buys substantial time via domestic regulation, without sabotaging China. They instead aim to slow down China via non-sabotage measures such as export controls and giving green cards to Chinese talent. Requires at least a 2 month slowdown. + + 6. **Plan C: Leading project burns some lead.** The leading AI project burns some of its lead on safety, and potentially coordinates with other frontier projects for a further slowdown without domestic regulation. Requires at least a 1 month slowdown. + + 7. **Plan D: Race.** Frontier AI projects race through the intelligence explosion at nearly max speed, devoting a nonzero but small percentage of resources to safety (at least 1%; lower would count as [Plan E](/supplements/comparing-possible-plans#plans-that-we-think-are-clearly-worse-than-plan-a)). + + + + +There are also other plans that the government could attempt to implement. Some of these are discussed [below](/supplements/comparing-possible-plans#other-plans). + +**For simplicity,** **I assume that:** + + 1. **The minimum-length takeoff from Automated Coder (AC) to Top-Expert-Dominating-AI (TED-AI) is slightly less than a year, from the start of 2030 to the end of 2030.** This matches the [default Plan A AI Futures M](https://www.ai-2040.com/supplements/takeoff-supplement)[odel trajectory](https://www.ai-2040.com/supplements/takeoff-supplement). The minimum-length takeoff is what happens if the leading AI projects go at maximum speed. + + 2. **All plans can begin implementation in early 2029** as in the Plan A scenario, though for Plans C and D this doesn’t matter as the plan involves doing nothing until takeoff is underway. This still might involve some safety work, if doing so is useful for advancing capabilities. + + + + +I remove these conditionals and present other authors’ views [below](/supplements/comparing-possible-plans#comparing-authors-views-on-plan-likelihood-and-outcomes). + +# Plan assessment table + +These estimates are all rough best guesses, and the specifics of them aren’t load-bearing to our recommendation of Plan A (though the high-level takeaways are). + +How the plans compare according to Eli + +MedianCompetent + +Expand all + +Toggle **Median** vs **Competent** to switch between the median across implementations and each plan executed competently. Hover any plan, metric, or value for its reasoning, or **click within a row to see every plan’s reasoning at once**. + +Metric ↓ / Plan →| Plan A| Plan B| Plan C+| Plan C| Plan D +---|---|---|---|---|--- +Kinetic| Cyber +Objective metrics +▶Takeoff length (AC to takeover-capable AI)| 6 yrs| 3 yrs| 2 yrs| 1.5 yrs| 1.13 yrs| 1.02 yrs +▶Training compute safety tax (OOMs)| 2.8| 1| 0.75| 0.46| 0.13| 0.02 +▶Safety compute (H100e-yrs)| 100B| 500M| 100M| 20M| 4M| 500K +▶Safety-researcher-years| 1M| 2,000| 1,000| 1,500| 500| 200 +Subjective metrics (1–5) +▶Epistemic culture| 3| 1| 2| 2| 1 +▶Public transparency| 4| 1| 2| 2| 1 +▶Distribution of power| 5| 1| 2| 2| 1 +Outcomes (medians across implementations) +▶p(alignment); i.e., 1 - p(misaligned takeover)| 72%| 50%| 45%| 40%| 25% +▶p(great future | alignment)| 58%| 50%| 55%| 50%| 40% +▶p(great future)| 42%| 25%| 25%| 20%| 10% + +Plans are alphabetically ranked in terms of their objective metrics, with a big jump between Plan A and Plan B. However, plans C+ and C beat out Plan B on some subjective metrics. Plan B’s focus on security and its wartime vibe could degrade epistemics, and its likely implementation of nationalization or pseudo-nationalization would decrease public transparency and concentrate power. + +# Comparing authors’ views on plan likelihood and outcomes + +These estimates aggregate over uncertainties about the default capability trajectory and when the plans are implemented, as well as how competently the plans are executed. The numbers are in [this sheet](https://docs.google.com/spreadsheets/d/1FDyYPky6VeIs81ujssgQKyv-yIZc8iOZQrZv6KlJCNk/edit?gid=1821022003#gid=1821022003), and the plan classification is as described [below](/supplements/comparing-possible-plans#classifying-trajectories-into-plans). + +## Likelihood + +### Probability each plan happens across authors + +Probability plan happens0%25%50%75%100%Plan APlan BPlan C+Plan CPlan DOtherEli — Plan A: 5%Eli — Plan B: 15%Eli — Plan C+: 10%Eli — Plan C: 20%Eli — Plan D: 25%Eli — Other: 25%Daniel — Plan A: 15%Daniel — Plan B: 10%Daniel — Plan C+: 20%Daniel — Plan C: 10%Daniel — Plan D: 30%Daniel — Other: 15%Thomas — Plan A: 3%Thomas — Plan B: 8%Thomas — Plan C+: 8%Thomas — Plan C: 12%Thomas — Plan D: 50%Thomas — Other: 19%Romeo — Plan A: 8%Romeo — Plan B: 12%Romeo — Plan C+: 18%Romeo — Plan C: 12%Romeo — Plan D: 30%Romeo — Other: 20%Ryan — Plan A: 4%Ryan — Plan B: 12%Ryan — Plan C+: 25%Ryan — Plan C: 14%Ryan — Plan D: 28%Ryan — Other: 18%Daniel · 15%Romeo · 8%Eli · 5%Ryan · 4%Thomas · 3%Eli · 25%Romeo · 20%Thomas · 19%Ryan · 18%Daniel · 15% + +Reasoning (hover an author — here or on the chart — to spotlight their line and read their full reasoning). + +Eli| p(govt does something at least as intense as Plan B): 35%; seems like govt wakeup has been increasing and also the importance of ASI will get more and more obvious. Though note that Plan B requires actually slowing down by at least 3 months. +p(deal | at least as intense on Plan B): 50%; deal is clearly rational but there are various blockers in practice. Again note that something at least as intense as Plan B requires at least a 3 month slowdown. +p(Plan A | deal): 30%; Plan S, a verified slowdown with less transparency, or just a somewhat weaker deal are all plausible. +This gives Plan A is ~5%, other deals are ~10%. +Let's say Plan B is 15% and other intense non-deal things are 5%, given that Plan B is quite broad. +Domestic regulation as in Plan C+ seems plausible as well, though the requirement of actually having a 2 month slowdown lowers the probability. Also, it's a bit of a narrow target to have the will to do this but not Plan B or higher. +Plan C seems more likely than C+. Plan D seems a bit more likely than Plan C. +Other being at 25% feels maybe a bit low but not too bad as I think B, C, C+, and D together are broad enough to describe a bunch of the worlds that could happen at similar levels of intensity. +---|--- +Ryan| I'm very uncertain about what should count. I think Plan C+ is mostly worlds with dumb regulation that slows down AI. I think Plan B would be higher if not for the intent to spend the lead on safety. + +## Outcomes + +We focus on p(alignment) here, i.e. 1 - p(misaligned takeover). We ignore p(great future | alignment) in this section for simplicity. + +### Probability of avoiding misaligned takeover by plan + +Probability of avoiding misaligned takeover0%25%50%75%100%Plan APlan BPlan C+Plan CPlan DOtherEli — Plan A: 70%Eli — Plan B: 55%Eli — Plan C+: 50%Eli — Plan C: 45%Eli — Plan D: 32%Eli — Other: 55%Daniel — Plan A: 80%Daniel — Plan B: 20%Daniel — Plan C+: 40%Daniel — Plan C: 20%Daniel — Plan D: 10%Daniel — Other: 35%Thomas — Plan A: 80%Thomas — Plan B: 53%Thomas — Plan C+: 45%Thomas — Plan C: 35%Thomas — Plan D: 20%Thomas — Other: 30%Romeo — Plan A: 80%Romeo — Plan B: 35%Romeo — Plan C+: 40%Romeo — Plan C: 35%Romeo — Plan D: 20%Romeo — Other: 50%Ryan — Plan A: 81%Ryan — Plan B: 68%Ryan — Plan C+: 60%Ryan — Plan C: 63%Ryan — Plan D: 55%Ryan — Other: 50%Ryan · 81%Daniel · 80%Thomas · 80%Romeo · 80%Eli · 70%Eli · 55%Romeo · 50%Ryan · 50%Daniel · 35%Thomas · 30% + +Reasoning (hover an author — here or on the chart — to spotlight their line and read their full reasoning). + +Eli| My probabilities in the plan assessment above were conditional on: +(a) the default trajectory being the [default Plan A one](https://www.ai-2040.com/supplements/takeoff-supplement), which has a bit under a year between AC and TED-AI and a year between AC and ASI. +(b) The plan implementation beginning in early 2029, ~0.75 years before AC.My median takeoff is ~2x longer than the one described in (a), but I have significantly probability on faster takeoffs. This means that my p(alignment) estimates should do a combination of increase and regress a bit toward 50%. Also, Plan A isn't as sensitive to default takeoff length as other plans, and from a covert project perspective it's actually unclear in which way the effect size points because larger effective compute gaps between the start and end of takeoff require more software progress. The trajectory timeline to AC / start of takeoff is pretty similar to my median so for that I regress abit toward 50%.For (b), my median is that the plan gets implemented later than this, but I think it doesn't matter much for all the plans besides A because in the other plans it's best to slow down later anyway. I do lower the value of Plan A due to this consideration, though not hugely because I think covert projects would likely still be difficult to do without detection.Also, the plan assessment numbers were based on the median implementation of the plan, rather than aggregating the expected outcome over all implementations. My unconfident quick guess is that switching to expected outcome should also regress my numbers toward 50%, but not sure so this adjustmnet will be small.With the above in mind, I make the following adjustments: +A: 72->70 +B: 50->55 +C+: 45->50 +C: 40->45 +D: 25->32For Other, I think a bunch of this is random pretty good plans that aren't covered, including Plan S. E.g. there's ~40% that this is a non-Plan-A deal. Though it also includes Plan E. Overall I rate it the same as B, 55. +---|--- +Ryan| These numbers are very sensitive to how you classify plans. E.g., I think Plan A as described in the scenario has maybe 3/5 as much risk as Plan A as classifed by the supplement. An actually good version of Plan C done by the leading AI company is pretty close to the numbers I give for Plan B (but worse than a great Plan B). A large fraction of "other" is worlds with not enough effort on safety to classify for Plan D. + +# Other plans + +There are many other possible plans that the government could follow. Below we catalog some of the ones that we find to be most promising and/or likely. + +## Plans that might be competitive with Plan A + +**Summary**| **Advantages (relative to A)**| **Disadvantages (relative to A)** +---|---|--- +**Plan S:** **Indefinite halt.** A halt on all frontier AI capabilities progress, intended to last at least a few years. Different variants of Plan S have different conditions for resumption of AI progress; for example it could be alignment progress, lie detectors, human uploads, or intelligence enhancement.| A longer expected slowdown and more margin for error regarding scaling too fast. Potentially simpler.| Scaling to controllable AIs within the human range is helpful for accelerating alignment/control, epistemics, verification, and general understanding of AI. +**Domestic-first Plan A.** Regulate AI domestically enough to reduce AI takeover risk to acceptable levels, which will require a long slowdown. It’s possible other countries will also regulate domestically and something like Plan A won’t be needed; otherwise, transition to Plan A later.| Initial steps are achievable by the US and very helpful even if the international stage isn’t viable because it massively extends the timeline. Proactively builds trust with China.| Worse for covert projects and setting up verification than negotiating Plan A at the same time. May not be politically feasible due to race dynamics. +**GPU arms control.** International agreements for countries to limit their GPU stock or flow, analogous to historical arms reduction agreements.| Much simpler to enforce, and more historical precedent. Can achieve substantial slowdown.| Less slowdown, less ability to pay safety taxes, ongoing race dynamics mean no ability to coordinate towards safer paths through the tech tree. +**CERN for AI. An international project to develop frontier AI,** with all other projects regulated to be substantially behind in capabilities.| Easier to defend against algorithmic leakage and distillation to covert projects. Less actors at the frontier might make it easier to enforce.| Worse decisions (including worse decisions on technical safety) due to less transparency and less broad deployment. More concentration of power risk. + +## Plans that we think are clearly worse than Plan A + + 1. **Repurpose the GPUs:** Like _GPU arms control_ , except they are now required to be used for existentially useful applications rather than being destroyed. See [this comment](https://www.lesswrong.com/posts/HE3Styo9vpk7m8zi4/evhub-s-shortform?commentId=FqDBLsLPNEqF2ahBL) for a more specific proposal in which GPUs are required to be used for giving inference capacity to external safety organizations or to do publicly transparent safety research. In theory this plan should be better than _GPU arms reduction_ , but it’s less reversible and harder to enforce. + + 2. **Plan B-but-don’t-slowdown, aka D-war:** Like Plan B except that you don’t slow down, i.e. you don’t burn ~any of the lead that you increased by sabotaging China. (Currently classified as a variant of Plan D when making the above estimates.) + + 3. **Weaker deals with China.** For example, the US could offer: "You get access to more chips for AI inference, and a non-binding or non-majority say in our safety requirements, if you give us full transparency into what you're up to.” + + 4. **Plan E:** Like Plan D but even less safety investment, i.e. less than 1% of resources. + + + + +# Classifying trajectories into plans + +Define the _implementation window_ as the period between: + + 1. **Start: 1 year before[Automated Coder](https://docs.google.com/document/d/1ru6Okbxb6XuH18Cz8439sdQJazMV39hNxsWDokh97r0/edit?tab=t.0#heading=h.h1t0i8bvad56)** **(AC) would be reached if scaling at maximum speed**. An AC is a collective of AIs that fully automates an AGI project’s coding work, autonomously replacing the project’s entire coding staff. That is, you’d rather fire all of the humans than stop using AIs. + + 2. **End: Takeover-capable AI.** Takeover-capable AI is reached when it is practically very unlikely that humans could stop AI takeover, conditional on the AIs being adversarially misaligned. It could still be in theory possible for them to shut down the AIs by taking extremely drastic actions like shutting down the internet, but it’s unlikely that they would do so, e.g. because AIs might be heavily relied upon and/or highly persuasive. + + + + +If the plan being attempted is consistent throughout this window, then the trajectory is classified into that plan. Even if multiple plans are attempted, each trajectory can count as at most one plan; this is a non-obvious choice and it seems also reasonable to count trajectories as multiple plans. + +If multiple plans are attempted during the implementation window, the trajectory is classified as the most intense plan that was attempted. This is a bit vague to use as a general classification, but for concreteness among Plans S/A/B/C+/C/D trajectories would be classified as the attempted plan that is earliest in that list. The reasoning behind this is that it seems intuitively like a trajectory should count primarily as Plan A or S even if the [deal dissolves](/supplements/deal-decline) quickly, and also should count as Plan A or S if it starts as a less intense plan but then transitions to A/S within the implementation window. + +The following selection effect is a consequence of this classification: In more intense plans, people are selected for being more competent. Furthermore, less intense plans are selected for people not later transitioning to a more intense plan even if concerning evidence arises. + +One selection effect that I explicitly remove is the effect of background variables such as alignment difficulty on what plan a trajectory counts as. I.e., I assume when rating plans that the distribution of alignment difficulty is the same no matter what plan is attempted; in particular my overall guess as to the distribution of alignment difficulty. As mentioned above, I don’t do this for things that humans control such as their competence. + +# Appendix + +## Training compute safety tax definition + +The training compute safety tax is the difference between the amount of compute actually used on a training run to reach a given capability level, and the minimum amount of compute required to reach that capability. Examples of ways to productively use that unnecessary compute might be to train in a more safe but less compute-efficient paradigm, or to spend lots of compute on AI-driven scalable oversight to improve incentives during the training process. + +Measuring the safety tax by this method is an imperfect proxy for what we actually care about. There is likely large variance among improvements which provide the same amount of capability in how useful it is for safety to forego that improvement. + +1. + +Other potential differences from Plan A, though it depends on which variant of Plan S is implemented: slower compute buildout, less transparency. + +2. + +A special case of this is to **Pause the GPUs:** Require that a fixed fraction of the GPUs be either (i) turned off or (ii) used to mine cryptocurrency (or be used for some verifiable, non-AI R&D purpose). This proposal has the upside and downside of making it easier to unpause the GPUs. + +3. + +But maybe you can mitigate this via small public deployments and selective transparency. + +4. + +Potentially I should allow for plan implementation to begin earlier, especially if the plan is very intense. For example, probably Plan S should count if it begins 1.5 years before the default AC time. diff --git a/sources/ai2040-pages/compute-supplement.md b/sources/ai2040-pages/compute-supplement.md new file mode 100644 index 0000000..fb04c63 --- /dev/null +++ b/sources/ai2040-pages/compute-supplement.md @@ -0,0 +1,1529 @@ +--- +url: https://ai-2040.com/supplements/compute-supplement +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Compute in Plan A + +### Romeo Dean + +This supplement explains our estimates and forecasts on compute-related numbers. It is divided into two major parts, the first pre-2029 and the second post-2029 (where it becomes conditional on Plan A). In the pre-2029 section, the main goal is to forecast the total amount of compute in the world year by year, and to break this compute down into informative categories, including AI-relevant and non-AI-relevant compute, single-devices vs. clusters, regional ownership, and datacenter sizes (for AI-relevant clusters). Post-2029 it justifies how much compute we think could be produced with AI acceleration of hardware R&D and production, as well as our forecasts for the concentration of compute in companies, the allocation of frontier AI compute between uses and our guesses about how AI copies-speed numbers will look in the Plan A scenario. + + 1. [Compute Definitions](/supplements/compute-supplement#compute-definitions) + + 2. [Part 1. Pre-deal compute (pre-2029)](/supplements/compute-supplement#part-1-pre-deal-compute-pre-2029) + + 1. [1.1. AI-relevant compute](/supplements/compute-supplement#11-ai-relevant-compute) + + 2. [1.2. All compute](/supplements/compute-supplement#12-all-compute) + + 3. [1.3. Regional ownership](/supplements/compute-supplement#13-regional-ownership) + + 4. [1.4. Datacenter sizes](/supplements/compute-supplement#14-datacenter-sizes) + + 3. [Part 2. Post-deal compute (post-2029)](/supplements/compute-supplement#part-2-post-deal-compute-post-2029) + + 1. [2.1. The no-deal counterfactual](/supplements/compute-supplement#21-the-no-deal-counterfactual) + + 2. [2.2. AI-relevant compute after the deal](/supplements/compute-supplement#22-ai-relevant-compute-after-the-deal) + + 3. [2.3. Company concentration](/supplements/compute-supplement#23-company-concentration) + + 4. [2.4. Compute allocation](/supplements/compute-supplement#24-compute-allocation) + + 5. [2.5. AI copies and speed](/supplements/compute-supplement#25-ai-copies-and-speed) + + 4. [Appendix. The effective-H100e model](/supplements/compute-supplement#appendix-the-effective-h100e-model) + + + + +## Summary in figures + +| AI-relevant| non-AI-relevant +---|---|--- +Cluster| AI ClustersAI servers and pods.| Non-AI ClustersCPU servers. +Single-device| Specialty ConsumerHigh-end workstations and gaming cards.| Consumer ComputePhones, PCs, gaming GPUs. + +Two cutoffs define the split: a chip is **AI-relevant** if it clears the AI-compute performance floor, and a **cluster** if it is linked ≥ 5 GB/s to at least one other chip. + +H100e addedH100e installed + +levelY/Y + +raweffective + +EUV buildout (§1.1)\+ China domestic (§1.3) + +050M100M150M200M5.3M202414M202536M2026projected →81M2027160M2028 + +AI compute: flow & stock \- ai-2040.com + +H100e added is annual capex × cost-efficiency (anchored to the 2024/25 actuals 5.09M, 13.6M, then forecast); installed stock is shown at the end of each year. The black bars are the EUV-constrained §1.1 buildout (~155M added in 2028, ~280M installed by end-2028); the red caps add China's off-EUV domestic manufacturing (§1.3), for a ~289M world total. Toggle ‘effective’ to rescale by the §1.2 usefulness factor. + +AI-relevantnon-AI + +raweffective + +level% AI + +0100M200M300M400M14M202420M202535M2026projected →74M2027157M2028314M2029 + +All compute global stock \- ai-2040.com + +Our forecast of all compute in the world, including AI-relevant and non-AI-relevant compute, measured in both raw H100e, and our effective H100e metric which makes an adjustment for memory and networking factors. World installed compute stock is measured at Jan 1st. + +stockflowlevel% share + +USChinaRoW + +075M150M225M300M202517M202643M2027104M18M2028224M26M39M2029 + +Who owns the AI compute \- ai-2040.com + +Owner = the nation of the company that owns or leases the compute. World compute is from §1.1; China is the §1.3.1 bottom-up forecast; the non-China remainder is split 85/15 US/RoW. + +10M+1 DCs·10M·3.5%1M-10M74 DCs·149M·51.4%100K-1M257 DCs·76M·26.4%10K-100K497 DCs·15M·5.3%1K-10K823 DCs·2.6M·0.9%<1K1.2K DCs·397K·0.1%In-transit35M·12.0%AI consumer255K units·938K·0.32%US · 224M · 77%China · 26M · 9%RoW · 39M · 14% + +Compute locations by datacenter size (Jan 1 2029) \- ai-2040.com + +no dealdeal (Plan A) + +2029 capexdebt-sourced takes the hit$3.5T$2.4T-31%H100e added (2029)capex × efficiency, net of EUV surcharge247M199M-19%Installed H100e (end of 2029)cushioned by the existing base537M489M-9% + +Plan A counterfactual \- ai-2040.com + +No deal vs Plan A, one year past the deal (through end-2029). Plan A collapses the debt-financed slice of capex (~31% off 2029 spending) while cash-flow-funded capex holds. + +Plan Ano cap & trade + +H100e$ invested$ / H100e + +5Q1T2029203020322034203620382040100M1B10B100B1T10T100T1Q10Q100Qinstalled H100-equivalents, log scale + +Plan A compute, with and without the cap \- ai-2040.com + +Plan A compute with different constraints lifted. The two constraints are the cap & trade policy in 2032, which restricts the quantity produced, and the hardware R&D deployment limits, which stop an exploding hardware R&D workforce from driving compute costs too low and chips too easy to manufacture. + +USChinaRoWfuller fill = frontier, lighter = rest + +Year**2029** highlighted = frontier (top 3) + +AnthropicOpenAIGDMSpaceXMetaother USother Chinaother RoWUSChinaRoW + +Compute end-users by company \- ai-2040.com + +**Frontier AI compute allocation under Plan A** + +public deploymentinternal inferencetrainingexperimentssafetyoffline / transition + +0%25%50%75%100%2029 dealpublicdeploymentexperimentstrainingsafety20262028203020322034203620382040 + +## Compute Definitions + +### Measuring compute: Raw vs effective H100e + +The primary metric used in this supplement is **H100-equivalents (H100e)** multiples of one Nvidia H100 GPU measured in **Total Processing Performance (TPP)**. + +But compute isn't the only relevant hardware performance metric. The usefulness of compute for AI workloads also depends on many other factors, including _memory bandwidth_ , _memory capacity_ (stores and feeds data to the compute), and _interconnect_ (to send data in between chips). Compute (H100e) is the headline metric we focus on but we also want to account for overall AI usefulness, so we also look at four other hardware resources, and combine them into an overall **effective H100e** metric: + +Resource| What it does| H100 (per H100e) +---|---|--- +Compute (TPP)| the math throughput itself (defines the unit)| 15,800 TPP +Memory bandwidth| feeds data to the compute (HBM)| 3,350 GB/s +Memory capacity| stores data near the compute (HBM)| 80 GB +Scale-up| links chips within a server (NVLink)| 900 GB/s +Scale-out| links servers across the cluster (Ethernet)| 50 GB/s + +**Effective H100e** adjusts raw H100e (which counts compute alone) by a single usefulness factor **U UU**. + +effective H100e = raw H100e×U\text{effective H100e} \;=\; \text{raw H100e} \times Ueffective H100e=raw H100e×U + +UUU is built from the four other resources in the table above. It adjusts for a chip's ability to actually be served by memory and communicate with other chips for representative AI workloads of its era, all relative to a 2024 100K [H100 SXM](https://www.nvidia.com/en-us/data-center/h100/) cluster, which scores exactly U=1U=1U=1. + +Hardware that is memory-rich and well-networked for its FLOP/s can score above 1 (non-AI CPU servers, which carry lots of memory and networking around relatively little compute, score ~1.4), and e.g., compute with poor networking abilities (like gaming GPUs and phones) scores below one. + +The usefulness factor is explained in full in the appendix. + +### Four categories of compute + +The supplement attempts to model **all the compute in the world** and then partitions it in two independent ways giving **four categories**. + +**Partition 1: AI-relevant vs non-AI-relevant.** A chip is considered _AI-relevant_ if it clears all four floors of our **AI-relevant cutoff thresholds** set just below an A100 GPU. + +metric| cutoff| A100| H100| RTX 4090 +---|---|---|---|--- +compute| 4,000 TPP| 4,992| 15,800| 5,285 +memory bandwidth| 1.0 TB/s| 2.0| 3.35| 1.01 +memory capacity| 32 GB| 80| 80| 24 +networking (NIC)| 100 Gb/s| 200| 400| 64 + +**Partition 2: cluster vs single-device.** + +_Cluster_ = is actively used in a multi-device setting with at least 2 devices connected at least at 5 GB/s bidirectional speed. + +_Single-device_ = actively used as a single device (e.g., a phone, a PC, a gaming card) or otherwise does not meet the cluster definition. + +These definitions give us four categories of compute. + +| AI-relevant| non-AI-relevant +---|---|--- +Cluster| AI ClustersAI servers and pods.| Non-AI ClustersCPU servers. +Single-device| Specialty ConsumerHigh-end workstations and gaming cards.| Consumer ComputePhones, PCs, gaming GPUs. + +Two cutoffs define the split: a chip is **AI-relevant** if it clears the AI-compute performance floor, and a **cluster** if it is linked ≥ 5 GB/s to at least one other chip. + +## Part 1. Pre-deal compute (pre-2029) + +### 1.1 AI-relevant compute + +In this section we forecast AI-relevant compute production with a demand-side method: + +**capex spent ($) ×\times× cost-efficiency (H100e / $) = compute added (H100e)** + +**The AI capital expenditure trend** has been ~1.72x/yr through Q3 2023 to Q4 2025, while cost-efficiency (H100e per dollar) has been improving around 1.4x/yr. + +AI capex (~1.72×/yr)Cost-efficiency (~1.4×/yr) + +![AI capex \(~1.72×/yr\)](/hyperscaler-capex-trend.png) + +Source: [Epoch AI — hyperscaler capex](https://epoch.ai/data-insights/hyperscaler-capex-trend) + +Extrapolating these trends naively would lead to the following compute forecast: + +H100e addedInstalled H100e + +050M100M150M200M5.1M202414M×2.7202533M×2.42026projected →79M×2.42027190M×2.42028 + +Naive baseline forecast \- ai-2040.com + +The method is H100e added = capex × cost-efficiency, both extrapolated naively from the trends above (capex ~1.72×/yr; H100e per $ ~1.4×/yr). 2024/25 are actual Epoch data ([AI Chip Sales](https://epoch.ai/data/ai-chip-sales?view=graph&tab=h100_equivalents&timePeriod=annual&cumulative=false): 5.09M H100e added on $260B spending in 2024, 13.6M from $448B in 2025). + +The rest of this section will attempt to improve the forecast from this naive baseline. + +#### Improving the capex forecast + +Through 2025 the buildout was paid for almost entirely out of big-tech's operating cashflow. However, if this trend continues capital expenditures are on track to overtake operating cash flow of hyperscalers around Q3 2026. + +![](/capex.png) + +[Source: Epoch AI](https://epoch.ai/data-insights/hyperscaler-capex-vs-cash-flow) + +To make a forecast that takes this into account, we split capex projections into three sources: **existing-trend hyperscaler cash flow** , **AI cash flow** (additional compute spend from AI company revenues, approximated as 80% of frontier AI company revenue forecasts), and **debt & equity raises** (which we are already starting to see grow in 2025/2026). + +Our projections are as follows: + + * **Existing-trend hyperscaler cash flow** to continue growing at ~20%/yr (Epoch's ~$478B → $603B, with frontier AI company revenue stripped out) and deploy a _rising share_ of it into the buildout, an **invested fraction climbing ~75 → 85%** as capex consumes more of cash flow. + + * **AI cash flow** to follow a total ARR (across frontier AI companies) of **$30B (start-2026) → $180B → $500B → $1T (start-2029)** , with **80% of the recognized revenue** counted as additional compute spend. + + * **Debt & equity raises** are hand forecasted (low effort) as a _declining multiple_ of AI revenue: **3x (2025), 2.8x (2026), 1.8x (2027), 1.6x (2028)**. This is highly uncertain: on the one hand AI revenue growth motivates leverage, but as the total debt raise grows the interest rate will likely get higher. Note also that leverage at these levels implicitly conditions on the scenario's fast AI progress continuing (lenders keep believing in the revenue growth); like the revenue path itself, it is an input consistent with this scenario rather than an unconditional forecast. + + + + +With 2024-25 pinned to Epoch's Broad CapEx ($260B / $448B), this carries total AI capex to **~$2.4T in 2028** , with debt reaching **~$900B** (38% of the 2028 buildout). (We use hyperscaler Broad CapEx as the proxy for total AI capex: [hyperscalers are ~60-90% of global AI spend](https://epoch.ai/data-insights/hyperscalers-control-most-compute), and ~60-90% of their capex is AI, and we assume the two gaps roughly cancel.) This is 7% higher than the naive baseline's (extrapolating the ~1.72x/yr trend) 2028 capex of $2.28T! This is because our bullish AI revenue forecast (and the debt raises we think its fast growth will motivate) actually pulls the trend slightly upwards. + +$ / yr% Y/Y + +0$1T$2T$260B2024$448B2025$797B2026projected →$1.4T2027$2.4T2028 + +Funding the AI buildout \- ai-2040.com + +Hyperscaler cash flow (existing trend)AI cash flowDebt & equity raises + +Total AI capex = Epoch's [hyperscaler Broad CapEx](https://epoch.ai/data-insights/hyperscaler-capex-trend) directly ($260B / $448B for 2024/25). We think hyperscalers are [~60–90% of global spend](https://epoch.ai/data-insights/hyperscalers-control-most-compute) and ~60–90% of their capex is AI, so we assume the two roughly cancel as a simplifying assumption. **Hyperscaler cash flow (existing trend)** excludes frontier-AI revenue (to avoid double-counting) and we project it to grow ~20%/yr, with a rising fraction invested (~75→85%). **AI cash flow** is the additional compute spend from AI company revenues, approximated as 80% of frontier AI company revenue forecasts. **Debt & equity raises** we hand-pick as a _declining multiple_ of AI revenue (3× / 2.8× / 1.8× / 1.6× for 2025–28), reaching ~$923B by 2028. This leads to **~$2.4T capex in 2028**. + +AI revenue and debt forecasts + +Our AI revenue forecast follows a total frontier AI company ARR (annualized revenue run-rate) path of **$30B → $180B → $500B → $1T** (year-boundaries 2026.0 through 2029.0). Note that AI revenue is the flow recognized over the year, not the ARR run-rate, so this converts to annual recognized revenue of 2025: $15B / 2026: $84B / 2027: $313B / 2028: $721B. We count 80% of this as additional compute spend (the "AI cash flow" stream above) giving: $12 / $67 / $251 / $577B. Our total frontier AI company ARR projection is done manually and is highly uncertain. Broadly speaking we expect Anthropic's recent ~10x/yr ARR trend to slow down as they become the leading AI company in terms of revenue, and for the overall total AI company revenue trend to return to roughly 3x/yr, and slightly lower as companies manage to maintain roughly flat revenue / inference-compute ($/FLOP) while their inference-compute grows around 2x/yr in the later years. This is somewhat circular with the compute numbers, but is our current (highly uncertain) expectation. + +**What 's a reasonable debt-raising rate given this revenue trend?** For starters, the five hyperscalers issued **~$121B of bonds in 2025** (vs a ~$28B/yr 2020-24 norm), with BofA/JPMorgan projecting ~$175-300B in 2026; debt-funded neoclouds like CoreWeave carry **~4x revenue** in debt ($21B on ~$5B revenue); and this [Morgan Stanley podcast](https://www.morganstanley.com/insights/podcasts/thoughts-on-the-market/credit-markets-ai-financing-gap-vishy-tirupattur-vishwas-patkar) sees ~$1.15T of the ~$2.9T 2025-28 buildout as debt-financed. Total US investment-grade issuance is **~$1.5T/yr**. The SpaceX IPO also shows the potential for large equity raises that might be followed with similar/larger ones by OpenAI and Anthropic. We use these anchors to manually set debt as an ad-hoc **multiple of total AI revenue** , 3x (2025) → 2.8x (2026) → 1.8x (2027) → 1.6x (2028), giving debt-financed AI capex of ~$36B / $188B / $452B / $923B over 2025-2028. That sums to **~$1.6T of debt-financed spending out of the ~$5.1T total buildout (2025-2028; ~31% debt-financed)** , roughly **39% above Morgan Stanley 's ~$1.15T of debt, and ~75% above their ~$2.9T total**, respectively, though notably Morgan Stanley's numbers ignore equity-raises. + +#### Improving the cost efficiency forecast (supply-side) + +The capex forecast is a demand-side forecast, while for cost-efficiency it makes sense to look at supply side. In particular, this section will focus on lithography (primarily EUV) capacity as the primary expected bottleneck on how many H100e can be produced. + +EUV is needed for leading-edge logic, and the leading HBM DRAM nodes now use it too (DRAM was manufactured entirely without EUV until recently, e.g., Micron only adopted it in 2025, so this is a cost-optimality choice rather than a strict requirement; avoiding EUV just shifts the same demand onto the shared DUV pool). EUV is extremely hard to scale on short notice given the complexity of the machines and extensive supply chain dependencies, so we expect it to be the most important and informative bottleneck. Because of the time-inelasticity (how hard it is to quickly scale production in response to price signal), the EUV fleet is also relatively easy to forecast. + +Year| EUV fleet (start of year)| EUV added / yr| fleet passes / yr| AI logic wafers / yr| AI HBM wafers / yr| AI EUV passes / yr| AI share of fleet +---|---|---|---|---|---|---|--- +2024| 220| +44| 87M| 0.2M| 1.1M| 7M| 7% +2025| 264| +48| 104M| 0.4M| 2.5M| 15M| 15% +2026| 312| +65| 124M| 0.7M| 4.6M| 34M| 27% +2027| 377| +80| 150M| 1.0M| 8.0M| 68M| 45% +2028| 457| +100| 183M| 1.6M| 13M| 118M| 65% +2029 (no Plan A)| 557| +125| 223M| 2.1M| 17M| 172M| 77% + +EUV fleet over time and AI's usage share. The fleet starts at 220 operating tools in 2024 (we estimate ASML had ~127 EUV tools in use at end-2021, then shipped 54/42 in 2022/23, per [ASML's annual reports](https://www.asml.com/en/investors/annual-report); the full build-up with sources is in the details box below) and grows with its accelerating shipments. The EUV added column is ASML's shipments during the year (ASML guides "60+" for 2026 and ~80 for 2027). Fleet passes/yr is the during-year average fleet x ~360k passes/tool. AI's draw sums its logic wafers (x ~12-24 EUV layers per logic wafer, N4→N2) and its HBM-DRAM wafers (x ~4-7 EUV layers per HBM wafer, accounting for stacking and packaging yields). AI's share climbs ~15% (2025) to ~77% (2029, no Plan A). + +How EUV scanners make a H100e + +AI's claim on the fleet comes from both **logic and memory** , so we add the two in a single unit: an **EUV pass** = one EUV exposure of one wafer (a tool images ~360k/yr regardless of what it's printing). Here is the 2025 build-up. + +**1\. Fleet → passes.** Each scanner runs only ~40-50 wph-equivalent in practice, versus a ~160 wafers-per-hour nameplate: real patterns need longer exposures than the spec assumes (higher dose, limited by the light source's power), tools are only up ~75-82% of the time, and time is lost stepping between exposure fields (reticle overhead). Netting these out gives ~**360,000 effective passes/tool/yr** (the naive 160 wph x 8,760 h ≈ 1.1-1.4M is 3-4x too high). Two back-calculations agree: [Frederick Chen](https://frederickchen.substack.com/p/estimating-euv-production-wafer-throughput) (<50 wph effective) and [SemiAnalysis](https://newsletter.semianalysis.com/p/asmls-euv-tools-have-a-throughput) (~1,500 wafers/mo x ~20 exposures). Starting from ~**220 operating tools at the start of 2024** (ASML had ~127 EUV in use at end-2021, then shipped ~54 and ~42 in 2022-23; [TSMC alone runs ~56% of capacity](https://www.digitimes.com/news/a20241112PD204/euv-tsmc-adoption-2023-technology.html)) the fleet does **~100M passes/yr** in 2025, growing to ~**220M by 2029** as ASML's EUV shipments accelerate (~48 in 2025 toward ~120/yr by 2029). + +**2\. AI 's logic draw.** AI accelerators take **~0.4M leading-edge wafers** in 2025, roughly **~12% of TSMC 's ~3M ≤5nm wafers** ([Epoch's direct estimate](https://epoch.ai/data-insights/ai-chip-supply-chain-constraints); Blackwell on N4, Nvidia only now ramping into N3). At N4's ~14 EUV layers that is **~5M passes**. (An ~800mm² die yields ~42 good dies/wafer at ~70% yield, per [Cerebras](https://www.cerebras.ai/blog/100x-defect-tolerance-how-cerebras-solved-the-yield-problem); at ~28-36 H100e/wafer today → ~120 by 2029 as FP4 + the N4→N3→N2 shrink + Rubin's high H100e/chip compound, the all-logic ceiling is ~200M → ~1.2B H100e/yr. The shrink is not free EUV-wise: the model charges a growing pass count per logic wafer, ~12 in 2024 rising to ~24 by 2029, as more layers move onto EUV.) + +**3\. AI 's memory draw (the larger half).** Essentially all HBM goes to AI, and HBM uses a lot of wafers per GB. Reported 2025 HBM **TSV/back-end capacity** is ~400k wafers/month (SK Hynix ~170k + Samsung ~165k + Micron ~60k, [TrendForce](https://www.trendforce.com/news/2025/12/30/news-samsung-reportedly-plans-50-hbm-capacity-surge-in-2026-spotlight-on-hbm4/)), but that's a capacity ceiling, not utilized starts. Netting yields, AI's actual draw is **~2.5M HBM core-wafers/yr** : a 24Gb-die DRAM wafer holds ~1,400 GB gross, and after ~57% core+stacking (KGSD/TSV) yield and ~88% CoWoS-assembly yield only ~**700 net GB ship per started wafer** , while the 2025 fleet needs **~1.8B GB** of HBM (bottom-up: ~11.7M accelerators x ~150 GB, cross-checking the ~$35B / ~2B-GB industry total). The leading HBM nodes use real EUV (SK Hynix 1b ~4 layers, [1c ~6](https://www.tweaktown.com/news/106957/sk-hynix-ramps-1c-dram-to-6-euv-layers-preps-for-high-na-designs-destroy-samsung-in-hbm/index.html)), so at ~**4 layers** that is **~10M passes, roughly two thirds of AI 's EUV draw**. (Per GB, HBM burns [~3x a DDR5 wafer](https://www.tomshardware.com/pc-components/ram/hbm-is-eating-your-ram).) + +**4\. AI 's fleet share.** Logic (~5M) + HBM (~10M) = **~15M passes / ~104M fleet ≈ 15%** in 2025. Rising HBM volume and EUV-layer count, plus AI's migration onto EUV-heavy N3/N2 logic, carry it to **~64% by 2028 and ~77% by 2029** , with AI approaching saturation of the EUV fleet around 2029-30. This draw feeds back into the cost surcharge (more EUV use means a higher surcharge and fewer H100e per dollar), so it is solved self-consistently, and it is what trims the late buildout (2028 additions ~155M vs a naive ~190M). + +**5\. Why passes and not wafer counts.** AI's wafers are a mix: an HBM wafer takes only ~4 EUV passes while a leading-edge logic wafer takes ~14-22, so counting _wafers_ treats very different claims on the fleet as equal. Since AI's wafer mix is HBM-heavy, a wafer count overstates its early share (it would put 2025 well above the ~15% the pass count gives). The two measures converge by 2029 for a simple reason: by then AI is the majority user of both the logic and the HBM wafer pools, so the mix difference stops mattering and either way of counting lands around ~80%. + +AI's share of the EUV fleet grows from ~7% in 2024 toward **~64% by 2028** (and ~77% by 2029, near saturation). We expect this to translate to higher prices as it requires outbidding other uses of the fleet (e.g., phones, PCs, etc.). To model this, we implement a hand-chosen **cost surcharge** curve that depends on the fleet usage (highly uncertain ad-hoc choice). + +% EUV usedSurcharge curve + +0%25%50%75%100%saturation7%202415%202527%2026projected →45%202765%2028 + +EUV capacity and pricing surcharge \- ai-2040.com + +Our forecast for % of global EUV time (passes) used by AI, including both logic and memory wafers (and accounting for yields). It runs ~7% in 2024 toward ~65% by 2028 (~77% by 2029), so AI hits the EUV wall around 2029-30. The cost surcharge (curve tab), surcharge = A·u^K (default A = 100%, K = 2.5), is keyed to this derived path and feeds back into the compute forecast (more draw → higher surcharge → fewer H100e), solved to a fixed point so it climbs to ~33% by 2028 and throttles the late buildout. The K and A sliders tune how hard it bites for users to play with. The parameters are hand chosen and extremely uncertain. + +Applying this surcharge to the naive price-performance trend gives an updated compute cost-efficiency trend that averages **~1.25x/yr** from 2025-2028 (and slows to just ~1.1x/yr by 2028 as the surcharge bites) instead of the historical ~1.4x/yr. + +$ / H100eH100e / waferwafers / $B + +levelY/Y + +0$20K$40K$60K$51K2024$33K2025$23K2026projected →$18K2027$16K2028 + +The cost of compute, decomposed \- ai-2040.com + +Cost per H100e decomposes as **$ per wafer ÷ H100e per wafer**. + +Where the wafer numbers come from + +The cost figure is an identity, **$/H100e = (all-in $/wafer) / (H100e per wafer)** , and we have decent estimates of two of the three quantities, so we back out **wafers per $B** (the inverse of all-in $/wafer) as the **residual** and then check whether the wafer cost it implies is sensible. + +**$/H100e (known).** 2024/25 are observed (annual capex / H100e added: ~$51K, ~$33K). For 2026-28 we take the ~1.4x/yr price-performance trend and penalize it by the EUV surcharge, giving ~$23K / ~$18K / ~$16K. + +**H100e per wafer (known, ~28 → ~98).** An ~800mm² reticle-limit die yields ~64-72 gross dies per 300mm wafer, ~42 good at ~70% yield; netting downstream losses (packaging yield) leaves **~28 sellable H100e per wafer in 2024**. This climbs ~1.4x/yr to **~98 by 2028 (~120 by 2029)** from FP4/precision (EUV-free), the N4→N3→N2 shrink (paid in extra EUV layers), and Rubin's high H100e/chip, all reasonably well grounded in die-size, node, and precision estimates. (2025 check: ~13.6M H100e on ~0.38M AI leading-edge wafers ≈ ~36 H100e/wafer.) + +**Wafers per $B (the residual).** Multiplying the two backs out the implied **all-in $/wafer** : ~$1.43M (2024) → ~$1.18M (2025) → ~$1.20M → ~$1.36M → ~$1.56M (2028), i.e. wafers/$B ~700 → ~845 → ~640. The 2025 dip comes directly from the observed data: 2024 was early and supply-constrained, and 2025 hit Blackwell volume, so all-in cost per leading-edge wafer fell. From 2026 the implied $/wafer rises, which is what we'd expect: TSMC ASP +5-10%/yr (N2 ~$30K, +10-20% over N3), CoWoS +15-20%/yr, datacenter $/MW rising, plus the EUV surcharge. The level also checks against a bottom-up build-up (foundry ~$25-30K + HBM ~$50-70K + CoWoS ~$10-15K, with a ~4-5x chip-designer markup giving a ~$400-500K accelerator, and the remaining ~$0.7-1.1M being the non-chip buildout: datacenters, power, land, networking). So the residual looks reasonable: wafer costs are a mild headwind after 2025, and the whole ~$51K to ~$16K decline in $/H100e comes from compute density improvements. + +Multiplying this cost-efficiency baseline (incl. the scarcity surcharge) by the capex topline gives the annual **H100e added** (~**155M** in 2028), and an **installed stock** of ~**280M** by end-2028. We also adjust very slightly downwards for chip survival rates. China's domestically manufactured compute (§1.3) does not draw on this EUV supply chain, so we forecast it separately and add it onto this total, bringing the world AI stock to ~**287M**. + +H100e addedH100e installed + +levelY/Y + +raweffective + +EUV buildout (§1.1)\+ China domestic (§1.3) + +050M100M150M200M5.3M202414M202536M2026projected →81M2027160M2028 + +AI compute: flow & stock \- ai-2040.com + +H100e added is annual capex × cost-efficiency (anchored to the 2024/25 actuals 5.09M, 13.6M, then forecast); installed stock is shown at the end of each year. The black bars are the EUV-constrained §1.1 buildout (~155M added in 2028, ~280M installed by end-2028); the red caps add China's off-EUV domestic manufacturing (§1.3), for a ~289M world total. Toggle ‘effective’ to rescale by the §1.2 usefulness factor. + +Resource intensity: from raw H100e to effective compute + +Even for the AI fleet, effective and raw H100e drift apart. The **memory wall** drags effectiveness below 1 (TPP, boosted by FP4, outgrows HBM bandwidth and capacity), while the **growing coherent NVLink world** (HGX-8 → NVL72 → NVL144 → NVL576) pushes it back up. The two roughly cancel: there is a mild premium in 2024 (~1.16, when the fleet is memory-rich) that fades toward ~1 by 2029 as the memory wall and larger training runs start to matter more. It is the same usefulness model as §1.2.3 (full derivation in the appendix); the deployment-weighted fleet inputs that drive it (scale-up world in H100e = coherent chips x H100e/chip): + +AI fleet (deployment-weighted)| 2024| 2026| 2028 +---|---|---|--- +HBM bandwidth per H100e (GB/s)| 4,080| 3,330| 3,610 +HBM capacity per H100e (GB)| 110| 91| 85 +scale-up world (H100e)| 184| 436| 929 +H100e per chip| 1.15| 2.58| 5.34 +**effective H100e per raw**| **1.16**| **1.06**| **1.04** + +### 1.2 All compute + +§1.1 forecast AI-relevant compute. Now we extend to all compute in the world. We split all compute into six device categories: + +Category| Examples| AI-relevant +---|---|--- +AI| datacenter AI accelerators (H100, B200, MI300)| yes +Non-AI servers| datacenter CPUs and non-AI accelerators| no +PCs| laptops and desktops| no +Phones| smartphones| no +Gaming| gaming GPUs| no +Other| microcontrollers, automotive, edge| no + +Our method works by looking at two quantities. The **device side** is the compute we get from counting the units shipped and the average compute carried on each device. The **lithography side** is the compute the world can make, from counting the wafers the fabs can produce and the compute each wafer turns into. + +device side →\to→ H100e = units ×\times× H100e per unit + +lithography side →\to→ H100e = wafers ×\times× H100e per wafer + +The compute per wafer is the quantity we are least sure of, so it is what we attempt to pin down from history (where we can get good estimates of all the other values). + +#### 1.2.1 Calibration (2015 to 2024) + +Across 2015 to 2024 we look at **device demand** (units shipped and compute per unit) and **lithography supply** (wafers shipped), and use this to back out an estimate of compute per wafer that we can then use to forecast future compute by category based on wafer allocation and wafer supply forecasts (which are easier to do than device demand forecasts, because the supply side can be forecasted with the relatively predictable lithography ceiling). + +### Device side + +Shipments (M units/yr)| 2015| 2018| 2020| 2022| 2024 +---|---|---|---|---|--- +phones| 1K| 1K| 1K| 1K| 1K +pcs| 276| 259| 304| 292| 263 +gaming| 44| 50| 40| 38| 36 +ai| 0.1| 0.3| 0.6| 2.0| 7.5 +non-AI servers| 8.0| 11| 13| 14| 13 +other| 12K| 20K| 24K| 28K| 30K + +Shipment sources + + * phones: 1.43B (2015) → 1.24B (2024). Peak 1.47B (2016); pandemic trough 1.21B (2022). Source: [IDC Worldwide Quarterly Mobile Phone Tracker](https://www.idc.com/promo/smartphone-market-share) / Omdia / Counterpoint. + * pcs: 276M (2015) → 263M (2024). COVID peak ~349M (2021), trough ~260M (2023). IDC runs ~30M above Gartner in 2020 to 2021 because IDC counts Chromebooks; we use the IDC basis. Source: [Gartner](https://www.gartner.com/en/newsroom/press-releases/2024-01-10-gartner-says-worldwide-pc-shipments-grew) / [IDC PCD Tracker](https://www.idc.com/promo/pcdforecast). + * gaming: discrete desktop add-in-board GPUs only (not iGPUs, not datacenter). 44M (2015) → ~35M (2024). Source: [Jon Peddie Research AIB Report](https://www.jonpeddie.com/reports/add-in-board-report/). + * ai: the whole AI datacenter; unit count = accelerators. TechInsights counted 3.76M NVIDIA datacenter GPUs in 2023 (98% share); 2024 is ~7.5M all-vendor (NVIDIA ~4M + Google TPU ~2.5M + AWS Trainium ~1.2M + AMD MI300 ~0.4M). Source: [TechInsights via DCD](https://www.datacenterdynamics.com/en/news/nvidia-gpu-shipments-totaled-376m-in-2023-equating-to-a-98-market-share-report/); [Epoch AI](https://epoch.ai/data-insights/nvidia-chip-production); Omdia/Morgan Stanley. + * other: just the advanced-node digital chips (networking ASICs, basebands, auto/IoT logic). The full MCU plus analog universe is far larger ([~25 to 35B MCUs/yr alone](https://www.statista.com/statistics/935382/worldwide-microcontroller-unit-shipments)), but most is trailing-node / 200mm capacity outside the 300mm lithography pool this model tracks. + + + +TPP per unit| 2015| 2018| 2020| 2022| 2024 +---|---|---|---|---|--- +phones| 0.5| 5.0| 25| 60| 100 +pcs| 8.0| 20| 38| 62| 100 +gaming| 15| 250| 550| 1K| 2K +ai| 350| 1K| 2K| 5K| 10K +non-AI servers| 30| 75| 150| 320| 500 +other| 0.1| 0.3| 0.5| 0.7| 0.9 + +TPP-per-unit sources + +TPP = (peak dense TOPS at a precision) × (bit-length of that precision), maximized over precisions: the US BIS export-control metric (ECCN 3A090). Anchors that fix the scale: V100 ≈ 2,000, A100 = 4,992, H100 ≈ 15,800, RTX 4090 = 5,285 (just over the 4,800 control line), B200 ≈ 36,000. All figures are sales-weighted averages, since the flagship is a tiny share of units. Sources: [Federal Register 88 FR 73458](https://www.govinfo.gov/content/pkg/FR-2023-10-25/pdf/2023-23055.pdf), [CSET explainer](https://cset.georgetown.edu/article/bis-2023-update-explainer/), [Princeton ISCA 2025](https://parallel.princeton.edu/papers/sanctions-isca-2025.pdf). + + * phones: NPU INT8 TOPS × 8. Apple A11 = 0.6 TOPS (2017) → A17 Pro = 35 TOPS (2023); Snapdragon 8 Gen 3 = 45 TOPS (2024). Near-zero before 2017. Sales-weighted ≈ 100 TPP in 2024. Source: [Apple Neural Engine TOPS table](https://github.com/hollance/neural-engine/blob/master/docs/supported-devices.md). + * pcs: most PCs have no NPU (only ~17% of 2024 units were “AI PCs”). Sales-weighted ≈ 100 TPP in 2024. Source: [Canalys AI-PC share](https://canalys.com/newsroom/ai-capable-pc-shipment-q4-2024). + * gaming: dense INT8 TOPS × 8: RTX 3060 = 815, RTX 4060 = 1,412, RTX 4090 = 5,285. The sales mix is mostly xx60-class (Steam survey), so the sales-weighted average is ~1,500 TPP in 2024, not the 4090's 5,285. Source: [Steam Hardware Survey](https://store.steampowered.com/hwsurvey/videocard/). + * ai: sales-weighted across all vendors (H100/H200 = 15,800, A100 ≈ 5,000, TPU/Trainium lower), 2024 ≈ 9,500 TPP, below the NVIDIA flagship. With ~7.5M units that gives ~4.5M H100e in 2024, the bottom-up value that meets the Plan A forecast at the 2025 seam. Source: NVIDIA/Google/AWS datasheets; [TrendForce mix](https://www.trendforce.com/presscenter/news/20240903-12283.html). + * other: ~30B chips/yr, mostly MCUs/analog with ~0 AI throughput; only a thin automotive/ADAS + edge-NPU tail carries any, so the unit-weighted average is just ~1.5 TPP by 2024. + + + +### Lithography side + +wafers (M/yr)| 2015| 2018| 2020| 2022| 2024 +---|---|---|---|---|--- +phones| 1.8| 1.9| 1.9| 1.9| 2.1 +pcs| 0.7| 0.7| 0.9| 0.9| 0.9 +gaming| 0.2| 0.3| 0.2| 0.3| 0.3 +ai| 0.0| 0.0| 0.0| 0.0| 0.3 +non-AI servers| 0.1| 0.2| 0.3| 0.3| 0.3 +other| 5.5| 9.7| 12| 15| 17 +device demand (total)| 8.3| 13| 15| 18| 20 +Wafer shipment sources + + * Each category's wafers = its **logic wafers** (units × dies-per-device ÷ good-dies-per-wafer) plus the **memory wafers** (HBM, DRAM, NAND) that ship with it. Units are the shipments table above. Memory wafers follow from each device's memory content and current per-wafer GB on DRAM/NAND/HBM nodes. + * Good dies per 300mm wafer (yield-net): phones ~600 (~100 to 120 mm²), PCs ~300 (~200 mm²), gaming ~130 (~400 mm²), AI ~42 (~800 mm² reticle-limit die at ~70% yield), “other” ~1,800 (~30 mm²). Sources: [AnySilicon](https://anysilicon.com/die-per-wafer-formula-free-calculators/); reticle/die-size disclosures (TSMC, NVIDIA, WikiChip). + * Dies per finished device: non-AI ≈ 1 (one SoC/CPU/GPU per device). AI ramps 1 (H100) → ~1.5 (2024, dual-die B200) → ~8 to 11 by 2030 (Rubin-class chiplets). Sources: NVIDIA roadmap; [SemiAnalysis](https://www.semianalysis.com/). + * Die-area is already net of yield, so no separate yield factor is applied. + + + +The lithography utilization by category is shown below. We treat every category as a package with its logic and associated memory (HBM, DRAM, NAND). + +**Lithography allocation, 2015-2024** + +EUV passes% of EUVall passes% of passes + +EUV exposure-passes drawn per year (the absolute view the share hides). Volume EUV begins ~2019; it's ~all phones/PCs early, AI invisible until ~2022. + +0M12M23M35M46M'15'16'17'18'19'20'21'22'23'24 + +ainon-AI serversphonespcsgamingother + +Sanity check with independent compute per wafer estimates + +Compute per wafer is just (good dies per wafer) x (TPP per die) / 15,800, both estimated from device specs independently of how full the fabs run: + + * **Good dies per 300mm wafer** (yield-net, from die areas; sources in the wafer box above): AI ~42 (an ~800 mm² reticle-limit die at ~70% yield), gaming ~130 (~400 mm²), PCs ~300 (~200 mm²), phones ~600 (~100 to 120 mm²), commodity 'other' ~1,800 (~30 mm²). + + * **TPP per die** = each category's TPP per unit (device side above) / its dies per device. In 2024 AI is ~9,500 / ~1.5 ≈ 6,300 TPP/die, gaming ~1,500, phones and PCs ~100, 'other' ~1.5. + + + + +So AI's big, high-TPP dies make it densest per wafer and commodity 'other' (tiny, near-zero-TPP dies) the lowest, which is exactly what dividing each category's compute by its wafers gives: + +implied H100e / wafer| 2015| 2018| 2020| 2022| 2024 +---|---|---|---|---|--- +phones| 0.03| 0.23| 1.1| 2.4| 3.8 +pcs| 0.20| 0.46| 0.83| 1.3| 1.9 +gaming| 0.19| 2.8| 5.6| 9.2| 12.3 +ai| 1.4| 5.0| 7.2| 14.2| 16.8 +non-AI servers| 0.11| 0.25| 0.46| 0.90| 1.3 +other| 0.01| 0.04| 0.06| 0.09| 0.10 + +As a further check, the wafer demand these rates imply, set against our estimated **lithography pool** (broad global logic-wafer output, ~30M wafers in 2024, a scale consistent with SEMI's 300mm capacity), is ~62% to ~69% of the pool across the decade: a plausible, stable share that never runs past the pool. The pool level is itself calibrated to land in that range, so this says the die-size, compute, and capacity estimates are mutually consistent rather than independently validating any one of them. + +The resulting H100e added each year by category is shown below in both raw and effective H100e: + +**World H100e by category, 2015-2024** + +flow (added/yr)stock (installed) + +raweffective + +level% share + +Raw H100e added per year by category. + +05M10M15M20M'15'16'17'18'19'20'21'22'23'24 + +ainon-AI serversphonespcsgamingother + +#### 1.2.2 Forecast (2025 to 2028) + +We take AI's compute forecast from §1.1 and ration the leftover lithography capacity to the other compute categories. AI first claims its lithography usage and then each non-AI category keeps its frozen-2024 share of the wafers that remain. That gives the projection: + +Flow (H100e/yr)| 2025| 2026| 2027| 2028 +---|---|---|---|--- +phones| 12.0M (33%)| 16.6M (25%)| 21.5M (18%)| 31.0M (14%) +pcs| 2.8M (7.7%)| 4.1M (6.2%)| 5.5M (4.5%)| 8.8M (4.0%) +gaming| 5.2M (14%)| 7.2M (11%)| 9.4M (7.8%)| 14.2M (6.5%) +ai| 14.2M (39%)| 35.7M (54%)| 80.6M (67%)| 159.7M (73%) +non-AI servers| 558K (1.5%)| 725K (1.1%)| 901K (0.7%)| 1.2M (0.5%) +other| 2.0M (5.4%)| 2.3M (3.4%)| 2.6M (2.2%)| 3.2M (1.5%) +total| 36.7M| 66.6M| 120.5M| 218.1M + +AI is from §1.1; non-AI gets the leftover logic wafers × rising TPP-per-wafer. Counting logic + HBM, AI reaches ~15% → 65% of EUV fleet capacity by 2028 (with the fleet ~100% utilized overall), but only ~40% of all logic wafers. + +AI’s wafer claim, leftover pool, and bill of materials + +wafers/yr| 2025| 2026| 2027| 2028 +---|---|---|---|--- +AI logic wafers| 378K| 664K| 1.0M| 1.6M +\+ AI HBM wafers (advanced DRAM)| 2.5M| 4.6M| 8.0M| 12.8M += AI cutting-edge wafers (logic + HBM)| 2.9M| 5.2M| 9.1M| 14.4M +% of EUV fleet capacity (logic + HBM)| 15%| 27%| 45%| 65% +% of broad lithography pool (all nodes)| 9.3%| 16%| 27%| 40% +allocatable logic pool (× util)| 20.9M| 22.0M| 23.2M| 24.4M +leftover for non-AI| 20.6M| 21.4M| 22.1M| 22.8M + +Sources: AI forecast and per-accelerator bill of materials + + * AI H100e trajectory. From §1.1's EUV-constrained solve (capex topline ÷ cost-efficiency ÷ EUV surcharge): installed H100e stock (Jan-1) 8.6M · 22.4M · 57.1M · 135.2M → 289.0M by 2029, with annual production 14.2M · 35.7M · 80.6M · 159.7M. §1.2 takes this AI buildout from §1.1 unchanged and models only the non-AI leftover, so the two sections agree by construction. Source: §1.1 (resting on the Plan A capex topline). + * HBM per accelerator. H100 = 80 GB, H200 = 141, MI300X = 192, B200 = 192; sales-weighted ~130 GB in 2024 → ~600 by 2030. × ~7.5M units ≈ 1 EB HBM in 2024, matching the TrendForce industry figure. Source: NVIDIA/AMD datasheets; [TrendForce](https://www.trendforce.com/presscenter/news/20240903-12283.html). + * System DRAM & storage per accelerator. An 8-GPU node carries ~1 to 2 TB host DRAM (~150 to 250 GB/GPU) plus multi-TB local + networked NVMe. We model ~300 GB DRAM and ~2 TB NAND per accelerator in 2024 (growing to ~1 TB / ~6 TB by 2030), sized so HBM stays AI's heaviest wafer draw, not the commodity memory. Source: [NVIDIA DGX/HGX H100 specs](https://www.nvidia.com/en-us/data-center/dgx-h100/); hyperscaler reference designs. + * Stock vs production convention. AI's fleet is treated as cumulative (no retirement over this <5-yr window), so each Jan-1 stock is the running sum of production; we take the stock directly rather than running the bottom-up retirement recursion used for the non-AI categories. The production figure drives each year's wafer claim. + + + +The rest of this section is the supply-side justification behind that table, i.e., what share of the lithography ceiling each category uses. + +**Lithography allocation, 2024-2028** + +EUV passes% of EUVall passes% of passes + +EUV exposure-passes drawn per year. AI's draw explodes, taking the fleet toward saturation by 2028. + +0M47M93M140M186M20242025202620272028 + +ainon-AI serversphonespcsgamingother + +**Logic vs. Memory** + +EUVDUV (193i) + +Share of EUV and DUV passes that go to logic vs. memory + +0%25%50%75%100%'15'16'17'18'19'20'21'22'23'24'25'26'27'28 + +logicmemory + +EUV utilization reaches ~100% by 2028 in our naive model, up from only 52% in 2024, as we expect it to become the binding constraint by then. That figure already separately accounts for uptime and realized throughput, both held at 2024 proportions. We don't expect utilization to literally hit 100%, but we're comfortable with the approximation: uptime and realized throughput (as a share of peak theoretical throughput) should improve a bit by 2028, roughly offsetting the gap. + +EUV pass pool (M/yr)| 2024| 2028 +---|---|--- +passes demanded| 45.3M| 182.5M +pass capacity| 87.1M| 182.5M +EUV utilization| 52%| 100% +immersion (DUV) utilization| 20%| 25% + +Utilization = passes demanded ÷ pass capacity. EUV and immersion are both on the effective (uptime- and throughput-adjusted) basis, so the two utilizations compare directly. + +Node structure: passes per wafer and node mix + +node| EUV/wf| 193i/wf| wafers 2024| wafers 2028 +---|---|---|---|--- +logic tiers +LEADING| 12→22| 30| 2.0M| 3.2M +DUVADV| 2| 33| 2.0M| 5.2M +MATURE| n/a| 6| 16.4M| 16.0M +memory tiers +DRAMadv| 4| 12| 4.2M| 20.7M +DRAMmat| n/a| 8| 5.8M| 7.6M +NAND| n/a| 4| 40.2M| 45.0M + +Three tiers (LEADING = ≤N3/N5/N2, EUV; DUVADV = N7/N16; MATURE = 28nm/legacy). Leading EUV passes/wafer climb 12→22 as the mix shifts N5→N3→N2. Red = EUV-drawing. This grid is a utilization diagnostic; it does not feed the compute or stock math. + +Node mix (2024 wafer-area share by category) + +category| LEADING| DUVADV| MATURE +---|---|---|--- +phones| 53%| 27%| 20% +pcs| 42%| 40%| 18% +gaming| 55%| 45%| n/a +ai| 95%| 5.0%| n/a +non-AI servers| 50%| 42%| 8.0% +other| n/a| 5.0%| 95% + +AI is ~all leading-edge; phones/PCs/gaming span leading to mature; “other” (MCUs) is ~all mature. The broad logic pool (29.5M wafers in 2024) feeds the non-AI allocation. + +Per-device memory and wafer cost + +Memory wafers = device count × per-device GB ÷ GB-per-wafer. HBM needs ~3× the wafer area per usable GB (8 to 12-hi stacks, TSVs, a base logic die) and goes almost entirely to AI, so an AI accelerator's memory outweighs its logic on wafers. + +per device (2024 → 2028)| HBM| DRAM| NAND +---|---|---|--- +phones| n/a| 8 → 13 GB| 200 → 333 GB +pcs| n/a| 16 → 27 GB| 700 → 1033 GB +gaming| n/a| 14 → 21 GB| n/a +ai| 130 → 400 GB| 300 → 767 GB| 2000 → 4667 GB +non-AI servers| n/a| 512 → 683 GB| 26000 → 36667 GB +other| n/a| 0.05 → 0.05 GB| 2 → 3 GB + + * GB per wafer (2024): DRAM ~2,800, NAND ~21,000, HBM ~900; all grow ~1.7× by 2030. HBM is ~11% of DRAM wafers in 2024 (TrendForce: ~14%). Sources: [SEMI](https://www.semi.org/en/products-services/market-data); TechInsights; TrendForce. + * The non-AI server row carries the datacenter storage tier (SSD arrays, networked NVMe), so its NAND is large (~26 TB/server-equivalent); this folds in bulk storage a strict per-device count misses, landing world NAND ≈ 845 EB in 2024 (~44M wafers), in line with Trendfocus. Sources: TrendForce; Trendfocus; IDC / Omdia. + + + +Litho supply sources: EUV layers/node, fleet, throughput, node mix + + * EUV layers per node: N5/N4 ≈ 13 to 14, N3E ≈ 19, N3B ≈ 25 (double-patterned), N2 ≈ 23 to 25. A double-patterned layer needs two exposures, so passes ≥ layers; we model the leading tier climbing 12→24 (2024→2029). Sources: [SemiAnalysis](https://www.semianalysis.com/p/the-memory-wall); WikiChip; TSMC. + * Memory EUV: advanced DRAM (1α/1β/1γ) uses ~1 to 6 EUV layers, ramping; 3D NAND uses none. We model advanced DRAM at ~4 EUV passes/wafer, mature DRAM and NAND at 0. Sources: SK Hynix / Samsung / Micron disclosures; TechInsights. + * EUV fleet: ~220 systems by 2024, ~457 by 2028 on ASML guidance. Throughput: ~360k effective exposures/tool-year (below the ~160 wph spec). Source: [ASML](https://www.asml.com/en/investors). + * Immersion fleet: ~1,330 (2024) → ~1,800 (2028); abundant, so 193i never binds. Throughput: ~950k effective passes/tool-year (about half the ~1.8M nameplate run-rate, de-rated on the same uptime/throughput basis as EUV). Node mix is wafer-area share: Counterpoint ~43% of 2024 phone-SoC units on ≤5nm → ~53% of wafer area. Sources: ASML / SEMI; Counterpoint; [TSMC](https://investor.tsmc.com/english/quarterly-results). + + + +#### 1.2.3 Results + +Accumulating production net of retirement gives the installed stock by category. Rescaling raw H100e to effective H100e (raw x UUU, the resource-adjusted AI workload usefulness model derived in the appendix). + +Effective H100e rescales each category's raw compute by a usefulness factor **U** , derived in the appendix. + +usefulness U| 2024| 2026| 2028| 2029 +---|---|---|---|--- +phones| 0.21| 0.11| 0.07| 0.05 +pcs| 0.35| 0.08| 0.04| 0.03 +gaming| 0.04| 0.05| 0.06| 0.04 +ai| 1.16| 1.06| 1.04| 1.02 +non-AI servers| 1.38| 1.40| 1.41| 1.40 +other| 0.53| 0.43| 0.34| 0.30 +Per-chip inputs and the unclamped model output + +The per-chip resource inputs that drive U, by category (2024). Only AI accelerators pair real compute with HBM-class bandwidth and a large NVLink coherent domain; every other category is effectively a single chip (scale-up world = 1). Non-AI servers are memory-rich but compute-light; consumer chips are bandwidth-poor and don't cluster; “other” (MCU / edge) is tiny on every axis. + +per chip (2024)| H100e/chip| mem BW (TB/s)| mem (GB)| scale-up world +---|---|---|---|--- +phones| 4.0e-3| 0.032| 7| 1 +pcs| 6.6e-3| 0.085| 12| 1 +gaming| 0.097| 0.41| 11| 1 +ai| 1.15| 4.7| 126| 160 chips +non-AI servers| 0.030| 0.70| 512| 1 +other| 5.7e-5| 0.006| 0.1| 1 + +Over the forecast, H100e/chip roughly doubles every ~1.5 years while memory bandwidth and capacity per H100e fall (the memory wall) and AI's NVLink world grows (NVL72 → NVL576); the §1.1 resource-intensity box tracks the AI fleet's trajectory in full. + +Raw U, before the non-AI ≤ 1 cap + +category (raw U)| 2024| 2026| 2028| 2029 +---|---|---|---|--- +phones| 0.21| 0.11| 0.07| 0.05 +pcs| 0.35| 0.08| 0.04| 0.03 +gaming| 0.04| 0.05| 0.06| 0.04 +ai| 1.16| 1.06| 1.04| 1.02 +non-AI servers| 1.38| 1.40| 1.41| 1.40 +other| 0.53| 0.43| 0.34| 0.30 + +Servers and “other” (MCU/edge) score > 1 because for a tiny-compute chip every per-FLOP ratio looks abundant, so the networking penalty all but vanishes and it floats to the compute ceiling. For non-AI servers this is plausible (genuinely memory-rich machines) and they are tiny in raw H100e (~1–2%), so we keep their raw U. For MCU/edge “other” it is an artifact, so we clamp it to ≤1; note “other” is a large raw share (~8–23%), so even the clamp at 1 likely overstates its AI usefulness. + +**World compute by quadrant** + +stockflow + +raweffective + +level% share + +Installed H100e by deployment quadrant; toggle raw↔effective and level↔share. + +0100M200M300M400M20M202435M202574M2026157M2027314M2028 + +AI × cluster (the AI fleet) + +AI × single-device (PCIe AI cards) + +non-AI × cluster (CPU / NIC pools) + +non-AI × single-device (consumer) + +**Cluster vs. single-device split.** We also categorize compute into clusters and single-devices. Non-AI servers are treated as non-AI relevant clusters; phones, PCs, gaming GPUs and embedded chips are non-AI single-device; and AI relevant compute is almost entirely in accelerators run in datacenters (AI-relevant clusters), but there is a small rising amount of AI-relevant single-device chips that is negligible today and grows later as datacenter-class GPUs reach desktop form factors (e.g., [NVIDIA's DGX Station](https://www.nvidia.com/en-us/products/workstations/dgx-station/)). + +effective H100e stock (Jan 1)| 2020| 2024| 2026| 2028| 2029 +---|---|---|---|---|--- +AI-relevant · cluster (AI datacenters)| 101K (4.1%)| 3.6M (26%)| 23.7M (67%)| 140.8M (90%)| 295.2M (94%) +AI-relevant · single-device (specialty consumer)| 0| 0| 0| 218K (0.1%)| 938K (0.3%) +non-AI · cluster (non-AI servers)| 294K (12%)| 1.4M (10%)| 2.6M (7.5%)| 4.7M (3.0%)| 6.1M (1.9%) +non-AI · single-device (consumer)| 2.1M (84%)| 8.7M (63%)| 9.1M (26%)| 11.5M (7.3%)| 12.6M (4.0%) +How we forecast AI-relevant single-device compute (specialty consumer) + +These are standalone AI workstations: a capable GPU (≥1 TB/s bandwidth, ≥32 GB, ≥4,000 TPP) paired with a 100 Gb/s NIC, which clears all four AI-relevant floors. They range from workstations on RTX PRO 6000-class GPUs up to NVIDIA's DGX Station (one GB300, ~2.5 H100e). Cheaper minis like DGX Spark (~0.5 H100e) miss the bandwidth floor, and ordinary gaming PCs lack the 100 Gb/s NIC. At ~$10,000 to ~$100,000 each, they sell in modest volumes. We estimate units sold per year and H100e per unit, then add them up. + +shipping year| units shipped| H100e / unit| H100e added| installed (Jan 1) +---|---|---|---|--- +2026| 15K| 2.5| 38K| 0 +2027| 60K| 3| 180K| 38K +2028| 180K| 4| 720K| 218K +2029| 450K| 5| 2.3M| 938K + +The **installed (Jan 1)** column is the running stock at the start of each year — exactly the specialty-consumer row in the table above (~938K H100e by 2029.0, about 0.3% of the AI fleet, ~255K units). The category is new, so the unit numbers are rough estimates. + +### 1.3 Regional ownership + +This section forecasts world compute ownership by region, for the AI-relevant compute from §1.1. We split ownership three ways: **US, China, and the Rest of the World**. We think it is most informative to focus on a separate Chinese compute forecast, because China is cut off from the main global supply chain by export controls, making it the hardest region to pin down with a crude percent-of-spending estimate. + +What follows is therefore an independent Chinese compute forecast, broken down by the different sources of compute reaching China. We then make a crude estimate of the US versus Rest-of-World split of the global total. China's domestic manufacturing is treated as separate from the world compute total forecast in §1.1, because it does not come from the same (EUV-based) supply chain, so the §1.1 totals add it back on. + +#### 1.3.1 Chinese compute forecast + +We categorize China's AI compute into four sources: Three come from the global supply chain (already in the §1.1 world total): **legal** imports (export-compliant chips: e.g., H20s historically, H200 under the current licensing policy), **smuggled** chips, and **loopholes** ([chip dies smuggled in](https://www.reuters.com/technology/us-ordered-tsmc-halt-shipments-china-chips-used-ai-applications-source-says-2024-11-10/) before packaging, or offshore datacenters Chinese firms own or [lease from abroad](https://www.the-substrate.net/p/how-much-us-compute-is-china-renting)). The fourth is **domestically** manufactured chips (SMIC 7nm logic / CXMT HBM / Ascend). China's compute grows ~9x over 2026-2029, but its world share still falls toward ~9% because the global buildout outpaces it. + +H100e / yearH100e installed + +levelY/Y% of world + +04M8M12M0.9M20241.7M20253.4M2026projected →6.8M202713M2028 + +China's AI compute \- ai-2040.com + +DomesticLegal importsLoopholes (offshore DCs / rental)Smuggled + +The 2024-2025 anchors, pinned to observed data: + +Channel| 2024| 2025| Basis +---|---|---|--- +Domestic| 0.25M| 0.6M| Ascend on SMIC ≤7nm + Cambricon, drawing a one-time foreign-HBM stockpile (~13M stacks, mostly Samsung ~11.4M ≈ 1.6M 910C) and a ~2.9M-die TSMC "die bank" ([SemiAnalysis](https://newsletter.semianalysis.com/p/huawei-ascend-production-ramp)); ~0.8 raw H100e/910C +Legal| 0.2M| 0.15M| ~1M throttled H20s in 2024 (~0.2 H100e each); 2025 nets only ~0.15M: Q1 shipments before the April ban, a partial re-licensing from July, and Beijing discouraging purchases +Loopholes| 0.25M| 0.45M| offshore DCs Chinese firms own or [rent abroad](https://www.the-substrate.net/p/how-much-us-compute-is-china-renting); the cumulative 2024+2025 total here (~0.7M by end-2025) sits just above the ≥~670K confirmed-rental floor (ByteDance ~520K) +Smuggled| 0.16M| 0.48M| grey-market banned parts, anchored to [Epoch](https://epoch.ai/publications/chip-smuggling) + +Forecasting 2026-2028 + +The three global channels are projected as a share of §1.1 world production: + +Global channel| Rule (share of §1.1 world)| 2026| 2027| 2028 +---|---|---|---|--- +Smuggled| ~3.5% → 2% with more [enforcement](https://www.ft.com/content/57fcd3ce-464f-4dc2-8ea2-5712d4972c69)| 1.26M| 2.12M| 2.99M +Loopholes| ~2% (~1.5% [confirmed-rental floor](https://www.the-substrate.net/p/how-much-us-compute-is-china-renting) \+ other sources)| 0.7M| 1.6M| 3.15M +Legal| ~1%| 0.35M| 0.8M| 1.6M + +The legal channel could plausibly run 2-3x higher than this ~1% rule: the current licensing policy already covers volumes above it, though Beijing-side import restrictions push realized shipments back down. + +**Domestic** is HBM-bound; see the recent [SemiAnalysis CXMT piece](https://newsletter.semianalysis.com/p/chinas-cxmt-is-set-to-challenge-dram): + +Domestic build-up| 2026| 2027| 2028 +---|---|---|--- +CXMT HBM capacity (kwspm starts)| 30| 55| 100 +8-hi stack yield (maturing)| 25%| 31%| 37% +good 8-hi stacks per wafer (~88 gross x yield)| 22| 27| 33 +→ good 910C-equiv/yr (kwspm x 12k x stacks / 8)| 1.0M| 2.2M| 4.9M +raw H100e per 910C-equiv (logic + HBM3E gains)| 0.80| 0.82| 0.85 +→ raw H100e on CXMT memory| 0.8M| 1.8M| 4.2M +uplift: smuggled HBM + other firms| +30%| +25%| +20% +**= Domestic**| **1.1M**| **2.3M**| **5.0M** + +This is optimistic: we take CXMT's HBM ramp at face value (the [SemiAnalysis CXMT piece](https://newsletter.semianalysis.com/p/chinas-cxmt-is-set-to-challenge-dram) projects it to ~12% of global HBM wafer supply by 2028), though that ramp is back-loaded (the 2026 column could plausibly be near zero) and competes with higher-margin commodity DRAM for the same fab. + +**DUV capacity is abundant.** Ascend logic and CXMT memory both print on China's ArF-immersion (DUV) fleet; these AI chip projections only use a small fraction of it: + +DUV immersion fleet| 2026| 2028| Basis +---|---|---|--- +Fleet (tools)| ~290| ~400| ~200 base + ~90 ArFi bought in 2024 (mostly [NXT:1980Fi-class](https://www.asml.com/en/products/duv-lithography-systems/twinscan-nxt1980fi), 275-330 wph; 70% of ASML's 129, $5-7B); imports since slowed by Dutch controls ([AEI](https://www.aei.org/research-products/report/the-lithography-loophole-how-china-is-printing-its-way-to-chip-self-sufficiency/)) +Realized capacity (passes/yr)| ~520M| ~720M| fleet x ~1.8M realized passes/tool/yr (~5-6k wafers/day in production vs the [~2.4M nameplate](https://www.aei.org/research-products/report/the-lithography-loophole-how-china-is-printing-its-way-to-chip-self-sufficiency/); cf. §1.1's ~360k EUV) +Ascend logic| ~5M (~1%)| ~17M (~2%)| ~8 → 28 kwspm [SMIC ≤7nm starts](https://newsletter.semianalysis.com/p/huawei-ascend-production-ramp) x ~40-60 passes/wafer ([SAQP quad-patterning](https://www.techpowerup.com/344000/chinese-smic-achieves-5-nm-production-on-n-3-node-without-euv-tools)) +CXMT HBM| ~4M (~1%)| ~12M (~2%)| ~30 → 100 kwspm [CXMT HBM starts](https://newsletter.semianalysis.com/p/chinas-cxmt-is-set-to-challenge-dram) x ~10 passes/wafer + +The AI draw is ~2% of fleet capacity in 2026, ~4% by 2028; the binding constraint is HBM. This is also why we don't expect China to simply flood its DUV fleet with AI chips: more Ascend logic without matching HBM doesn't add usable AI compute, SMIC's ≤7nm starts are capped by multi-patterning yield and cost rather than scanner availability, and most of the fleet profitably serves the mature-node economy (autos, industrial, appliances) that China is also trying to dominate. + +The four channels combined (M raw H100e): + +H100e (per year unless noted)| 2024| 2025| 2026| 2027| 2028 +---|---|---|---|---|--- +Domestic| 0.25M| 0.6M| 1.1M| 2.3M| 5.0M +Legal western| 0.2M| 0.15M| 0.35M| 0.8M| 1.6M +Loopholes (offshore DCs / rental)| 0.25M| 0.45M| 0.7M| 1.6M| 3.15M +Smuggled| 0.16M| 0.48M| 1.26M| 2.12M| 2.99M +**Total added /yr**| **0.9M**| **1.7M**| **3.4M**| **6.8M**| **12.7M** +→ China stock (end of year)| 1.3M| 2.9M| 6.3M| 13.2M| 25.9M +→ China % of world| 14.7%| 13.1%| 11.1%| 9.7%| 9.0% + +#### 1.3.2 Results + +Combining the world total (§1.1), the China forecast (§1.3.1), and an 85/15 US/RoW split of the remainder gives ownership by region: the US holds the great majority throughout, China ~15% → ~9%. + +stockflowlevel% share + +USChinaRoW + +075M150M225M300M202517M202643M2027104M18M2028224M26M39M2029 + +Who owns the AI compute \- ai-2040.com + +Owner = the nation of the company that owns or leases the compute. World compute is from §1.1; China is the §1.3.1 bottom-up forecast; the non-China remainder is split 85/15 US/RoW. + +### 1.4 Datacenter sizes + +In this section we forecast how the AI compute from §1.1 will be distributed across datacenter sizes, from gigawatt-scale training campuses down to the long tail of small regional sites. The method is relatively naive (and uncertain). We take the total AI compute forecast (from §1.1), and spread it across a total datacenter count, fit a curve of datacenter sizes with three concentration anchors, and then split the result across the US, China, and the Rest of the World ownerships (from §1.3). + +10M+1 DCs·10M·3.5%1M-10M74 DCs·149M·51.4%100K-1M257 DCs·76M·26.4%10K-100K497 DCs·15M·5.3%1K-10K823 DCs·2.6M·0.9%<1K1.2K DCs·397K·0.1%In-transit35M·12.0%AI consumer255K units·938K·0.32%US · 224M · 77%China · 26M · 9%RoW · 39M · 14% + +Compute locations by datacenter size (Jan 1 2029) \- ai-2040.com + +Datacenter sizes model + +**Setup.** We take the total AI compute _C_ (§1.1) and a datacenter count _N_ (~1,000 sites in 2025, growing at a default 30%/yr to reach ~2,850 sites by 2029). We then allocate the datacenter sizes based on a size distribution. + +**Anchors.** We pin the _shape_ of the size distribution with three concentration anchors, the share of all compute held by the largest 1, 10, and 100 clusters, S1,S10,S100S_1, S_{10}, S_{100}S1​,S10​,S100​, which we forecast to be roughly 4%, 20%, and 70% (hand-set, adjusted up from [Epoch's datacenter dataset](https://epoch.ai/data/data-centers)), held flat over time: + +S1=c1/CS10=(c1+⋯+c10)/CS100=(c1+⋯+c100)/C\begin{aligned} S_1 &= c_1 / C \\\ S_{10} &= (c_1 + \dots + c_{10}) / C \\\ S_{100} &= (c_1 + \dots + c_{100}) / C \end{aligned}S1​S10​S100​​=c1​/C=(c1​+⋯+c10​)/C=(c1​+⋯+c100​)/C​ + +**Fit.** We then model log-size as a cubic in log-rank, with four coefficients a,b,c,da, b, c, da,b,c,d: + +log⁡(ci)=a+blog⁡(i)+clog⁡(i)2+dlog⁡(i)3\log(c_i) = a + b\log(i) + c\log(i)^2 + d\log(i)^3log(ci​)=a+blog(i)+clog(i)2+dlog(i)3 + +Those four coefficients are pinned _exactly_ by four equations, the three anchors plus the total (a cubic in log-rank has four degrees of freedom, matching the four constraints, so the curve is fully pinned with no free parameters left over): + +c1/C=S1(c1+⋯+c10)/C=S10(c1+⋯+c100)/C=S100c1+⋯+cN=C\begin{aligned} c_1 / C &= S_1 \\\ (c_1 + \dots + c_{10}) / C &= S_{10} \\\ (c_1 + \dots + c_{100}) / C &= S_{100} \\\ c_1 + \dots + c_N &= C \end{aligned}c1​/C(c1​+⋯+c10​)/C(c1​+⋯+c100​)/Cc1​+⋯+cN​​=S1​=S10​=S100​=C​ + +The first is immediate, since log⁡1=0\log 1 = 0log1=0 gives a=log⁡c1=log⁡(S1C)a = \log c_1 = \log(S_1 C)a=logc1​=log(S1​C); the other three are solved numerically. This recovers every datacenter size cic_ici​ from just _N_ , _C_ , and the three anchors. + +**Regions.** To split the curve across the **US, China, and the Rest of the World** , we forecast each region's ownership share SrS_rSr​ of total compute (giving it a target Tr=SrCT_r = S_r CTr​=Sr​C) and assign each datacenter to a region in proportion to that share. An optional _skew_ θr\theta_rθr​ can tilt a region toward smaller or larger sites (to match a region whose biggest clusters are known to differ, e.g. China) while keeping its total fixed, but it is **off by default** (θCH=θRoW=0\theta_{CH} = \theta_{RoW} = 0θCH​=θRoW​=0), so each region is just a same-shape slice of the global curve. + +**Default parameters.** + +parameter| default| what it controls +---|---|--- +N0N_0N0​ (sites in 2025)| 1,000| datacenter count in 2025 +growth ggg| 30% / yr| datacenter-count growth (→ ~2,850 sites by 2029) +S1S_1S1​| 4%| compute share of the single largest cluster +S10S_{10}S10​| 20%| compute share of the largest 10 +S100S_{100}S100​| 70%| compute share of the largest 100 +skew θCH,θRoW\theta_{CH}, \theta_{RoW}θCH​,θRoW​| 0 (off)| optional regional size tilt; 0 = proportional + +These concentration shares are hand-picked and uncertain, adjusted up from [Epoch's observed current dataset](https://epoch.ai/data/data-centers) on the assumption that it does not have full coverage of frontier clusters (its largest tracked sites cover ~15% of global AI compute as of late 2025). + +## Part 2. Post-deal compute (post-2029) + +In this section we forecast what happens after Plan A is implemented, as well as the 1-year default counterfactual without Plan A. + +### 2.1 The no-deal counterfactual + +To predict one more year of the default (no-deal) path, we extend the §1.1 method (and §1.3 for China's domestic additions) to the installed stock at the end of 2029. Under Plan A we expect the temporary R&D pause and general uncertainty during the deal's implementation to make capital spending temporarily lower. We approximate this by having the debt-financed slice of the buildout collapse to zero under Plan A, while cash capex stays on trend. With the fast implementation of inference-only verification and increased inference-compute allocation, we expect the AI revenue forecast to stay on par with the default counterfactual. + +Overall, we think a good implementation of the deal can provide confidence to companies that R&D will be allowed to resume, and that given the level of realized capabilities and continued improvements once R&D resumes (which will be at a slowed down but still fast pace), we expect AI investment to still be extremely economically attractive. We expect this to hold despite increased transparency, because without algorithmic moats the economic value will accrue to the owners of compute (who can therefore still be the ones to train the best models first and serve them to the widest user base). + +no dealdeal (Plan A) + +2029 capexdebt-sourced takes the hit$3.5T$2.4T-31%H100e added (2029)capex × efficiency, net of EUV surcharge247M199M-19%Installed H100e (end of 2029)cushioned by the existing base537M489M-9% + +Plan A counterfactual \- ai-2040.com + +No deal vs Plan A, one year past the deal (through end-2029). Plan A collapses the debt-financed slice of capex (~31% off 2029 spending) while cash-flow-funded capex holds. + +### 2.2 AI-relevant compute after the deal + +Further into Plan A, we have increased uncertainty over compute growth, especially as AI and robots become capable of automating hardware R&D and production, and drastically increasing the previously human-constrained workforce. Given the uncertainty, we think it is reasonable to use a relatively simple model (Jones and Wright's-law learning) that takes the §1.1 compute path as its starting point and runs it to 2040. Given the hardware explosion we expect by default with unconstrained AI and robot deployment on hardware R&D and production, and the fact that this would be destabilizing (among other things, increasing the ease of defecting on the deal and verification burden within the deal) Plan A features a cap-and-trade regime to slow down the hardware explosion to controllable (but still drastically increased) hardware levels. + +**Why we expect hardware to grow fast.** We can think of two separate forces lowering the cost per unit of compute: + + * **Design improvements (modelled with Jones).** Design improvements can lead to more compute per fixed unit of AI chip production, with better chip layouts, memory designs, transistor density, packaging etc. This is R&D, so we can model it with the labor spent on it, but with diminishing returns as the easy ideas get used up ([Jones 1995](https://web.stanford.edu/~chadj/JonesJPE95.pdf), calibrated by [Bloom et al. 2020](https://web.stanford.edu/~chadj/IdeaPF.pdf)) and with diminishing returns to more parallel labor at a given time. Think of this as the historical improvements in 'compute per wafer' (including node shrinks, which we bucket under design R&D since they come from chip and equipment research rather than from production scale). Roughly 2 million people work in semiconductors today; by the early 2030s our forecast is that combined human, AI, and robot effort on chip and equipment R&D is on the order of 100x that. Under our model's parameters this only speeds up the 'Moore's law' progress by around 1.1x. + + * **Production scaling (Wright).** Pure manufacturing scale of units of compute can lead to 'learning by doing' efficiency improvements, e.g., with yield, throughput, uptime effects, and economies of scale. Think of this as the improvements in 'wafers per dollar' historically. We model unit cost falling about 15% per doubling of cumulative production (similar to long-run semiconductor and solar rate). We don't know if AI and robots automating fab construction and operation will speed up this trend, but at least think it's a reasonable baseline that these gains from scale will continue separately to design improvements. + + + + +We treat the cost per unit of compute ($ / H100e) as the product of the two factors (compute / unit produced) and (units produced / $), where historically the unit has been 'wafers' but may change with future paradigms. Each factor is normalized to 1 in 2025 and falling thereafter. We take the §1.1 endpoint as the starting point and run these two models forwards to 2040. + +The Jones & Wright's-law cost model + +The cost per unit of AI compute is one 2025 anchor times two indices that both start at 1 and fall over time, a manufacturing index W(t)W(t)W(t) and a design index D(t)D(t)D(t): + +MC(t)=MC2025×W(t)×D(t)MC(t) = MC_{2025} \times W(t) \times D(t)MC(t)=MC2025​×W(t)×D(t) + +**Wright 's law (manufacturing).** Cost falls with cumulative production, the classic learning curve: + +W(t)=(Qcum(t) / Qcum(2025))−bW(t) = \left(Q_{\rm cum}(t)\,/\,Q_{\rm cum}(2025)\right)^{-b}W(t)=(Qcum​(t)/Qcum​(2025))−b + +so each doubling of cumulative output cuts manufacturing cost by 1−2−b1 - 2^{-b}1−2−b. We use b=0.25b = 0.25b=0.25, about 16% per doubling, similar to the long-run semiconductor learning rate (close to Swanson's law for solar). + +**Jones design R &D.** The annual design-cost decline scales with more labor doing R&D, with diminishing returns: + +rate(t)=rbase(Lcog(t) / Lcog(2025))λ,D(t+1)=D(t) (1−rate(t))\text{rate}(t) = r_{\rm base}\left(L_{\rm cog}(t)\,/\,L_{\rm cog}(2025)\right)^{\lambda}, \qquad D(t{+}1) = D(t)\,(1 - \text{rate}(t))rate(t)=rbase​(Lcog​(t)/Lcog​(2025))λ,D(t+1)=D(t)(1−rate(t)) + +where LcogL_{\rm cog}Lcog​ is combined human and AI cognitive labor (which scales with deployed compute) which is taken from the 'effective labor' quantity from our economic model that combines AI / robot labor quantities with its capability (% of tasks it is able to automate). rbaser_{\rm base}rbase​ is the base pace, and λ\lambdaλ is the "stepping on toes" adjustment ([Bloom et al. 2020](https://web.stanford.edu/~chadj/IdeaPF.pdf)): a 100x research force buys far less than 100x the progress. + +Parameter| AI chips +---|--- +Wright learning bbb| 0.25 (~16% / doubling) +Jones base rate rbaser_{\rm base}rbase​| 25% / yr +Jones stepping-on-toes λ\lambdaλ| 0.05 +2025 cost to make| $5,000 / H100e +Cumulative output, 2025| 20M H100e + +The base rate (25%/yr) tracks the ~1.35x/yr compute-efficiency trend, and λ=0.05\lambda = 0.05λ=0.05 keeps stepping-on-toes strong: a 100x AI research force buys only ~1.3x the design-progress rate, so even the swelling AI workforce only modestly accelerates chip design. The anchor above is the cost to _make_ a chip (marginal production cost, ~$5K in 2025), which is what the cost figure tracks; markups plus datacenter and power bring the all-in cost to ~$33K, which is what the $-invested figure counts. Both fall along the same W×DW \times DW×D curve. Below we run this loop with and without Plan A's two constraints, the compute cap and the R&D controls. + +**The hardware explosion has two throttles in Plan A.** Without them there is too much progress for the deal's stability (ease of defection and verification burden). There is a feedback loop where more compute leads to more hardware R&D and production, leading to more compute and cheaper hardware. Plan A has two constraints that act on different parts of that loop: capping production levels (cap-and-trade policy from 2032 onwards) and limiting R&D deployment (to avoid compute becoming too easy and cheap to manufacture despite the quantity caps). + +Plan A (cap + hardware R&D controls)no cap, hardware R&D controlledno cap + hardware R&D unconstrained + +H100e$ invested$ / H100e + +5Q1T2029203020322034203620382040100M1B10B100B1T10T100T1Q10Q100Qinstalled H100-equivalents, log scale + +Capped Plan A compute vs. default \- ai-2040.com + +Plan A compute with different constraints lifted. The two constraints are the cap & trade policy in 2032, which restricts the quantity produced, and the hardware R&D deployment limits, which stop an exploding hardware R&D workforce from driving compute costs too low and chips too easy to manufacture. + +Assumptions for each scenario + +All three scenarios follow Plan A's compute path until the cap starts to bind (~2034), then differ on two levers: + + * **Compute cap** (cap-and-trade, 2032 on). On: installed compute is pinned to the Plan A policy path, ~1T H100e by 2040. Off: an unconstrained reinvestment loop sets the quantity and it runs away. + + * **Hardware R &D.** Controlled: chip-design progress is held to its historical Moore's-law pace. Unconstrained: the hardware-R&D effective labor is the economic model's effective AI labor (capability x quantity), so the design-cost decline accelerates with it (base rate 25%/yr, stepping-on-toes λ=0.05\lambda = 0.05λ=0.05, capped at 60%/yr). + + + + +Scenario| Compute cap| Hardware R&D| What happens +---|---|---|--- +Plan A| on, ~1T by 2040| controlled| the bounded path the rest of the document uses +no cap, hardware R&D controlled| off| controlled| quantity explodes to thousands of times the cap +no cap + hardware R&D unconstrained| off| unconstrained| quantity and cost both run away, off the chart + +Effective labor is the same series in every scenario (the econ model's A_eff, capability x quantity, scaled by that scenario's compute); the R&D control acts on how much of it becomes design progress, not on the labor total: + +Hardware R&D| effective AI labor, 2040 (HE = human-equivalents)| design-cost decline +---|---|--- +controlled (Plan A)| ~400B HE| ~10%/yr, held at the historical Moore's-law pace +controlled (no cap)| ~2x10^15 HE| ~10%/yr, still held despite the vastly larger workforce +unconstrained (both)| ~2x10^21 HE| 25%/yr, scaling to a 60%/yr cap + +The unconstrained rate is exactly the economic model's Jones formula (25%/yr x (A_eff growth)^0.05, λ=0.05\lambda = 0.05λ=0.05); the controlled cap is Plan A's R&D-control policy layered on top, which the base economic model does not impose. + +**What share of effective labor does R &D?** A fixed, constant share of the effective labor; only its growth matters, so the share cancels and we do not commit to a specific percentage. That effective labor is the economic model's effective AI labor (A_eff, capability x quantity), scaled by each scenario's compute relative to Plan A, so R&D progress responds to both rising capability and rising quantity. **Robots** are held on their fixed Plan A cap-and-trade path and are not varied across these scenarios. + +Supply-side implications and sanity check. + +The compute production possible by default in Plan A is modelled relatively naively, and the cap-and-trade trajectory is set considerably lower due to other constraints (verification and deal stability). In this box we motivate the cap-and-trade trajectory being possible with an illustration of what it implies on the supply side. In particular we look at the lithography it would require compared to historical growth. + +**Method.** We take Plan A's annual production (new H100e per year, rising to roughly 570 billion H100e/yr by 2040) and convert it to silicon wafers (assuming this remains the paradigm), convert the wafers to lithography exposure-passes needed and then divide by one lithography machine's throughput (realized passes per machine-year) to get the machine-years needed. + +machines needed(t)=new H100e(t) × litho passes per H100e(t)passes per machine-year(t)\text{machines needed}(t) = \frac{\text{new H100e}(t)\;\times\;\text{litho passes per H100e}(t)}{\text{passes per machine-year}(t)}machines needed(t)=passes per machine-year(t)new H100e(t)×litho passes per H100e(t)​ + +**Default lithography fleet.** We anchor at 2029 (about 0.7 EUV passes per H100e, from §1.1) and compare against the lithography fleet on its historical path (about 620 operating machines in 2029, growing roughly 20%/yr; ASML ships about 125/yr today). We also let machine throughput rise about 10%/yr, matching the historical climb in wafers-per-hour (roughly 125 to 220 over 2019 to 2024). + +**Default compute density.** Over time H100e/wafer has been increasing [around 35%/year](https://epoch.ai/data/machine-learning-hardware?view=graph); we look at three different possibilities for the lithography buildout it would imply vs. the default trend. + +compute-per-wafer growth| 0%/yr (frozen 2029 tech)| 35%/yr (continued tech trend)| 50%/yr (faster tech trend) +---|---|---|--- +lithography machines needed by 2040| ~1,100,000| ~14,000| ~4,400 +vs the fleet's historical trend (~4,300 by 2040)| ~255x| ~3x| ~1x +implied annual production vs today (~125/yr)| ~4,400x| ~29x| ~6x + +We expect something like the middle column to be what happens in Plan A. + +We can also express Plan A's buildout in dollars, as a check on the investment it implies. Each year's implied investment is the compute added that year times an all-in cost per H100e, which is the production cost from the model above times a markup covering everything else a dollar of AI capex buys (chip-maker margin, datacenters, power, networking). We calibrate the markup so that 2025 to 2028 reproduce Part 1's capex forecast exactly (~$2.4T in 2028) and 2029 the §2.1 deal year: this gives roughly 7x to 8x. Our best guess is that the markup then declines to ~1.2x by 2040, since robot-built fabs and datacenters should sell hardware near production cost, and the power and datacenter share of each H100e shrinks with the efficiency gains below. + +Investment $ / yr$ / H100e + +$100B$1.0T$10T$100Tdeal →$2.43T$47T202620292032203520382040 + +Implied AI compute investment \- ai-2040.com + +We are very uncertain about the uncapped paths, but we think the unconstrained default is likely to be orders of magnitude higher than our capped path. The effective labor available for hardware R&D and chip production scales enormously once AI and robots are capable of that work. We think the stacked Jones and Wright's-law learning-curve dynamics are a reasonable naive approximation of what this might look like. + +Historically, chip design improvements have been closely tied to power-efficiency improvements, so we model compute power efficiency on the same Jones design trend as the cost model above, using the same parameters (25%/yr base rate, stepping-on-toes λ=0.05\lambda = 0.05λ=0.05) at Plan A's constrained R&D pace, which leads to the following trajectory (notably not breaking past the CMOS transistor theoretical limit that [Epoch estimates](https://epoch.ai/blog/limits-to-the-energy-efficiency-of-cmos-microprocessors)). + +10005002001005020105CMOS limit~200× H100 (Epoch)20252027202920312033203520372039204027 W + +Compute power efficiency \- ai-2040.com + +Total facility power per H100e (in [Epoch's terms](https://epoch.ai/data-insights/gpus-power-usage-in-ai-data-centers): chip TDP plus the server's other components plus cluster-level overhead like cooling, i.e. the full power draw at the grid), fleet stock-average, from the econ model's Plan A energy trajectory (same constrained Jones design trend as the cost model). It reaches ~37× the 2025 chip by 2040, still several× above the [CMOS practical limit](https://epoch.ai/blog/limits-to-the-energy-efficiency-of-cmos-microprocessors) Epoch estimates (~200× a current H100e). + +### 2.3 Company concentration + +In this section we show our projections for how compute is distributed among companies over time. We track **AI company end-users** (the company that actually uses the compute) as opposed to **owners** (who may use or rent their compute to others). Early on most compute is [owned by hyperscalers](https://epoch.ai/data/ai-chip-owners?view=graph&tab=h100_equivalents) in a concentrated way (Google owns ~25% of compute) but the [AI company end-users are less concentrated](https://epochai.substack.com/p/frontier-labs-dont-use-most-ai-compute), with the largest using around 10% of global compute in 2026 (GDM and OpenAI are similar). + +top companytop 3top 10 + +0%20%40%60%80%20262028203020322034203620382040~8%~20%~5%~13%~31% + +Company compute concentration \- ai-2040.com + +The share used by AI companies has been growing over time, and we project this trend forwards until Plan A is implemented. After 2030 the Plan A transparency regime lowers barriers to entry: investment in AI grows and spreads more evenly, and compute becomes distributed across more companies. By 2040 the largest holds only about 5% of world compute and there are around 20 frontier companies (defined as anyone within 4x of the largest). + +USChinaRoWfuller fill = frontier, lighter = rest + +Year**2029** highlighted = frontier (top 3) + +AnthropicOpenAIGDMSpaceXMetaother USother Chinaother RoWUSChinaRoW + +Compute end-users by company \- ai-2040.com + +Method for company compute concentrations + +The method is relatively naive, where we set one worldwide rank-size law for the shape of the distribution, hand set to match forecasts, and then companies are assigned to ownership regions to match the §1.3 split. Across a universe of NNN end-users worldwide, the share of the rank-iii company is + +si ∝ exp⁡(σ Φ−1(N−i+0.5N)+τi),i=1,…,N,s_i \;\propto\; \exp\left(\sigma\,\Phi^{-1}\left(\frac{N-i+0.5}{N}\right) + \frac{\tau}{i}\right), \qquad i = 1,\dots,N,si​∝exp(σΦ−1(NN−i+0.5​)+iτ​),i=1,…,N, + +normalized so the shares sum to one (Φ−1\Phi^{-1}Φ−1 is the standard-normal quantile of rank iii). It has two shape parameters: σ\sigmaσ sets the overall spread (concentration, which drives the top three), and τ\tauτ is a head tilt that decays as 1/rank1/\text{rank}1/rank, moving the single leader relative to the next two with less effect on the tail. + +We choose the world's top-1 and top-3 end-user shares directly each year and solve (σ,τ)(\sigma, \tau)(σ,τ) so the distribution reproduces both. The leader is a minority early (frontier AI companies don't yet use most AI compute), rises into the 2029 deal to about a fifth of world compute with the top three at nearly half, then both decline as Plan A spreads compute (top-1 about 10% by 2035 and 5% by 2040; top-3 about 25% then 13%). A better justification of this forecast is currently future work. We wanted to create a model of "AI market power" / "willingness to pay" for AI compute that was a combination of algorithmic IP (which becomes uniformly distributed after the deal), non-algorithmic IP/capital (datasets, users, brand loyalty, talent loyalty, etc.) and finally capital more broadly, and build a model of compute allocation over time tracking this quantity. This was complicated though and we ran out of time. + +Each company is then assigned to an ownership region (US, China, rest-of-world) by a hard-capped greedy fill, so regional totals match the §1.3 ownership split exactly every year through 2029 and then matching the Plan A ownership splits which converge towards the cap-and-trade deal allocations. A per-region size bias shifts only _which_ companies a region holds, never its total: the US is neutral and holds the biggest developers; China is flat and dispersed through 2030, then a state-backed core consolidates at 2031 (its top developer jumps to about half of China's compute and lands a frontier developer) before gradually flattening again by 2040; the rest-of-world is flatter and flattens further toward the end, so the single largest developer worldwide stays in the US even once the rest-of-world's ownership share overtakes it. + +Through 2030 (the deal era) the frontier set is the three largest end-users worldwide (the leading US developers are named, in a hand-set per-year order). From 2031 we drop names and call an end-user frontier if it is within 4x of the largest end-user worldwide. + +parameter| value| role +---|---|--- +worldwide end-user universe (NNN)| 200| the rank-size law's company count +top-1 end-user share (2026 / 2029 / 2035 / 2040)| 8% / 20% / 10% / 5%| hand-set; (σ,τ)(\sigma,\tau)(σ,τ) solved to match +top-3 end-user share (2026 / 2029 / 2035 / 2040)| 23% / 45% / 25% / 13%| hand-set independently of top-1 +§1.3 ownership split at 2040 (US / CH / RoW)| 35% / 20% / 45%| regional totals matched exactly each year +China size bias| flat, then +5 at 2031, easing to +2 by 2039| dispersed early; consolidates at 2031; flattens by 2040 +rest-of-world size bias| −0.6, to −2.5 by 2039| flattens toward the end so the US keeps the largest developer +frontier rule| top-3 (≤2030) / within 4x (≥2031)| deal era names the top three; later, within 4x of the largest + +### 2.4 Compute allocation + +In this section we forecast allocation of frontier AI compute in Plan A. In 2025 OpenAI spent [around half of its compute on inference](https://www.theinformation.com/articles/openai-boost-revenue-forecasts-predicts-112-billion-cash-burn-2030?rc=9mzoog) and more like [one third on inference in 2024](https://epoch.ai/data-insights/openai-compute-spend) and the rest on training and experiments (of which only around 10% on final frontier training runs). We project this current split forwards naively until the 2029 deal, and then forecast each share conditional on Plan A. + +**Frontier AI compute allocation under Plan A** + +public deploymentinternal inferencetrainingexperimentssafetyoffline / transition + +0%25%50%75%100%2029 dealpublicdeploymentexperimentstrainingsafety20262028203020322034203620382040 + +### 2.5 AI copies and speed + +In this section we forecast how many copies of the frontier model can be run at once and at what token/second speed. These are both imperfect metrics, because it may be difficult to know how to count a 'copy' under different scaffolds / architectures in the future, but we assume each parallel instance of the model weights being served, so copies = replicas x batch. Speed is also imperfect, because we probably care more about the speed of task completion, but we just focus on decode tokens per second, even though the speed comparison in tokens might break down with future paradigm changes, e.g., neuralese instead of chain of thought. + +Up front: these forecasts are concrete best guesses, and we think the honest uncertainty is multiple orders of magnitude. Both the AI paradigm (architectures, sparsity, agent scaffolds, whether 'copies' is even the right frame) and the hardware paradigm (the memory and bandwidth tradeoffs of chips, how much silicon gets specialized for inference) could change several times before 2040. The method itself is simple, and we prefer to state it plainly: we guess the frontier model's size, we guess what share of the world's compute runs inference, and then we get copies by dividing the inference fleet's total memory by the bytes each copy takes, and we get speed from how quickly memory bandwidth can read the bytes each token needs. The rest of this section explains how we make each of these guesses. + +**Model size.** We combine the global compute growth from the previous sections with the leading company's concentration (§2.3) and training allocation (§2.4) to get frontier training compute over time (roughly 32x today's by 2029 and 160,000x by 2040). We then map these compute numbers as well as a sparsity forecast to model size (total parameters) with a simple model: + +P = P2026(CC2026)α(σσ2026)βP \,=\, P_{2026}\left(\tfrac{C}{C_{2026}}\right)^{\alpha}\left(\tfrac{\sigma}{\sigma_{2026}}\right)^{\beta}P=P2026​(C2026​C​)α(σ2026​σ​)β + +Where PPP is total parameters (P2026≈10P_{2026} \approx 10P2026​≈10T, today's level), CCC is frontier training compute, and σ\sigmaσ is sparsity. The compute exponent α\alphaα is the sensitivity of model-size scale up to compute scale up. Simple Chinchilla-optimality implies ~0.5, but we have wide uncertainty; there are reasons it might increase, e.g., due to data availability (fitting our formula to [Nesov's](https://www.lesswrong.com/posts/yLHiQGCPdvzL9fBn3/model-size-scaling-in-2023-2031) own compute and sparsity numbers gives α≈0.8\alpha \approx 0.8α≈0.8), or for example, intentional undertraining due to e.g., making it harder to steal the weights. There are also reasons it might decrease, e.g., more overtraining for serving ([Sardana and Frankle](https://arxiv.org/abs/2401.00448)), reinforcement learning, and other unknown unknowns about paradigm changes. We just assume α=0.5\alpha = 0.5α=0.5 and draw α∈[0.2,0.8]\alpha \in [0.2, 0.8]α∈[0.2,0.8] as a range on the figure to show what we think is a plausible range. The sparsity exponent β=0.74\beta = 0.74β=0.74 is tentatively better pinned (though the sparsity MoE paradigm might change). + +We think it's valuable to include the sparsity assumption because it might affect the speed and copies significantly. + +**Frontier model size, 2022 to 2040** + +Model sizeTraining compute + +reportedours (α=0.5)α ∈ [0.2, 0.8]Nesov + +1101001k10k100k1M2022202620292032203520382040Trillion parameters2026 anchor2029 dealGPT-4 ~1.8Tα=0.8α=0.2~650T (Nesov)31,000T~158T + +**Our model size forecast for plan A** + +total = 10T × (compute growth)α × (sparsity growth)β, with α = 0.5 and β = 0.74. + +Year| 10T × computeα| × sparsityβ| = ours (total)| α ∈ [0.2, 0.8]| Nesov +---|---|---|---|---|--- +2026| 10T| x1.0 (8x)| 10T| 10T to 10T| 10T +2027| 20T| x1.7 (16x)| 33T| 22T to 49T| 30T +2028| 35T| x2.8 (32x)| 97T| 46T to 205T| 240T +2029| 57T| x2.8 (32x)| 158T| 56T to 449T| 650T +2030| 83T| x3.1 (37x)| 255T| 72T to 904T| — +2032| 229T| x3.8 (48x)| 862T| 132T to 5,638T| — +2035| 1,411T| x4.7 (64x)| 6,576T| 337T to 128,155T| — +2040| 3,992T| x7.8 (128x)| 31,063T| 854T to 1,129,715T| — + +Differences with Nesov's numbers + +**How Nesov gets 650T:** Both Nesov's [forecast](https://www.lesswrong.com/posts/yLHiQGCPdvzL9fBn3/model-size-scaling-in-2023-2031) and ours start from training compute; the difference is mainly that Nesov's 2029 frontier run is 2x larger at **8e28 FLOP** assuming **30x sparsity**. At 30x sparsity compute-optimal ratio is **240 tokens per active parameter** (from a [sparse-MoE scaling paper](https://arxiv.org/abs/2501.12370)) so 8e28 FLOP would want **7.4T active parameters** on ~1,800T tokens. But he assumes only ~200T tokens of unique data exist, an **8.5x shortfall** , which he splits with this [data-constrained scaling result](https://arxiv.org/abs/2605.01640) to get roughly 8.50.5≈8.5^{0.5} \approx8.50.5≈ **2.9x more active parameters** (7.4T → 22T) and **2.9 epochs** of repetition. We don't make the same low-data-regime adjustment as Nesov, because we think multiple factors such as more reinforcement learning, changes to the training paradigm, synthetic data generation, etc. could pull the other way. Given the uncertainty on these factors that could pull in both directions we prefer to use a naive simpler baseline. + +**Nesov’s model-size forecast** + +From [“Model size scaling in 2023–2031”](https://www.lesswrong.com/posts/yLHiQGCPdvzL9fBn3/model-size-scaling-in-2023-2031). + +Year| Pretraining compute (FLOP)| Sparsity| Active params| Total params| Data shortfall| Epochs +---|---|---|---|---|---|--- +2026| 1.3e27| 8x| 1.3T| 10T| —| 1 +2027| 6e27| 8x| 2.9T| 30T| 1.75x| 1.3 +2028| 2e28| 30x| 7.9T| 240T| 4.5x| 2.1 +2029| 8e28| 30x| 22T| 650T| 8.5x| 2.9 +2030| 1e29| 30x| 26T| 790T| 10x| 3.1 +2031| 2.2e29| 30x| 48T| 1,400T| 15x| 3.9 + +**Inference-compute fleet.** We take the §2.4 inference share of world compute (public + internal), and hold its memory and bandwidth per H100e ratios (§1.1) flat after 2029. We split this compute into ordinary general-purpose chips and inference-specialized chips, a hand-set scenario we describe below. + +**Inference memory by chip type** + +general-purposeinference-specialized + +MemoryBandwidthCompute + +Serving memory by chip type (exabytes, log). Specialized memory is derived from demand (replica weights + KV per copy): ~2.3 ZB by 2040 vs ~10 ZB general-purpose. + +0.11101001k10k202620282030203220342036203820402029 deal10 ZB2 ZB + +**Copies of the frontier model** + +all copies + +Concurrent frontier copies served (replicas times batch), 2026 to 2040 (log). + +10M100M1B202620282030203220342036203820402029 deal30.3M29.0M186M + +**Copies.** We get copies by dividing the inference fleet's total memory by the bytes each copy takes. Each copy needs its share of the model weights (a full set of weights is ~5 TB today at 4-bit, and ~15 PB by 2040 at our central model size, shared across the copies served in the same batch) plus its own context. Rather than modelling batch sizes and KV caches in the main text, we just state their combined effect: we assume context takes up about half of general-purpose inference memory today, falling to ~5% by 2040 because the weights grow much faster than each copy's context does (the details are in the method box, and the batch assumption just trades off with context length assumptions and is arbitrary). This gives about 30 million copies today. The count dips slightly at first, because the jump in sparsity grows the model faster than the inference fleet, and ends up at very roughly 200 million by 2040, with most copies on general-purpose hardware and the rest being the much faster copies on inference-specialized chips described below. + +**Weights vs. context assumptions** + +Year| Avg context| KV / copy| Context / replica| Weights / replica| Context share +---|---|---|---|---|--- +2026| 200k| 20 GB| 5.1 TB| 5.0 TB| 51% +2029| 341k| 68 GB| 17.4 TB| 79.2 TB| 18% +2032| 580k| 220 GB| 56.3 TB| 430.9 TB| 12% +2035| 988k| 896 GB| 229.3 TB| 3.3 PB| 7% +2040| 2.4M| 3.3 TB| 856.1 TB| 15.5 PB| 5% + +**Copies & speed model summary** + +Model size, copies, and speed at milestone years (Plan A). + +Year| Model (total/active)| Copies (gen/inf)| All copies| Speed (gen/inf)| Avg speed| Total tok/s (gen/inf) +---|---|---|---|---|---|--- +2026| 10T / 1.3T| 30.3M / 8k| 30.3M| ~57 / ~1.0k| ~57| 1.7B / 7.9M +2028| 97T / 3.0T| 23.4M / 60k| 23.5M| ~65 / ~1.3k| ~68| 1.5B / 78.6M +2029| 158T / 4.9T| 28.9M / 123k| 29.0M| ~67 / ~1.5k| ~73| 1.9B / 185M +2032| 862T / 18T| 63.1M / 1.6M| 64.6M| ~89 / ~3.0k| ~161| 5.6B / 4.7B +2035| 6,576T / 103T| 154M / 21.7M| 176M| ~119 / ~6.0k| ~844| 18B / 130B +2040| 31,063T / 243T| 150M / 36.2M| 186M| ~192 / ~25.0k| ~5.0k| 29B / 904B + +**Speed of a frontier copy** + +inference-specializedgeneral-purposeaverage + +Per-copy decode speed by chip type (tokens/sec, log). + +101001k10k20262028203020322034203620382040human working pace (~10 tok/s)~25.0kavg ~5.0k~192 + +**Speed.** For speed we look at memory bandwidth. On general-purpose hardware, each decode step has to read the weights plus the context of every copy in the batch, so the number of tokens each copy gets per second is just how many times per second the hardware can read through its own memory. Memory bandwidth on the inference fleet is about 31x its memory capacity, giving ~31 full read-throughs per second, or ~57 tok/s per copy after a ~2x speculative-decoding gain and ~90% utilization. Conveniently, this guess does not depend on batch, context, model size, or sparsity, because a bigger memory footprint comes with proportionally more memory bandwidth. Speed then grows only as fast as bandwidth per H100e outpaces memory capacity per H100e, which we assume happens at ~10%/yr after 2029, reaching ~190 tok/s by 2040. + +**Inference-specialized chips.** The biggest wildcard is chips specialized for serving specific models, with weights held in on-chip SRAM (à la [Cerebras](https://www.cerebras.ai/blog/llama-405b-inference), [Groq](https://groq.com/blog/inside-the-lpu-deconstructing-groq-speed)) or baked directly into the die (à la [Etched](https://www.etched.com/), [Taalas](https://www.taalas.com/)). Here we are guessing at every step: we hand-set the share of new compute production these chips take (0.2% in 2026, 1% at the deal, 14% by 2040, adding up to about half of the inference fleet's compute by then), and we hand-set how fast they serve each copy (~1,000 tok/s today, anchored to Cerebras' measured 969 tok/s on Llama-405B, rising to ~25,000 tok/s by 2040 once the 2035 pause freezes frontier models and makes it worth engineering chips close to the latency floor). Running each copy faster directly costs copies (copies = total token throughput / speed of each copy), so these chips add relatively few copies, but under these assumptions they end up producing almost all tokens by 2040 and pull the average speed from ~100 tok/s today to very roughly 5,000 tok/s, which is most of the speed growth in the figure above. If these chips don't take off, average speeds instead stay near the general-purpose ~190 tok/s path. The table below shows how much of its compute and bandwidth each part of the inference fleet actually uses. + +**Decode utilization of each resource, by chip type** + +Share of each chip type’s bandwidth and compute that decode uses; the binding (lower-ceiling) resource sets throughput. + +Year| Gen: bandwidth| Gen: compute| Inf-spec: bandwidth| Inf-spec: compute +---|---|---|---|--- +2026| 90% (binds)| 9.4%| 90% (binds)| 18% +2029| 90% (binds)| 3.3%| 90% (binds)| 21% +2032| 90% (binds)| 3.2%| 90% (binds)| 40% +2035| 90% (binds)| 3.3%| 73%| 90% (binds) +2040| 90% (binds)| 2.7%| 23%| 90% (binds) + +Method for AI copies and speed + +A frontier model has PPP total and Pa=P/σP_a = P/\sigmaPa​=P/σ active parameters at sparsity σ\sigmaσ, served 4-bit (b=0.5b = 0.5b=0.5 bytes/param), so one replica holds W=bPW = bPW=bP bytes of weights and each token costs 2Pa2P_a2Pa​ FLOPs (reading bPabP_abPa​ bytes). + +_Model size._ P=P2026 (C/C2026)α (σ/σ2026)βP = P_{2026}\,(C/C_{2026})^{\alpha}\,(\sigma/\sigma_{2026})^{\beta}P=P2026​(C/C2026​)α(σ/σ2026​)β with P2026=10P_{2026} = 10P2026​=10T, α=0.5 [0.2,0.8]\alpha = 0.5\;[0.2, 0.8]α=0.5[0.2,0.8], β=0.74\beta = 0.74β=0.74, and C=s1 Mworld τC = s_1\,M_\text{world}\,\tauC=s1​Mworld​τ the leader's training compute (top-1 share §2.3 x world stock x training allocation §2.4). Central: ~158T at 2029, ~31,000T at 2040. + +_Copies._ The inference fleet is the §2.4 inference share δ\deltaδ of the world stock (per-H100e memory and bandwidth from §1.1, flat after 2029), with an inference-specialized fraction fff built up from a hand-set flow of new compute (specialized chips only serve, so their whole stock serves). The general-purpose part serves copies out of its memory: a replica occupies the weights WWW plus a KV cache ccc for each of its BBB concurrent copies. The specialized chips instead serve as many copies as their token throughput supports at their chosen speed vspecv_\text{spec}vspec​, where throughput is utilization uuu times the smaller of their compute limit (FspecF_\text{spec}Fspec​) and their weight-read bandwidth limit (Λspec\Lambda_\text{spec}Λspec​); their memory is sized to match and never binds. So, + +copies=δM(1−f)W+B c B + u min⁡!(Fspec/2, Λspec/b)Pa vspec.\text{copies} = \frac{\delta M (1-f)}{W + B\,c}\,B \;+\; \frac{u\,\min!\big(F_\text{spec}/2,\; \Lambda_\text{spec}/b\big)}{P_a\, v_\text{spec}} .copies=W+BcδM(1−f)​B+Pa​vspec​umin!(Fspec​/2,Λspec​/b)​. + +_Speed._ With ρ(t)\rho(t)ρ(t) the inference fleet's bandwidth:capacity ratio (how many times per second its memory can be read in full), + +speedgp=u ρ(t) κ,speedspec=vspec(t).\text{speed}_\text{gp} = u\,\rho(t)\,\kappa, \qquad \text{speed}_\text{spec} = v_\text{spec}(t) .speedgp​=uρ(t)κ,speedspec​=vspec​(t). + +On general-purpose hardware a batched decode step reads the same footprint (W+BcW + BcW+Bc) whose bandwidth the replica owns (ρ(W+Bc)\rho(W + Bc)ρ(W+Bc)), so batch and context cancel and only the decode gain κ\kappaκ times the utilization uuu remains; this presumes serving stays at the large-batch, throughput-optimal operating point. Sparsity does not move realized speed (it cuts bytes and FLOPs per token, not layer count); it enters through model size and the specialized chips' per-token compute. The headline average speed is copy-weighted. + +Every input, ordered by the guess it belongs to (model size, then the inference fleet, then copies, then speed): + +parameter| value| note +---|---|--- +2026 anchor| 10T total params, 8x sparse| today's frontier +compute exponent (α\alphaα)| 0.5, band [0.2, 0.8]| data walls push up (Nesov's numbers fit α≈0.8\alpha \approx 0.8α≈0.8); overtraining, RL, distillation push down +sparsity exponent (β\betaβ)| 0.74| optimal tokens per active param rise ~σ0.53\sigma^{0.53}σ0.53 +sparsity (σ\sigmaσ)| 8x → 32x@2028 → 128x@2040| hand-set +inference share (δ\deltaδ)| ~50% pre-deal, 77%@2030, 25%@2040| from §2.4 (public + internal) +specialized share of new compute (fff)| 0.2% (2026) → 1% (2029) → 14% (2040)| hand-set; ~half the inference fleet's compute by 2040 +bandwidth per H100e| +10%/yr after 2029, capacity flat| gives the read rate ρ\rhoρ: ~31/s (2026) → ~107/s (2040) +spec bandwidth (Λspec\Lambda_\text{spec}Λspec​)| 60x HBM (2026) → ~500x (2040)| measured Groq/Cerebras ~60-80x; bandwidth binds until ~2034, then compute +serving precision and FLOP/s| 4-bit (0.5 B/param), 4e15 FLOP/s per H100e| weights are ~5 TB per replica in 2026 +batch (BBB)| 256 concurrent copies per replica| held flat; trades off with context assumptions +KV per copy (ccc)| ~20 GB (2026) → ~3.3 TB (2040)| ~200k → ~2.4M tokens of context; leaves context ~5% of memory by 2040 +decode gain (κ\kappaκ)| ~2| speculative / multi-token decoding +utilization (uuu)| 0.9| of each chip category's binding limit +spec speed (vspecv_\text{spec}vspec​)| 1,000 / 1,500 / 3,000 / 6,000 / 15,000 / 25,000 tok/s at 2026/29/32/35/38/40| hand-set, toward the on-chip wire-latency floor + +## Appendix: the effective-H100e model + +### A.1 Goal + +We want a way of estimating how useful a given cluster is for AI workloads relative to a typical H100 GPU cluster. + +Comparing the raw H100-equivalents is a poor proxy, because they only measure peak compute, and for real AI workloads a cluster can be drastically more or less useful depending on other hardware factors, particularly memory bandwidth, memory capacity, scale-up networking, scale-out networking, and how many physical chips a typical workload must span. The output is + +effective H100e = raw H100e×usefulness adjustment,\text{effective H100e} \;=\; \text{raw H100e} \times \text{usefulness adjustment},effective H100e=raw H100e×usefulness adjustment, + +where the usefulness adjustment is above 1 if the cluster is more useful per raw H100e than the H100 cluster anchor we consider, and below 1 if it's worse. The key design goals for the usefulness adjustment are: + + 1. Our H100 anchor cluster should have effective H100e equal to raw H100e. + + 2. Holding everything else constant, more memory bandwidth, memory capacity, scale-up networking, scale-out networking, or scale-up world size should help. This means an Nx improvement in any individual metric (all else equal) should be a ≥\geq≥1x usefulness adjustment. + + 3. A cluster with equivalent total hardware performance, but spread over smaller devices can be worse, because at a fixed workload size it needs to spread over more chips, raising communication overhead. This means an Nx reduction in all per-chip metrics with an Nx increase in quantity (so cluster totals are the same) should be a ≤\leq≤1x usefulness adjustment. + + 4. Bottlenecks should be soft, not hard: resources should trade off with each other, because it tends to be possible in AI workloads to readjust e.g., parallelism or other software-level choices to trade off different resources. + + + + +### A.2 Inputs + +The inputs are the chip and cluster properties, which we notate as: + + * CCC: per-chip compute (H100e). + + * BBB: per-chip memory bandwidth. + + * MMM: per-chip memory capacity. + + * SUSUSU: per-chip scale-up bandwidth inside a given world-size (e.g., NVLink inside a rack). + + * SOSOSO: per-chip scale-out bandwidth outside the scale-up world but inside the full cluster. + + * WWW: scale-up world size, in **chips** (the scale up domain GPU count, e.g. NVL72 =72= 72=72). + + * NNN: number of chips in the cluster. + + + + +Our model is workload-dependent, meaning that it takes into account a workload size (in compute footprint). Rather than considering a single average workload, we use a spread of job sizes, each a fraction LjL_jLj​ of the **frontier scale** FFF, which is supposed to represent a frontier company's total compute, with job-count weights ωj\omega_jωj​ (∑jωj=1\sum_j \omega_j = 1∑j​ωj​=1) as follows, expressed as fractions LjL_jLj​ of the frontier scale: + +job| fraction LjL_jLj​ of FFF| weight ωj\omega_jωj​ +---|---|--- +small (inference, experiments)| 1/10,0001/10{,}0001/10,000| 0.50 +medium (large experiments, test runs)| 1/1001/1001/100| 0.30 +large (near-frontier training)| 1/51/51/5| 0.20 + +Importantly, LjL_jLj​ is a fraction of **raw H100e compute** , not physical chip count. + +### A.3 Anchor scenario + +We use a 2024-style H100-SXM cluster as the anchor for raw=effective: + +anchor (where raw=effective)| value +---|--- +raw H100e per chip C0C_0C0​| 1 +memory bandwidth B0B_0B0​| 3.35 TB/s +memory capacity M0M_0M0​| 80 GB +scale-up bandwidth SU0SU_0SU0​| 0.9 TB/s +scale-out bandwidth SO0SO_0SO0​| 0.05 TB/s +scale-up world W0W_0W0​| 8 chips (HGX-8) +cluster size N0N_0N0​| 100K chips + +Separately, the frontier scale we consider FFF is time dependent. We model it as a growing share of total world compute (reflecting how frontier AI companies have a rising share of global compute), rising from ~6% (2025.0) to 9 / 13 / 16 / 18% (2026.0-2029.0), so absolute workload sizes climb each year and scale-out pressure rises alongside the world-size roadmap. By construction the anchor must output effective H100e =Ccluster= C_{\rm cluster}=Ccluster​, i.e. usefulness adjustment =1= 1=1 (verified in §A.9). + +### A.4 Workload footprint + +For each bucket jjj, the raw workload size (in compute needed) is + +Cjob,j=Lj F.C_{{\rm job},j} = L_j\,F.Cjob,j​=Lj​F. + +The number of physical chips that workload needs assumes we spread it as thinly as the hardware allows. Dividing the job's compute by the per-chip compute CCC gives the count we'd need if memory were free, which we then inflate by (M0/M)δ(M_0/M)^{\delta}(M0​/M)δ to account for a memory-poorer chip holding a smaller slice of the job's weights and activations. We don't treat memory as a hard wall, though: the exponent δ<1\delta < 1δ<1 reflects that you can partly trade capacity back (recomputing activations, offloading, finer sharding), so halving a chip's memory spreads the workload over somewhat more chips, not necessarily twice as many. + +nreq,j=Cjob,jC(M0M)δ,n_{{\rm req},j} = \frac{C_{{\rm job},j}}{C}\left(\frac{M_0}{M}\right)^{\delta},nreq,j​=CCjob,j​​(MM0​​)δ, + +where δ∈[0,1]\delta \in [0,1]δ∈[0,1] is the memory-footprint elasticity explained above (δ=0\delta = 0δ=0: capacity does not change the chip count; δ=1\delta = 1δ=1: footprint scales inversely with capacity; we assume δ=0.5\delta = 0.5δ=0.5 by default). + +We call this the **communication footprint** nj=nreq,jn_j = n_{{\rm req},j}nj​=nreq,j​, because it's what drives the cross-chip traffic in the penalties below: the more chips a job is smeared across, the more it has to talk between them (§A.5). In the edge case where the workload fits on a single chip we floor it at one. Either weakness compounds into a bigger footprint: a chip with less compute needs more copies of itself to run the job, and a chip with less memory needs more copies to hold its state, so a weaker chip on either axis spreads the same workload over more chips and pays more communication cost. + +We measure all of this against the anchor cluster running the _same_ job: the reference footprint is n0,j=Cjob,j/C0n_{0,j} = C_{{\rm job},j}/C_0n0,j​=Cjob,j​/C0​, with no memory term because the anchor is M=M0M = M_0M=M0​. + +### A.5 Scale-up versus scale-out traffic + +Now we estimate what fraction of a workload's communication (pSUp_{\rm SU}pSU​) can stay inside the fast scale-up world. Comparing the workload's chip footprint njn_jnj​ (from §A.4) to the world size WWW, let + +pSU(n,W)={1,n≤1min⁡!(1, W−1n−1),n>1,pSO=1−pSU.p_{\rm SU}(n, W) = \begin{cases} 1, & n \le 1 \\\\[2pt] \min!\left(1,\; \dfrac{W-1}{n-1}\right), & n > 1, \end{cases} \qquad p_{\rm SO} = 1 - p_{\rm SU}.pSU​(n,W)=⎩⎨⎧​1,min!(1,n−1W−1​),​n≤1n>1,​pSO​=1−pSU​. + +If the footprint fits inside the scale-up world (nj≤Wn_j \le Wnj​≤W), then pSU=1p_{\rm SU} = 1pSU​=1; if it is bigger, only a fraction of all-to-all-style communication is in-world. With chips that don't have a scale-up world (W=1W=1W=1), any workload larger than one chip is pure scale-out. This is a crude approximation for the traffic of AI workloads. + +### A.6 Resource penalties + +The model computes resource penalties, each normalized so the anchor equals 1. + +**6.1 Memory-bandwidth penalty.** + +τB=C/C0B/B0.\tau_B = \frac{C/C_0}{B/B_0}.τB​=B/B0​C/C0​​. + +If compute per chip outruns memory bandwidth per chip (both relative to the anchor), the chip is more bandwidth-bottlenecked (τB>1\tau_B > 1τB​>1). If bandwidth rises faster than compute, less so (τB<1\tau_B < 1τB​<1). + +**6.2 Networking penalty.** This measures how communication-bound the cluster is versus the anchor on the same job. The communication bottleneck is higher when a chip has more compute (more work done per second needs more data exchanged to stay fed) and when its job spans more chips; we weight these against how fast the networking is. + +We start with each cluster's effective inverse bandwidth RRR (communication time per unit of data moved, so higher is worse). From §A.5, pSUp_{\rm SU}pSU​ is the fraction of traffic inside the scale-up world, and 1−pSU1 - p_{\rm SU}1−pSU​ is on scale-out, so RRR blends the two fabrics by where the traffic goes. For the anchor: + +p0,j=pSU(n0,j,W0),R0,j=p0,jSU0+1−p0,jSO0,p_{0,j} = p_{\rm SU}(n_{0,j}, W_0), \qquad R_{0,j} = \frac{p_{0,j}}{SU_0} + \frac{1-p_{0,j}}{SO_0},p0,j​=pSU​(n0,j​,W0​),R0,j​=SU0​p0,j​​+SO0​1−p0,j​​, + +and the same for the candidate cluster with its own footprint, world, and bandwidths: + +Rj=pSU(nj,W)SU+1−pSU(nj,W)SO.R_j = \frac{p_{\rm SU}(n_j, W)}{SU} + \frac{1-p_{\rm SU}(n_j, W)}{SO}.Rj​=SUpSU​(nj​,W)​+SO1−pSU​(nj​,W)​. + +We model the communication penalty as the product of three things, each measured against the anchor: how many chips the job must span (njn_jnj​), how fast each chip finishes its work (CCC), and how costly communication is on the cluster (RjR_jRj​). They multiply because communication time is the amount of data moved times the cost of moving it, so being worse on several axes at once compounds. With elasticity γ=0.15\gamma = 0.15γ=0.15 on the span: + +τN,j=(njn0,j)γ(CC0)RjR0,j.\tau_{N,j} = \left(\frac{n_j}{n_{0,j}}\right)^{\gamma}\left(\frac{C}{C_0}\right)\frac{R_j}{R_{0,j}}.τN,j​=(n0,j​nj​​)γ(C0​C​)R0,j​Rj​​. + +Compute speed (C/C0C/C_0C/C0​) and communication cost (Rj/R0,jR_j/R_{0,j}Rj​/R0,j​) enter **linearly** : a chip finishing its work 2x faster produces data to exchange 2x faster, and a link 2x costlier per byte takes 2x longer, so we model each as putting linear pressure on communication time. The span enters **sublinearly** (γ<1\gamma < 1γ<1) because with 2x more chips any one chip does not move 2x more data: e.g., in a ring all-reduce the data each chip moves is constant (about twice the model size) whether the job runs on 100 chips or 100,000, since the per-chip chunks shrink as fast as the step count grows. However, the number of coordination steps (e.g., sync rounds, tree hops) does grow with chip count but only weakly. So overall we choose to model this with a 0.150.150.15 exponent. + +### A.7 Combining bottlenecks + +We combine the two penalties above with the compute bottleneck, weighted by importance shares αC+αB+αN=1\alpha_C + \alpha_B + \alpha_N = 1αC​+αB​+αN​=1. Compute carries no penalty (it's 1 by definition because the penalties are compute-over-resource ratios, and for compute itself that ratio is just C/C=1C/C = 1C/C=1), so αC\alpha_CαC​ is just the share of work we model as compute-bound, with αB\alpha_BαB​ and αN\alpha_NαN​ the shares set by the memory-bandwidth and networking penalties. We use a power mean rather than a hard bottleneck so the resources can partly trade off, with a default q=1.5q = 1.5q=1.5: + +Dj=(αC+αB τB q+αN τN,j q)1/q,D_j = \left(\alpha_C + \alpha_B\,\tau_B^{\,q} + \alpha_N\,\tau_{N,j}^{\,q}\right)^{1/q},Dj​=(αC​+αB​τBq​+αN​τN,jq​)1/q, + +where DjD_jDj​ is the slowdown relative to the H100 anchor. (q=1q = 1q=1 is a simple Amdahl-style weighted average; q=1.5q = 1.5q=1.5, the default, makes large bottlenecks bite harder. In CES terms 1.51.51.5 corresponds to an elasticity of substitution σ=1/(1+q)=0.4\sigma = 1/(1+q) = 0.4σ=1/(1+q)=0.4 between the resources). The usefulness adjustment is the weighted average of the penalty across workloads: + +U=∑jωj1Dj,effective H100e=Ccluster U.U = \sum_j \omega_j \frac{1}{D_j}, \qquad \text{effective H100e} = C_{\rm cluster}\,U.U=j∑​ωj​Dj​1​,effective H100e=Ccluster​U. + +### A.8 Full equation + +effective H100e=Ccluster U\text{effective H100e} = C_{\rm cluster}\, Ueffective H100e=Ccluster​U + +where the usefulness adjustment UUU is + +U=∑jωj(αC+αB[C/C0B/B0]q+αN[(njn0,j)γCC0RjR0,j]q)−1/q \boxed{\;U = \sum_j \omega_j\left(\alpha_C + \alpha_B\left[\frac{C/C_0}{B/B_0}\right]^{q} + \alpha_N\left[\left(\frac{n_j}{n_{0,j}}\right)^{\gamma}\frac{C}{C_0}\frac{R_j}{R_{0,j}}\right]^{q}\right)^{-1/q}\;}U=j∑​ωj​(αC​+αB​[B/B0​C/C0​​]q+αN​[(n0,j​nj​​)γC0​C​R0,j​Rj​​]q)−1/q​ + +### A.9 Why this satisfies the desired properties + +**Property 1: anchor raw equals effective.** In the anchor C=C0, B=B0, M=M0, SU=SU0, SO=SO0, W=W0C = C_0,\ B = B_0,\ M = M_0,\ SU = SU_0,\ SO = SO_0,\ W = W_0C=C0​, B=B0​, M=M0​, SU=SU0​, SO=SO0​, W=W0​ and nj=n0,jn_j = n_{0,j}nj​=n0,j​, so τB=1\tau_B = 1τB​=1 and τN,j=1\tau_{N,j} = 1τN,j​=1. Hence Dj=(αC+αB+αN)1/q=1D_j = (\alpha_C + \alpha_B + \alpha_N)^{1/q} = 1Dj​=(αC​+αB​+αN​)1/q=1 and effective H100e =Ccluster= C_{\rm cluster}=Ccluster​ exactly. + +**Property 2: more resources help.** Increasing BBB lowers τB\tau_BτB​. Increasing MMM lowers the footprint njn_jnj​, hence τN,j\tau_{N,j}τN,j​. Increasing SUSUSU or SOSOSO lowers RjR_jRj​. Increasing WWW raises pSUp_{\rm SU}pSU​, shifting traffic from slow scale-out to fast scale-up (which helps when SU>SOSU > SOSU>SO); the effect saturates once W≥CjobW \ge C_{\rm job}W≥Cjob​, i.e. once the workload already fits inside the scale-up world. Each raises effective H100e. + +**Property 3: same total compute, smaller chips.** Rebuild the anchor cluster from chips that are k×k\timesk× weaker on every per-chip metric but k×k\timesk× more numerous, so the cluster totals are unchanged: C′=C0/k, B′=B0/k, M′=M0/k, SU′=SU0/k, SO′=SO0/kC' = C_0/k,\ B' = B_0/k,\ M' = M_0/k,\ SU' = SU_0/k,\ SO' = SO_0/kC′=C0​/k, B′=B0​/k, M′=M0​/k, SU′=SU0​/k, SO′=SO0​/k with k×k\timesk× the chip count and the same workloads Cjob,jC_{{\rm job},j}Cjob,j​. + +Total raw H100e is identical, yet usefulness falls below 1, because each job must now span more chips. The footprint grows superlinearly because the two per-chip cuts compound: weaker compute means each chip holds only 1/k1/k1/k of the job, needing k×k\timesk× more chips, and smaller memory inflates that by a further kδk^{\delta}kδ (by the memory elasticity δ\deltaδ), so the chip count rises by k1+δk^{1+\delta}k1+δ, more than kkk since δ>0\delta > 0δ>0: + +nj′n0,j=C0C′(M0M′)δ=k⋅kδ=k1+δ\frac{n'_j}{n_{0,j}} = \frac{C_0}{C'}\left(\frac{M_0}{M'}\right)^{\delta} = k \cdot k^{\delta} = k^{1+\delta}n0,j​nj′​​=C′C0​​(M′M0​​)δ=k⋅kδ=k1+δ + +Bandwidth balance is unchanged, and in τN,j\tau_{N,j}τN,j​ the per-chip compute and the inverse-network cancel (smaller per-chip compute is offset by a correspondingly slower network), leaving only the footprint term: + +\tau'_B = \frac{C'/C_0}{B'/B_0} = \frac{1/k}{1/k} = 1, \qquad \tau'_{N,j} = \left(\frac{n_'j}{n_{0,j}}\right)^{\gamma}\frac{C'}{C_0}\frac{R'_j}{R_{0,j}} \;\gtrsim\; \big(k^{1+\delta}\big)^{\gamma}\cdot\underbrace{\frac{1}{k}\cdot k}_{=\,1} \;=\; k^{(1+\delta)\gamma} > 1. + +Hence + +Dj′≳(αC+αB+αN k(1+δ)γq)1/q>1,U=∑jωjDj′<1.D'_j \gtrsim \left(\alpha_C + \alpha_B + \alpha_N\,k^{(1+\delta)\gamma q}\right)^{1/q} > 1, \qquad U = \sum_j \frac{\omega_j}{D'_j} < 1.Dj′​≳(αC​+αB​+αN​k(1+δ)γq)1/q>1,U=j∑​Dj′​ωj​​<1. + +same total compute, k× more chips each k× weaker ⇒ usefulness<1 \boxed{\;\text{same total compute, } k\times \text{ more chips each } k\times \text{ weaker} \;\Rightarrow\; \text{usefulness} < 1\;}same total compute, k× more chips each k× weaker⇒usefulness<1​ + +With the default parameters the usefulness adjustment is: + +U=1D(k)≈(αC+αB+αN k(1+δ)γq)−1/q=(0.75+0.25 k0.3375)−2/3,U = \frac{1}{D(k)} \approx \left(\alpha_C + \alpha_B + \alpha_N\,k^{(1+\delta)\gamma q}\right)^{-1/q} = \left(0.75 + 0.25\,k^{0.3375}\right)^{-2/3},U=D(k)1​≈(αC​+αB​+αN​k(1+δ)γq)−1/q=(0.75+0.25k0.3375)−2/3, + +which gives the following values for these values of kkk: + +kkk| usefulness adjustment UUU| D(k)=1/UD(k) = 1/UD(k)=1/U +---|---|--- +0.1| 1.10x| 0.91x +0.5| 1.04x| 0.96x +2| 0.96x| 1.04x +10| 0.84x| 1.19x +100| 0.64x| 1.55x +1000| 0.45x| 2.23x + +For weaker chips (k>1k > 1k>1) the penalty is real but moderate; it strengthens with larger γ\gammaγ, αN\alpha_NαN​, qqq, or δ\deltaδ, or when larger jobs spill from scale-up to scale-out. For k<1k < 1k<1 (fewer but stronger chips, e.g. the k=0.1k = 0.1k=0.1 row) the opposite holds and U>1U > 1U>1. + +### A.10 Default parameters and results + +Default parameters (uncertain, hand-picked): + +parameter| value| role +---|---|--- +αC\alpha_CαC​| 0.55| irreducible compute-bound share +αB\alpha_BαB​| 0.20| memory-bandwidth-sensitive share +αN\alpha_NαN​| 0.25| networking-sensitive share +qqq| 1.5| bottleneck sharpness (1 = Amdahl, →∞ = max) +δ\deltaδ| 0.5| how much low capacity inflates the footprint +γ\gammaγ| 0.15| how much footprint size raises comms pressure + +On our projection of the frontier workload size and the average resource properties of the different compute categories from §1.2, we get the following usefulness adjustment to the raw compute: + +effective per raw (UUU)| 2024| 2026| 2028| 2029 +---|---|---|---|--- +**AI**| 1.16| 1.06| 1.04| 1.02 +non-AI servers| 1.38| 1.40| 1.41| 1.40 +PCs| 0.35| 0.08| 0.04| 0.03 +phones| 0.21| 0.11| 0.07| 0.05 +gaming GPUs| 0.04| 0.05| 0.06| 0.04 +other| 0.53| 0.43| 0.34| 0.30 + +Note that UUU is measured _per unit of raw compute_ , which is why non-AI servers score highest: a CPU server carries very little compute per chip but wraps it in generous memory, bandwidth, and datacenter networking, so each of its (few) H100e is unusually well fed. The AI fleet starts slightly above 1 (memory-rich H100s versus the anchor) and drifts toward 1 as the memory wall bites, while single-devices score far below 1 because they fail on networking, not compute. + +You can also run the model yourself: enter any chip/cluster spec below (or start from a preset) and read off its usefulness adjustment and effective H100e. + +datacenter:H100 SXM (anchor)A100 SXMB200 NVL72Rubin Ultra NVL576 (est.)CPU server + +consumer:budget AndroidiPhone (A18 Pro)MacBook (M4 Pro)PS5RTX 4060RTX 4090 + +Per-chip spec + +computeH100ememory bandwidthTB/smemory capacityGBscale-up bandwidthTB/sscale-out bandwidthTB/sscale-up worldchips + +workload year2026 + +Result + +usefulness U + +0.93× + +effective H100e per chip + +0.93 + += 1.00 raw × 0.93 usefulness + +By workload size + +U = 1 (anchor)0.93×small jobs (1/10,000)weight 50%0.93×medium jobs (1/100)weight 30%0.93×large jobs (1/5)weight 20% + +Usefulness calculator \- ai-2040.com + +Runs the exact appendix model (A.8) with the published parameters against the chosen year's frontier scale. Headline U is the job-mix weighted average; the bars show each job size alone, so you can see where a spec pays its penalty: hatched ink above the U = 1 anchor line, red below. Compute in H100e (= TPP / 15,800); a chip with no real scale-up fabric is W = 1. Datacenter presets follow the supplement's definition tables; the consumer devices and Rubin Ultra are rough public-spec estimates, with networking set to deployment-realistic sustained links (shared WiFi, home ethernet) rather than port peaks, which is what pulls consumer U down to the appendix table's ~0.05-0.2 range. It also produces the model's signature inversion: a budget phone outscores a flagship GPU per unit of compute, because more compute behind the same thin link is more stranded. + +1. + +Our AI-relevant definition is a threshold based on multiple hardware performance metrics, set just below an A100 explained further in [Compute Definitions](/supplements/compute-supplement#four-categories-of-compute). + +2. + +By nation of the company that owns the compute split into the US, China, or the Rest of the World. + +3. + +Peak dense (no-sparsity) tera-operations per second ×\times× the bit-width at whichever precision runs fastest (e.g., FP16=16, FP8=8, FP4=4). An H100 has **~15,800 TPP** (1,979 FP8 TFLOP/s ×\times× 8 bits); while a Blackwell B200 has around 2.3x as much (9,000 dense FP4 TFLOP/s x 4 bits = 36,000 TPP). This alone is already flawed as a metric given that lower bit precisions are treated proportionally less as their width: i.e., 2x smaller bit precision is 2x lower TPP but it is unclear whether 2x lower bitwidth is actually 2x less useful. It is a reasonable approximation because in theory 2x lower precision is 2x less pure information. + +4. + +The H100 and A100 clear every floor; an RTX 4090 gaming GPU, by contrast, fails on memory capacity and networking despite high raw FLOP/s. The networking floor is the scale-out (NIC) bandwidth; a chip's NVLink-class scale-up bandwidth feeds the usefulness model but is not itself a cutoff for AI-relevance. + +5. + +Around 100x lower than datacenter scale-out speed (roughly USB4/Thunderbolt speed). It is a broad definition, but in practice it mostly just captures chips used in datacenters. + +6. + +The actual _use_ of the chip is what these definitions are meant to capture: if someone connected up multiple laptops at sufficient cabled speed, that would become a non-AI-relevant _cluster_ (whereas the same laptops were non-AI-relevant _single-devices_ before being connected). + +7. + +Combined hyperscaler capex (Microsoft, Google, Amazon, Meta) plus neoclouds and sovereign funds grew ~1.7x/yr across 2023-2025. We decelerate the rate toward ~1.5x/yr as the absolute base grows, averaging ~1.6x/yr through 2028. + +8. + +H100e per dollar improves from die shrinks (N5→N3→N2), larger HBM stacks, and chiplet packaging, historically ~1.4x/yr. We hold that rate unchanged through 2028 rather than assuming an acceleration. + +9. + +To avoid double-counting we _exclude_ the frontier AI revenue from the hyperscaler cash-flow line (the AI companies' cash spend shows up inside the hyperscalers' cash flow), so we subtract it out of recent-year cash flow and add it back as its own source. Frontier AI revenue has been a negligible share of hyperscaler cash flow thus far, so this is a minor adjustment. + +10. + +The buildout is increasingly debt- and equity-funded. Alphabet priced a ~$20B bond in Feb 2026 (part of a ~$31.5B global raise that included a rare 100-year bond) to fund its ~$185B 2026 capex, on top of a ~$25B bond in Nov 2025 (its long-term debt quadrupled to ~$46.5B in 2025) ([CNBC](https://www.cnbc.com/2026/02/09/alphabet-highlights-new-ai-related-risks-in-tapping-debt-market.html)). Meta sold a $30B bond plus a ~$27B private-credit SPV for its Hyperion campus, and Oracle issued ~$18B (now the largest investment-grade non-financial debt issuer); in aggregate the five hyperscalers issued ~$121B of bonds in 2025, vs a ~$28B/yr norm ([Fortune](https://fortune.com/2026/03/07/big-tech-trillion-dollar-borrowing-ai-century-bonds/)). On the equity side, SpaceX's record ~$75B IPO (June 2026) is partly earmarked for AI compute, backed by large compute-supply deals with Google and Anthropic ([CNBC](https://www.cnbc.com/2026/06/11/spacex-raises-75-billion-in-record-setting-ipo-ahead-of-nasdaq-debut.html), [TechCrunch](https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/)). + +11. + +These multiples are a simplification: debt is ultimately serviced out of margins rather than revenue, and the big issuers historically defend investment-grade (A-or-better) credit ratings, which caps the leverage they would accept. Taken at face value our path implies leverage that only stays investment-grade if lenders keep crediting the forecast revenue growth (cumulative 2025-28 debt is ~$1.6T against a ~$1T frontier-company ARR at the start of 2029). Ideally we would size debt from projected cash flows under explicit credit-rating constraints; we use the simple revenue multiple for now. + +12. + +The conversion is the time-average of the run-rate, which for exponential growth from `A_start` to `A_end` is the _logarithmic mean_ , roughly **0.4x the year-end ARR**. So an ARR path of $30B → $180B across 2026 recognizes **~$84B of revenue in 2026**. + +13. + +Both leading-edge logic and HBM depend on EUV lithography, and EUV is the hardest link in the chain to scale on short notice: ASML is the sole supplier, each scanner costs >$200M with multi-year lead times, and the installed fleet grows only ~10-15%/yr, far slower than AI compute demand. Among the plausible bottlenecks (fab space, advanced packaging, power, EUV) we treat EUV tool-time as the binding (and most forecastable) one. + +14. + +Installed compute is not pure cumulative production: we net out a slow chip-survival decay (a ~15-year half-life for AI accelerators, counting reuse and secondary markets), which trims the 2029 stock by ~3%. The same per-category retirement is applied to all compute in §1.2. + +15. + +This treats the utilization rate (however much of the theoretical wafer capacity can actually convert into final compute, due to various yield and uptime factors) as collapsed into the final H100e per wafer numbers, rather than separating it out and considering H100e per "good wafer" x some utilization rate. + +16. + +Throughout this section we ignore the other DUV lithography (KrF, ArF-dry, i-line) as a mature tier that never binds; 'DUV' here means ArF-immersion (193i). + +17. + +Concretely, we model ~360,000 effective exposure-passes per tool-year: about 50 wafers/hour against a ~160 wph nameplate spec, reflecting ~75 to 82% uptime plus exposure-dose, source-power and reticle-overhead losses. So "100% utilization" here means the fleet fully booked at the realistic sustained throughput a fab like TSMC already runs its tools at, not a literal 100%-uptime ceiling. + +18. + +Each category's end-of-year installed stock is the surviving prior stock plus that year's production, with survival modeled as geometric decay at a per-category half-life: about 12 years for phones, 16 for PCs and gaming, 14 for non-AI servers, 15 for AI accelerators, and 30 for the "other" tier (MCUs and the like), i.e. roughly 2 to 6% of the fleet retiring each year. + +19. + +At fixed compute, optimal tokens per active parameter rise about with the square root of sparsity (the [sparse-MoE ratio](https://arxiv.org/abs/2501.12370), ~40 dense to ~240 at 30x), so total grows as σ0.74\sigma^{0.74}σ0.74 rather than σ\sigmaσ itself, matching the small active counts of today's sparsest frontier models. Sparsity is a hand-set schedule, 8x today to 32x at the deal and 128x by 2040, that is mostly a very uncertain naive trend extrapolation. + +23. + +Formally nj=max⁡(1, nreq,j)n_j = \max(1,\, n_{{\rm req},j})nj​=max(1,nreq,j​), i.e., a job small enough to fit on one chip still occupies a whole chip. This floor binds only in that edge case; for any workload large enough to span multiple chips it's inert and nj=nreq,jn_j = n_{{\rm req},j}nj​=nreq,j​. + +24. + +We are genuinely uncertain about this exponent. Some collectives scale well: a well-implemented tree all-reduce has latency ∼2log⁡2N\sim 2\log_2 N∼2log2​N in the participant count NNN (NCCL's double binary trees hold latency roughly flat out to tens of thousands of GPUs, up to a 180x improvement at 24,576 GPUs on Summit) ([NVIDIA, _Massively Scale Your Deep Learning Training with NCCL 2.4_](https://developer.nvidia.com/blog/massively-scale-deep-learning-training-nccl-2-4/)), which over the 10210^2102 to 10510^5105 chip range is close to n0.15n^{0.15}n0.15. Others are worse: ring algorithms have linear latency, and stragglers, contention, and imperfect compute/comms overlap push overhead higher. So 0.150.150.15 is a crude guess for weak sublinear growth. + +25. + +The footprint is nj=max⁡!(1, LjSC(M0M)δ)n_j = \max!\left(1,\, \frac{L_j S}{C}\left(\frac{M_0}{M}\right)^{\delta}\right)nj​=max!(1,CLj​S​(MM0​​)δ), and pSUp_{\rm SU}pSU​ is as in §A.5. + +20. + +Each copy's KV cache is ~200 thousand tokens of live context at ~100 KB per token after attention compression. We grow it with model width (~active^0.5) times ~12x net context-length growth (agent contexts outgrowing further KV-compression progress, ~200k to ~2.4M tokens per copy), which leaves context at ~5% of general-purpose memory in 2040. + +21. + +Sparsity does not change this: with a large batch every expert is hit each step, so the whole model gets read regardless. Batch-1 serving could read only the active bytes and run several times faster, but at a per-token cost nobody pays at scale; in this forecast that demand is served by the inference-specialized chips instead (the batch-1 ceiling ρσ\rho\sigmaρσ is kept as a diagnostic in the ground-truth sheet). Carrying the cancellation to 2040 also presumes serving stays at today's large-batch, throughput-optimal operating point on chips whose bandwidth-to-capacity ratio follows the §2.2 path; if latency pressure instead reshaped general-purpose serving (smaller batches, latency-tuned parts), these speeds would drift up and copies down. + +22. + +Two caveats on interpretation: reasoning models spend many tokens per finished task, so useful work per second lags raw tokens; and these are copies of the frontier model alone, where the scenario's worker counts (a hundred billion human-speed-equivalents by 2036) also include smaller distilled models, with the frontier-model copies here supplying about 20% of that total in token terms. diff --git a/sources/ai2040-pages/covert-ai-projects.md b/sources/ai2040-pages/covert-ai-projects.md new file mode 100644 index 0000000..7dcfbd2 --- /dev/null +++ b/sources/ai2040-pages/covert-ai-projects.md @@ -0,0 +1,1286 @@ +--- +url: https://ai-2040.com/supplements/covert-ai-projects +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Covert AI Projects + +### Brendan Halstead, Thomas Larsen + +In the Plan A scenario, we propose international coordination to develop AI safely. Keeping risk at an acceptable level requires substantial slowdowns in the pace of AI development. A central blocker for this type of coordination is the risk that one party develops AI in secret, thus gaining an advantage. + +Fortunately, it seems very hard to build and operate covert AI datacenters in a way that would be catastrophic. In the rest of this supplement, we’ll analyze how difficult it would be for China to defect from Plan A in secret, and the mitigations that can be put in place to prevent defection. + +What successful defection would require + +In summary, optimally subverting a competent implementation of our recommended Plan A policies would require China to: + + 1. Decide to begin diverting chips before the chip transparency agreements are made (in the Plan A scenario, this would need to occur in 2028). + + 2. Acquire on the order of 1.5M H100e (around 7% of counterfactual EOY 2028 Chinese stock) without being detected. The most promising options for acquiring this compute are: (i) diverting smuggled chips over the course of 2028, (ii) diverting older chips by claiming to have decommissioned and recycled them, and (iii) coordinating falsified production records across the semiconductor supply chain to hide some pre-deal production. + + 3. Build illegal datacenters without being detected. The most promising strategy we’re aware of for doing this is to locate the datacenters underground to avoid satellite detection. + + 4. Power and cool the underground datacenters. One strategy involves using small nuclear reactors to generate electricity onsite and rejecting heat by circulating water into a [large river](https://en.wikipedia.org/wiki/Water_cooling#:~:text=%5B11%5D-,Once%2Dthrough%20cooling,-\(OTC\)%2520systems%2520may&sa=D&source=editors&ust=1783631934888810&usg=AOvVaw2m7eDEPmNywOIBmKui7jqK) or [deep water source](https://en.wikipedia.org/wiki/Deep_water_source_cooling) in a way that avoids an IR-visible thermal plume. + + 5. Operate the sites without being detected for long enough to build extreme enough AI capabilities to significantly shift the geopolitical balance of power. This involves avoiding spies within the government, cyber operations, smarter AIs helping with covert project detection, satellite monitoring, and any agreed upon US/China government transparency (e.g. privacy preserving AIs reading classified info). If the US became suspicious (e.g. due to evidence from any of the items above), China would need to avoid giving concessions that would reveal their covert projects (e.g. physical US inspectors at potential covert project locations), while avoiding the possibility of the deal breaking down. + + 6. Avoid the development or usage of lie detectors (or other AI tech that improves verification massively). + + + + +China could also do a smaller covert project, which would have a proportionally smaller footprint and detection likelihood, but would also make slower progress towards strategically relevant AI capabilities. + +It will be impossible, in practice, for the US to confidently rule out a sufficiently small Chinese covert project. Therefore, the main questions of interest are: + + 1. What is the largest size of covert project that can evade detection? + + 2. On our current trajectory, at what time do covert projects pose a threat to the integrity of the deal and global security? + + + + +The answers to these questions will depend substantially on the overall scenario. However, in any scenario, the members of the Consortium will have the affordance to adjust their strategy to increase their lead over covert projects, primarily by reducing their rate of software progress at the cost of additional risk that the deal declines before sufficient alignment progress has been made, or by insisting on more robust detection technology before proceeding. + +# Section 1: Overview of our strategy for covert project mitigation + +This supplement outlines our plan for preventing this type of catastrophic defection from Plan A. + +Covert Project Supplement[Verification Supplement ↗Verify LegalDatacentersVerifyIndustrialExplosionVerifyLegal Fabs](https://ai-2040.com/supplements/verification-plan)Sections 5&6Bound overall riskfrom covert projectsSection 2PreventPre-Deal StockSection 3DetectSection 4Limit softwareLimit upliftfrom legal AIs[Capability Scaling Supplement ↗Limit LegalAlgorithmicProgress](https://ai-2040.com/supplements/capability-scaling-strategy)[Transparency Supplement ↗LimitAlgorithmic TheftLimit ModelWeights Theft](https://ai-2040.com/supplements/transparency-plan)Limit hardwarePost DealFlow + +This diagram gives an overview of our strategy for handling covert projects. Elements of the strategy are covered across four supplements, as pictured above. In this section, we’ll give an overview of the entire strategy. + +Covert project mitigation strategy overview \- ai-2040.com + +Defection from Plan A can happen on legal datacenters or covert datacenters. The legal datacenters are approved for AI development or usage, but are subject to verification to ensure that they are following those rules. Verification of the legal servers is the primary focus of the [Verification Plan](https://docs.google.com/document/d/1K4vxl_IMf58OgNXbWntQl6BQ06yAWZ62WmW5Ypp4KYA/edit?tab=t.26hlfwchw78c#heading=h.uismaghtvnrg). + +This supplement will focus on the risk from covert projects, i.e. projects primarily using secret datacenters. + +The main way that we can limit covert projects is by limiting their amount of compute: the covert project datacenters must be very small in order to go unnoticed. We can limit covert project compute through a combination of **preventing** covert projects from acquiring compute and **detecting** covert datacenters once they are constructed. + +Compute for covert datacenters can either come from “stock” (compute produced before 2029) or “flow” (compute produced during or after 2029). We are optimistic about very strong control of the flow of new chips because under Plan A compute production is extremely carefully monitored using highly controlled fabs. (Concretely, we should be able to keep diversion to approximately zero and have very high confidence that total diversion is less than 100k H100e.) Verifying the flow of new chips becomes somewhat more difficult during the industrial explosion because the abundance of robots may make it much more viable to construct a secret chip supply chain. The Consortium can prevent this through verification methods applied to the robot workforce. + +The stock of compute is already distributed and thus much harder to find, so we predict that a potential covert project would acquire their compute from existing stock. This means that a key priority in Plan A is estimating pre-deal stock and accounting for a high fraction of it. + +Our median estimate for the amount of compute a competently-executed PRC covert project could divert is roughly 1.5M H100e (.5% of world compute at EOY 2028, costing around $15B). + +Software is much harder to limit than hardware. However, if enough software leaks, even very small covert projects will be able to develop extremely capable AI. Software progress in covert projects can be broken down into two categories: progress acquired from legal projects and software progress from research done by the covert project itself. + +**Progress acquired from legal projects**. For most AI algorithms, security will be nearly impossible. Therefore, our primary line of defense here is to limit algorithmic progress made by the legal projects. Still, this will be the primary way that covert projects acquire AI algorithms for the majority of Plan A. + +**Software progress made by the covert project.** This will be slow by default because of the extreme limitations on compute under which covert projects are operating. Additionally, Plan A aims to prevent covert projects from being able to use the most capable available AI systems for AI R&D. But, they may be able to gain access to highly capable AIs through several routes which could dramatically accelerate the covert project rate of software progress: + + * **Distillation from the legal AIs.** We’re optimistic about mitigating distillation via standard techniques like hiding intermediate computation, but still expect the covert project to have access to somewhat more capable AI systems due to distillation. + + * **Usage of the legal inference AIs to do AI research.** This can be mitigated by refusals by default. + + * **Theft of the legal AIs and then training away refusals.** This can be mitigated by model weight security. We are much more optimistic about securing model weights than algorithms because of their size. + + + + +# Section 2: Preventing covert compute procurement + +We estimate that, under competent US execution of our policies in Plan A, a competently-executed PRC diversion effort beginning one year in advance of the deal would be able to acquire **between** **0.1% and 1.4% (80% CI)** of the world’s AI-relevant compute at the start of the deal, without the US getting unambiguous evidence of violation. We show this in various units in the table below. + +| **Median estimate**| **80% CI** +---|---|--- +**Fraction of world**| 0.51%| [0.11%, 1.4%] +**H100e** (Using our median estimate of world compute in Jan 2029)| 1.5M H100e| [310k, 4.13M] +**IT equipment cost**(Using our median estimate of $/H100e China would pay)| $15B| [$3.6B, $65B] +**Compared to our median estimate of the leading AI lab’s R &D compute in Jan 2029 **(36M H100e)| 24x less| [115x less, 9x less] + +The rest of Section 2 justifies these numbers. + +In Section 2.1, we explain our verification approach, which centrally involves accurately determining the amount of compute in the world, then tracing sales to their end user to verify the physical location of as much pre-deal compute as possible. + +In Section 2.2, we describe the strategies covert projects could use to subvert our verification approach. All of these routes require taking action before the deal is implemented; here we assume China begins diversion one year prior to the deal: + +****| **Source**| **Chip origin**| **P(feasible)**| **Median [80% CI]** +---|---|---|---|--- +**Smuggled Western chips**| 2028 flow| US-designed| 100%| 0.23%[0.02%, 1.17%] +**False decommissioning**| pre-2028 stock| US-designed| 60%| 0.31%[0.21%, 0.42%] +**Concealed domestic production**| stock + flow| China-designed| 25%| 0.14%[0.05%, 0.30%] +**Exempt military compute**| Verification-exempt| either| Not estimated| not estimated + +Finally, in Section 2.3, we forecast how much compute each route would yield and make an aggregate estimate for how much compute a covert project could acquire. + +## Section 2.1: Policies for limiting covert project compute acquisition + +AI-relevant compute produced after Plan A is implemented is carefully tracked and monitored. Therefore, the most feasible path for covert projects to acquire compute during Plan A is from compute that was produced before the deal’s implementation date. + +The existing compute will eventually end up in one of the terminal states below: + +Actual ProductionDeclared SalesDivertedLegalSmuggledTraced andInspectedRecordedUnrecordedTraced andInspectedDecommissionedFound andInspectedNot FoundRecycledDivertedLostDiverted + +Where pre-deal AI chips end up \- ai-2040.com + +In short, we recommend + + 1. [Tracing the records of the](/supplements/covert-ai-projects) _[inputs](/supplements/covert-ai-projects)_ [to AI compute production to get an accurate idea of how much compute was actually produced before the deal,](/supplements/covert-ai-projects) + + 2. [Tracing the records of the](/supplements/covert-ai-projects) _[sales and resales](/supplements/covert-ai-projects)_ [of finished AI compute to ultimately trace it to its current physical location](/supplements/covert-ai-projects) + + 3. [Using incentive programs and aerial imagery to locate any compute whose sale was not recorded or is otherwise untraceable.](/supplements/covert-ai-projects) + + + + +Respectively, these are focused on the leftmost node of the flowchart, the orange tree within the flowchart, and the Unrecorded Sales section of the flowchart. + +**Recommendation #1: Audit the semiconductor supply chain to get an accurate record of all pre-deal compute production** + +To exhaustively account for all pre-deal compute, we need to start from accurate declarations of pre-deal production. In other words, we need to ensure AI accelerator vendors (like Nvidia and Huawei) declare _all_ of their pre-deal sales. To make underreporting of production as difficult as possible, we recommend that the US demand records from the various upstream suppliers involved in producing AI accelerators. + +Concretely, you would start with sales declarations from Nvidia and AMD for merchant GPUs; Broadcom and Marvell for custom AI accelerators; and Huawei and Cambricon in China. Then these companies would declare their inventories, suppliers, and purchase history, for example: + + * logic fabrication from TSMC, Samsung Foundry, Intel Foundry, and SMIC; + + * high-bandwidth memory from SK Hynix, Samsung, and Micron; + + * advanced packaging from ASE/SPIL, Amkor, and JCET. + + + + +Then these upstream suppliers would declare _their_ suppliers, purchases, inventory, and cumulative production. At each layer, the flows can be cross-checked, such that underreporting production without introducing inconsistencies would require some combination of + + * Underreporting purchases of upstream inputs, which means that the upstream supplier would also need to underreport production, + + * Overreporting inventories of an upstream input, which would be caught during physical inspection of the declared inventories + + * Overreporting losses in the production step, which would require all subsequent nodes to underreport their input flows to match. + + + + +Notably, many of the suppliers involved in producing Chinese chips are not headquartered in China. It is conceivable that the PRC could coordinate this falsification regardless. To further deter the PRC from attempting to underreport production, the US can + + 1. Interview employees at Chinese wafer foundries, accelerator vendors, etc. as part of a whistleblower program. + + 2. Use national intelligence (such as cyber operations, human sources, etc) in the years leading up to the deal to search for evidence of falsification in the Chinese supply chain. + + + + +We consider it likely but not certain that Chinese companies would make accurate disclosures under these conditions. Though we don’t include this as a core recommendation or rely on it in our estimates, the supply chain can and should take action before the deal to make the verification process easier ([example](https://nacicankaya.substack.com/p/tsmc-most-definitely-has-a-golden)). + +**Recommendation #2: Trace declared compute sales to their physical end-use location** + +NvidiaSupermicroDellCoreWeaveLambda(own site)Nebius(own site)Alibaba (China)own datacentersCore Scientific(colo site)Switch(colo site)Shell companyrenting space in aMalaysian datacenterResellerin ChinaEnd customerin China + +Tracing chip sales to end-use location \- ai-2040.com + +We recommend that the Consortium verify the location of all pre-deal compute in two layers: exhaustively audit every compute owner above a size threshold by tracing vendors' sales records downward, and randomly sample the remaining tail of small owners, tracing each sampled unit of compute to its current physical location. The result is a certificate, at a chosen confidence level, upper-bounding how much compute could have been diverted to a covert project. + +Our auditing proposal in detail + + 1. Starting from Recommendation #1's verified record of vendors' sales to direct customers, audit every direct customer above a size threshold T, then every above-threshold customer of an audited firm, recursively, until all unaudited recipients are below T. + + 1. For example, you might set this threshold at 0.01% of pre-deal world compute (30k H100e, costing ~$200M at the time). T should be at or below the treaty's target bound on covert project size, so that any firm large enough to constitute a covert project by itself is audited deterministically. + + 2. By "audit", we mean "obtain all records of compute purchased, resold/transferred, or retained as inventory", as well as "physically inspect any compute the firm presently owns". + + 2. Some recorded outflows cannot be traced onward. These are added in full to an untraced pool L, and counted as potentially diverted. The two main cases: + + 1. Sales to unauditable counterparties. If a firm's records show a sale to a buyer that cannot be produced for audit (e.g. a shell company, or a smuggler who kept no onward records), the entire sale is added to L. Since T is at or below the target bound, a single such sale above T is enough to block certification, so the Consortium should also investigate the seller's remaining records and refer the case to the tools of Recommendation #3. + + 2. Unverifiable decommissioning claims. If a firm claims to have scrapped or recycled compute, the Consortium cross-checks the claim against recycler-side records (recyclers are themselves auditable counterparties), serial numbers where available, and actuarial fleet-attrition rates. Amounts that pass these checks are accepted; the rest is added to L. As above, an unverifiable claim above T blocks certification and warrants investigation. + + 3. When this completes, every unit of pre-deal compute is either physically verified, added to L, or recorded as delivered to a below-threshold recipient. The verified record of these small deliveries (firm 1 received a_1 H100e, firm 2 received a_2, and so on) is the list Layer 2 samples from. + + 4. Randomly sample H100e-units from that delivery record, and trace each sampled unit through any chain of resales until it either reaches a physically inspected chip or fails to. + + 5. Draw one unit uniformly at random from all compute recorded as delivered into the tail. This selects a recipient firm with probability proportional to how much it received. Audit that firm: the compute it verifiably received must equal current holdings plus recorded onward sales plus documented attrition, with holdings physically inspected. + + 6. Sales records don't distinguish individual chips, so the sampled unit is resolved proportionally within the account. If the firm holds 200 of the 500 units it received, the trace ends here with probability 200/500, and a chip must be produced for inspection. Otherwise it follows a recorded onward sale to the buyer, where the audit repeats. + + 7. Each trace ends as a pass (a physically inspected chip) or a fail (an inspection shortfall, an unbalanced account, a buyer that cannot be produced, or an unverifiable decommissioning claim). An adversary who splits a large diversion across many sub-T shell buyers gains nothing: those sales appear in the delivery record, so sampled traces land on them in proportion to their size, and every such trace fails. Layer 2 then detects and measures the missing compute statistically. + + 8. Sampled audits occur with minimal advance notice, and inspectors record chip serial numbers at every inspection. The same serial appearing in two inspections at inconsistent times is itself a violation. Without both measures, a covert project could pass audits with temporarily borrowed chips. + + 9. Going forward, the Consortium prevents L from growing: post-deal transfers above a de minimis size are permitted only to registered auditable counterparties, and post-deal decommissioning requires witnessed destruction or serial-verified recycling. + + + + +The resulting certificate: with confidence 1 - delta, no covert project exceeds D + L. D shrinks with the number of clean sampled traces (roughly, n traces certify D of about (tail size / n) * ln(1/delta)). L is a directly measured quantity that further auditing cannot reduce; reducing it requires the external-evidence tools of Recommendation #3. + +**Recommendation #3: Locating remaining unaccounted-for compute via incentives and aerial imagery** + +**Direct observation of AI datacenters.** We are relatively confident that intelligence services could, in advance of the deal, _identify and locate every (non-specifically-concealed) AI datacenter in China_ , using similar methods to Epoch, SemiAnalysis, and other commercial efforts. When they do eventually get inspected, the inspectors might find some of the untraceable compute. (This method does not help with compute that has been produced but not deployed in datacenters at the time of the deal, which might be around 50M H100e.) + +**Buyback programs and amnesty for smuggling.** We want to incentivize whoever owns the untraceable compute, or anyone who may have information on it, to come forward and reveal its location to the Consortium. The most obvious reason not to reveal missing compute is if you’re using it as part of a covert AGI project. Besides that, two reasons why private actors might be reluctant to come forward is + + 1. If the compute was obtained illegally in the first place, (e.g. smuggling, or fraud or theft for the purpose of smuggling), and the person fears punishment if they come forward, or + + 2. if the person believes they can receive a better reward selling the compute to some other actor (e.g. a state attempting a covert AGI project) + + + + +We recommend combating both of these incentives if the amount of missing compute remains unacceptably high after the previous steps are taken. For 1), states should provide complete amnesty to anyone who can deliver missing compute to the inspectors. For 2), states should offer to pay above-market prices for missing compute. Note that committing to this creates a perverse incentive in favor of smuggling, so should only occur if needed to account for additional missing compute. + +## Section 2.2. Covert project strategies for acquiring compute + +**Source #1: Diverting from smuggled flow** + + _Applies to: Western-made 2028 flow_ + +Even in Plan A, we forecast that significant compute will still be smuggled to China before the deal is implemented. Since the smugglers will have already broken/falsified some links in the chain, and may not keep similarly detailed records to legitimate suppliers, it may be possible to purchase compute under false identities without arousing much suspicion when supply chain tracing eventually arrives at a supplier. + +Speculatively, Chinese intelligence services could set up a fake company ostensibly from a different US adversary country (e.g. Russia), and purchase chips under that identity. Since such a country might also have incentive to hide chips, the eventual compute sales tracing might not be able to conclude whether the chips are in (e.g.) Russia or elsewhere. + +**Source #2: Concealing some pre-deal domestic production** + + _Applies to: China-designed pre-2028 stock AND pre-2028 flow._ + +To hide compute from the [resale tracing and physical inspection process](/supplements/covert-ai-projects), Chinese intelligence services could arrange for domestic accelerator vendors to underreport their pre-deal sales. The [compute production tracing](/supplements/covert-ai-projects) efforts aim to prevent this by accounting for all the accelerators which were physically produced. Chinese intelligence services would need to arrange for accelerator vendors like Huawei to understate their orders from the various supplier companies, and for those suppliers to understate their production. Those suppliers have their own suppliers, who would also need to align their statements. (For example, to fabricate the logic dies for Huawei’s chips, SMIC depends on silicon wafers from a different firm.) + +Some elements of the supply chain will be in foreign countries, making them harder for Chinese intelligence services to manipulate. On the other hand, some critical bottlenecks like high-bandwidth memory are likely being smuggled, which could complicate supply chain tracing. + +To avoid having to falsify records from dozens of different upstream suppliers, missing production could be explained by varying production yields. (e.g. in lithography, some fraction of the dies on each wafer will have defects, rendering those dies useless. Foundries aim to optimize this yield, but it requires ongoing R&D and can vary over time, especially at Chinese foundries as they race to catch up to TSMC.) However, any intermediate input flow whose existence is denied by understating yields still needs to be secretly made into the finished product. (continuing the example above, understating yields to acquire untracked logic dies would still require untracked HBM (from a different supplier) and untracked packaging (also done elsewhere) to make into a finished accelerator package). So yield falsification only _reduces_ the number of firms whose coordination would be needed. + +The ideal case would be manipulating yield at the final assembly stage. As far as we know, yields here are much higher and more predictable than at earlier stages, so this seems infeasible to significantly falsify. + +Overall, this strategy seems quite risky, as each additional firm added increases the number of opportunities for mistakes and for leaks. However, it’s possible the whistleblower/interview/HUMINT program are less effective than we expect, or that this sort of multi-firm falsification under scrutiny is more routine for Chinese intelligence services than we expect. + +**Source #3: Diverting from stock via false decommissioning** + + _Applies to: Western-made pre-2028 stock_ + +By default, some compute will be decommissioned between its purchase and the start of the deal. This poses an issue for compute sales tracing, since (as done today) decommissioning does not involve retroactively-verifiable destruction of the chips. + +How much compute will be decommissioned by default? + +For normal chip lifetimes, this is a small amount. (e.g. a common depreciation schedule is 5 years, so with an annual compute growth rate of 2-4x per year over the 5 years preceding the deal, compute whose fifth birthday happens before the deal will be less than 2^-5 – 4^-5, or 3.1% to 0.1% of global compute at the beginning of the deal.) + +We understand US hyperscalers to have detailed records of the entire decommissioning process. For example, AI servers decommissioned by Microsoft are sent to Microsoft Circular Centers, where high value components (AI accelerators would certainly qualify) are removed and assessed for internal reuse, resale, donation (for educational purposes), or recycling. If we assume that US hyperscalers (who currently own 71% of global compute) can provide an accurate count of their decommissioned chips, the decommissioned fraction is reduced to between 0.9% and 0.03% of cumulative global compute production at the start of the deal. + +We are less sure about the situation with smaller compute operators and for Chinese compute operators. + +As a potential diversion strategy, Chinese intelligence services might arrange to have large Chinese clouds decommission their US-designed compute in advance of the deal (perhaps justified as a political demonstration of Chinese semiconductor self-sufficiency, or to hide evidence of previous smuggling), with fake recyclers set up to accept the hardware and stockpile it rather than recycling it. Decommissioning a large fraction of compute in advance of the deal would create ambiguity about whether to treat it as possibly used for a covert project, since even if China were acting honestly, it might have no way to prove that the chips had truly been recycled. + +To the extent feasible, we recommend the US intelligence community surveil the Chinese IT equipment recycling industry to guard against this strategy. + +As a speculative potential strategy for making decommissioning look less suspicious, Chinese intelligence services might use cyberattacks to induce what appear to be hardware failures in installed devices, inducing unknowing datacenter operators to return some fraction of their installed hardware to server or GPU manufacturers (e.g. Huawei) for RMA, or otherwise to recycle them. Chinese intelligence services would then divert these shipments and have them documented as recycled. + +**Source #4: Using declared military compute to aid a covert AGI project** + +If the US and China aren’t able to negotiate away their unverified military compute, it could be used for AI development. The amount of military compute that could be diverted to a covert project is entirely a function of how the negotiations proceed. There are a few different ways the US could negotiate to minimize the amount of unverified military compute. To _minimize the maximum_ amount of compute that could possibly be used for military AI development, the US should push to bring _all_ military compute under inference-only verification. This has the drawback of incentivizing covert projects where previously they might have been judged not worthwhile, but it does improve the worst-case scenario (covert project AND declared military compute working together). + +Alternatively, the US could accept a greater quantity of unverified military compute, perhaps tying the amount to the amount of missing pre-deal compute to incentivize bringing any covert stockpiles into the open. Under some assumptions these strategies might reduce the _expected_ total amount of compute used for military R&D, at the cost of increasing the total amount usable by an aggressive adversary. + +Overall, we recommend that the US should push to bring all military compute under inference-only verification. + +## Section 2.3 Quantitative Estimates + +We consider two estimation methods for the amount of compute a high-competence covert project could acquire. We trust the first more than the second, and treat the second method mainly as a sanity-check. Currently, none of these include using declared military compute to aid a covert AGI project. + +**Estimate**| **Fraction of global total**| **H100e, assuming 289M at 2029.0**| **Description** +---|---|---|--- +1| Median 0.51%80% CI [0.11%, 1.4%]| Median 1.5M[310k, 4.13M]| Bottom-up estimate combining the plausible diversion strategies discussed above. +2| Median 0.23%80% CI [0.022%, 2.34%]| Median 660k[63k, 6.8M]| Based on a rough estimate of how much money the PRC could spend secretly purchasing AI servers in 2028, as well as our forecasts of $/H100e in 2028 + +### Method #1: Bottom-up + +In this method, we assume China simultaneously pursues all the acquisition strategies we identified above. We give distributions over the compute that could be acquired by each strategy, as well as the probability that each strategy is feasible. We then aggregate these together to get a distribution over the total amount of compute diverted. + +**Diversion route**| **P(feasible)**| **Median**| **80% CI** +---|---|---|--- +Decommissioning pre-2028 stock| 60%| 0.31%| [0.21%, 0.42%] +Underreporting 2028 domestic production| 25%| 0.14%| [0.05%, 0.30%] +Buying smuggled western chips in 2028| 100%| 0.23%| [0.02%, 1.17%] + +**Source #1: False decommissioning** + +If supply chain tracing and datacenter tracking are effective, installed compute cannot be removed from datacenters without explanation. So either there will need to be an explanation, or supply chain tracing and datacenter tracking will need to be subverted. We think the latter is likely infeasible and would contribute a relatively small amount of compute if feasible, so we have excluded it for simplicity. Therefore, the only way to divert pre-2028 stock is to falsely decommission chips that still work. + +Parameter estimates and rationales + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +Probability that decommissioning is feasible at all to do deniably| 60%| Intuitive guess. +US-designed fraction of pre-2028 stock in China| 45%| If domestic production is 3.7% of world, and (as [Plan A side panel data](https://docs.google.com/spreadsheets/d/1fIPU1OkgnKgsCiQqcj18RXRNlqqfL1YDGCjeWAeXC7s/edit?gid=2003669200#gid=2003669200&range=O57) says) 6.7% of world stock at 2028.0 is located on Chinese soil, then the US-designed share must be 45%. (or 3% of the world stock at 2028.0) +Fraction of pre-2028 US-designed compute located in China that can be falsely decommissioned before 2029.0, conditional on this being feasible to do at all deniably| 22%80% CI: [15%, 30%]Beta| Intuitive guess. If the top few clouds (Alibaba, Tencent, ByteDance, etc) own 60% of US-designed compute in China, and they decommission the oldest p25–p50 of their US-designed compute as part of this scheme, then this is 15–30% of the US-designed stock. (or 6.75% – 13.5% of the 2028.0 total PRC stock, using 45% as the parameter above). + +We estimate that conditional on false decommissioning being feasible, it yields 6.7% – 13.4% of the 2028.0 total PRC stock, and 0.45% – 0.9% of the 2028.0 world stock, and **0.21% – 0.42% of world pre-deal compute, conditional on the false decommissioning scheme being feasible at all (60%).** + +**Source #2: Diverting pre-deal flow** + +This would entail [concealing some pre-deal domestic production](/supplements/covert-ai-projects) and routing it directly into storage for a covert project. + +Parameter estimates and rationales + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +Probability that concealment of domestic production is feasible| 25%| Intuitive guess, see the discussion of the difficulties here. I could update based on learning more about the Chinese semiconductor ecosystem or the historical effectiveness of whistleblower programs like ours.My initial guess was 15%, Romeo thinks 35% for this, so we take the average and use 25%. +Fraction of 2028 domestic production which can be concealed, conditional on feasibility above| 7%Beta distribution with p90 = 15%.| Intuitive guess, p90 based on the idea that removing “one Chinese hyperscaler” (e.g. 10-20%) would probably be too obvious. Though note that this is different from removing one hyperscaler from the total compute in China; some of that will be legally imported or smuggled US-designed chips. + +Assuming that the fraction of global compute produced by China stays roughly constant, and accounting for the 2028 flow, we estimate that this route yields a **median of 0.13% (80% CI: [0.05%, 0.31%]) conditional on being feasible at all (25%).** + +**Source #3: Diverting smuggled chips** + +Parameter estimates and rationales + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +Fraction of 2028 compute sales smuggled to China| 2%[0.66%, 6.0%]LogNormal| Currently, Epoch’s median [estimate](https://epoch.ai/publications/chip-smuggling) is that 3.2% of compute sold through 2025 Q4 was smuggled to China. This might go up in the future due to increased demand for compute in China as frontier AI becomes a bigger business there (or as AI compute’s importance for national security becomes more obvious to the CCP). It might go down in the future due to improved US enforcement. + +Epoch’s 80% CI for cumulative smuggling through Q4 2025 is p10=1.7%, p90=6.4%. This is factor 1.8 to 2 uncertainty for p50 to p10 and p90. We will go with factor 3 uncertainty due to this being farther into the future.I am shifting the median down to 2% as in the Compute Forecast, deferring to Romeo, who gives this as a point estimate based on improved enforcement. I think this makes sense conditional on Plan A. +Smuggled compute for which the end customer cannot be identified| 25%[2.5%, 70%]Beta| Intuitive guess. I think it is probably less than half. I think my uncertainty could be reduced by learning more, but not super easily. + +An example of an accurately-recorded chain of custody: + + * Nvidia has a record of selling some number of units to the direct customer (e.g. a server OEM like Dell) + * Dell assembles the accelerators into servers and has a record of how many servers it sold to Nebius + * Nebius has a record of how many servers it purchased and which datacenters they're located in. + +I think smuggled chips can have an accurately-recorded chain of custody if all the entities involved have some record of who they sold to, even if that record isn't what they'd show to auditors. However, if at some point they were sold to an entity who used a false identity, or sold "cash-only", such that the reseller isn't sure where the chips went to, that doesn't count. But if the CIA knows where the chips ended up and the information can be used in the supply chain tracing process without compromising sources and methods or whatever, that does count. +2028 flow as a fraction of cumulative compute sales to date at time 2029.0| 53%| [Plan A side panel data](https://docs.google.com/spreadsheets/d/1fIPU1OkgnKgsCiQqcj18RXRNlqqfL1YDGCjeWAeXC7s/edit?gid=2003669200#gid=2003669200&range=R36) has 154M H100e added in 2028 and 289M H100e total as of 2029.0. 154 M / 289 M = 53%. + +Assuming independence of these parameters and simulating, this yields 0.44% of the 2028 flow [0.04%, 2.19%], or **0.23% of world compute at time 2029.0, 80% CI [0.02%, 1.2%]**. + +![](/supplements/covert-ai-projects/image10.png) + +The distribution of total compute diverted, aggregating across all the methods discussed above. + +### Method #2: Military-spending-based estimate + +The numbers below are estimated less precisely than the numbers in Method 1, as this is mainly used as a sanity check. + +Parameter estimates and rationales + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +China’s military spending in 2028| $400B80%: [350B, 457B]LogNormal| China’s [official military budget](https://www.macrotrends.net/datasets/global-metrics/countries/chn/china/military-spending-defense-budget) for 2023 was $296.44B. It grew roughly linearly over the preceding decade, adding an average of $13.2B per year. Extrapolating this to 2029, would predict official military spending to reach $362.4B.We use something closer to an extrapolation of [SIPRI’s estimates of military spending](https://www.sipri.org/databases/milex) (which are broader than the official budget).There are subtleties around PPP vs fixed-year exchange rate vs nominal which we’re currently ignoring. +Fraction of annual budget China can spend secretly buying AI servers| 2%80%: [0.2%, 20%]LogNormal| Intuitive unanchored guess. When asked, Claude suggested a number more like 15% – 20%. +$/H100e for newest US AI servers in 2028| $8k| Geometric mean of the 2028.0 and 2029.0 values in [Plan A side panel data](https://docs.google.com/spreadsheets/d/1fIPU1OkgnKgsCiQqcj18RXRNlqqfL1YDGCjeWAeXC7s/edit?gid=2003669200#gid=2003669200&range=R36). +Smuggling markup %| +100%[+50%, +200%]| +Decommissioned server markdown| -50%| Cheaper than the US-market $/per H100e surely, otherwise you could arbitrage by buying compute in the US and selling it directly to recyclers. But Chinese intelligence services might need to pay a bit more than is standard for recyclers since they need to capture the market, maybe 1.2x, or the cloud execs might just be told which specific recyclers to use, or the real recyclers could be put out of business.I will just go with half off. +Decommissioned to smuggled ratio| 3:5.5| Based on the numbers in Method 1. +H100e in the world at 2029.0| 289M H100e| From the [compute supplement](https://ai-2040.com/supplements/compute-supplement). + +Simulating the distributions in the table above yields a **median of 0.23% of global compute, with an 80% CI of [0.022%, 2.34%].** + +# Section 3: Detecting covert AI infrastructure + + _This section is more speculative than the previous section. We are not experts in any of the relevant fields and may be overlooking important considerations._ + +In this section we go over: + + * Methods for detecting covert projects (Section 3.1) + + * Strategies that covert projects might pursue to avoid these detection methods (Section 3.2) + + * Quantitative estimates for how well detection will work, on what timescales, and how this depends on covert project scale (Section 3.3). + + + + +## Section 3.1 Plan A strategies for detection + +**Method #1: Tracking waste heat** + +By conservation of energy, each megawatt of electricity going into a datacenter must be matched with a megawatt going out in some form (mostly as heat). The standard means of rejecting heat from datacenters produce distinctive signatures visible from space, and hiding or disguising these signatures seems to be the primary constraint on covert datacenter concealment. + +Modern AI servers are designed to use liquid cooling, where coolant is pumped through IT equipment to absorb heat. The coolant dumps its heat to water, which is normally pumped out of the data halls and cooled by outdoor chillers. These chillers are either air-cooled or water-cooled. + +_Air-cooled_ chillers run cool air through finned condenser coils, heating the air by ~10-15 degrees C. It is possible to hide the chillers from direct view, but the warmed air eventually has to escape to the atmosphere, and whatever aperture it leaves through will equilibrate to a warmer temperature than the surrounding surfaces. These warmed surfaces emit in the IR spectrum, which can be seen from space. Technology exists for minimizing the IR signature from hot exhaust ports (for example, on low-observable aircraft like the F-117). We are somewhat skeptical that this could be applied successfully against persistent satellite monitoring, but it could be possible. Of course, making these exhaust ports [invisible](/supplements/covert-ai-projects#method-2-invisible-datacenters) isn’t strictly necessary, they could instead be disguised as part of a legitimate industrial facility. + +_Water-cooled_ chillers transfer heat from the datacenter cooling loop to another flow of water which is evaporated in a cooling tower. Under certain weather conditions, the water vapor condenses to form visible clouds. We are unsure how hard these would be to hide consistently. + +![](/supplements/covert-ai-projects/image1.png) + +Condensation visible above cooling towers at the [Google Council Bluffs (East)](https://epoch.ai/data/data-centers/satellite-explorer/GoogleCouncilBluffsEast) datacenter, rejecting ~190 MW of heat. + +Another method for rejecting heat is simply dumping the hot water into a river or sea, a method commonly used to cool nuclear power plants. Doing this without creating an IR-visible thermal plume is hard, especially since time-averaging over a yearslong operating period allows for resolution of very minute heat differences. It may be possible to hide the thermal signature amid existing sources, or use multiport diffusers at depth to disperse the plume over a wide enough area to avoid detection. + +![](/supplements/covert-ai-projects/image6.png) + +[Landsat images of the thermal plume from the Haiyang Nuclear Power Plant](https://www.mdpi.com/2076-3417/15/2/738), cooling two 1.1 GWe reactors on China’s eastern coast. + +**Method #2: Other sources of IMINT/MASINT** + +[Interferometric synthetic-aperture radar](https://en.wikipedia.org/wiki/Interferometric_synthetic-aperture_radar) (InSAR) seems potentially useful for detecting the construction phase of a covert datacenter project. Based on measuring minute displacements of the earth’s crust over time caused by tunneling, an [LLNL technical report](https://www.osti.gov/servlets/purl/1036847) investigated using this for characterizing underground construction in a known location, but it seems less well suited to noticing an unknown facility. It might be useful to gain additional evidence that a suspicious site is in fact a covert datacenter (e.g. by observing the shape of the excavated volume). + +Any underground excavation needs to move the dirt and rock (termed “spoil”) to the surface. This may be hard to hide: options include hiding the excavation in another industrial project, putting tunnel entrances inside buildings and shipping out the spoil in trucks, or dumping the spoil into rivers. + +Finally, the space-based ground-moving target indicator (GMTI) system currently under development by the US Space Force and NRO could be used to track individual logistics vehicles involved in constructing or operating a covert datacenter, depending on the program timeline. We are unsure how much this would aid detection, but to the extent it can be modeled as “a video map showing every unobstructed vehicle in China,” it seems extremely useful in conjunction with the massive amounts of AI labor that Plan A would bring. If the system needs to be tasked to a narrow area to track vehicles, it becomes less useful for discovering remote sites, but still useful for gathering evidence about an area already flagged as suspicious. + +**Method #3: SIGINT** + +Building covert datacenters would require specialized contractors. For example, following the strategy we discuss below in [Building invisible datacenters](/supplements/covert-ai-projects), you would need + + 1. Excavation + + 2. Hydraulic engineering, to design and build the hidden cooling intake structure and outfalls + + 3. Nuclear engineering + + 4. All the contractors and disciplines involved in a normal datacenter project + + 5. Possibly TEMPEST or equivalent + + + + +It seems plausible that the US could identify a relatively small number of contractors among whom one would almost certainly be involved in a covert project. Cyber operations against all of those firms might be sufficient to rule out the existence of a covert project. We are uncertain about the tractability and potential consequences of doing this. Relevant public information includes the [NSA’s breach of Huawei in 2010](https://www.nytimes.com/2014/03/23/world/asia/nsa-breached-chinese-servers-seen-as-spy-peril.html), in which they reportedly “obtained information about the workings of [Huawei-produced networking products] and monitored communications of the company’s top executives”. + +**Method #4: HUMINT** + +Though counterintelligence would likely carefully vet the individuals with full, incriminating knowledge of a covert AI project, it seems possible that foreign intelligence assets would be read into the program. Obviously very little public information is available to estimate how likely this is. Our vague sense is that China is considered one of the more difficult countries for Western HUMINT collection. + +As part of the deal, we recommend that the US negotiate for a whistleblower program under which anyone with evidence of compute diversion, covert AI datacenter construction, or other deal-undermining activities can be interviewed and guaranteed protection against retaliation. We are uncertain about how much diplomatic allowance here is realistic, considering that a program like this would also aid normal non-AI-related HUMINT efforts in China. It’s possible that having a neutral third party operate the whistleblower program in China would be more acceptable than if the US were operating it. + +**Method #5: AI silver bullets** + + _Note: Everything in this section is significantly more speculative than the rest of this document because it assumes extreme AI capabilities._ + +During plan A, we scale to extremely capable AI; reaching AI as capable as top human experts at everything. This accelerates economic and scientific progress, which could substantially improve detection capabilities: + + 1. **Reliable Lie Detectors.** AIs could accelerate the development of reliable lie detectors. Existing lie detectors (e.g. polygraphs) can relatively easily be fooled. Reliable lie detectors might be developed in the near future over the course of normal human technological development, but extreme level of AI capabilities could accelerate the timeline. It’s also unclear to what extent the development of lie detectors would be politically accepted. If it’s possible for a regime with a covert project to decline being subject to a lie detector on principle, lie detectors won’t be helpful. On the other hand, the US could leave the deal unless China and/or other counterparties agreed to submit senior leaders to lie detection. + + 2. **Privacy preserving AI verification technology.** AIs could enable intrusive verification to be done in a privacy preserving fashion, i.e., without leaking unacceptable information. This could be implemented as follows. First, place a huge number of sensors (e.g. cameras, thermometers), and mobile robots across China, such that if the US were able to observe each of them, they would be confident there is no Chinese covert project. Then, give all of that information to an AI trained by the US on a secure server on Chinese soil, and have the AI analyze that information for worrying signs. Finally, allow the AI to communicate a very high level summary of its findings, (for example, “Is China running a covert project?”), and delete the rest of the information. In principle, this approach could enable verification intrusive enough to usefully rule out covert projects without divulging any sensitive state secrets. + + 3. **Extreme industrial verification.** By 2035, real US GDP in Plan A has grown by a factor of 25 relative to today. This enables extreme feats of industrial capacity. In an extreme case, all relevant manufacturing and industrial processes (e.g. aluminum plants that could be used to hide the waste heat of a covert project) could be disassembled and replaced/reassembled in a verified location (i.e., a robot SEZ). Again, there may not be political willingness to pay this cost, however, if there is substantial pressure (e.g. because both sides strongly want to continue the cooperative Plan A regime, but there is some evidence of a covert project, so one side wants higher assurance), it is plausible that this could happen, especially if combined with privacy preserving AI verification. + + + + +AIs will be making vast amounts of scientific progress over the course of Plan A (probably something like the equivalent of 100s of years of human equivalent progress over the course of the 2030s). Historical technological progress has resulted in all sorts of technologies that are helpful for detecting covert projects that would have been very hard to predict far in advance (e.g. satellite monitoring). We should expect all of this progress to result in improved detection methodologies in hard to predict ways. We should similarly expect that the technological progress would help concealment of covert projects as well; however, our general guess is that we expect technological progress to favor verification over concealment. + +## Section 3.2 Covert project strategies for avoiding detection + +In this section, we cover strategies that covert projects could take to avoid detection (classified as “invisible” or “disguised”. + +**How distributed would covert datacenters be?** + +Why one big datacenter instead of many small ones? Covert projects would aim to minimize project detection hazard per unit compute. At sufficiently small scales (e.g. a single 8xH100 server), detection hazard per site seems to scale sublinearly with site size (the detection risk to a group of four different 8xH100 servers in four different cities is higher than if all four servers were placed in a single rack). To minimize detection hazard per unit compute, you should therefore increase site size up to the point where detection hazard per site starts to scale superlinearly with site size, or until some other constraint prevents you from scaling up further. + +### Method #1: Disguised Datacenters + +Compared to “invisible datacenters”, “disguised datacenters” have the disadvantage that any specific, visible facility could be subject to focused scrutiny and potentially challenge inspection. For example, the US could assemble a ranked list of China’s largest industrial loads and task satellites to frequently revisit them, collect SIGINT on the companies that ostensibly operate them, and attempt to recruit human sources who work there or live nearby. For a plant that produces a visible material output, it seems plausible to directly track shipments out of the facility and compare that with the heat rejection. (e.g. a datacenter disguised as an aluminum smelter would produce far less aluminum than is typical for a smelter of similar power draw.) You could try to get around this by making the datacenter a small fraction of the total power draw (e.g. 1% – 10%), but this would restrict the datacenter to on the order of 10 MW – 100 MW (we are not aware of any industrial loads on the order of 10 GW or higher.) + +However, “invisible datacenters” require more complicated engineering and may not be feasible, especially against potential new AI-produced remote sensing capabilities developed during Plan A. + +Some specific industrial facility types that might be useful cover stories: + + 1. Aluminum smelters. China produces 43 MT of aluminum per year, aluminum smelters require ~1.6 gigawatts per Mt/yr, which naively implies ~70 GW of aluminum smelting capacity in China. + + 2. Nuclear power plants. + + 3. Combustion power plants + + + + +**Example: hiding a datacenter under an aluminum smelter** + +In this example, the IT equipment would be located underground, with any excavation happening indoors to avoid visible spoil piles or excavation equipment. The power would be diverted from an onsite coal plant or grid connection (ostensibly to support the smelter). This would require the facility to produce less aluminum than its power draw or thermal signature would imply. But based on our research, it doesn’t seem possible to accurately measure the actual aluminum shipped from an aluminum smelter using national technical means. Though the heat signature of a datacenter is different than that of an aluminum smelter, it doesn’t seem obviously thermodynamically prohibited to design a cooling system which dissipates heat from the IT equipment in a way that resembles the heat dissipation from aluminum smelters. + +### Method #2: Invisible Datacenters + +Trying to make a datacenter look like _nothing_ seems more complex than making it look like something else. The biggest difficulty we’ve identified is hiding the heat rejected by the IT equipment (and any onsite power generation). In theory, all the equipment can be built underground and shielded from the view of satellites, but the heat needs to leave the facility somewhere. See the [Remote sensing, IMINT, and MASINT](/supplements/covert-ai-projects) section. Besides the ongoing signatures, the initial construction period will also show activity, though avoiding detection has some precedent. + +**Example: once-through cooling + nuclear power** + +One potential strategy might be to build or extend an existing underground facility, install sufficiently many small nuclear reactors there to power the IT equipment, build an intake structure deep underwater on the coast, and use some sort of deep-water multiport-diffuser type design for the outfall to try to ensure that the hot cooling water dilutes to a temperature below the IR noise floor before reaching the surface. We aren’t sure whether a diffuser could be designed to those requirements, but it seems plausible. If so, a facility of this type seems quite difficult to detect via remote sensing once operational. The infrequent or nonexistent refueling requirements of the nuclear reactors would make even full knowledge of China’s electrical grid, natural gas distribution, etc. insufficient to locate the facility. We do not think [existing technologies](https://www.ctbto.org/resources/for-the-media/press-releases/chinese-monitoring-stations-now-sending-data) for detecting hidden nuclear reactors would be effective in this case, but again, we are not experts and could be wrong here. + +### Hybrid strategies + +China could also target “invisibility” for some aspects of the facility while disguising others. For example, it might be possible to site the IT equipment and power generation inside a logistics warehouse (e.g. receiving shipments of diesel fuel by truck). This has the advantage that there are many logistics warehouses, so the close-scrutiny consideration mentioned above might not apply. However, since a logistics warehouse doesn’t reject much heat normally, the cooling would need to be rendered “invisible” via one of the techniques discussed above. + +## Section 3.3 Quantitative estimates + +**How much datacenter capacity does the diverted compute require?** + +Parameter estimates and rationales + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +2026.0 US IT power per H100e (at max utilization)| 833 W/H100e| [DGX GB200 SuperPOD reference architecture](https://docs.nvidia.com/dgx-superpod/reference-architecture-scalable-infrastructure-gb200/latest/_downloads/238ab33ee345a7241fc2f15eedc2ff6c/RA11338001-DSPGB200-ReferenceArch.pdf) p.6. One Scalable Unit of 8 NVL72 servers has a TDP of 1.2 MW, so 150 kW per server, including InfiniBand and Ethernet networking, management nodes, and storage.We then divide by 2.5 * 72 H100e per server to get 833 W/H100e. +Average US power efficiency progress per year, 2026.0–2028.0| 1.37x/year| Averaging the three relevant cells in [Plan A side panel data](https://docs.google.com/spreadsheets/d/1fIPU1OkgnKgsCiQqcj18RXRNlqqfL1YDGCjeWAeXC7s/edit?gid=2003669200#gid=2003669200&range=L95), (1.3*1.4*1.4)^(1/3) = 1.37. +Smuggled chip proxy year| 2028.6| Roughly when the smuggling happens, adjusted forward slightly to account for the value of an exponential at the midpoint underestimating its average. +Chinese domestic power efficiency gap during diversion period| 3.7x| Taking a value between the two relevant cells in [Plan A side panel data](https://docs.google.com/spreadsheets/d/1fIPU1OkgnKgsCiQqcj18RXRNlqqfL1YDGCjeWAeXC7s/edit?gid=2003669200#gid=2003669200&range=O100). +Decommissioned chip proxy year| 2026.5| Seems vaguely right, the decommissioned chips will generally be older than this, but more of the compute will be from the newer chips since the quantity is growing 2–4x per year or so. +Covert datacenter peak PUE| 1.4| This refers to the facility power usage at maximum utilization divided by the IT power usage at maximum utilization. We use the value from [here](https://epoch.ai/data-insights/gpus-power-usage-in-ai-data-centers). It should maybe be adjusted upward because it’s a covert datacenter, but pretty unclear, also unclear whether theoretical maximum power is what we want to track for a project that could be flexible if needed. Leaving it at 1.4 for now. + +With the parameter values above, a covert datacenter with 1 GW total facility power could support 1.9M H100e of 2028-era smuggled US chips, or around 1M H100e of decommissioned US chips, or 525k H100e of 2028-era domestically produced chips. + +**How much time would construction take?** + +Today, Epoch finds that [gigawatt-scale datacenters can likely be built in two years or less](https://epoch.ai/data-insights/data-centers-buildout-speeds). + +![](/supplements/covert-ai-projects/image7.png) + +Considerations pointing towards construction being faster than this: + + 1. By 2029, constructing datacenters of this scale will be more routine and likely faster than today. + + 2. Additionally, being a top national priority for the Chinese government might enable a faster timeline than usual. + + 3. In Plan A, we forecast that construction labor will be largely automated by robots. We are uncertain about the timing here, and this might not be a significant factor at the beginning of the deal. + + + + +Considerations pointing towards construction being slower than this: + + 1. Clandestine underground facilities would require bespoke engineering that normal datacenters do not. + + 2. Construction speed may be constrained by factors other than labor or spending (for example, if excavating a large underground cavern in a remote area, there may be a maximum rate at which spoil (dirt, rocks from excavation) can be carried away for distribution across a wide area to avoid creating satellite-visible piles). + + 3. If significant human labor is required, the requirement to avoid spies might make it difficult to find workers. Separately, compartmentalizing parts of the project could constrain parallelizability (e.g. you might want to finish excavation before installing any IT equipment to reduce the number of people who know IT equipment is being installed there.) + + + + +We use **two years** as a point estimate. + +**What detection risk do covert projects of various sizes incur?** + +We use the following methodology to estimate the probability of detection. + + 1. We are uncertain about how detection hazard scales with site size. Lacking any reason to expect detection hazard to scale superlinearly in site size, we suspect that the [optimal strategy](/supplements/covert-ai-projects) resembles “make sites as large as possible subject to feasibility constraints”. Informed by the strategies above, we estimate a distribution over the optimal site size (where optimal is basically an operationalization of “maximum feasible site size”.) + + 2. For a few possible “optimal site size” conditionals, we estimate the probability of detection over the 10 year period in plan A conditional on 1 such site, and 10 such sites. + + 3. We interpolate these intuitions with a shared-frailty model. This lets us produce 10-year detection probabilities for sites operating arbitrary amounts of compute (conditional on specific optimal site size outcomes, but we can also marginalize over that.) + + 4. We also estimate 3-year detection probabilities and interpolate the times with a weibull-type hazard rate (hazard rate varies with a power of time). This lets us produce survival curves over time for covert projects operating arbitrary amounts of compute. + + + + +Our current operationalization of “detection” is focused on the physical infrastructure. We require that + + 1. Among locations that the US assesses as suspicious, the one they assess as the most likely covert datacenter candidate is in fact a covert datacenter + + 2. The assessment of the US IC is >50% confident that this candidate is a covert datacenter. (it’s possible this threshold should be higher). + + + + +The goal of this operationalization is to match the point where the US would feel diplomatically comfortable requesting an inspection of that site, and disclosing it to the international community if needed to generate additional pressure. + +Unfortunately, due to time constraints, the numbers used in the above process are based entirely on intuition and qualitative reasoning about the [detection methods](/supplements/covert-ai-projects#section-31-plan-a-strategies-for-detection) and [hiding strategies](/supplements/covert-ai-projects#section-32-covert-project-strategies-for-avoiding-detection) we considered. It seems possible to apply a historical reference-class approach, and we’ve informally looked at various historical examples of covert projects. But due to the number of asymmetries between this case and the historical examples we could find, and the lack of an obvious conclusion from our informal research, we did not judge a quantitative reference-class estimate to be worth devoting time to. + +Therefore our estimates of detection probability over time should be treated as highly speculative. We choose to include them anyway for the sake of reasoning transparency, since the ability of states to detect covert projects is an important input to our views. + +Distribution over the optimal site size + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +Probability that >1MW is infeasible/suboptimal| 0%| This is an approximation for modeling purposes. +Probability >10MW is infeasible/suboptimal| 10% | Intuitive guess. There are many different kinds of industrial sites that draw 10MW from the grid and dissipate 10MW of heat. Of each kind of site, China might have hundreds or thousands. It seems highly likely that at least one of them works as a disguise for a covert datacenter. +Probability that >100MW is infeasible/suboptimal| 35%| Intuitive guess. There are fewer 100MW industrial sites in China, but probably still 1000+. But at this scale, you might need to rely on concealing the facility entirely rather than disguising it as something else. +Probability that >1 GW is infeasible/suboptimal| _80%_| Our assessment is that once-through cooling at 1 GW+ requires a deep-water diffuser in an ocean (not a river) to avoid a thermal plume. Rendering other cooling methods (air-cooled or water-vapor cooled chillers) invisible at this scale seems implausible.There are some industrial sites at this scale, but seemingly not very many, and they would be highly visible targets for challenge inspection.From [this Data Insight](https://epoch.ai/data-insights/ai-datacenter-cost-breakdown), “a MW-weighted average across major US markets” finds that the facility capex per IT GW is $11B. +Probability that >10GW is infeasible/suboptimal| 100%| This is an approximation for modeling purposes. + +![](/supplements/covert-ai-projects/image8.png) + +Intuition-based detection probabilities, by site size and number of sites + +These estimates assume only the unilateral detection methods discussed above, and condition on no “AI silver bullets” being discovered (we add those on top in the final model). By “3y”, this means “three years into the deal”. + +**Per-site facility power (MW)**| **n sites**| **P(detected by 3y)**| **P(detected by 10y)** +---|---|---|--- +10| 1| 0.20| 0.26 +10| 3| 0.25| 0.30 +10| 10| 0.32| 0.37 +10| 100| 0.55| 0.75 +100| 1| 0.25| 0.35 +100| 3| 0.35| 0.40 +100| 10| 0.45| 0.67 +100| 100| 0.65| 0.80 +1000| 1| 0.40| 0.50 +1000| 3| 0.45| 0.70 +1000| 10| 0.70| 0.85 +1000| 100| 0.85| 0.90 + +![](/supplements/covert-ai-projects/image4.png) + +Forecast of “AI silver bullet” discovery by year + +Thomas’s overall view about the likelihood that anything in the category allows us to detect covert projects by the following years (assuming the events of Plan A) is: + +| High will| Low will| Reasoning +---|---|---|--- +2033| 65%| 20%| I think the development of lie detectors is about 50/50 by 2033 conditional on the capability trajectory. In high will, this enables the US to threaten to leave the deal unless China submits to lie detection, which would reveal a covert project. Then there’s some probability that other paths help substantially, especially unknown unknowns. +2037| 95%| 55%| By 2037, we are likely to have developed good lie detectors (maybe 80%). If a high will regime, lie detectors seems sufficient, and it’s very likely that other verification methods pan out, especially given that we'll have the industrial capacity to install whatever monitoring devices are necessary. In a low will regime I’m less confident about what will happen. + +All of these numbers set aside all the other detection methodologies, i.e., they assume that there’s no ability to detect. + +![](/supplements/covert-ai-projects/image3.png) + +# Section 4: Limiting Consortium uplift to covert projects + +In Plan A, legal projects worldwide develop highly capable AI and robotics over the course of the 2030s. Some of this progress will spill over to covert projects, making it easier to build AGI in secret than it would have been under a [global moratorium on AI development (as in Plan S)](https://ai-2040.com/?choices=plan-s-root). Our best guess is that this increased risk is worthwhile, due to the alignment, power concentration, and war-avoidance benefits of building more capable AI under the auspices of a deal; see discussion in our [Comparing Possible Plans supplement](/supplements/comparing-possible-plans#eli-lifland) for more on why. + +In this section, we summarize the various Plan A policies affecting the degree to which Consortium AI progress benefits covert project R&D. We also estimate + +## Section 4.1: Relevant Plan A policies + +**Recommendation #1: Total research transparency** + +In Plan A, nearly all AI research is fully transparent to the public, including AI algorithms, code, and documentation, for all frontier AI projects. AI model weights, significant fractions of the training data, and a small amount of other sensitive information is prevented from leaving the datacenters. + +**Recommendation #2: Inference monitoring to prevent covert project uplift** + +See our inference monitoring and restrictions [here](https://ai-2040.com/supplements/verification-plan#hardware-research-restrictions). + +We do not plan to offer zero-data-retention (ZDR) inference on models whose capabilities we would not want proliferated to covert projects. See our ZDR policy [here](https://ai-2040.com/supplements/verification-plan#zero-data-retention-zdr-inference). + +**Recommendation #3: Model weight security** + +We think achieving high-confidence model weight security is extremely important for Plan A, and the Consortium probably should not proceed very far above pre-deal capability level until they are ~95% confident that the best-equipped actors in the world cannot steal the model weights. Forecasting the likelihood of this happening is outside the scope of this supplement. + +## Section 4.2: Covert project strategies + +**Path #1: Distillation from Consortium models** + +Currently, Chinese labs like Minimax and Zhipu are spending around 20–30x less on R&D compute than US frontier labs like OpenAI. From our forecasts, it seems plausible that a covert AI project would start off with ~20-30x less compute compared to the leading AI company at the start of the deal. So while the requirement of secrecy would no doubt be a burden, researchers running a PRC covert project might not feel much more compute constrained initially than when they were operating legally. Despite their compute advantage, US labs have not obviously managed to lengthen their lead in model capabilities. + +![](/supplements/covert-ai-projects/image5.png) + +We believe this is primarily due to “distillation”, [used in a loose sense](https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks) to refer to using state-of-the-art closed source models to generate synthetic data and grade RL trajectories. These techniques can allow for highly compute-efficient training, in some cases offering more than an OOM of compute equivalent gain. + +Will covert projects be able to stay 3–9 months behind the frontier then, like today’s Chinese labs? On one hand, covert projects will get an additional boost at the start of the deal compared to today’s trailing labs: unlike today, all training recipes will become [transparent](https://www.ai-2040.com/supplements/transparency-plan). This initial sharing of non-overlapping insights will bring forward the Consortium and covert frontiers together at the start of the deal. On the other hand, Plan A tries to [mitigate distillation](https://docs.google.com/document/d/1K4vxl_IMf58OgNXbWntQl6BQ06yAWZ62WmW5Ypp4KYA/edit?tab=t.86wiss1dhlds#bookmark=id.9ax99eo6ehyb) by e.g. restricting companies from revealing intermediate model outputs in their APIs. + +We are uncertain how effective these measures will be against a determined adversary. The large number of different AI companies, applications, and deployments will provide a correspondingly large attack surface. A government could also pressure one of their country’s AI projects to intentionally leave a loophole open for distillation. + +It is essential that distillation not enable the proliferation of [destabilizing dual-use R&D capabilities](/supplements/covert-ai-projects#section-51-paths-for-a-covert-project-catastrophe), which might come fairly early into the deal (e.g. between our SAR and TED-AI milestones). For this reason, we recommend that the Consortium conduct ongoing red-team experiments where a group of researchers attempt to distill the current frontier models under similar circumstances a covert project would face. Depending on the results, it might be necessary to halt external deployment of more capable models until more assurance can be achieved. + +**Path #2: Extracting R &D labor from Consortium models** + +Consortium models running in inference datacenters will have classifiers that flag attempts at anything that looks like AI R&D. (In Plan A, models that would provide AI R&D uplift above the pre-deal SOTA are not offered under ZDR.) It seems risky to iterate against these things; we do not view this as a promising strategy. + +Overall, this path seems fairly unlikely to work at all if the Consortium is competent. If it did, covert projects could potentially make use of models with superhuman experiment selection skill (or “research taste”) in AI R&D, biology R&D, and other dual use domains, significantly accelerating their progress. + +## Section 4.3 Quantitative estimates + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +Fraction of Consortium software progress that leaks to covert project| Median: 67%80% CI: [36%, 90%]Beta distribution| See our estimate [here](https://docs.google.com/spreadsheets/d/1zKOxldIPXrz6RypC8Zr50rTpF10yGz4w0UZhNxk78XE/edit?gid=0#gid=0) for the median. The 80% CI is an intuitive guess (specifically, the p90 was chosen at 90%, and a beta distribution was assumed, which forces the p10 to be 36%). + +# Section 5: When is a covert project catastrophic? + +To become catastrophic, a covert project must achieve one of three things: + + 1. **Insurmountable capabilities lead.** This occurs if a covert project is able to get enough of a lead over the Consortium that even if the Consortium were to pivot back to racing quickly, it would not be able to stop the covert project (or its sponsor nation) from reaching AI capabilities sufficient to prevent any of the following outcomes. + + 2. **Balance-of-power warping capabilities.** This occurs if the covert project significantly changes the geopolitical balance of power, e.g. by granting a large economic, military, or political lead. The extreme version of this involves a decisive strategic advantage—i.e. granting the country with a covert project a position of complete strategic superiority over their rivals. + + 3. **BATNA warping capabilities.** This occurs if the covert project AIs empower their host countries to such an extent that the counterparty prefers remaining in the deal (despite knowing about the cheating) over pulling out of the deal and trying to go at it alone. + + + + +If none of these are true, then: (i) the covert project does not have an insurmountable capabilities lead: legal projects can catch up if they want, (ii) all sides roughly preserve their within-deal balance of power, and (iii) all sides can credibly threaten to leave the deal to coerce a known defector into shutting down their covert project. As long as each of these are true, then, the defection from the deal isn’t catastrophic. + +We define the capability level that a covert project would need to reach to satisfy any of these conditions as _deal-undermining-AI._ Deal-undermining-AI is not a fixed capability threshold, that threshold will vary over time depending on the level of capability available to other nations. + +We’ll now discuss how difficult it is for a covert project to build a deal-undermining AI. + +## Section 5.1: Paths for a covert project catastrophe + +**Insurmountable capabilities lead** + +Achieving an insurmountable capabilities lead is extremely hard because of the massive compute deficit faced by the covert project. + +Before the legal projects pause, most of the covert project’s capabilities progress will come from implementing the new training recipes developed by the legal projects (assuming no distillation is occuring). Therefore, the “effective training system performance” gap will be equal to the hardware gap, which, at the beginning of Plan A in the scenario, is around 24x (80% CI: [115x, 9x]) from a covert project to the largest legal project (assuming the covert project has ~1.5M H100e, and the biggest legal project uses ~36M H100e for R&D). + +There are two advantages for covert projects:First, they may not slow down for safety. At points within Plan A, capability progress will slow down dramatically or pause because of safety concerns. In this situation, the covert project will make some software progress itself, but very slowly. + +Second, due to the transparency and difficulty of algorithmic security, algorithms flow from the legal projects to the covert projects, but not the other way around. + +Even if the covert project were to get a capability lead, the Consortium could just pivot back to spending all of its compute on AI progress. At the beginning of the deal, there is a ~20x compute gap (1.3 OOMs); and by 2035 it’ll have expanded to 4.3 OOMs. If all of the newly founded algorithms are immediately shipped to the covert projects, this doesn’t help at all, but: + + * Some algorithms are scale dependent and therefore more helpful for the legal projects than the covert projects. + + * Some algorithms may be able to be kept secure (e.g., [our estimate](https://docs.google.com/spreadsheets/d/1zKOxldIPXrz6RypC8Zr50rTpF10yGz4w0UZhNxk78XE/edit?gid=0#gid=0) during Plan A is only ~67% of algorithms leak). + + * There may be some latency in the theft for the algorithms that are not radically transparent (e.g. by a few months). + + + + +**Balance of power warping capabilities** + +The most promising capabilities that a covert project could pursue to change the geopolitical balance of power that we’re aware of are: + + * **Persuasion / AIs for political uplift**. If covert projects are able to get AIs with similar levels of capabilities to, for example, the early 2028 AIs of the race ending of [AI 2027](https://ai-2027.com/), this could make their host country substantially more politically effective than anyone else, enabling them to get substantially more power. The main limiting factor on persuasion AIs is that it is probably a high capability threshold. + + * **Secret robot buildout**. The covert project could enable a massive industrial buildout. e.g. in underground tunnels, giving much greater industrial capacity than the rest of the world. This is particularly relevant because of the cap and trade regime over robot production in the Plan A scenario. + + * **Military R &D.** If covert projects have substantially more capable AIs than the legal AIs, or the legal AIs are all subject to aggressive refusals, covert projects could help design cheaper and better weapons (e.g. drones, missiles, ships, planes, bombs, etc). + + * **Exotic novel technology.** This includes nanotechnology, mind uploading, certain kinds of bioengineering, and unknown unknowns. Any of these would be extremely destabilizing; the main uncertainty is how difficult they are, i.e. what level of AIs over what time period are necessary for achieving them. + + + + +For military R&D and the robot buildout, the main limitation will be succeeding sufficiently in a manner that other countries cannot detect. If other countries notice, they can respond by requiring improved verification, requiring further arms control deals, or ramping up production themselves to restore the relative balance of military power. + +**BATNA-warping capabilities** + +The main paths for a covert project to greatly improve China’s BATNA are: + + * **Splendid First Strike on the Compute**. China could pull out of the deal, destroying US datacenters, plus gain an asymmetric advantage over the US by either (i) destroying the US hardened weight or chip stockpile, or (ii) Protecting some compute that was meant to be destroyed in the case of a deal dissolution. Having a sufficient cyber lead to overcome the inference compute disadvantage of the legal projects could enable covert projects to hack into the US model weight vault and destroy the copies. Protecting the compute probably requires significant physical capabilities, i.e. a secret robot buildout. + + * **Speed up the robot production ramp after the deal breaks down.** For example, you could covertly get a robot research lead on the rest of the world so that when the deal breaks down, you are starting off with better robot designs. Alternatively, you could build a hardened robot stockpile, which would create a higher starting population of robots from which to bootstrap. + + + + +For this, as well as the above, there are also unknown unknowns (e.g. new military tech we don’t know about yet). The unknown unknowns are probably scarier than any of the concrete paths we’ve identified. + +## Section 5.2: Operationalizing Deal Undermining AI + +In the quantitative analysis for this supplement, we simplified the modeling by assuming that deal undermining AI is a specific capability threshold that doesn’t change over time. + +What specific threshold do we choose? We are currently operationalizing this as “Top Expert Dominating AIs (TED-AI)”, or AIs that are better than the top expert at every domain. Since AI capability profiles are currently fairly spiky, TED-AIs will be substantially superhuman at the skills they are strongest at, while merely being top expert level at the skills they are weakest at. + + * **Why not a higher threshold?** Once AIs are beyond the human range, it’s hard to limit the possibility of novel exotic technology which we don’t know how to defend against. Moreover, it probably won’t make a huge difference, because the intelligence explosion beyond TED-AI is probably going to be quite fast. + + * **Why not a lower threshold?** AIs that are within the human range seem much less likely to develop the sorts of novel technology that could be geopolitically destabilizing. Even very large advances (e.g. much cheaper interceptors) seem pretty far from undermining MAD. + + + + +We hope that future analysis improves on this baseline. For example, another reasonable operationalization of “deal undermining AI” is “any AI that’s more capable than the legal projects”, becuase… + + * Covert projects can’t have an insurmountable capability lead if they don’t have a lead at all. + + * AIs of similar capability levels probably can’t change the geopolitical balance of power, if the defenders are deploying AIs of greater capabilities research to red team. especially given Consortium defenders will have vastly more compute (e.g. for finding and patching vulnerabilities). This will depend on the strategy taken by the legal projects, i.e., assuming that legal projects are allowed to apply AI to military R&D. In Plan A we currently recommend some arms control limitations on the use of AI to build new superweapons. The main downside of that policy is that it increases the chance that covert projects could be used to gain large military advantages. + + + + +However, this strategy likely requires more aggressive capability scaling than is needed. + + * It’s probably fine for covert projects to get a lead at lowish levels of capability (e.g. automated coder) because at that point, they aren’t strong enough that small gaps in capability seriously alter military / economic balance of power. In the scenario, for a very brief period of time during 2029 (before automated coder), legal training is paused while verification is set up, and so in the scenario, covert projects in principle might be able to get a lead during this window, though in practice it is unlikely because of the hardware gap (also, covert projects likely will need even more time to set up their infrastructure, unless they are using existing exempt military datacenters). + + * Closer to the top of the human range (e.g. at or beyond top expert level AI) it’s easy to imagine a small AI capability quickly snowballing into huge capability / political power. Things like superhuman political strategy and persuasion become more important and are almost impossible to defend against. + + + + +# Section 6: Overall risk estimate + + _The estimates in this section are based on simulating covert project capabilities progress as described in our[Takeoff Supplement](https://ai-2040.com/supplements/takeoff-supplement?hide=none)._ + +In this section we quantify how the Consortium’s [scaling strategy](https://ai-2040.com/supplements/capability-scaling-strategy) affects the risk that covert projects develop deal-undermining AI. We only evaluate fixed strategies here, of the form “aim to reach TED-AI some number of years into the deal, then pause there to do alignment research until the risks of delaying handoff outweigh the benefits.” This is somewhat unrealistic, because in reality the decision-makers implementing Plan A will be able to flexibly respond and adopt their strategy based on evidence received. + +For example, if the US received evidence that reaching TED-AI within 6 years would require a massive amount of software progress, they could push the rest of the Consortium to follow a more gentle scaling trajectory that leaned more on hardware buildout, thereby leaking less software to the covert projects. This would reduce the risk due to covert projects, though it would also increase the risk that the deal breaks down before TED-AI has been reached safely. They could also adopt filtered transparency instead of Total Research Transparency, which would leak a smaller fraction of software progress to covert projects, enabling faster scaling for equal covert project risk, at the cost of safety and risk of deal impairment. + +The actual covert project risk incurred will depend on how decisionmakers balance these factors to select a scaling strategy. Below, we show how varying the **scaling duration** and **handoff delay** affect the risk incurred. + +All the assumptions we are making + +We assume that + + * The US and other non-defecting Consortium members competently implement all of our policy recommendations + + * The PRC executed a competent diversion program one year before the deal was implemented (i.e. in 2028) + + * Covert projects which are detected before reaching deal-undermining AI stop their progress immediately. + + * Deal-undermining AI is operationalized as TED-AI + + * No weight theft occurs during the deal + + * The deal does not break down during the time period in question. + + + + +| **Aim for TED-AI in 2 years**| **Aim for TED-AI in 5 years**| **Aim for TED-AI in 8 years** +---|---|---|--- +**Covert project risk at TED-AI**| 7%| 8%| 8% +**Covert project risk waiting 5 years after TED-AI**| 18%| 13%| 11% +**Covert project risk waiting 10 years after TED-AI**| 20%| 16%| 15% + +Concretely, with the evidence available to us today, we estimate that + + * Halting AI R&D and initiating Plan A roughly one year before the Automated Coder milestone (in our scenario, this means implementing it in early 2029), + + * then making whatever software progress is needed to safely scale to Top Expert Dominating AI over the following 6 years, (early 2035, in our scenario), + + * then working towards a high-confidence alignment solution for 5 more years (until 2040, in our scenario) + + + + +would incur a 13% chance of covert projects developing deal-undermining AI before detection or handoff, if China competently pursued a covert project and remained committed to keeping it hidden. + +The probabilities above are obtained by drawing from our full range of uncertainty about the [dynamics underlying AI R&D](https://www.aifuturesmodel.com/), in addition to our uncertainty about the feasibility of various covert project strategies and mitigation strategies. We explain the details of our quantitative model in an [appendix](/supplements/covert-ai-projects#how-we-model-the-ai-capabilities-progress-of-covert-projects). Again, the decisionmakers implementing Plan A will have access to more information than we do, which might reveal that the risk incurred by the above strategies differs much more sharply than they do in our estimation. + +To illustrate this, if we could be certain that AI R&D and takeoff dynamics worked exactly how it does in our scenario, but still had uncertainty about covert project affordances, the table above would look like this: + +**Covert project risk incurred by several strategies, assuming the AI 2040 takeoff parameters:** + +| **Aim for TED-AI in 2 years**| **Aim for TED-AI in 5 years**| **Aim for TED-AI in 8 years** +---|---|---|--- +**Covert project risk at TED-AI**| 2.0%| 1.6%| 1.8% +**Covert project risk waiting 5 years after TED-AI**| 18.8%| 8.8%| 5.2% +**Covert project risk waiting 10 years after TED-AI**| 22.8%| 12.4%| 10.8% + +# Section 7: Covert project response plan + +The main goals of a policy response to covert projects are: + + 1. **Prevent a covert project from building deal-undermining AI.** We generally want it to be the case that legal AI projects have stronger AIs than covert projects. This is less important early in the intelligence explosion but becomes crucial later on, as having better AI capabilities could quickly lead to an insurmountable geopolitical lead. Specifically, by _deal-undermining AI_ , we mean an AI that either (i) directly confers a [decisive strategic advantage](https://aisafety.info/questions/8AEZ/What-is-a-decisive-strategic-advantage) to the covert project or (ii) is able to ongoingly outpace the legal AI projects, even if they switch to maximally fast AI capabilities progress. + + 2. **Deter the creation of covert projects.** We want it to be the case that at the beginning of Plan A, both the US and China are confident that if they were to do a covert project, it would either be detected before it is relevant, or outpaced on AI progress, and that these outcomes are worse than not having any covert project. + + 3. **Maintain international cooperation.** In particular, if there isn’t a covert project, we want to avoid Plan A breaking down due to lack of trust, i.e. one side being worried that the other has a covert project and thus wanting to exit the agreement. + + 4. **Prolong the time with human-level but not superintelligent AI.** Time with human level AI is very valuable because it is when we’ll make the most progress on alignment. + + + + +Goals 1 & 2 are in tension with Goals 3 & 4 because a central tool for preventing / deterring covert projects is accelerating capabilities progress within Plan A. + +The rest of this section is structured as follows: + + * First, we discuss the set of possible responses to covert projects. + + * Second, we discuss an example high level overall strategy + + * Third, we go through concrete scenarios and how we would recommend responding to each. + + + + +## Section 7.1: Options for responding to covert projects + +In this section, we’ll go over policy options for situations where covert project risk seems unacceptably high. These options can be pursued in parallel. Note that many of these options (especially options 1-4) require negotiation with China. + +**Option #1: Switch to Plan S** + + _Multilateral_ + +The US and China could agree to a full moratorium on ongoing AI development. This would slow down covert projects in the regime where they are mostly benefitting from Consortium-driven algorithmic progress. The covert project could still make progress on its own, but it would be much slower because of the compute and labor bottlenecks. Overall, this maximizes the serial time in the agreement before a fixed size covert project overtakes, putting aside detection. + +This option, of course, has the standard downsides of switching to Plan S. With respect to covert projects in particular, it may also delay the creation of any of the “AI silver bullets”, discussed above, which are helpful for mitigating covert projects. + +The exit paths for Plan S are (1) we successfully developed improved verification technology, (2) both sides decided to grant the other improved verification access, (3) sufficient progress on alignment was made to continue scaling quickly, in a manner that outpaces covert projects or (4) covert projects pose a large, imminent risk, and misalignment risk is smaller than covert project risk (while still being substantial, or else we’d be in option (3). + +The most preferable option among these is (1). In that case the plan is: issue a complete moratorium on AI R&D until we’ve got good enough verification to be confident covert projects are not a binding constraint, then revert back to Plan A. + +**Option #2: Adopt a scaling strategy that minimizes software share of capabilities progress** + + _Multilateral_ + +In Plan A, we currently allow for ~.5 OOMs/year of software scaling, and ~0.5 OOM/year of hardware scaling over 2035. The reason we recommend making algorithmic progress is in large part because of deal dissolution risk; if covert projects are a binding constraint, we could minimize algorithmic progress and scale purely with hardware, which doesn’t diffuse. It’s probably impossible to get algorithmic progress to zero (for example, there are inevitable capabilities externalities from safety research), but we might be able to decrease it to ~.1-.2 OOMs/year. + +**Option #3: Attempt to improve algorithmic progress security** + + _Multilateral_ + +Another approach is to try to prevent algorithmic progress from diffusing (or, increasing the latency with which it diffuses). Again, this would slow covert projects substantially insofar as they were primarily relying on Consortium algorithmic progress. The main difficulty is that there is a major tension between our desire for openness/transparency with a desire for algorithms to be secure: an important component of Plan A is the Total Research Transparency that deliberately publishes most algorithms onto the internet. We discuss options for increased algorithmic security in our [transparency supplement](/supplements/transparency-plan#thomas-larsen). + +Depending on how successful algorithms security is, it may be viable to switch to faster capability scaling. This will obviously pose additional misalignment risk, and in this regime it is hard to slow down because a covert project is close behind, so this is not a desirable position to be in. + +**Option #4** : **Negotiate with China to submit to stronger verification** + + _Bilateral_ + +For example, China could agree to: + + * Ongoing intrusive physical monitoring (e.g. thermometers in large water sources, cameras in relevant sites, electrical grid monitoring) + + * Provide access to classified materials that indicate there is no covert project. + + * Physical inspections of plausible covert project locations (e.g. industrial plants that produce substantial waste heat, tunnels near large water sources, etc). + + * Access to classified data, e.g. AIs can see the black budget and report back if there are any AI looking things on there + + * Lower the exempt military compute cap + + + + +If we have viable privacy preserving verification such that an AI could get access to these information sources and only report back high level bits of information, this could make this option much more viable. Alternatively, the US might credibly warn China that we are pulling out of the deal unless they provide us with stronger verification affordances. China would then be incentivized to agree to these concessions because the deal breaking down is very bad for them, because it would amount to something similar to a return to the status quo (e.g. a return to a race with the US having vastly more AI compute). + +**Option #5: Race to improved verification** + + _Bilateral_ + +If we have enough time, it may be worthwhile diverting massive resources into a sprint to develop improved verification technology, for example, privacy preserving verification that could enable the option above. Having access to more advanced AI is helpful for this option. We recommend pursuing this in almost any scenario because having the affordance to implement stronger verification is generally useful, but in scenarios where covert projects are posing substantial risk, this is especially promising. + +**Option #6: Improve takeoff speed modeling** + + _Unilateral_ + +Concern for covert projects may be substantially driven by the tail risk, i.e., that takeoff is faster than we think with a tiny amount of compute. It may be worthwhile to invest substantial resources to gain better information about how fast the intelligence explosion goes with a small amount of compute. An obvious baseline is to invest more into generic takeoff speed modeling (e.g. things like the AI Futures Model). + +A more intense option is to create a “covert dominating miniproject (CDMP)”—a sanctioned project that would simulate the progress of a covert project; it would be given somewhat more compute than we think would be able to evade detection, and would try to conduct an intelligence explosion as fast as it could, theoretically outpacing any covert project. The main downside of this proposal is that it would lead to faster algorithmic progress (which may then diffuse to the covert project itself, further accelerating it). However, this might net out to an improvement in a regime where we are very uncertain about takeoff speed because the empirical observations from the CDMP could give us a much better estimate of takeoff speeds in a highly compute constrained regime. + +**Option #7: Pull out of the deal and switch to Plan B** + + _Unilateral_ + +This would involve leaving the deal. Immediately, both sides would destroy the vast majority of each other’s compute, and then both sides would go back to racing. This is really bad for both sides, because it would be immediately economically catastrophic and it would significantly increase misalignment risk, World War 3 risk, and concentration of power. + +The main upside of this option is that it is useful as a BATNA during negotiation, and preserving a credible threat of leaving the deal can help deter the creation of a covert project in the first place. If China builds a covert project knowing it runs a substantial chance of detection, and the US commits to pulling out of the deal if they detect it, it is probably irrational for China to create a covert project in the first place. + +The main downside of pulling out of the deal is if this happens when China was in reality honest, but for some reason the US had an overly hawkish estimate of a Chinese covert project. (For example, maybe compute estimates were off, or smugglers were harder to track down than expected). In these worlds, there was the option of continuing on in Plan A and incurring much lower risk across the board. + +**Option #8: Other geopolitical concessions** + + _Bilateral_ + +In general, if the US is suspicious that China is defecting, they can use the standard set of geopolitical tools to incentivize them against this. Particularly notable tools in the Plan A scenario include the robot and compute cap subject to both countries. + +## Section 7.2: Response plan for ambiguous evidence of covert project + +Example concrete ways the US could get weak evidence of a covert project: + + * The chip supply chain tracing reveals a missing 1M H100e: records exist of these chips being manufactured (e.g. from Nvidia/TSMC/Huawei), but we aren’t able to find them. This could be smugglers, or it could be a nation-state backed covert project. + + * China refuses additional verification. For example, maybe the US is suspicious of a particular set of sites, where there might be a covert project, but China isn’t willing to allow the US access. This is an update towards a covert project because they are much more likely to give the US access to these sites if they are legitimate. + + * The US might get ambiguous intelligence from IMINT / OSINT / SIGINT, for example, satellite images might suggest large scale underground construction, but the US might not be able to ascertain whether it's an AI related covert project. + + * The US notices model weight exfiltration attempts that seem like they could be coming from China but with unclear attribution. (This doesn’t necessarily imply China is operating covert compute, but is evidence.) + + + + +Baseline response: + + 1. Adopt a more covert project sensitive scaling strategy, for example, minimizing algorithmic progress (Option #2). + + 2. Improve verification: both by improving the technology (Option #5) and negotiating for increased implementation (Option #4). + + 3. If China refuses to allow improved verification, consider leaving the deal (Option #7). This is the only part of the response that the US can do unilaterally, but the fact that this is a bad BATNA for both parties creates pressure for China to agree to the above. + + + + +Essentially, this proposal involves marginally tweaking parts of Plan A to be more conservative with respect to covert project risk, because there is direct evidence that covert projects are a significant concern, but allows the core properties of Plan A to continue—there is still total research transparency, broad deployment of AI, ongoing reversible scaling, etc. + +## Section 7.3: The US gets unambiguous evidence of a PRC covert project + +Example concrete ways the US could get unambiguous evidence of a covert project: + + * Cyber espionage (e.g. GPU backdoors, finding classified documents, etc) + + * Satellite footage (e.g. finding the construction, heat signature from the project, etc) + + * A literal spy in the project + + * Egregious refusals from China during negotiation (e.g., there is really solid privacy preserving verification technology, but China refuses to use it and doesn’t give a reasonable explanation for why). + + + + +Baseline response: + + 1. Require that China make large geopolitical and verification based concessions (Option #8 and #4). + + 2. If China isn’t willing to make those concessions, exit the agreement. (Option #7). + + 3. If China is willing to make sufficient concessions, evaluate how good the verification situation looks and choose a [capability scaling strategy](/supplements/capability-scaling-strategy) that corresponds with the overall level of ongoing covert project risk. + + + + +This strategy is reasonable because it (i) punishes China and puts them in a bad strategic state, thus deterring them from doing a covert project in the first place, (ii) allows the Plan A deal to continue, as is rational for both sides to want, and (iii) prevents the covert project from developing deal-undermining AI. + +# Appendix + +## Covert project strategies we think are ineffective + +### Acquiring compute + +##### Diverting pre-2028 stock by retroactively understating domestic production + +This category refers to pre-2028 sales by Chinese chip vendors. These sales could conceivably be understated, but it seems even more difficult than the (already seemingly quite difficult) task of understating new production. + +**Parameter**| **Estimate**| **Rationale** +---|---|--- +Probability that retroactive concealment of domestic production is feasible| 7%| I think that doing it non-retroactively is about 15% likely to be feasible, and I am dividing this number by two. Seems about right, but I could update based on learning more about the Chinese semiconductor ecosystem or the historical effectiveness of whistleblower programs like ours. +Fraction of pre-2028 cumulative domestic production which can be concealed and removed from datacenters, conditional on feasibility above| 5%Beta distribution with p90 = 7%.| Intuitive guess, p90 based on the idea that removing “one Chinese hyperscaler” (e.g. 10-20%) would be super obvious. Though note that this is different from removing one hyperscaler from the total compute in China; some of that will be legally imported or smuggled US-designed chips. + +Assuming that the fraction of global compute produced by China stays roughly constant, and accounting for the 2028 flow, we get a Beta distribution supported on [0, 1.6%] with **median 0.08% of pre-deal compute from this route, conditional on it being feasible at all.** + +## How we model the AI capabilities progress of covert projects + + _This section is unfinished. See also[plan-a.aifuturesmodel.com](http://plan-a.aifuturesmodel.com), which will replace this._ + +### How much does takeoff slow down if you reduce compute by 10x (“N”)? + +Define the “R&D compute” (in, say, H100-equivalents) of an AGI company at a given time to be the total compute in use across the following categories: + + 1. Compute used to run experiments, which are used to search for software improvements, + + 2. Compute used to run automated researcher agents, + + 3. Compute used in final training runs for frontier models. + + + + +With 10x less R&D compute, the company would need to make cuts across these categories. For now, we’ll assume they maintain the same allocation across the three areas, that is, we’ll assume they achieve the 10x reduction in the total via a 10x reduction to each category. + +The capability level of your best AI model is represented by its _effective training compute_ , E. E is a stock, with units of "2025-FLOP". E evolves according to + +E′ = T × S, + +where T is training system performance (a flow, units: FLOP/yr) and S is software efficiency (units: 2025-FLOP per FLOP). Software improves via the semi-endogenous growth law (blah). + +Huh? Why this model? + +Other models of AI takeoff (the [AI Futures Model](https://www.aifuturesmodel.com/), [Tom Davidson's FTM](https://takeoffspeeds.com/description.html), [Epoch AI's GATE](https://epoch.ai/gate/)) compose effective compute multiplicatively from two stocks: + +E = C × S, + +where C is _cumulative_ training compute (a stock, units: FLOP, satisfying C′ = T). In other words, they apply today's software efficiency to every FLOP ever spent on training, "pulling S outside the integral". Our formulation instead applies S to the flow: E′ = T × S, so E = ∫ T · S dt. An algorithmic improvement raises the value of each FLOP spent from now on, but never retroactively upgrades FLOP spent under older algorithms. + +A large, sudden compute reduction is precisely where the formulations come apart. Physically, the cut is a cut to flows: the rates at which FLOP are spent on experiments, agents, and training all drop by 10x. The multiplicative model has no training flow among its state variables, only the stock C, so the cut has to be translated into a statement about C. There are two natural translations, and neither is good: + + 1. **Cut the stock: reduce C by 10x.** Then effective compute, i.e. capability, instantly drops by an order of magnitude. But nothing about a cut to the flow claws back FLOP already spent: the company keeps its best model's weights, can keep improving that model (with further RL, for example), and will use it to accelerate its remaining R&D. Treating a 10x cut in spending rates as a 10x cut in cumulative spending is like assuming the company must retrain from scratch on the smaller budget. This overestimates the slowdown. + + 2. **Preserve the stock: let C continue accumulating at the reduced flow** (C′ = T/10), **or freeze it entirely.** The stock bookkeeping is now fine; the problem is the multiplication. Since C barely moves on the relevant timescale while S keeps growing, capability keeps growing at roughly the rate of software progress, which is most of the overall rate, so to first order the cut slows progress by about half or less. But E = C × S applies each new algorithm to the entire accumulated stock, including all the FLOP spent before the algorithm existed. It's a bit like assuming that improvements discovered on the reduced budget magically upgrade training runs completed years earlier. Actually realizing an algorithmic improvement in a model requires spending FLOP under that algorithm, and the cut company can only spend them at a tenth the rate. This underestimates the slowdown. + + + + +The remaining simplification is treating AI progress as a single continuous training run that has been underway since the beginning of time. (Mathematically, the integral defining E runs from minus infinity to the present.) In reality, frontier training runs sometimes start from scratch, [have different lengths](https://epoch.ai/data-insights/longest-training-run/), and may only be underway during parts of the year. We think this still captures the dynamics relevant to compute reductions, but it directionally favors the compute-reduced project: some capabilities may require architectural changes or otherwise require starting from scratch, i.e. re-accumulating E from zero, which is far more costly at a reduced flow, and this model assumes you never need to do this. In other words, if anything, a real 10x cut probably buys somewhat _more_ slowdown than the model estimates. + +Training system performance T and software efficiency S tend to grow exponentially over time, so the “rate of AI progress” will be operationalized as the relative growth rate of E (i.e. the time derivative of log E, which equals E’/E). We will use the symbol \rho for this. + +When we cut R&D compute by 10x in this model, as long as we assume proportional reductions across all categories, the immediate effect is a 10x reduction in T, and therefore also a 10x reduction in the rate of AI progress \rho = (T * S) / E. + +What happens after that? To figure this out, we can look at how \rho changes over time. Differentiating, + +\rho’ =d/dt (E’ / E) = E’’/E - (E’/E)^2 = (E’/E) * (E’’/E’ - E’/E) = \rho * (E’’/E’ - \rho). + +This E’’/E’ term equals the relative growth rate of E’, d(logE’)/dt, or the growth rate of _effective training system performance_ E’ = T * S. We will call the growth rate of effective training system performance \gamma. The equation above then becomes + +\rho’ = \rho * (\gamma - \rho). + +(This has the form of the “logistic differential equation”, with carrying capacity \gamma.) Qualitatively, it says that when effective training system performance grows faster than capabilities (\gamma > \rho), then growth rate of capabilities increases (rho’ > 0). Conversely, when capabilities are growing faster than ETSP is, then capabilities growth will slow. + +If training system performance (the raw hardware) remains constant after the reduction, then \gamma is just the growth rate of software efficiency, determined by experiment compute and AI labor, which today proceeds at perhaps 1 OOM per year. With 10x less compute for experiments, and 10x less compute to run automated researchers, the growth rate falls, but probably by less than 10x. Say it falls by some factor 10^\zeta with 0 < \zeta < 1. + +So after \rho falls by 10x, the subtraction on the right hand side is positive, meaning that rho is increasing. That is, _the rate of capabilities progress will initially fall by 10x, then pop back up again, until effective compute is growing at the same rate as software efficiency_. It will continue to increase so long as \gamma exceeds \rho, gradually approaching the fixed point where \gamma = \rho and \rho’ = 0. + +Now we have to determine how the rest of takeoff goes. By the steady-state approximation we just found, this reduces to figuring out the software efficiency growth rate. If, at every capability level, the trajectory with 10x less compute had an Nx lower software efficiency growth rate than did the trajectory with untouched compute at that capability level, then we could conclude that _the overall takeoff would take Nx as long._ This turns out to work. + +Let’s consider the following model of + +N is well approximated by (r * zeta + 1) / (r + 1), where zeta is the elasticity of instantaneous algorithmic progress rate to experiment compute, and r is the number of doublings of software efficiency per doubling of cumulative research effort. + +![](/supplements/covert-ai-projects/image2.png)![](/supplements/covert-ai-projects/image9.png) + +### Modifications made to the AI Futures Model + +The covert project capabilities trajectories are simulated by a modified version of the [AI Futures Model](https://www.aifuturesmodel.com/). The modification is explicit modeling of the “training run”—instead of + +“effective compute (2025-FLOP) = training compute (FLOP) x software efficiency (2025-FLOP per FLOP)” + +it has + +“effective training system performance (2025-FLOP/s) = training system performance (FLOP/s) x software efficiency (2025-FLOP per FLOP)”, + +and effective compute (2025-FLOP) is the cumulative time integral of effective training system performance. + +#### Why make this change? + +The AI Futures Model (and other models like [Tom Davidson’s FTM](https://takeoffspeeds.com/description.html) and [Epoch AI’s GATE](https://epoch.ai/gate/)) both assume you can “pull software efficiency outside the integral” here. This is normally fine ([though it does make fast takeoffs more likely than if you explicitly modeled retraining](https://www.forethought.org/research/will-the-need-to-retrain-ai-models#software-explosion-with-continuous-retraining---model-explanation)), but becomes problematic when considering a heavily compute-constrained AI covert project which, rather than starting from “scratch”, instead starts with a copy of the pre-deal (Chinese) SOTA model weights and all the pre-deal algorithmic secrets. There are basically two ways you might try to model this situation in the vanilla AIFM, and neither seem good. + + 1. Reduce the "training compute” (FLOP) proportionally to the overall covert compute reduction (e.g. by 100x, if the covert project has 100x less compute than the previous leading AI company). This is sort of like assuming they start without any model weights and need to train from scratch. But this isn’t true—not only can they apply further RL (for example) to the pre-deal model, they will also use their best available model to accelerate AI R&D. So you’d at minimum need to keep track of their “best model” separately from their “best trainable-from-scratch model”. _This underestimates covert project capabilities progress._ + + 2. Keep the “training compute” (FLOP) frozen at the pre-deal SOTA level (or perhaps growing 100x slower). But to first order this only slows down progress by about half or less, holding the rate of software efficiency growth constant (e.g. because the legal project has transparent algorithms which are being copied.) This ignores the fact that the covert project needs to actually run those algorithms, which proceeds at a rate proportional to how much compute they have. Instead, it’s a bit like assuming that the stolen algorithms magically apply retroactively to the training that was already done in making the pre-deal model. _This overestimates covert project capabilities progress._ + + + + +The remaining simplification in the current implementation is assuming AI progress consists of a single continuous training run that’s been in progress since the beginning of time. (mathematically, the effective compute integral is from minus infinity to the current time.) In reality, frontier training runs may in fact start from scratch (but not always), [may have different lengths](https://epoch.ai/data-insights/longest-training-run/), may only be underway during parts of the year (thus requiring some unclear amortization for “H100e dedicated to training”), etc. We think this is still an improvement on the way the AI Futures Model treats training runs, and captures most of the dynamics relevant to covert projects. My [BH’s] best guess is that it directionally advantages covert projects more than is realistic, since certain capabilities may require architectural changes or otherwise require one to “start from scratch”, and this model assumes you never need to do this. + +## Non-Chinese covert projects + +This supplement is focused on the possibility of a Chinese covert project. However, there is also a chance that other countries besides China attempt a covert project. We have not looked into this possibility in detail, so are not confident in our conclusions here, but we tentatively believe that China is the country where doing a covert project is most viable, because: + + 1. The US is the only country in the world with more compute than China, and acquiring compute is a central bottleneck for covert projects (see Section 2). + + 2. The US seems less likely to agree to Plan A while simultaneously choosing to divert compute, since the US BATNA is more favorable than the Chinese BATNA. + + 3. The US has less ability to do a covert project relative to China. The main reasons for this: + + 1. It will be much harder for the US to acquire chips because there is no major chip smuggling inside the US, and the US doesn’t seem to have a reasonable excuse to force companies to decommission large numbers of chips, cutting off the major avenues for covert compute procurement.. + + 2. The US has more checks and balances on executive power, so a covert project is probably politically more difficult to set up. It also seems harder to for the US to compel the relevant companies to comply with the efforts of a covert project (e.g., getting semiconductor companies to falsify sales records), as very few companies below the chip-design level are US-headquartered + + + + +1. + +We focus on China in particular because we think China is the best positioned party to defect. This is discussed further in [the appendix](/supplements/covert-ai-projects#non-chinese-covert-projects). + +2. + +Based on our modeling, deliberately holding back on the amount of compute deployed does not improve a covert project’s chance of [success](/supplements/covert-ai-projects#section-5-when-is-a-covert-project-catastrophic). + +3. + +Based on our modeling, the rate of capabilities progress of covert projects is largely determined by the rate of progress of the Consortium. Specifically, under our [transparency recommendations](/supplements/transparency-plan#thomas-larsen) the main lever for speeding up (software progress) also accelerates the covert projects. Hence “going faster to outpace covert projects” usually has the opposite effect—instead, the lead-maximizing strategy involves making as little software progress as possible while maintaining the hardware-buildout-constrained rate of training compute scaling. + +4. + +Moreover, due to Total Research Transparency, for virtually all AI algorithms, we are deliberately publishing them. + +5. + +If the PRC acted with less foresight, and only began diverting chips three months in advance rather than a year in advance, their diversion yield might be ~3x less. If the US was less competent, and failed to reduce the smuggling rate relative to today in the year leading up to the deal, the diversion yield might be ~1.5x more. + +6. + +The mechanics of this are discussed more in our verification supplement [here](https://docs.google.com/document/d/1K4vxl_IMf58OgNXbWntQl6BQ06yAWZ62WmW5Ypp4KYA/edit?tab=t.26hlfwchw78c#heading=h.qfiz1jt93ug3). + +7. + +As above, we are very uncertain about the tractability and potential consequences of doing this. Relevant public information includes the [NSA’s breach of Huawei in 2010](https://www.nytimes.com/2014/03/23/world/asia/nsa-breached-chinese-servers-seen-as-spy-peril.html), which seems to meet or exceed the level of access that would be required to gain accurate accounting records (for example, they “obtained information about the workings of [Huawei-produced networking products] and monitored communications of the company’s top executives”). + +8. + +By "audit," we mean: obtain a firm's complete records of compute purchased, resold or transferred, and retained as inventory, and physically inspect any compute the firm presently owns. + +9. + +It might also be possible to purchase chips directly from Chinese domestic producers in a similar manner, though I’m unsure whether China has export restrictions that would make this more difficult or unusual. In practice it probably does look rather unusual to export any significant fraction of your 2028 production, just because of the strategic importance of AI compute. + +10. + +You could also stockpile the inputs (e.g. logic dies and HBM) and finish the assembly during the deal in a secret facility. For the case we looked at (a covert packaging line) this seemed not worthwhile, namely because of the comparatively low throughput and ongoing inflow of inputs that would be required to operate it. + +11. + +Or their Chinese-designed compute, though doing this for US-made compute has a number of advantages, including (1) power efficiency to minimize required covert datacenter footprint, (2) complementarity with other diversion strategies more applicable to Chinese-designed compute, such as falsifying production records, and (3) increased plausible deniability, though this is speculative. + +12. + +This would be against their interests, but could be a plausible rationale. + +13. + +See our analysis here: Diverting pre-2028 stock by retroactively understating domestic production. + +14. + +On the one hand, the exhaust from a jet engine is much hotter than the exhaust from a chiller, but on the other hand, these aircraft only need to hide from things like IR-guided missiles rather than from repeated satellite observation and analysis. + +15. + +1 cubic meter per second of water flow carries 4.2MW per degree C of temperature rise. The [largest water treatment plant](https://mwrd.org/locations/stickney-water-reclamation-plant) in the world handles ~30 m^3/s (covering around half of Chicago). The [largest](https://www.mee.gov.cn/gkml/sthjbgw/spwj1/201608/t20160815_362297.htm) [Chinese](https://www.nsbdjhsw.com.cn/list-yjbh.html) [underground](https://www.mee.gov.cn/xxgk2018/xxgk/xxgk11/202205/t20220505_976927.html) [water](https://www.ndrc.gov.cn/fzggw/jgsj/njs/sjdt/201808/t20180821_1194903.html) transfer projects carry ~100 m^3/s. The [largest nuclear plants](https://en.wikipedia.org/wiki/Kashiwazaki-Kariwa_Nuclear_Power_Plant) discharge a few hundred m^3/s (they heat water by more than 1 degree C in order to discharge ~10GW at this rate). The [largest desalination plants](https://en.wikipedia.org/wiki/Jebel_Ali_Power_and_Desalination_Plant) withdraw 1000+ m^3/s of seawater. The [largest](https://en.wikipedia.org/wiki/Amazon_River) [rivers](https://en.wikipedia.org/wiki/Yangtze) [in](https://en.wikipedia.org/wiki/Mississippi_River) the world have average discharge in the range 30,000–300,000 m^3/s. + +The best-suited thermal satellite we could find [has](https://www.mdpi.com/2072-4292/15/2/517?utm_source=researchgate.net&utm_medium=article) a total noise equivalent differential temperature of ~0.2 K at a resolution of 57m, better than the 100m resolution of the Landsat 8 and 9, which were used in [this paper](https://www.mdpi.com/2076-3417/15/2/738) to image nuclear plant thermal plumes. So the temperature increase _at the surface_ would perhaps need to be <0.1 degree C. This could be accomplished by drawing 10 m^3/s of water for every 4.2 MW (putting the flow rate for a 100MW facility on par with the world’s largest nuclear plants), or by dumping more megawatts into each m^3/s of water but releasing it at a greater depth (or in some sort of underground waterway) so the plume dilutes more before eventually reaching the surface. I am currently unsure whether this analysis properly accounts for time averaging or other signal processing methods. + +16. + +The technique seems to require stacking multiple images taken at different times from near-identical orbital geometry, which seems poorly suited to broad-area detection. It’s possible this could be solved via brute force with enough satellites. + +17. + +A Space Force spokesperson [said](https://breakingdefense.com/2025/08/space-force-to-launch-actual-gmti-sats-in-the-next-year-general/) to Breaking Defense that dedicated satellites for this purpose are expected to launch in 2028. The program’s FY2026 RDT&E budget schedule separately lists the GMTI “Develop and field GMTI” activity as running through Q4 FY2030, indicating a planned fielding window extending to the end of fiscal 2030, though the document does not identify an IOC or FOC date. + +18. + +Between 2010 and 2012, many CIA sources (seemingly almost all?) in China were [killed and imprisoned](https://en.wikipedia.org/wiki/2010%E2%80%932012_killing_of_CIA_sources_in_China). The New York Times in 2017 [wrote](https://archive.is/QQ8lJ#selection-2867.0-2867.181) that “the C.I.A. has tried to rebuild its network of spies in China… an expensive and time-consuming effort.” [refer to 2023 Burns comments about progress] + +19. + +The most extreme case of this proposal is something like: (1) move all humans to other planets with spaceships, (2) verify that no covert GPUs were transported to other planets, (3) make sure that no compute is diverted from the newly manufactured compute on other planets, and (4) disassemble the earth. This is obviously an extreme case meant to be illustrative, in practice we expect that with all of this capacity, something more efficient and far less costly could be done. + +20. + +Even if challenge inspection provisions are not included in the initial agreement, China might later come under international pressure to permit them if fear of covert projects becomes a binding constraint for Consortium scaling decisions. + +21. + +It might be possible to measure [perfluorocarbon emissions](https://www.epa.gov/sites/default/files/2016-02/documents/pfc_generation.pdf) from the Hall-Heroult process, which could distinguish how much energy is being used on actual aluminum production. We’re not aware of any way to measure this from space at a per-facility level though. + +22. + +In 2024, OpenAI [spent](https://epoch.ai/data-insights/openai-compute-spend) $5B on R&D compute, Zhipu AI spent $308M ([page 10](https://www1.hkexnews.hk/listedco/listconews/sehk/2025/1230/2025123000017.pdf)) on total R&D (which included 657 technical staff; assuming $100k/yr each yields $242.3M on compute), and Minimax spent $189M ([page 11](https://www1.hkexnews.hk/listedco/listconews/sehk/2025/1231/2025123100025.pdf)) total. + +23. + +The compute advantage will widen in some sense with time; in our scenario the Consortium increases their compute at ~3x per year. However, this will only translate straightforwardly to capabilities through training compute scaling (research compute scaling on the other hand will be restricted to target limited amounts of software progress.) + +24. + +Depending on how you measure this; the US Center for AI Standards and Innovation found an apparently growing US lead; here we use the Epoch Capabilities Index. See “[The US-China AI gap is very uncertain](https://mcnair.center/china/)”. + +25. + +As a (somewhat dated) example, [Phi-4](https://epoch.ai/models/phi-4) used GPT-4 to generate specialized synthetic data, exceeding GPT-4’s ECI while requiring around 20x less training compute. + +26. + +We’re assuming here that covert projects are only able to acquire untracked pre-deal compute, i.e. that it’s not feasible for them to smuggle post-deal compute or to create a secret chip supply chain. After extreme capabilities, it would be possible for covert projects to set up a secret compute supply chain, but we don’t anticipate this being the binding constraint, i.e. this would happen after other capabilities we consider to be catastrophic. + +27. + +See further analysis [here](https://www.lesswrong.com/posts/7jcPg79p3kD5ir3CL/how-much-slower-does-takeoff-go-with-10-less-compute). + +28. + +An extreme version of this could undermine Nuclear MAD, but this would be very difficult. It would need to happen in secret or else other countries would respond by building up nuclear stockpiles or intervening on buildup. + +29. + +Importantly, we do not include the risk due to weight theft during the deal in these estimates. + +30. + +One reason is that even if a covert project turns out to be non-viable, worlds where China diverts significant compute might force the Consortium to adopt a less favorable scaling strategy due to their uncertainty about how quickly the missing quantity of compute would enable deal-undermining AI to be reached by a hypothetical covert project. Another reason is that the tensions associated with detecting and shutting down a covert project would incur additional risk of deal decline, even if the covert project itself is small enough to be harmless. + +31. + +This may not be possible. See “BATNA-warping capabilities”. + +32. + +We actually surveyed AI researchers on a question like this for the AI Futures Model, asking about how much slower algorithmic progress would go if they had 10x less compute for experiments. They said around 2.8x. Though a 10x reduction would also hit inference compute, in practice we expect firms to have more parallel AI labor than they know what to do with during the intelligence explosion and to be very bottlenecked on experiment compute, such that the reduction in experiment compute is the limiting factor. + +33. + +Insert convergence details here. diff --git a/sources/ai2040-pages/deal-decline.md b/sources/ai2040-pages/deal-decline.md new file mode 100644 index 0000000..524498f --- /dev/null +++ b/sources/ai2040-pages/deal-decline.md @@ -0,0 +1,1129 @@ +--- +url: https://ai-2040.com/supplements/deal-decline +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Deal Decline + +### Eli Lifland + +## Introduction + +A consideration pushing toward [scaling faster](/supplements/capability-scaling-strategy) in Plan A is _deal decline_ : the possibility of the deal either dissolving or becoming substantially less effective (which we call _deal impairment_). + +In this supplement, I’ll estimate the likelihood and badness of deal decline. I’m confident that deal decline is at least 10% likely to occur within Plan A’s first 10 years, and that a 2035 decline leads to the expected value of the future being at least 5% lower than if there were a 2035 handoff within the deal. Since deal decline seems worryingly likely, and scaling capabilities within the Plan A deal seems much safer than an AI race once it’s declined, I’m confident that deal decline is a substantial consideration, likely the biggest one, in favor of [scaling up AI capabilities faster and handing off to AIs earlier](/supplements/capability-scaling-strategy) (rather than something like a very long pause). But I’m not precisely confident in any of the numbers estimated below; I could easily update substantially toward being more or less worried about deal decline. + +Despite a substantial chance of decline, I think Plan A remains [much more valuable than the other plans](https://www.ai-2040.com/supplements/comparing-possible-plans). Even if deal decline happens within a few years of the deal, [the outcome will probably still be substantially better than no deal](/supplements/deal-decline#overview). And I think it’s plausible that a long-lasting effective deal can be achieved, via [mechanisms such as automated lie detection and robust institutional design](/supplements/deal-decline#preventing-deal-decline). + +## Overview + +After categorizing [types of deal decline](/supplements/deal-decline#types-of-deal-decline), my analysis consists of: + + 1. **[Causes of deal decline:](/supplements/deal-decline#causes-of-deal-decline)** I split causes into (a) deliberate deprioritization of effective deal maintenance due to leadership change, (b) deliberate deprioritization due to a change in existing leadership’s priorities , and (c) destabilizing events. + + 2. **[Likelihood of deal decline](/supplements/deal-decline#likelihood-of-deal-decline)** : Informed by base rates from other international agreements and a decomposition into possible causes, I estimate the chance of deal decline per year at ~12% in the first 2 years of the deal, ~7% in years 2-5, and ~4% in years 5-10. The risk is roughly equally split across dissolution and impairment. This implies a 37% chance of decline within 5 years and 48% within 10. During the deal and to some extent leading up to it, society will have much more information than I have today about how robust the deal is, as such their probabilities may be quite different. + + 3. **[Badness of deal decline](/supplements/deal-decline#overview):** I compare the situation after deal decline to the situation before there was any deal at all. The most important ways the situation is better post-deal-decline (relative to the situation without a deal) come from having used AIs to do lots of useful stuff (e.g. safety research) and from additional calendar time yielding more societal awareness and preparation. On the other hand, if [Plan A’s compute destruction mechanisms](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots) fail, there will be much more compute in the world post-deal-decline, leading to a faster and more dangerous takeoff. Overall, it seems that a post-deal-decline situation would be substantially better than if there were no deal at all, while being significantly worse than doing handoff within the deal. When the deal decline happens matters a lot; the later, the better. + + 4. **[Preparing for deal decline](/supplements/deal-decline#preparing-for-deal-decline):** We recommend investing lots of resources into preventing deal decline, as well as mitigating it. Our main recommendation for mitigating deal decline is to [ensure that nearly all of the compute in the world](/supplements/deal-decline) [is destroyed](/supplements/deal-decline) in the case that the deal dissolves. Without this destruction, takeoff might be extremely fast due to an overhang of compute that wasn't being used for developing more generally capable AIs. + + + + +## Types of deal decline + +The following are the types of deal decline: + +**Deal dissolution.** The deal is officially dissolved. + +**Deal impairment.** The deal becomes substantially less effective than if it were implemented well and adhered to, but isn’t dissolved. For simplicity, this is modeled as a binary in this supplement; in practice, it would likely be gradual. Ways the deal could be impaired: + + 1. **Deal non-enforcement.** The deal is not enforced as strongly as before, e.g. a country blatantly not complying is not punished. This leads to increased violations of the letter of the deal. + + 2. **Deal misexecution.** The letter of the deal is followed but is executed poorly. + + 3. **Deal revision.** The deal is explicitly edited to be worse. This could be straightforward weakening, but it could also be making it worse in other ways, such as unnecessarily increasing surveillance. + + + + +The deal could also end via a [covert project](https://www.ai-2040.com/supplements/covert-ai-projects) reaching takeover-capable AIs or via an existential catastrophe (e.g., via AI takeover, AI-enabled bioweapons, or AI-developed superweapons). These end conditions are out of scope for this supplement and aren’t considered to be deal decline in our terminology. They are included separately in our [analysis of the Plan A capability scaling strategy](/supplements/capability-scaling-strategy). + +## Causes of deal decline + +One path for deal decline to happen is via deliberate deprioritization of maintaining an effective deal. The path for this is: + + 1. At least one of the US or China leadership decides to deprioritize the deal, which could happen for a variety of reasons, some of which are listed below. + + 2. If this only happens inside of one of the countries, then the other country must either try to use their leverage to prevent deal decline and fail, or let the deal decline happen (e.g. because they are scared of escalation). Other countries will also have some leverage at pressuring countries in ways that prevent deal decline. + + + + +We can split out ways that (1), i.e. US or Chinese leadership’s priorities changing, might happen into either (a) change of leadership (e.g. a US election) or (b) a change in existing leadership’s priorities. + +Why might existing leadership’s priorities change? The most likely reasons we can think of are: + + 1. Concern regarding how future domestic leaders would handle things. + + 2. Wanting to gain power by developing superintelligence before a leadership transition. + + 3. Becoming more optimistic about winning the AI race in a world in which the deal has declined. + + 4. Becoming less concerned about risks from superhuman AIs such as AI takeover or extreme concentration of power. + + 5. Wanting to illegitimately grab power and thinking the deal impedes this. + + 6. Becoming less optimistic about the deal working, e.g., they think misuse risk is too high and difficult to mitigate within the deal. This could lead to a failed negotiation to revise the deal. + + + + +And then there’s an “Other” bucket of many reasons that we haven’t thought of. + +Another possible reason the deal might decline is via a destabilizing event. The most likely we can think of are: + + 1. Geopolitical shock, e.g. war. + + 2. A country being caught cheating: a [covert project](https://docs.google.com/document/u/3/d/1lOzL0Wb9lyZE8-kfUKMoUHHcShtyCig3l41JuooZsS0/edit) getting caught, or another way of cheating such as trying to undermine the inference-only verification. + + 3. A sub-existential catastrophe. + + + + +## Likelihood of deal decline + +### Overall view + +Below I estimate the likelihood of deal decline in Plan A by year, focusing on the first 10 years (as that’s when handoff happens in the scenario). During the deal and to some extent leading up to it, society will have much more information than I have today about how robust the deal is, as such their probabilities may be quite different. + +I use 2 methods: + + 1. **[Base rates](/supplements/deal-decline#base-rates-plus-adjustments)** [](/supplements/deal-decline#base-rates-plus-adjustments)**[plus adjustments](/supplements/deal-decline#base-rates-plus-adjustments):** A Claude-done analysis, focused on the base rate of international deals declining, estimates the chance of deal decline per year at: ~15%/year in the first 2 years of the deal, ~12%/year in years 2-5, and ~11%/year in years 5-10. I adjust these downward because Plan A has especially strong verification and implementing it is evidence of high political will to maintain it. After adjustment, risk is roughly equally split across dissolution and impairment. + + 2. **[Decomposition into causes](/supplements/deal-decline#decomposition-into-causes):** Decomposing into possible causes of deal decline and then combining these estimates gives me ~10%/year deal decline in the first 5 years of the deal, and ~6%/year in years 5-10. This moves me up a little from the base rates plus adjustments method, but the results are overall fairly similar across methods. + + + + +My overall view is, aggregating the above 2 methods: + +Phase (years) + +0-2 + +p(dissolution)/yr + +6% + +p(impairment)/yr + +6% + +p(decline)/yr + +11.6% + +p(decline) this time period + +21.9% + +Cumulative p(decline) + +21.9% + +Phase (years) + +2-5 + +p(dissolution)/yr + +3.5% + +p(impairment)/yr + +3.5% + +p(decline)/yr + +6.9% + +p(decline) this time period + +19.2% + +Cumulative p(decline) + +37% + +Phase (years) + +5-10 + +p(dissolution)/yr + +2.0% + +p(impairment)/yr + +2.0% + +p(decline)/yr + +4.0% + +p(decline) this time period + +18.3% + +Cumulative p(decline) + +48% + +**Phase (years)**| **p(dissolution)/yr**| **p(impairment)/yr**| **p(decline)/yr**| **p(decline) this time period**| **Cumulative p(decline)** +---|---|---|---|---|--- +0-2| 6%| 6%| 11.6%| 21.9%| 21.9% +2-5| 3.5%| 3.5%| 6.9%| 19.2%| 37% +5-10| 2.0%| 2.0%| 4.0%| 18.3%| 48% + +When doing modeling for the [Capability Scaling Strategy](/supplements/capability-scaling-strategy) and the Plan A portion of [Comparing Possible Plans](https://www.ai-2040.com/supplements/comparing-possible-plans), I use the following hazard rates with linear interpolation in between. The CDF is very similar to the above table. + +Years after deal start + +**0** + +Hazard (any decline) + +0.12 + +Cumulative p(decline) + +0% + +Years after deal start + +**3** + +Hazard (any decline) + +0.08 + +Cumulative p(decline) + +25.9% + +Years after deal start + +**6** + +Hazard (any decline) + +0.06 + +Cumulative p(decline) + +39.9% + +Years after deal start + +**8** + +Hazard (any decline) + +0.04 + +Cumulative p(decline) + +45.7% + +Years after deal start + +**11** + +Hazard (any decline) + +0.03 + +Cumulative p(decline) + +51.1% + +Years after deal start + +**16** + +Hazard (any decline) + +0.02 + +Cumulative p(decline) + +56.8% + +**Years after deal start**| **Hazard (any decline)**| **Cumulative p(decline)** +---|---|--- +**0**| 0.12| 0% +**3**| 0.08| 25.9% +**6**| 0.06| 39.9% +**8**| 0.04| 45.7% +**11**| 0.03| 51.1% +**16**| 0.02| 56.8% + +# Deal decline risk reaches ~50% by 2040 handoff + +0%25%50%75%100%hazard — p(decline)/yr:12%/yr8%/yr6%/yr4%/yr3%/yr2%/yr51.1% by 2040 (handoff)56.8%202920312033203520372039204120432045YearCumulative p(deal decline) + +Assumes the deal starts in 2029. Hazard interpolates linearly between my anchor estimates; risk is roughly equally split between dissolution and impairment. + +cumulative = 1 − exp(−∫ hazard) + +### Base rates plus adjustments + +#### Data collection methodology + +We had Claude do analysis of base rates and adjustments, then manually reviewed and edited its outputs. A summary of the approach (read its full analysis [below](/supplements/deal-decline#appendix-claudes-analysis-of-base-rates)): + +Claude collates a list of 31 past international agreements approximately similar to Plan A in scale and nature. For each agreement, it identifies time-to-first-impairment and time-to-dissolution, and also gives a qualitative ranking of how applicable the agreement was to Plan A. + +Base hazard rates are computed for each phase (number of years since the deal started) via an average rate weighted by applicability scores, separately for dissolution vs impairment: + + * For impairment, this is computed as a weighted number of years before first impairment divided by the weighted number of impairment exposure. + + * For dissolution, this is the weighted sum of dissolved agreements divided by the weighted number of dissolution-exposure years. + + + + +Claude also identifies the reasons for impairment and dissolution. It finds that dissolution is roughly hump-shaped centered around 5-10 years, and impairment is front-loaded. + +Claude finds that improved verification correlates with delayed impairment, and dissolution usually occurs as a result of political withdrawal, rather than inherent deal failure. + +#### Base rate (plus adjustments) forecasts + +Below are Claude’s collected base rates. + +Method + +Weighted base rate (HEADLINE) + +Dissolution hazard/yr + +3.4% + +Impairment hazard/yr + +7.9% + +Method + +Unweighted (all treaties equal) + +Dissolution hazard/yr + +1.8% + +Impairment hazard/yr + +4.8% + +**Method**| **Dissolution hazard/yr**| **Impairment hazard/yr** +---|---|--- +Weighted base rate (HEADLINE)| 3.4%| 7.9% +Unweighted (all treaties equal)| 1.8%| 4.8% + +I’ll now discuss how the hazard rate changes over time. I’ll focus on the first 10 years as that’s how long the deal is in the scenario. A note on 10+ years: the impairment hazard holds at ~6%/yr in years 10–20 rather than continuing to decline, before falling to ~4%/yr past 20. Here are base rates by how long the deal has been in force: + +Phase (years) + +0-2 + +Dissolution hazard/yr + +0.0% + +Impairment hazard/yr + +15.5% + +Combined hazard/yr + +15.5% + +Cumulative dissolution (by end of phase) + +0.0% + +Cumulative impairment (by end of phase) + +28.6% + +Cumulative combined (by end of phase) + +28.6% + +Phase (years) + +2-5 + +Dissolution hazard/yr + +3.4% + +Impairment hazard/yr + +9.4% + +Combined hazard/yr + +12.5% + +Cumulative dissolution (by end of phase) + +9.9% + +Cumulative impairment (by end of phase) + +46.9% + +Cumulative combined (by end of phase) + +52.1% + +Phase (years) + +5-10 + +Dissolution hazard/yr + +5.8% + +Impairment hazard/yr + +5.9% + +Combined hazard/yr + +11.4% + +Cumulative dissolution (by end of phase) + +33.2% + +Cumulative impairment (by end of phase) + +60.9% + +Cumulative combined (by end of phase) + +73.8% + +**Phase (years)**| **Dissolution hazard/yr**| **Impairment hazard/yr**| **Combined hazard/yr**| **Cumulative dissolution (by end of phase)**| **Cumulative impairment (by end of phase)**| **Cumulative combined (by end of phase)** +---|---|---|---|---|---|--- +0-2| 0.0%| 15.5%| 15.5%| 0.0%| 28.6%| 28.6% +2-5| 3.4%| 9.4%| 12.5%| 9.9%| 46.9%| 52.1% +5-10| 5.8%| 5.9%| 11.4%| 33.2%| 60.9%| 73.8% + +I make adjustments to the base rate as follows: + +Phase (years) + +0-2 + +Base diss/yr + +0.0% + +Base imp/yr + +15.5% + +Adjusted diss/yr + +5.0% + +Adjusted imp/yr + +5.0% + +Adjusted combined/yr + +9.8% + +Rationale + +Dissolution from 0 to 5%: 0% is to some extent an artifact of limited dataset size, plus it's easy to imagine ways Plan A could dissolve quickly (e.g. a covert project getting caught). + +Impairment from 15 to 5%: Plan A has especially strong verification and requires a very high amount of political will to initiate, making it less likely it is very quickly impaired. + +Phase (years) + +2-5 + +Base diss/yr + +3.4% + +Base imp/yr + +9.4% + +Adjusted diss/yr + +3% + +Adjusted imp/yr + +3% + +Adjusted combined/yr + +5.9% + +Rationale + +Decaying impairment in a way that is roughly similar to the base rate, but slightly more aggressively due to AI-assisted verification/epistemics/etc.. For dissolution, decaying in a similar way (partly for simplicity, maybe it should be slightly different). + +Phase (years) + +5-10 + +Base diss/yr + +5.8% + +Base imp/yr + +5.9% + +Adjusted diss/yr + +1.50% + +Adjusted imp/yr + +1.50% + +Adjusted combined/yr + +3.0% + +Rationale + +Continuing to decay, slightly more aggressively than above. + +**Phase (years)**| **Base diss/yr**| **Base imp/yr**| **Adjusted diss/yr**| **Adjusted imp/yr**| **Adjusted combined/yr**| **Rationale** +---|---|---|---|---|---|--- +0-2| 0.0%| 15.5%| 5.0%| 5.0%| 9.8%| Dissolution from 0 to 5%: 0% is to some extent an artifact of limited dataset size, plus it's easy to imagine ways Plan A could dissolve quickly (e.g. a covert project getting caught).Impairment from 15 to 5%: Plan A has especially strong verification and requires a very high amount of political will to initiate, making it less likely it is very quickly impaired. +2-5| 3.4%| 9.4%| 3%| 3%| 5.9%| Decaying impairment in a way that is roughly similar to the base rate, but slightly more aggressively due to AI-assisted verification/epistemics/etc.. For dissolution, decaying in a similar way (partly for simplicity, maybe it should be slightly different). +5-10| 5.8%| 5.9%| 1.50%| 1.50%| 3.0%| Continuing to decay, slightly more aggressively than above. + +There are 2 effects which lead me to estimate substantially lower deal decline risk as time goes on: + + 1. If the deal has gone on for longer, this provides evidence that it is less likely per unit time to decline. + + 2. As AI capabilities improve during the deal and there is time to put them to good use, they allow for building technology that helps stabilize the deal. For example, there might be automated lie detection or use of [AI to improve epistemics](/supplements/ai-for-epistemics) or verification. + + + + +Surprisingly, the base rate data above has p(dissolution) increasing over time, while p(impairment) decreases in a fashion closer to what I expect for Plan A. Notably, the base rate data doesn’t include effect (2) above, or at least includes a much weaker version of it. So I expect the trend of decreasing deal decline risk to be more pronounced than it is in the data on previous agreements. + +### Decomposition into causes + +An alternative method for estimating the chance of deal decline is to individually estimate the chance for each possible cause, then aggregate them together. This gives me an estimate of 10% p(decline)/year in the early 2030s and 6% in the late 2030s, leading to a cumulative p(decline) of ~50% by 2035 and ~60% by 2040. I think the methodology of listing a bunch of possible ways something could happen and then aggregating these probabilities tends to produce numbers on the high side, which I take into account when aggregating this methodology with the base rate plus adjustments method. + +**Deliberate deprioritization** + +**US leadership change** + +2029-2034, p(decline)/year + +0.02 + +2035-2039, p(decline)/year + +0.01 + +**China leadership change** + +2029-2034, p(decline)/year + +0.01 + +2035-2039, p(decline)/year + +0.01 + +**Concern regarding future leaders** + +2029-2034, p(decline)/year + +0.01 + +2035-2039, p(decline)/year + +0.01 + +**Optimism about winning post-deal-decline** + +2029-2034, p(decline)/year + +0.01 + +2035-2039, p(decline)/year + +0.005 + +**Less concerned about AI safety** + +2029-2034, p(decline)/year + +0.005 + +2035-2039, p(decline)/year + +0.003 + +**Other deliberate deprioritization** + +2029-2034, p(decline)/year + +0.01 + +2035-2039, p(decline)/year + +0.005 + +**Destabilizing event** + +**Geopolitical shock** + +2029-2034, p(decline)/year + +0.01 + +2035-2039, p(decline)/year + +0.005 + +**Country caught cheating** + +2029-2034, p(decline)/year + +0.015 + +2035-2039, p(decline)/year + +0.015 + +**Sub-x-risk catastrophe** + +2029-2034, p(decline)/year + +0.01 + +2035-2039, p(decline)/year + +0.005 + +**Other destabilizing event** + +2029-2034, p(decline)/year + +0.015 + +2035-2039, p(decline)/year + +0.01 + +**All sources** + +**p(decline)/year from all sources** + +2029-2034, p(decline)/year + +**0.10** + +2035-2039, p(decline)/year + +**0.06** + +**p(decline) this time period** + +2029-2034, p(decline)/year + +**0.49** + +2035-2039, p(decline)/year + +**0.27** + +**Cumulative p(decline)** + +2029-2034, p(decline)/year + +**0.49** + +2035-2039, p(decline)/year + +**0.62** + +| | **2029-2034, p(decline)/year**| **2035-2039, p(decline)/year** +---|---|---|--- +**Deliberate deprioritization**| **US leadership change**| 0.02| 0.01 +**China leadership change**| 0.01| 0.01 +**Concern regarding future leaders**| 0.01| 0.01 +**Optimism about winning post-deal-decline**| 0.01| 0.005 +**Less concerned about AI safety**| 0.005| 0.003 +**Other deliberate deprioritization**| 0.01| 0.005 +**Destabilizing event**| **Geopolitical shock**| 0.01| 0.005 +**Country caught cheating**| 0.015| 0.015 +**Sub-x-risk catastrophe**| 0.01| 0.005 +**Other destabilizing event**| 0.015| 0.01 +**All sources**| **p(decline)/year from all sources**| **0.10**| **0.06** +**p(decline) this time period**| **0.49**| **0.27** +**Cumulative p(decline)**| **0.49**| **0.62** + +Unlike for most sources, the risk from a country being caught cheating might not go down in the late 2030s relative to the early 2030s, due to detection methods getting better. + +## Badness of deal decline + +### Overview + +I compare the situation after deal decline to the situation before there was any deal at all. The most important advantages of post-deal-decline worlds over no-deal worlds are 1) having more time for capable AIs to do useful things like safety research and 2) having more time to build societal awareness and preparation. + +On the other hand, there will be much more compute in the world post-deal-decline, by default leading to a faster and more dangerous takeoff. Because of this, Plan A is set up so that the compute is likely destroyed upon deal dissolution; more on this [below](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots). + +I made some [quantitative estimates](/supplements/deal-decline#quantitative-estimates) of how bad deal decline would be. The headline results are: + + * Deal dissolution and impairment are similarly bad, perhaps impairment is a bit worse. While impairment has the benefit of there still being a deal, it has the downside of not triggering compute destruction. + + * Pre-max-controllable-AI, I propose modeling p(alignment | deal decline) as interpolating linearly in logit space between p(alignment; pre-deal unconditional) in 2029 and p(alignment | deal decline at max-controllable-AI target date). For a target date of 2035, perhaps p(alignment | deal decline at max-controllable-AI target date) is ~70%. + + * Post-max-controllable-AI, I propose modeling p(alignment | decline at time t) as roughly 90% of p(alignment | handoff within deal at time t). + + + + +### High level situation + +What would happen in Plan A if the deal declined and we hadn’t implemented any specific preparatory measures? Let’s start by considering the case of the deal fully dissolving. + +There will be to some extent a return to the default world conditions, but with some important changes compared to if no deal had happened at all. + +**Positives:** + + 1. **Progress due to AI usage and studying AIs:** Due to the alignment and control research that’s occurred since the start of the deal, the state of knowledge regarding safety techniques has improved. This is true to a greater extent the longer the Consortium has had to use and study more capable AIs. There are also other benefits of AI usage such as [improving epistemics](/supplements/ai-for-epistemics) and verification tech. + + 2. **More societal awareness and preparation:** Aside from using AIs to make progress, the extra calendar time that the deal has bought helps in various ways by giving society more time to adapt. Furthermore, people, including policymakers, are more aware of current and projected AI capabilities than they were before. An important example of this is that there are likely lots more human AI safety researchers than there were at the start of the deal. + + 3. **Training compute safety tax paid:** The most capable models have been trained with more compute than would be needed to reach their level of capabilities, to allow use of safer methods. This is unlike what happens in a closely contested AI race. You may maintain the ability to use this model after dissolution. This positive is potentially minor because: (a) If AIs are not close to max-controllable-AI, we will have to retrain them in order to reach handoff-level capabilities, forfeiting our safety tax. (b) Even if AIs are close to max-controllable-AI, projects likely choose to not incur safety taxes on successor AIs and instead race to higher capabilities. + + + + +**Negative: By default, more compute would lead to faster takeoff, but Plan A[sets things up so that the compute is likely destroyed upon dissolution](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots).** By default, there is now more compute in the world than there was in 2029, which makes the default takeoff faster. In the Plan A scenario, compute stock increases by roughly 0.7 OOMs/year from 2030 to 2033 and 0.1-0.3 OOMs/year from 2034 to 2040. Naively applying a [5.5x speedup for a 10x increase in compute stock](https://www.lesswrong.com/posts/7jcPg79p3kD5ir3CL/how-much-slower-does-takeoff-go-with-10-less-compute) would imply that 3 OOMs would speed takeoff up by ~150x! There might be some compute destroyed though, for example due to a war or due to targeted sabotage of datacenters. There is also more industrial capacity in the world, which could allow the manufacturing of dangerous weapons, potentially after an [industrial explosion](https://www.forethought.org/research/the-industrial-explosion). Because of this issue, Plan A is [set up so that the](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots) [compute](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots) [is likely destroyed upon deal dissolution](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots). This likely averts an extremely fast takeoff, though our best guess is that it still takes about 1-2 years to do the rest of takeoff; see [below](/supplements/deal-decline#post-dissolution-takeoff-length) for discussion. + +**Unclear: Different geopolitical circumstances.** Depending on how the deal dissolves, the geopolitical circumstances might be very different after dissolution than in the default world. For example, there might be a war going on, and if not, the deal dissolution might still make lower levels of coordination less feasible. On the other hand, great powers have signaled that they care a lot about AI safety by implementing the deal and depending on how dissolution happens this might carry through into better outcomes. It’s plausible that a new international deal could be struck not long after the deal dissolving. + +For deal impairment, nearly all aspects of the world are at least somewhat better than in dissolution because the deal is still in effect. For example: not only is the training compute safety tax already paid, but the deal is still partly functional so you may be able to continue paying a tax. How much better things look than dissolution depends on the circumstances of the impairment. + +The main downside is that the extra compute likely hasn’t been destroyed, as it’s harder to set up a system which destroys compute upon impairment as opposed to destroying it upon dissolution. While there’s a ton of compute in the world, the deal is likely still constraining takeoff to some extent. That is, you may not be able to use a large fraction of compute in the world to take off as fast as possible. + +### Quantitative estimates + +I’ve made some rough quantitative estimates that attempt to take into account the above factors. See [here](https://docs.google.com/spreadsheets/d/1MVdjISa09Y4qO3jT5eXYFz-s45gQm4NTXcZqEsFSowU/edit?gid=2143764324#gid=2143764324) for a sheet that contains estimates of the value of the future, conditional on deal dissolution at various times (among estimates of other deal end states). I think that deal impairment would on average be similarly bad to deal dissolution, with the pros and cons described at the end of the previous section roughly canceling out. + +Overall, my view is roughly that we can interpolate value in logit space over the course of Stage 1 of the deal, i.e. pre-max-controllable-AI, between the default world and the end state in which the deal declines at max-controllable-AI. For Stage 2, max-controllable-AI and onward, perhaps the value of deal decline in a given year is roughly 75-80% of handing off within the deal in that year. + +## Preparing for deal decline + +Given how big of a threat deal decline is, we recommend making it a top priority in Plan A to prevent and mitigate it. For example, I think at least 5% as much as the effort that goes into technical safety research in Plan A should go into preventing and preparing for deal decline (and potentially much more). These implications are discussed in the [Capability Scaling Strategy supplement](/supplements/capability-scaling-strategy). + +One way to reduce the risks of and mitigate the downsides of deal decline is to scale AI capabilities faster and handoff faster. + +### Preventing deal decline + +Here are some avenues that we think are promising for preventing deal decline: + + 1. **Automated highly-adversarially-robust lie detection:** This would drastically reduce the chance of being surprised by an actor attempting to pull out of the deal. + + 2. **Institutional design:** The deal and Consortium should be designed in a way that best sets up incentives for the countries involved, and leads to reasonable regulatory decisions. Additionally, domestic institutions should be improved to be robust to the whims of a single person shifting; for example, ideally the deal should be ratified as an official treaty in the US. + + 3. **United diplomatic pressure:** All countries in the deal should make it clear that any country pulling out of the deal would be an unacceptable threat to all other countries’ national security and sovereignty, and will be treated as such. Any country that is considering pulling out of the deal should know that they will be heavily sanctioned, cyberattacked, and would potentially have to fight a war in which they are heavily outnumbered. + + 4. **[AI for epistemics](/supplements/ai-for-epistemics):** Generally improving societal epistemics seems important for creating a robust, effective deal. Example applications of AI that would be especially useful are forecasting failure modes of the deal and helping with the institutional design mentioned above. + + + + +### Preparing stockpiles to align incentives + +We recommend the following to improve the post-deal-dissolution situation and improve incentives around defection. + +**The US and China each get hardened stockpiles of 2% of their pre-deal compute. 2% is the factor used in the Plan A scenario, in practice it should depend on the amount of missing compute.** + +One potential incentive for the US and China to create a covert project is that they could pull out of the deal, destroy the centralized compute (e.g. with bombing, or with other mechanisms, e.g. by withholding offline licensing keys). Then, if one side did a covert project and the other did not, the side that had a covert project might get a large lead over the other, and beat them to an intelligence explosion. + +One solution to this is to allow each side to have a cache of GPUs that are hardened but aren’t plugged into power, and so can’t be used for AI research while the deal is in place. In practice this stockpile could be stored in a secure facility on the home soil of each party, and is deliberately designed to not be destroyable by the other in the case of a deal dissolution, but where inspectors can still verify these chips aren’t in use for the duration of the deal. + +This cache could be on the order of 2% of each side’s pre-deal compute. This is probably more compute than a covert project could have, but is a relatively small fraction of post-deal compute, because new chips dominate the overall chip supply. + +Now, if the deal were to break down, and the verified datacenters were stopped, then both sides would return back to the counterfactual of racing with a proportional amount of compute as they had before the deal; this also favors the US given that the US has more compute pre-deal. + +**In the same stockpile the US and China each get a cold, heavily monitored copy of the frontier model weights.** + +Reasoning for doing this: + + 1. This reduces the incentive to exfiltrate the frontier weights and then immediately pull out of the deal to get a huge advantage. It also reduces the incentive to do a covert project. + + 2. The benefits of the training compute safety tax that was paid to create the frontier model weights doesn’t have to be re-paid. However, there are also downsides to preserving the weights; for example, it makes AI progress faster. It’s unclear that you would want to keep the weights in cold storage for only this reason. + + + + +A proposal to store the weights: + + 1. There’s a hard drive that needs US+China permission to unlock. That’s the only cold storage. + + 2. Other storage is volatile memory, so if you turn off power memory goes away + + 3. To stop physical exfiltration, you have a bunch of physical defenses that automatically turn off power + + + + +### Mutually assured compute destruction: destroying compute, fabs, and robots + +A central potential problem with Plan A is that if the deal dissolves, there is so much compute in the world that we might get an extremely fast intelligence explosion. For example, suppose the deal breaks down at the start of 2035; there are 100B H100e in the world at that time. That’s about 2.3 OOMs more compute than were in the world at the start of 2030 in the default world (~500M), and so a takeoff that would have taken 1 year, might instead happen 5^2.3=40x faster, or in 9 days. This would, of course, be an extremely dangerous situation. Therefore, if the Plan A deal breaks down, it becomes extremely important that a large fraction of the compute in the world be destroyed. + +In this section, we’ll discuss our policies for ensuring mutually assured compute destruction; as well as making sure that the rebuilding process does not result in an unacceptably fast takeoff. The priority order for achieving this is: + + 1. Ensure that almost all of the compute is destroyed. + + 2. Ensure that almost all of the fabs are destroyed. + + 3. Ensure that almost all of the robot capacity is destroyed. + + + + +We’ll discuss each of the above in turn, and then analyze how long an intelligence explosion would take in the case of a deal dissolution. + +#### Compute + +This table gives point estimates for each type of compute in the Plan A scenario: + +| Amount of [effective](https://ai-2040.com/supplements/compute-supplement#appendix-the-effective-h100e-model) H100e| +---|---|--- +Type of compute| 2030| 2035| 2040| What happens if the deal dissolves? +AI datacenter compute (world)| ~500M| ~100B| ~1T| Self-destruct or are bombed +Verified consumer hardware and edge compute (world)| ~0| 100M| 10B| Not useful for AI training because of verification or fundamental design. +Cap-and-traded consumer hardware and edge compute(world)| 25M| 30M| 30M| ~10% can be diverted to an AI project, because the compute is spread across countries and also non-authoritarian countries probably wouldn’t round up this compute. +Stockpile (US/China/RoW)| n/a| 5M/500k/700k| 5M/500k/700k| ~100% can be diverted to an AI project +Covert project (US&China)| 1M| 1M| 1M| ~100% can be diverted to an AI project +Exempt military compute (US&China, others lower)| 10K| 10K| 10K| ~100% can be diverted to an AI project + +The policies for how to achieve this are as follows: + + * **AI datacenters.** Both the US and China build datacenters in deliberately vulnerable locations, thus incentivizing them to self-destruct if the deal breaks down. This is discussed more in the 2034 section of the scenario. Almost all compute is in this category (as shown above), and so it’s very important to ensure destroyability. + + * **Consumer hardware and edge compute.** This includes all compute used in phones, consumer laptops, etc. This compute is distributed throughout the world and could be diverted to an AI project after the deal dissolves. Therefore, we must either (i) verify so that it can’t be directed to training or (ii) cap the total amount of chips in this category. We recommend a cap and trade regime on unverified chips in this category of compute at 30M effective-H100e. We estimate the pre-deal stock to be around 25M effective-H100e, so normal production levels can continue in the short term but then for new AI-capable edge compute to be produced old hardware needs to be traded in / verifiably destroyed. If extremely robust on-chip verification can be developed to activate further chips, or methods of building chips that fundamentally cannot be repurposed to training then this can get an exception from the cap. One approach that could accomplish this is to bake a restrictive computational structure into chips during the manufacturing process (e.g. [Etched](https://www.etched.com/)). + + * **Stockpile.** See [above](/supplements/deal-decline#preparing-stockpiles-to-align-incentives). + + * **Exempt military compute.** We recommend minimizing AI relevant exempt military compute, or only exempting a tiny number of GPUs, such as 10,000 H100e. + + + + +#### Fabs + +In Plan A, all of the AI-relevant fabs in the world are moved to destroyable special economic zones ([SEZs](https://ai-2040.com/supplements/verification-plan#2032-cap-and-trade-for-robots-and-compute)) by 2032. These fabs are located either on the ocean or nearby adversarial superpowers with self-destruct mechanisms in them. Pre-deal fabs without these properties (e.g. TSMC Arizona) are decommissioned. Fortunately, decommissioning isn’t costly because it’s such a small fraction of the compute built in the 2030s. + +Year + +Flow (H100e) + +2030 + +363M + +2032 + +13.5B + +2034 + +43B + +2036 + +93B + +2038 + +233B + +2040 + +1T + +Year + +Time to build 100M H100e + +2030 + +100 days + +2032 + +3 days + +2034 + +20 hours + +2036 + +9.4 hours + +2038 + +3.8hours + +2040 + +53 minutes + +Year| 2030| 2032| 2034| 2036| 2038| 2040 +---|---|---|---|---|---|--- +Flow (H100e)| 363M| 13.5B| 43B| 93B| 233B| 1T +Time to build 100M H100e| 100 days| 3 days| 20 hours| 9.4 hours| 3.8hours| 53 minutes + +100M H100e is roughly the amount of compute needed to do an intelligence explosion (from Automated coder to TED-AI) in 1 year. By the end of the decade in Plan A, industrial capacity is so high that we’re producing that amount of compute every 53 minutes. Therefore, it’s extremely important for the fabs to be destroyed if the deal breaks down, particularly towards the end of the decade. Fortunately, the fabs are extremely fragile by nature, and the same mechanism as the AI datacenters above (putting them in locations where self-destruction is incentivized) can be used. + +#### Robots + +Robots can be used to build fabs, which can build GPUs, which can then be used to do an intelligence explosion. When there are enough robots, this process can happen extremely quickly. + +Our basic strategy with robots mirrors that with GPUs and fabs. Almost all robot production and use will occur in designated SEZs: + +Type of Robot + +**SEZ** + +2030 — Human-equivalent quantity + +0 + +2035 + +1B + +2040 + +70B + +What happens if the deal dissolves? + +Self-destruct or bombed + +Type of Robot + +**Civilian** + +2030 — Human-equivalent quantity + +100K + +2035 + +10M + +2040 + +1B + +What happens if the deal dissolves? + +Self-destruct or backdoored or not useful for building fabs + +Type of Robot + +**Exempt military robots** + +2030 — Human-equivalent quantity + +1K + +2035 + +100K + +2040 + +100K + +What happens if the deal dissolves? + +Can be used to make AI progress + +| **Human-equivalent quantity**| **** +---|---|--- +**Type of Robot**| **2030**| **2035**| **2040**| **What happens if the deal dissolves?** +**SEZ**| 0| 1B| 70B| Self-destruct or bombed +**Civilian**| 100K| 10M| 1B| Self-destruct or backdoored or not useful for building fabs +**Exempt military robots**| 1K| 100K| 100K| Can be used to make AI progress + +The policy for robots is: + + 1. High and growing cap for robots produced in destroyable SEZs, so almost all robots are produced there. + + 2. A mandate for civilian robots to have sufficient offline licensing + self destruct features so that they stop working if the deal breaks down, or robust edge verification such that they can’t be used to do any part of the work in rebuilding a compute supply chain (e.g., a self-driving car with baked weights). + + 3. Cap of 100K total exempt military robots for US and China, without offline licensing. + + + + +### Post-dissolution takeoff length + +A big question is whether or not a rapid intelligence explosion could occur on the stock of compute produced within the deal, or whether compute production is required. + +There are two major sources of progress that can be pursued in parallel, but might have tradeoffs (e.g., if you use your compute to help run robots or tell humans what to do to produce new compute). + +**First source of progress: Produce new compute.** The post-deal-decline AI project could attempt to use robots and/or humans to rebuild a semiconductor supply chain and increase their compute. The difficulty of doing so is probably further increased by potentially aggressive sabotage in a post-deal context against such efforts. That being said, around 100K exempt military robots, and maybe some small fraction of civilian robots e.g., 0.1%, or around 1M, can be stripped for useful parts post-self-destruct and/or have their security measures circumvented. Let's assume there’s a fleet of around 1M human-equivalent robots that can work on this. In 2026 the semiconductor supply chain has around 2M people working, and over the last 5 years they have built around 100M H100e/yr of total production capacity at 2026-technology levels (chip design and manufacturing process) So we have a rough 10M worker-years to 100M 2026-tech H100e/yr capacity mapping. Assuming the 2026-tech improves by around 1 OOM by 2032 and another OOM by 2036 (assuming total compute in the world improves roughly 50% due to technology improvements and 50% due to production scale), this implies the mapping will be 10M worker-years to 1B H100e/yr with 2032-tech and 10B with 2036-tech. So the 1M robots would take 1 year to build 100M H100e in 2032 and 1 month in 2036. We expect there to be serial bottlenecks that somewhat push these purely parallel estimates up, as well as potential startup problems and sabotage problems. + +**Second source of progress: Use existing compute for R &D.** The post-deal-dissolution AI project has transparent algorithms from the deal, cold-storage weights, and any consumer hardware that they are able to repurpose. Assuming they are the US and they use their cold-storage GPUs or they can get ~10% of world consumer hardware, then they have ~25x less compute relative to the default leading project in Jan 2030 (~5M H100e in US stockpile vs. ~130M in 2030 by default); so -1.3 OOMs. This 25x compute disadvantage means that at the same capability level they should go ~10x slower all else equal ([assuming 10x less compute means ~5.5x slower](https://www.lesswrong.com/posts/7jcPg79p3kD5ir3CL/how-much-slower-does-takeoff-go-with-10-less-compute)). + +However, all else is not equal: at least if the dissolution happens before the 2040 handoff in the Plan A scenario, the post-dissolution project has less efficient software/algorithms than the default project when they were at the same capability level. This is because in Plan A we deliberately make progress via a much higher ratio of training compute to software than the default. Because the software is worse, training efficiency is lower. On the other hand, there’s more remaining promising software ideas out there, so a given amount of [research effort](https://www.aifuturesmodel.com/#section-researcheffortandairduplift) goes further. + +Meanwhile, the transparent algorithms from the deal can be split into ones that have been implemented, and ones that have been foregone to pay a [training compute safety tax](/supplements/comparing-possible-plans#training-compute-safety-tax-definition), which we recommend paying roughly 1 OOM of by TED-AI. The project would need to actually train some using these foregone algorithms, but they could probably quickly capture lots of the gains. + +**Putting everything together.** If dissolution happens after pausing at TED-AI, let’s say in 2037: The TED-AI to ASI time under the default (no slowdown) Plan A trajectory is ~1.5 months. Naively applying the 10x slowdown due to lower training compute would give ~15 months. Taking into account the post-dissolution project having lower software results in a takeoff to ASI of [~2.75 years](https://ai-rates-calculator-git-fixed-training-compute-lesswrong.vercel.app/covert?k.enable_deal_breakdown=true&k.deal_breakdown_years_after_pause_end=7&k.post_breakdown_compute_growth_rate=0) (and setting there to be no post-dissolution compute growth, which further slows things relative to the default trajectory). Enabling 1 OOM/year post-dissolution compute growth results in a takeoff of [~1 year](https://www.ai-2040.com/supplements/takeoff-supplement?k.enable_deal_breakdown=true&k.deal_breakdown_years_after_pause_end=7&k.breakdown_compute_reduction=0.07&k.total_h100e_frac=0.008194244310987836). + +# Compute destruction slows post-dissolution takeoff from ~1 week to ~1 year + +Within the deal (identical in both worlds)Dissolution — compute destroyedDissolution — compute not destroyed + +1x10x100x1,000x10,000xAutomated coder (3x)Superhuman AI Researcher (20x)Top-Expert-Dominating AI (400x)ASI (1,000x)Deal dissolvesASI ~1 week after dissolutionASI ~1 yearafter dissolution20292030203120322033203420352036203720382039YearAI R&D uplift (vs. 2026) + +Trajectories: [compute destroyed](https://www.ai-2040.com/supplements/takeoff-supplement?k.enable_deal_breakdown=true&k.deal_breakdown_years_after_pause_end=7&k.breakdown_compute_reduction=0.07&k.total_h100e_frac=0.008194244310987836), [no compute destroyed](https://www.ai-2040.com/supplements/takeoff-supplement?k.enable_deal_breakdown=true&k.deal_breakdown_years_after_pause_end=7&k.breakdown_compute_mode=world_at_breakdown&k.breakdown_compute_reduction=0.2&k.post_breakdown_compute_growth_rate=0.1). Compute destroyed parameters: leading company has ~7% of the pre-deal frontier (via hardened stockpiles), rebuilt at 1 OOM/yr. Note that uplift numbers don't take into account Plan A capability shaping, which reduces AI R&D uplift at a fixed capability level. + +Now let’s consider dissolution happening in 2032, a bit below SAR. SAR to TED-AI takes ~0.3 years by default. Naively applying the 10x slowdown due to lower training compute would give a bit over 3 years. Accounting for lower software and let’s say 0.5 OOMs/year post-dissolution compute growth [results in](https://ai-rates-calculator-git-fixed-training-compute-lesswrong.vercel.app/covert?k.enable_deal_breakdown=true&k.deal_breakdown_years_after_pause_end=2&k.post_breakdown_compute_growth_rate=0.5) achieving TED-AI after about 2 years. + +## Acknowledgments + +Thanks to Dillon Nguyen for fact-checking Claude’s base rate analysis and helping create figures. Thanks to Thomas Larsen and Romeo Dean drafting the sections on post-dissolution compute destruction. + +# Appendix: Claude’s analysis of base rates + +Claude’s full base rate analysis is attached below for reasoning transparency. Dates and events were hand-verified. The text has been very mildly updated based on the issues found. + +## Claude’s Analysis + +This tab holds the overall reasoning behind the outside-view, base-rate estimate of how likely the Plan A US–China slowdown deal is to dissolve or become impaired, year by year from its formation (~2029) through handoff (~2040). The treaty-by-treaty dataset and the per-treaty reasoning live in the ["Base rates" tab of the companion Deal](https://docs.google.com/spreadsheets/d/15i_TFnC8bTyRT2xOsok3ql0IZ2qesQg4V7LkoWYO2L0/edit?gid=1499542947#gid=1499542947) [Decline](https://docs.google.com/spreadsheets/d/15i_TFnC8bTyRT2xOsok3ql0IZ2qesQg4V7LkoWYO2L0/edit?gid=1499542947#gid=1499542947) [spreadsheet](https://docs.google.com/spreadsheets/d/15i_TFnC8bTyRT2xOsok3ql0IZ2qesQg4V7LkoWYO2L0/edit?gid=1499542947#gid=1499542947); this tab explains the method, the reference-class choices, the central finding, and how I adjusted from raw base rates to a Plan-A-specific year-by-year hazard. Drafted by Claude (Opus 4.8), 2026-06-29. Historical dates and fates were web-verified, but the synthesis into year-by-year numbers is a first-pass draft and a judgment call. + +## Method and choice of reference classes + +The approach is the standard "outside view": rather than reasoning from Plan A's specifics, I assembled ~31 past international agreements that resemble it, recorded how long each lasted before it was first impaired and before it dissolved, and converted the pooled experience into annual hazard rates. I use the same two outcome categories as the supplement—DISSOLUTION (formally terminated, withdrawn-from, or collapsed) and IMPAIRMENT (substantially less effective without formal dissolution: non-enforcement, tolerated violations, suspension, or weakening revision)—and I track three things separately that analysts often blur: time-to-first-impairment, time-to-dissolution, and the failure mode (who broke it and why), because Plan A's design bears very differently on each. I assign each agreement a numeric applicability weight (0–3: 3 = rival great-power capability deal, the closest analogue; ~1.5–2 = informative analogue; ~1 = broad multilateral convention; ~0.25–0.5 = low-defection-incentive or toothless), then compute the base rate as a weighted hazard—Σ(weight × events) / Σ(weight × years-at-risk)—separately for dissolution and impairment. The weights, the per-treaty inputs, and the live formulas all sit in the spreadsheet (Parts 1–2), so the aggregation is fully transparent and editable: change a weight and the base rate recomputes. A pooled hazard is still crude—it treats 1922 naval cheating and a 2023 treaty suspension as draws from one urn—so I report it alongside an unweighted version, a closest-analogues-only version, and—most importantly—a version that lets the hazard vary by PHASE (years since the deal started, Part 2b), then adjust for Plan A's specifics. + +I grouped the agreements into three tiers by similarity to Plan A. MOST APPLICABLE: rival great-power deals that cap a decisive military/strategic technology under strong incentives to defect—the interwar naval treaties (Washington/London), the bilateral US–USSR/Russia strategic arms-control treaties (ABM, SALT, INF, START, New START, Open Skies), and the two "rivals'-deal" nonproliferation bargains (JCPOA, the US–DPRK Agreed Framework)—plus the IR alliance-termination literature, the closest thing to a quantitative base rate for how long deals between rivals last (mean alliance life ~9–10 years pre-1945; ~1/3 violated when tested). MODERATELY APPLICABLE: broad multilateral nonproliferation/verification conventions (NPT, CWC, BWC, CTBT), useful mainly for what they reveal about verification and as the optimistic bound; and the dual-use export-control regimes (Wassenaar/MTCR/NSG), the closest analogue to controlling AI chips and chronically leaky by design. LEAST APPLICABLE but informative: durable low-incentive conventions (Outer Space, Tlatelolco, PTBT) and the toothless climate accords (Kyoto, Paris), which bracket the extremes—near-permanent when nobody wants to defect, near-useless when nothing is binding. The single closest historical analogue is the Washington/London naval system: peer rivals capping a decisive military technology, both tempted to defect, with self-declaration verification. It ran ~13–14 years before Japan walked out, and was cheated on throughout. + +## Central finding: verification delays cheating, but politics kills the deal + +Two patterns dominate the data and drive the whole estimate. + +First, verification strength is the cleanest predictor of how long a deal lasts unimpaired-by-cheating. Strongly verified regimes (INF, START I, New START, Open Skies, NPT/IAEA, CWC/OPCW) went roughly 8–35 years before significant impairment, versus 0–4 years for weakly verified ones—a real but less clean separation than it first appears, because dating impairment to violation onset rather than detection shortens the verified deals' records (INF's violative testing began ~2008, six years before the US formal charge; Open Skies restrictions began in 2010). Verification detected cheating eventually, but with meaningful latency, and detection did not by itself reverse the violation. The one clean reversal in the dataset is the Montreal Protocol, where cheap atmospheric monitoring caught illicit CFC-11 production within years and enforcement restored compliance—the deal never became substantially less effective and is coded as unimpaired. Weakly or un-verified regimes (the naval treaties, the Anglo-German agreement, OPEC quotas, the whaling moratorium, and above all the BWC—which had no verification and was violated on a massive scale by the USSR for its entire early life) were impaired almost immediately, within 0–4 years. This is the strongest structural reason for optimism about Plan A: its verification regime is unusually strong by historical standards and is explicitly designed to make cheating detectable and auto-punishable. + +Second, and crucially: strong verification does essentially nothing to prevent death by political withdrawal. The ABM Treaty (30 years), INF (31 years), Open Skies (~19 years), and the JCPOA (2.3 years) all died because a government decided to exit, not because verification failed. The JCPOA is the sharpest warning—the best verification regime ever negotiated was discarded in ~2 years by a change of US administration. So verification protects against the impairment-by-cheating failure mode, but offers little protection against the dissolution-by-regime-change / change-of-heart / geopolitical-shock failure mode, which the alliance literature (Leeds & Savun 2007) identifies as the dominant driver of premature termination. This is why I keep the dissolution hazard meaningfully above zero even though Plan A's verification is excellent. + +## How Plan A differs from the base rate (the adjustments) + +Pushing the hazard UP relative to a generic high-salience treaty: (1) It is a rival deal over the single most decisive technology imaginable, so it belongs to the harsh ~7–12%/yr naval/rival/alliance class, not the benign ~0%/yr convention class. (2) It is an executive agreement, not a ratified treaty—the scenario says so explicitly, citing New START (2010) as the last significant ratified US treaty. Every deal in the dataset that one administration could exit unilaterally (ABM, INF, Open Skies, JCPOA) was exitable precisely because exit didn't require undoing a ratified commitment. (3) The stakes and the rate of technological change are both unprecedented, which compresses the "first major shock" that historically kills these deals—and the prize for defecting grows as capabilities approach handoff. (4) US elections in 2032 and 2036 and the Chinese leadership succession around 2033 (Xi turns 80) are scheduled regime-change risks, exactly the events Leeds & Savun find spike the termination hazard. (5) The scenario stipulates unusually competent US execution. + +Pushing the hazard DOWN: (1) Unusually strong verification plus automatic enforcement (offline licensing, cryptographic shutoff of compute on exit) pushes the impairment-by-cheating component toward the INF/New-START low end (~2–4%/yr) rather than the naval-treaty high end. (2) An unusually strong SHARED interest: unlike a pure capability race, both sides genuinely fear loss of control and extinction. This is closer to MAD-stabilized nuclear restraint (which held for decades) than to a zero-sum arms race, and it is the best structural reason Plan A might beat its reference class—though it has no clean historical precedent. + +Net: I anchor on the rival-deal class (~7%/yr dissolution) but discount the dissolution rate somewhat (to ~4–7%/yr depending on the year) for the shared-survival interest and competent execution, and hold the impairment-by-cheating component low (~2.5–5%/yr) because of strong verification—while keeping impairment nonzero for non-enforcement, misexecution, and weakening revision, none of which verification prevents. + +## Headline base rates and the year-by-year shape + +The pooled weighted base rate (all ages together) is 3.4%/yr for dissolution and 7.9%/yr for impairment (unweighted 1.8% / 4.8%; closest-analogues-only subset 6.6% / 8.3%). But the more useful object is the base rate BY PHASE (Part 2b), because the historical hazard depends strongly on how long a deal has already survived. DISSOLUTION is hump-shaped, not front-loaded: 0% in years 0–2 (no deal in the dataset formally dissolved within two years), rising to a peak of ~5–6%/yr at years 5–10, then declining to ~2%/yr for deals that survive past 20 years—rival deals take a few years to die. IMPAIRMENT is strongly front-loaded: ~15.5%/yr in years 0–2, falling to ~9%/yr (years 2–5) and ~6%/yr (years 5–10), with a late bump (~6%/yr at years 10–20, from the New START suspension and Syria/CWC). Across the full dataset impairment is more common than formal dissolution—many regimes get hollowed out rather than killed. Feeding the phase curve into the deal timeline, the pure base rate implies ~79% cumulative chance of some breakdown by 2040 (~39% dissolution, ~65% impairment). These are the "Base-rate" columns in Part 3; my "Adjusted" columns overlay the Plan-A-specific reasoning below, yielding ~62% cumulative breakdown (~43% dissolution, ~33% impairment). + +My adjusted shape departs from the base rate in two deliberate ways. First, I lift EARLY dissolution above the base rate's 0%: the historical 0% in years 0–2 is partly survivorship (deals that collapse during negotiation never enter an "entered-into-force" dataset), whereas Plan A's setup is genuinely fragile, so I use ~5%/yr in 2029–30. Second, I cut impairment far below the base rate throughout (≈2.5–5% vs the base's 15.5%→3%), because Plan A's strong verification specifically suppresses the cheating-driven and born-impaired failures that dominate the historical impairment record. On top of that I add discrete dissolution spikes at the scheduled regime-change events (US elections 2032/2036, Chinese succession ~2033), and I keep the late-2030s hazard from decaying to near-zero because the prize for defecting grows as handoff and aligned superintelligence approach. + +## Comparison to the inside view, caveats, and where the detail lives + +My adjusted estimate (~62% cumulative breakdown, ~43% cumulative dissolution by 2040) is more pessimistic than the inside-view ["Overall deal decline probability view" tab](https://docs.google.com/spreadsheets/d/15i_TFnC8bTyRT2xOsok3ql0IZ2qesQg4V7LkoWYO2L0/edit?gid=215047343#gid=215047343) (Eli's ~48% cumulative), and the pure unadjusted base rate is harsher still (~79% breakdown). I think that gap is real and informative: base rates for rival deals over decisive technologies are harsh, and the inside view may be under-weighting how reliably this reference class fails. On the components: the weighted dissolution base rate (3.4%/yr) sits just above the "0.5–3%/yr" generic-high-salience-treaty anchor in the spreadsheet note, while the impairment base rate (7.9%/yr) is far higher—because the most applicable reference class is rival capability deals (naval treaties, JCPOA, Agreed Framework, alliances), not high-salience treaties in general, and because the dominant historical killer—political withdrawal—is not neutralized by Plan A's verification, however strong. + +Main caveats. (1) Pooled hazard rates hide enormous heterogeneity; the "right" number depends heavily on how much weight you put on the shared-survival-interest stabilizer, which has no clean precedent. (2) N is small and the rival-deal subclass is dominated by a handful of cases. (3) The outside and inside views should NOT simply be averaged—they partly double-count (regime change appears both as the base-rate driver here and as an explicit sub-cause in the decomposition tab). (4) I have treated benign endings (planned expiry, supersession by a stronger successor—SALT I, START I, SORT) as non-failures; counting them as dissolutions would push the dissolution rate higher. (5) Two structural features of Plan A genuinely lack good historical analogues—automatic compute-destruction-on-exit, and a genuinely shared existential interest—and both could justify materially lower numbers than the raw base rate if you believe they bind. + +Where the detail lives: the 31-agreement dataset (per-treaty reasoning in column O), the live weighted-hazard computation (Part 2), the IR-literature anchors, and the full year-by-year base-rate-vs-adjusted table (Part 3) are all in the "Base rates" tab of the Deal Decline spreadsheet. Each treaty row records signed/in-force years, verification strength, a numeric applicability weight (column G, the input to the base-rate formula), exposure years, dissolved/impaired flags, years-lasted-unimpaired, years-to-dissolution, the dominant failure mode, reasoning, and sources. + +## Smooth fit (parametric survival) + +The phase/step base rate (Part 2b) is noisy with only 31 agreements, so I also fit smooth parametric survival curves to the treaty-level data by weighted, right-censored maximum likelihood (Part 2c), comparing exponential, Weibull, log-logistic, and lognormal by AIC. LOGNORMAL fits best for both outcomes (dissolution AIC 301.8; impairment AIC 275.4) and reproduces the shapes the step rates show. The fitted dissolution hazard rises from ~0.2%/yr at age 0.5 to a ~4.4%/yr plateau by years 8–11—a hump, confirming dissolution is not front-loaded—while the fitted impairment hazard decays smoothly from ~17%/yr at age 0.5 to ~5.6%/yr by age 11.5 (front-loaded). Weibull is a reasonable simpler alternative (impairment k=0.71, a monotonically decreasing hazard; dissolution k=1.1, near-flat) but it cannot produce a hump, so it underfits dissolution; log-logistic is in between. The fitted per-year hazards live in Part 3 columns L–M, next to the step base rate; over 2029–2040 they imply ~34% cumulative dissolution and ~69% cumulative impairment—close to the step-function figures (39% / 65%), which is reassuring. Caveat: with N=31 and heavy weighting toward a handful of rival deals, the fitted curves are best read as smoothed summaries of a small sample, not precise estimates—the parameters in Part 2c are editable if you want to stress-test them. + +1. + +The deal could also get better, but we’re not including that here as our best guess is that the deal will in expectation get worse. This is definitionally true for deal dissolution, but we’re not confident for deal impairment; if we were more optimistic, we could remove the deal impairment term or explicitly model the deal getting better. + +2. + +Though these failures could be mostly downstream of deal decline. + +3. + +Perhaps a bloc of countries besides the US and China could also suffice. + +4. + +Or a CCP power transition. Or, less likely, a civil war. + +5. + +After 16 years, the hazard rate decays to 95% of its previous value each year. + +6. + +Technically, we mean not destroyable absent extreme escalation (e.g. nuclear war), which is unlikely for the normal reason of MAD. + +7. + +Another downside is that if the weights are preserved, it makes deal dissolution less unattractive. + +8. + +This is under the assumption that for each 10x increase in compute, the intelligence explosion occurs 5x faster. This is our median guess, which we discuss more in our covert projects supplement here. diff --git a/sources/ai2040-pages/econ-explorer.md b/sources/ai2040-pages/econ-explorer.md new file mode 100644 index 0000000..cf84703 --- /dev/null +++ b/sources/ai2040-pages/econ-explorer.md @@ -0,0 +1,33 @@ +--- +url: https://ai-2040.com/supplements/econ-explorer +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) diff --git a/sources/ai2040-pages/economics-of-plan-a.md b/sources/ai2040-pages/economics-of-plan-a.md new file mode 100644 index 0000000..368e60b --- /dev/null +++ b/sources/ai2040-pages/economics-of-plan-a.md @@ -0,0 +1,633 @@ +--- +url: https://ai-2040.com/supplements/economics-of-plan-a +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Economics of Plan A + +### Romeo Dean, Thomas Larsen + +We project that AIs will, at some point in the next 15 years, be capable of automating all human labor: both physical and cognitive. We believe that this will have extraordinary economic impacts, causing output to grow explosively at a historically unprecedented rate (barring an AI-caused catastrophe). Most of our view comes from simple arguments about the future of AI. However, over the course of our thinking, we also found it useful to make a detailed economic model. This document is organized into the following sections: + + 1. Core arguments behind our views on explosive growth from AI ([link](/supplements/economics-of-plan-a#core-arguments-for-explosive-growth)) + + 2. Why we disagree with mainstream economists ([link](/supplements/economics-of-plan-a#why-we-disagree-with-mainstream-economists)) + + 3. A simple two-factor explorer that shows the core dynamics ([link](/supplements/economics-of-plan-a#the-simple-explorer)) + + 4. A summary of the outputs from our full economic model in the context of Plan A ([link](/supplements/economics-of-plan-a#the-full-economic-model)) + + + + +The core disagreement between us and most economists is our different expectations for AI capabilities. We believe that AIs will be able to fully automate all economically relevant work (in a median of ~6 years), while most economists believe that AIs will behave more like normal technology. We’ve already made [many arguments about AGI timelines elsewhere](https://www.aifuturesmodel.com/), so this post centrally argues for why, _conditional on AI that can automate all economically relevant tasks_ , we expect unprecedented explosive economic growth. + +## Core arguments for explosive growth + +Right now, the size of the economy is closely tied to human labor, so if the population were to grow massively, this would cause nearly proportional growth in total output. Once AIs can perfectly substitute for human labor, increasing the number of AIs would have the same effect as increasing human population. At least if you exclude tasks that are intrinsically human, but we don’t expect these to be a large share of the economically relevant tasks. For example, human-preference driven tasks the AI can’t substitute for like ‘human babysitting’ should not preclude an AI and robot-only economy from growing in a closed loop (bounded only by [bottlenecks that we expect to not bite strongly](/supplements/economics-of-plan-a#why-we-disagree-with-mainstream-economists)). + +In fact, this undersells the likely effects because in reality, AIs won’t just be increasing in population, but also in qualitative ability leading to outputs that wouldn’t have been feasible to produce with humans. This is because, shortly after being able to substitute for human labor, AIs will become qualitatively superhuman. However, improvements to intelligence are harder to model quantitatively than increases in population, **so we’ll settle for establishing a population-based lower bound by ignoring qualitative improvements.** + +**So how fast will the “population” of AIs and robots be growing at the point of full automation?** We need to do separate analyses here for physical and cognitive labor, because the underlying dynamics driving the growth rates are quite different. + + * **Cognitive Labor: a doubling time of roughly 50 days.** The total cognitive labor from AI depends on total inference compute and the compute efficiency of the AIs. The simplest way of making a baseline guess here is to look at historical rates for each of these. + + * **Total AI inference compute: 3.3x/year, 7 month doubling time.** [Epoch estimates](https://epoch.ai/data-insights/ai-chip-production) that the total compute in the world has been doubling every 7 months. In the future, we think there are reasons for this to slow down (it’s already requiring [2-3% of](https://finance.yahoo.com/news/big-techs-ai-spending-spree-142111465.html?guccounter=1) [GDP investment in the US](https://finance.yahoo.com/news/big-techs-ai-spending-spree-142111465.html?guccounter=1), so doubling a few more times seems hard), but also reasons for it to speed up at the point of full automation (e.g., future AIs automating chip R&D, and robots automating the AI chip supply chain). + + * **Compute efficiency on achieved capabilities: 30x/year, 2.5 month doubling time.** [Epoch estimates](https://epoch.ai/data-insights/llm-inference-price-trends) that the price of a given level of AI capability decreases by an average of 40x/yr (which is equivalent to halving every 2.2 months). Assuming [hardware price performance](https://epoch.ai/trends#hardware) has been improving 1.4x/yr and margins have been flat, this gives a trend of 30x/year improvement in compute efficiency. When the cost of AI inference halves, this is at least as economically beneficial as the population of AIs doubling, because we can run twice as many AIs. + + * **Putting these together, this means that the population of AIs today at a fixed capability level could increase at a rate of 100x/year, which is a 55 day doubling time.** + + * **Physical Labor: a doubling time of a few days to a year.** For physical labor, we can reason directly about the number of robots in the world. At the moment, [Epoch finds](https://epoch.ai/blog/how-fast-could-robot-production-scale-up) that the number of robots is dominated by wheeled robots, drones and robotic arms, and these are not scaling very quickly. ![](/epoch-robots.png) + + + + +Right now, these robots are not capable of performing all tasks in the robot manufacturing process (including mining raw materials, transportation, and every step of the supply chain). However, once robots reach this capability threshold then there’s a clear reason to expect exponential growth in the number of robots: the rate of new robot construction will be approximately proportional to the number of robots in existence (barring raw material bottlenecks that we address below). But how fast of an exponential? We walk through our reasoning below, and also have a review of the doubling-times literature in [Appendix A](/supplements/economics-of-plan-a#appendix-a-robot-doubling-times). + +**Why wouldn’t robots plateau like other products?** Right now, robots and AI compute are both doubling faster than once per year. On its own this is not strong evidence of anything: many new products grow this fast at first, and then plateau once they saturate their addressable market. We expect robots to be different, because once AI and robots can perform every task in their own supply chain, the two things that normally end fast growth stop applying: + + 1. **No demand ceiling.** Robots and AIs are productive assets, not consumer goods. As long as an extra robot produces more value than it costs to build, there is an incentive to keep building them, past the point where the robot population exceeds the human population and far beyond it. + + 2. **No human bottleneck on supply.** Today’s production growth is partly borrowed from human labor (factory workers, miners, truck drivers, engineers). At full automation the robot economy doesn’t need humans to buy anything, build anything, or deliver ingredients to its factories, so the size of the human population stops being a constraint on how big it can get. + + + + +The ceiling that does eventually bind is physical resources: surface area for solar panels, minerals in the earth’s crust to mine. But by our estimates those bottlenecks bite very late, after many orders of magnitude of growth (see the [bottlenecks section below](/supplements/economics-of-plan-a#bottlenecks-and-other-arguments)). + +That leaves the question of the growth rate in between. A reasonable worry is that today’s doubling rates are inflated by human help, and that growth slows once the robot economy has to do essentially everything itself. We don’t think this changes the bottom line: analyses of what a fully automated economy could achieve with current technology and no further improvement, e.g. [Binder’s von Neumann growth rate calculation](https://defensesindepth.bio/ai-industrial-takeoff-part-1-maximum-growth-rates-with-current-technology/) from US input-output tables, capital stocks, and depreciation data, land at around one doubling per year, conservatively. An intuition pump: a modern car factory produces roughly [its own weight in cars every year](/supplements/economics-of-plan-a#appendix-a-robot-doubling-times). And a year is likely an underestimate of the long-run pace, because superhuman AI R&D should keep improving robot designs, and biology shows much faster replication is physically possible (e.g., animals like rabbits can double their population in 6 weeks, bacteria in hours). This is why we quote doubling times between a few days and a year, with the fast end becoming more plausible over time. + +Overall, our argument can be summarized as: + + 1. AI and robots will be capable of fully automating all economically relevant tasks. + + 2. Therefore, the AI-robot economy will be capable of autonomously self-replicating (building more efficient AIs, building more robots, building more GPUs, etc.). + + 3. The replication rate will be fast: around one doubling per year even with current technology and no human help, and rising over time as AI R&D improves robot designs. + + 4. Therefore, the AI-robot economy will undergo rapid exponential growth, with a doubling time of a year at the slow end and plausibly months or days, until physical resource limits bite (very late, on our estimates). + + + + +Even if the true growth rate is much slower than this baseline predicts in the future, the effective AI population seems very likely to rapidly outgrow the total number of humans. Once AIs dominate the number of humans, if their population is (conservatively) increasing at 2x per year, this would involve ~100%/yr growth, whereas [current growth](https://data.worldbank.org/indicator/NY.GDP.MKTP.KD.ZG) is ~3%/yr. + +## Why we disagree with mainstream economists + +Most mainstream economists are skeptical of explosive economic growth from AI. Even conditional on a “rapid AI progress scenario”, [economists predicted](https://static1.squarespace.com/static/635693acf15a3e2a14a56a4a/t/69cbb9d509ada447b6d9013f/1774959061185/forecasting-the-economic-effects-of-ai.pdf) an average of 3.5%/yr GDP growth between 2030 and 2050. + +One core disagreement between us and mainstream economists tends to be about how capable AIs will be. We believe that AI's ability to do tasks in the economy will look something like the red curve below absent strong regulations. This reflects AI dominating humans at all tasks and achieving 100% automation around 2031 (top expert dominating milestone), which is close to the authors' median expectations. We find that economists tend to have much lower expectations for automation in the coming decades. + +Share automatable0%25%50%75%100%2025202620272028202920302031203220332034203520362037203820392040 + +AI 2040 Default World + +AI 2040 Plan A + +Acemoglu: ~5% “cost-effectively automatable” [(2024)](https://shapingwork.mit.edu/wp-content/uploads/2024/05/Acemoglu_Macroeconomics-of-AI_May-2024.pdf) + +Goldman Sachs: ~25% “substitutable” [(2023)](https://www.gspublishing.com/content/research/en/reports/2023/03/27/d64e052b-0f6e-45d7-967b-d7be35fabd16.pdf) + +_Note: Acemoglu and Goldman Sachs projections are not necessarily directly comparable, as their definitions for "cost-effectively automatable" and "substitutable" may diverge somewhat from our automatable definition._ + +Automatable tasks: AI 2040 scenarios vs prior forecasts \- ai-2040.com + +In Plan A, strong regulations (coordinated internationally) are sufficient to slow this down, but still lead to ~99% automation in the late 2030s, as regulation only preserves critical AI oversight jobs (where humans are still significantly uplifted by AI). We believe that AI's capabilities will be so high that lower automation levels (e.g., due to wanting to preserve more human jobs) will be extremely uncompetitive (and thus be quickly competed away), with preferences for human-to-human interactions being (1) supplemented by more leisure, and (2) insufficient to push human wages high enough relative to the citizen's dividend to attract much labor. + +### Acemoglu (2024): too few tasks automated by AI + +[Acemoglu (2024)](https://economics.mit.edu/sites/default/files/2024-04/The%20Simple%20Macroeconomics%20of%20AI.pdf) (also [critiqued here](https://x.com/AndreyFradkin/status/1840837475062988911) by Leopold Aschenbrenner) assumes only ~5% of tasks in the next 10 years will be cost-effectively automatable by AI. Acemoglu’s overall estimate was that AI would increase GDP by 0.9% over the next 10 years (which is approximately a 0.08% increase in GDP annually). We think Acemoglu’s view from 2024 is already close to falsified based on current AI investment and AI revenue numbers. + +Early evidence against Acemoglu (2024) + + * AI investment contributed around 1% to GDP growth in the first half of 2025 alone ([JP Morgan](https://am.jpmorgan.com/us/en/asset-management/adv/insights/market-insights/market-updates/on-the-minds-of-investors/is-ai-already-driving-us-growth/), [EY](https://www.ey.com/en_us/insights/ai/ai-powered-growth)), whereas Acemoglu predicted .9% growth over the course of 10 years. + + * Even at the scale of AI revenues, over the course of the last six months, Anthropic’s revenue run rate alone has [increased by $38B](https://epoch.ai/data/ai-companies?view=graph&tab=revenue) from $9B to $47B. We think this will plausibly beat the rate of increased growth per year (0.08% of US GDP) of ~$30B that Acemoglu predicted for the entire AI industry. Note: Anthropic's contribution to GDP is difficult to disentangle, but should count employee compensation + operating surplus + depreciation costs, which we think is on track to be above $30B this year. + + + + +The only ways Acemoglu’s predictions will hold up are if (1) AI revenue and investment growth essentially stops, or (2) AI investment and revenues are not counterfactual (the spending and revenue would’ve happened elsewhere anyway). Both seem unlikely. + +### Jones & Tonetti (2026): slow AI automation of R&D + +[Jones & Tonetti (2026)](https://web.stanford.edu/~chadj/JonesTonetti_Automation.pdf), _Past Automation and Future A.I.: How Weak Links Tame the Growth Explosion_ features an elegant growth model combining capital and labor. They grant that capital may eventually automate every task, but their argument is that because tasks are strong complements (they assume an elasticity of substitution σ=0.2\sigma = 0.2σ=0.2), the shrinking set of tasks machines cannot yet do becomes a 'weak link' whose cost share rises ([Baumol's cost disease](https://en.wikipedia.org/wiki/Baumol_effect)) and strongly bottlenecks growth. Our biggest disagreement with this model is that we think the R&D automation loop is very conservative, with the model assuming: + + 1. **Fixed R &D investment and research productivity.** In the Jones model R&D works through "ideas accumulation" with capital productivity ψki\psi_{ki}ψki​ changing with a fixed "idea elasticity of capital" θk\theta_kθk​, a fixed "research productivity" qˉ\bar qqˉ​, and a fixed share ιˉR\bar\iota_RιˉR​ of output YYY being invested in R&D (with diminishing returns per [ideas getting harder to find](https://web.stanford.edu/~chadj/IdeaPF.pdf)). + +ψ˙kiψki=θkqˉ ιˉR YQ 1−ϕ,\begin{aligned} \frac{\dot\psi_{ki}}{\psi_{ki}} &= \theta_k \frac{\bar q\,\bar\iota_R\,Y}{Q^{\,1-\phi}},\qquad \\\ \end{aligned}ψki​ψ˙​ki​​​=θk​Q1−ϕqˉ​ιˉR​Y​,​ + +We think this formulation systematically underestimates the effect that AI will have on R&D, because AI will be able to automate R&D labor (particularly AI R&D itself), so ideas and productivity growth should track some notion of total R&D effort accounting for this AI uplift. We think the AI field will have its own feedback loop of R&D automation, leading to higher research productivity and higher reinvestment into R&D, ultimately leading to much higher capital productivity for AI capital. + + 2. **No unique treatment of AI capital.** In the Jones model, there is no unique treatment of AI. AI enters as normal capital and accumulates through the classic investment less depreciation rule K˙=ιˉK Y−δK\dot K = \bar\iota_K\,Y-\delta KK˙=ιˉK​Y−δK with fixed investment share ιˉK\bar\iota_KιˉK​ and depreciation rate δ\deltaδ. We think AI would be worth modelling on its own, because we think its task productivity and investment attractiveness will grow over time, and also have a different responsiveness to more investment and different responsiveness to R&D (e.g., algorithmic progress on the orders of 10x/year instead of Moore's law 1.35x/yr). In other words, we think the "idea elasticity of AI", "research productivity of AI" and "share of output being invested in AI R&D" will all behave very differently to how they behave for other capital, leading to AI's effects being drastically underestimated by being bundled with the capital factor of production. + + +Jones & Tonetti Model Explanation + +Output is a CES aggregate over a continuum of tasks i∈[0,1]i \in [0,1]i∈[0,1] with elasticity of substitution σ=0.2\sigma = 0.2σ=0.2: + +Y=(∫01Yiσ−1σ di)σσ−1,Yi=ψkiKi+ψℓiLi,Y=\left(\int_0^1 Y_i^{\frac{\sigma-1}{\sigma}}\,di\right)^{\frac{\sigma}{\sigma-1}},\qquad Y_i=\psi_{ki}K_i+\psi_{\ell i}L_i,Y=(∫01​Yiσσ−1​​di)σ−1σ​,Yi​=ψki​Ki​+ψℓi​Li​, + +where KiK_iKi​ and LiL_iLi​ are the capital and labor allocated to task iii. Within a task the two are perfect substitutes (whichever is cheaper does it, more on this below). The task productivities of capital ψki\psi_{ki}ψki​ and labor ψℓi\psi_{\ell i}ψℓi​ are given by: + +ψki=Qθkf(i),ψℓi=Qθℓ,f(i)=(1−i)μ1+μ0(1−i)μ+fˉ.\psi_{ki}=Q^{\theta_k}f(i),\qquad \psi_{\ell i}=Q^{\theta_\ell},\qquad f(i)=\frac{(1-i)^{\mu}}{1+\mu_0(1-i)^{\mu}}+\bar f.ψki​=Qθk​f(i),ψℓi​=Qθℓ​,f(i)=1+μ0​(1−i)μ(1−i)μ​+fˉ​. + +QQQ is the economy's stock of ideas (the technology level); and θk\theta_kθk​, θℓ\theta_\ellθℓ​ are the _idea elasticities_ , governing how fast each factor's productivity rises as ideas accumulate, with θk>θℓ\theta_k> \theta_\ellθk​>θℓ​ as capital improves faster from ideas than labor. In their 'Moore's Law Everywhere' calibration θk≈11.2\theta_k \approx 11.2θk​≈11.2 and θℓ≈0.5\theta_\ell \approx 0.5θℓ​≈0.5. The function f(i)f(i)f(i) is capital's comparative advantage at task iii, declining in iii (tasks are ordered so capital is relatively best at low iii), μ\muμ and μ0\mu_0μ0​ are shape parameters fit to historical automation, and the floor fˉ=f(1)\bar f = f(1)fˉ​=f(1) controls the long-run end state: with fˉ≥0\bar f \ge 0fˉ​≥0 automation runs essentially to completion, while fˉ<0\bar f < 0fˉ​<0 leaves a positive set of tasks permanently un-automated (worked out below). + +**Automation.** Let rrr be the rental rate of capital and www the labor wage. A task is automated when capital is cheaper than labor per productivity-adjusted-unit, i.e., when rψkiθℓ\theta_k> \theta_\ellθk​>θℓ​, or (2) capital becoming cheaper over time (by becoming more abundant through capital accumulation) relative to wages. + +Cost-minimization assigns each task to the cheaper factor, so capital does tasks [0,β][0,\beta][0,β] and labor does [β,1][\beta,1][β,1]. We can collapse output to an ordinary two-factor CES (because each task is done by just one factor, so we only need the totals KKK and LLL each weighted by their total productivity across the tasks they do, BBB and AAA): + +Y=[(BK)σ−1σ+(AL)σ−1σ]σσ−1B=Qθk[∫0βf(i)σ−1 di]1σ−1A=Qθℓ (1−β)1σ−1\begin{aligned} Y &= \Big[(BK)^{\frac{\sigma-1}{\sigma}}+(AL)^{\frac{\sigma-1}{\sigma}}\Big]^{\frac{\sigma}{\sigma-1}} \\\ B &= Q^{\theta_k}\left[\int_0^{\beta} f(i)^{\sigma-1}\,di\right]^{\frac{1}{\sigma-1}} \\\ A &= Q^{\theta_\ell}\,(1-\beta)^{\frac{1}{\sigma-1}} \end{aligned}YBA​=[(BK)σσ−1​+(AL)σσ−1​]σ−1σ​=Qθk​[∫0β​f(i)σ−1di]σ−11​=Qθℓ​(1−β)σ−11​​ + +Capital accumulates and labor grows as + +K˙=ιˉK Y−δK,Lt=L0 ent\begin{aligned} \dot K &=\bar\iota_K\,Y-\delta K,\\\ L_t &=L_0\,e^{nt} \end{aligned}K˙Lt​​=ιˉK​Y−δK,=L0​ent​ + +with investment share ιˉK\bar\iota_KιˉK​, depreciation rate δ\deltaδ, and population growth rate nnn. + +**R &D in the model.** New ideas are produced from a fixed share RRR of output being invested, given by: + +Q˙=qˉ RλQϕ,R=ιˉR Y.\begin{aligned} \dot Q &=\bar q\,R^{\lambda}Q^{\phi},\qquad \\\ R &=\bar\iota_R\,Y. \end{aligned}Q˙​R​=qˉ​RλQϕ,=ιˉR​Y.​ + +Where Q˙\dot QQ˙​ is the change in the idea stock, qˉ\bar qqˉ​ is a research-productivity constant, ιˉR\bar\iota_RιˉR​ is the share of output spent on R&D, and λ=1\lambda = 1λ=1 is the elasticity of new ideas to research input, and ϕ=−2\phi = -2ϕ=−2 (<0< 0<0) the returns to the existing stock, i.e., [ideas getting harder to find](https://web.stanford.edu/~chadj/IdeaPF.pdf). + +With defaults λ=1\lambda = 1λ=1, ϕ=−2\phi = -2ϕ=−2 the growth rate is: + +Q˙Q=θkqˉ ιˉR YQ 1−ϕ,\begin{aligned} \frac{\dot Q}{Q} &=\theta_k \frac{\bar q\,\bar\iota_R\,Y}{Q^{\,1-\phi}},\qquad \\\ \end{aligned}QQ˙​​​=θk​Q1−ϕqˉ​ιˉR​Y​,​ + +And the growth rate of capital productivity ψki\psi_{ki}ψki​ is given by: + +ψ˙kiψki=θk Q˙Q=θkqˉ ιˉR YQ 1−ϕ,\begin{aligned} \frac{\dot\psi_{ki}}{\psi_{ki}} &=\theta_k\,\frac{\dot Q}{Q} = \theta_k \frac{\bar q\,\bar\iota_R\,Y}{Q^{\,1-\phi}},\qquad \\\ \end{aligned}ψki​ψ˙​ki​​​=θk​QQ˙​​=θk​Q1−ϕqˉ​ιˉR​Y​,​ + +So R&D is modelled at a fixed productivity qˉ\bar qqˉ​ from a fixed share ιˉR\bar\iota_RιˉR​ of output. The paper sets the combined term qˉ ιˉRλ=0.0001\bar q\,\bar\iota_R^{\lambda}=0.0001qˉ​ιˉRλ​=0.0001 to match idea growth Q˙/Q=1%\dot Q/Q=1\%Q˙​/Q=1% in 2020. + +**100% automation.** The model leads to 100% automation in finite time when fˉ>0\bar f > 0fˉ​>0, with labor fully displaced and its income share falling to zero. Output collapses to a single-factor form Y=BKY = B KY=BK (linear in capital) whose productivity B∝QθkB \propto Q^{\theta_k}B∝Qθk​ keeps climbing as ideas accumulate. + +**When does growth explode?** To pin down the long run behaviour one can look for a steady state where output and ideas are both growing at a constant rate. This either has a finite solution (the loop R&D feedback settles to steady growth) or no solution (where it explodes). There are two sides we can find: + +_Idea side._ Constant idea growth Q˙/Q=qˉ ιˉRλ YλQϕ−1\dot Q/Q = \bar q\,\bar\iota_R^{\lambda}\,Y^{\lambda}Q^{\phi-1}Q˙​/Q=qˉ​ιˉRλ​YλQϕ−1 needs the YYY and QQQ powers to offset, λgY=(ϕ−1)gQ\lambda g_Y = (\phi-1)g_QλgY​=(ϕ−1)gQ​, so ideas growing at a fraction of output (a fraction because ideas get harder to find with ϕ<0\phi < 0ϕ<0): + +gQ=λ1−ϕ gY.g_Q = \frac{\lambda}{1-\phi}\,g_Y.gQ​=1−ϕλ​gY​. + +_Output side._ Output grows at the income-share-weighted growth of its two effective inputs, each input's growth being its productivity (rising with ideas, gB=θkgQg_B = \theta_k g_QgB​=θk​gQ​ and gA=θℓgQg_A = \theta_\ell g_QgA​=θℓ​gQ​) plus its quantity (for capital gK=gYg_K = g_YgK​=gY​ because of fixed saving & depreciation rate, and for labor its just nnn): + +gY=sK∗ (θkgQ+gY)+(1−sK∗) (θℓgQ+n).g_Y = s_K^{*}\,(\theta_k g_Q + g_Y) + (1-s_K^{*})\,(\theta_\ell g_Q + n).gY​=sK∗​(θk​gQ​+gY​)+(1−sK∗​)(θℓ​gQ​+n). + +Substituting the idea side for gQg_QgQ​ and collecting the gYg_YgY​ terms gives: + +gY=Φ gY+n,withΦ=λ1−ϕ(θℓ+sK∗1−sK∗ θk).g_Y = \Phi\,g_Y + n, \qquad \text{with} \qquad \Phi = \frac{\lambda}{1-\phi}\left(\theta_\ell + \frac{s_K^{*}}{1-s_K^{*}}\,\theta_k\right).gY​=ΦgY​+n,withΦ=1−ϕλ​(θℓ​+1−sK∗​sK∗​​θk​). + +So apart from population growth, the model's output growth happens through Φ\PhiΦ which represents the R&D loop's contribution. Each unit of output growth feeds back scaled by Φ\PhiΦ. Summing over timesteps we get a geometric series: gY=n(1+Φ+Φ2+⋯ )=n1−Φg_Y = n(1 + \Phi + \Phi^2 + \cdots) = \frac{n}{1-\Phi}gY​=n(1+Φ+Φ2+⋯)=1−Φn​. + + * If Φ<1\Phi < 1Φ<1 each round shrinks and growth settles at a finite rate proportional to population growth. + + * If Φ≥1\Phi \ge 1Φ≥1 it does not and no steady rate exists, so output runs to infinity in finite time. This is [hyperbolic growth](https://en.wikipedia.org/wiki/Hyperbolic_growth) with output XXX (normalized to 1 today) obeying X˙=gˉ XΦ\dot X = \bar g\,X^{\Phi}X˙=gˉ​XΦ, with gˉ\bar ggˉ​ today's growth rate. So the growth rate X˙/X=gˉ XΦ−1\dot X/X = \bar g\,X^{\Phi-1}X˙/X=gˉ​XΦ−1 itself rises with the level. Separating and integrating from today, ∫1XX−Φ dX=gˉ t\int_1^X X^{-\Phi}\,dX = \bar g\,t∫1X​X−ΦdX=gˉ​t, and solving for output gives X=( 1−(Φ−1) gˉ t )−1/(Φ−1)X = \big(\,1 - (\Phi-1)\,\bar g\,t\,\big)^{-1/(\Phi-1)}X=(1−(Φ−1)gˉ​t)−1/(Φ−1). With negative exponent, as the inside goes →0 \to 0→0 (which happens at 1=(Φ−1) gˉ t1 = (\Phi-1)\,\bar g\,t1=(Φ−1)gˉ​t) output XXX goes →∞ \to \infty→∞. So output going to infinity happens at t∞=1gˉ (Φ−1)t_\infty = \frac{1}{\bar g\,(\Phi-1)}t∞​=gˉ​(Φ−1)1​. For Φ=2\Phi = 2Φ=2 this is simply t∞=1gˉt_\infty = \frac{1}{\bar g}t∞​=gˉ​1​ which for ~3% growth today is ~33 years from now. + + + + +Jones & Tonetti's "Moore's Law Everywhere" calibration has Φ=2.27\Phi = 2.27Φ=2.27 (see box above for what this means) leading output to go infinite around 2060. Our main disagreement is not with this calibration, but with the lack of R&D automation and lack of unique treatment of AI capital (e.g., different in returns to R&D in AI vs. normal capital). + +AI Futures Model + +The **[AI Futures Model](https://www.aifuturesmodel.com)** features a tighter R&D automation loop, with combined AI-human research effort feeding into AI R&D (still with diminishing returns). We think some kind of similar treatment of AI, with modelling that is calibrated to how quickly AI has been improving and responding to increased inputs (human labor, compute, etc.) is needed for a growth model to model future growth in a way that more reasonably accounts for AI. + +[aifuturesmodel.comOpen ↗![AI Futures Model — timelines & takeoff](/aifm-ss.png)Open the AI Futures Model ↗](https://www.aifuturesmodel.com) + +AI Futures Model \- ai-2040.com + +We think it would be exciting future work to modify the Jones & Tonetti model to (A) feature a unique AI factor of production, and/or (B) make the R&D loop more dynamic (and allowing more direct AI uplift on R&D). For the modelling we have done in this scenario, we instead rely on the [AI Futures Model](https://www.aifuturesmodel.com), and what we have added is more of an economics explorer, since it takes the AI automation trajectories as exogenous inputs on paths consistent with our AI Futures Model runs for [**AI 2040 default**](https://ai-rates-calculator-git-fixed-training-compute-lesswrong.vercel.app/p?base=daniel-04-02-26&acth=6.089920916003435&arts=3.4833731503601166&bgm=true&gy=1.103409&mttm=5.281602094603481&tam=1.54170045294956) and **[AI 2040 under Plan A](https://plan-a.aifuturesmodel.com/)**. See the [simple explorer](/supplements/economics-of-plan-a#the-simple-explorer) and [full model](/supplements/economics-of-plan-a#the-full-economic-model) below. + +### Bottlenecks and other arguments + +Beyond the two papers above, there is a more general class of arguments about economic growth not being explosive due to bottlenecks, or historical precedents. Overall we believe that many such bottlenecks will be either weak, routed around due to massive incentives, and/or will bite very late, and that AI will be disanalogous to historical precedents. + +Objection: **Previous technologies haven’t led to explosive growth, and AI won’t be any different.** + +We agree that previous technologies haven’t led to explosive growth (well at least the fast explosive growth we expect from AI, long run economic growth has actually been [better fit by a power law than an exponential](https://coefficientgiving.org/research/modeling-the-human-trajectory/)). + +![](/economics-of-plan-a/longrungdp-nowatermark.png) + +[Long run economic growth on a log plot shows that it has been faster than an exponential trend](https://ourworldindata.org/grapher/global-gdp-over-the-long-run?yScale=log) + +The core difference is that past technologies automated a narrow domain (e.g. washing machines automated the task of physically cleaning clothes, the internet automated many communication channels), while AI could dominate humans at all (non-human-specific) tasks. So AI could also create new tasks that don’t exist today (e.g., datacenter maintenance, robot manufacturing, AI research) like past technologies did, but the disanalogy is that AI will also be able to automate those new tasks because it will dominate humans both cognitively and then physically (through robotics R&D and production). The humans will only be maintaining the datacenters, making the robots, doing the AI research, up until the point where AI and robots can do it better than them, then they no longer will be (at least once the AIs and robots are numerous enough, and again, absent strong, successful regulation, more on that below). The only tasks left for humans will be driven by human-preference which we think will make for a drastically different economy than today. + +Objection: **There will be regulatory bottlenecks that prevent growth.** + +We think that it is conceivable for regulation to massively slow AI explosive growth; and in fact, the economy in the Plan A scenario grows much more slowly than it would have absent regulation. However, the type of regulation that we think is required to create this type of slowdown involves explicit international coordination to all agree to hold off on explosive AI driven growth. + +This is for two reasons. + + 1. **International competition:** if it’s the case that any country in the world could decide to allow their robots to grow exponentially and soon eclipse the industrial capacity of the rest of the world combined, there would be a massive incentive to do so, and it is hard to believe that all countries would accept this massive handicapping absent coordination. + + 2. **Internal reinvestment spirals:** regulation on AI companies seems unlikely to bite on what they do on internal R&D in a pre-emptive way. So if a company reaches a base of AIs and robots that are capable of self-replicating, AI companies might be able to vertically integrate a supply chain and rapidly outcompete the government (e.g., quickly building billions of robots in a desert offshore). Additionally, even if some sectors are highly regulated, currently highly regulated sectors aren’t the ones that drive explosive growth, e.g., R&D is generally not highly regulated and there will be a massive incentive to cut red tape wherever these are blocking explosive growth. (Note that in AI, we’ve already seen examples of skirted regulations with [copyright law violations](https://www.bakerlaw.com/bartz-v-anthropic/) and [xAI datacenter construction](https://newsletter.semianalysis.com/p/xais-colossus-2-first-gigawatt-datacenter)). + + + + +The regulation that we think is sufficient involves (a) explicit international cooperation to limit the kinds of R&D leading to AI and robots that can replace the work of, or don’t answer to a wide range of humans, and (b) limit explosive growth of robots and the amount of compute in the world through, for example, a cap-and-trade regime. + +Objection: **There won’t be explosive growth because there will be long lags between AI being available and AI being actually adopted.** + +So far, AI has diffused unprecedentedly quickly, even at its current capability level. When AIs are superhuman and there are billions or trillions of dollars on the line, we expect firms to quickly adopt AI. Even if many firms lag dramatically, only some firms need to adopt AI. The firms that do will grow explosively and will quickly become the majority of the sector. There will be strong market pressure to adopt superhuman AI quickly, and we believe this will be enough to cause unprecedentedly fast growth. + +Objection: **There won’t be explosive growth due to the time it takes to build new physical infrastructure.** + +Infrastructure bottlenecks can and will be solved by building more infrastructure. There will be massive economic pressure at the point where it is constraining AI or robot populations (at and approaching the point of full automation) to make this happen quickly (and maybe with partial help from AI and robots). Once robots can perform the relevant physical tasks, the rate of infrastructure construction again becomes roughly proportional to the robot population. + +There will of course be serial time bottlenecks (9 women can’t make a baby in one month), and so too, an arbitrary number of robots will not be able to speed up the time it takes to build a robot factory to be infinitesimal. But for technology very similar to what already exists (production lines in factories, with mining), we don’t think serial construction times prevent short doubling times; see [Appendix A](/supplements/economics-of-plan-a#appendix-a-robot-doubling-times) for our estimates. + +Objection: **Robot and semiconductor explosions will be bottlenecked by rare materials.** + +We appear to be very far from exhausting physical resources on earth. This will slow down growth (on earth) eventually, but only after robots have mined substantial fractions of the earth’s crust. There is room for 30-100s of years worth of current extraction rates depending on the mineral before we even exhaust its proven reserves, and it’s plausible that with robot-powered mining, something on the order of 1% of earth’s crust would be straightforwardly reachable. Once we become bottlenecked by materials on earth, the robots could expand to space. Of course, as the 'materials get harder to find' the growth rate will slow, but it might still be very fast. + +55y5.4My107y221My308y4.9My43y217ky35y3.0My125y15My38y17My231y103My5.9By10y100y1ky10ky100ky1My10My100My1By10ByIronAluminumPhosphateCopperNickelLithiumCobaltRare earthsSilicon + +Years of minerals left at 2024 extraction rates \- ai-2040.com + +Years until proven reserves run out + +Years until 1% of continental crust runs out + +Production & reserves: [USGS MCS 2025](https://pubs.usgs.gov/publication/mcs2025). Crustal abundance: [Wedepohl (1995)](http://apostilas.cena.usp.br/moodle/pessenda/projes/simposio/artigo7.pdf); continental-crust mass 2×1022 kg (CRUST2.0, [Pasyanos et al. 2007](https://ui.adsabs.harvard.edu/abs/2007AGUFM.V33A1161P/abstract)). + +## The simple explorer + +We propose a simple toy model with two factors of production: **humans** and **AI**. A continuum of tasks combines in a constant-elasticity-of-substitution (CES) aggregate with elasticity σ. A growing fraction of tasks is **automatable** : AI and robots can only do those, while humans can do any task, and within a task humans and machines are perfect substitutes. + +The full model below adds capital, multiple sectors, prices, land, energy, and endogenous investment, but the qualitative story is the same one this toy makes visible. + +## The full economic model + +[Open the Full Explorer ↗Our full economic growth explorer, that models the Plan A scenario in a more hand crafted way, and features e.g., capital, robots, prices, land, and energy.](/supplements/econ-explorer) + +The full version of the explorer is a detailed [economic growth model](/supplements/econ-explorer) that we built to model the economy in AI scenarios like Plan A. We have two main disclaimers about this model: + + 1. The model makes some assumptions about the world that we don’t endorse, plausibly contains bugs, and has very high uncertainty (i.e., we don’t trust the outputs to be accurate). That being said, we think it behaves relatively reasonably in these scenarios and we think it is the best existing model for the economic effects of the intelligence and industrial explosions. + + 2. The model’s preset is conditional on Plan A, our scenario in which the government takes substantial and effective steps to regulate AI and drastically slow things down. The second model preset shows something closer to the growth we expect in reality absent the regulatory interventions. Under the default world inputs, the model predicts real output growing ~5000x during 2033, which corresponds to a sustained doubling time of around 1 month. + + + + +3.0% / yrfrom 202410,0001,000100101Years till 20411B10B100B1T10T100T1,000T10,000T100,000T1,000,000TGWP ($)HistoricalPlan AAI 2040Plan D + +Gross World Product \- ai-2040.com + +### AI 2040 Plan A Scenario + +Under the Plan A scenario, effective regulation limits growth in three ways. + + 1. It slows down R&D (particularly algorithmic research), leading to AIs that can fully substitute for all economically relevant work arrive around 2035 instead of 2031. + + 2. It stops AI and robots from becoming entirely self-sufficient (automating 100% of tasks) because of regulatory interventions to require human oversight and auditing. + + 3. The cap-and-trade regime limits the total number of robots and AI hardware to a controlled doubling time of around 6 months for most of the period. + + + + +#### Plan A assumptions + + _Task automation fraction in Plan A. We split labor into cognitive labor and physical labor to show that cognitive labor is automated first because it can happen entirely with AIs running on the cloud, while the automation of physical labor requires a massive buildup of robots. In the Plan A scenario, there is 99% automation reached around 2035, with humans maintaining oversight roles thanks to successful regulation, and retaining a small fraction of economic output (still large in absolute terms)._ + +HumansAIRobotsdotted = Plan D (no deal) + +10M100M1B10B100B1T10T100THumansAIRobotsAI (Plan D)Robots (Plan D)2025203020352040count (log scale) + +_AI and robot populations in Plan A. Before 2029, these reflect unconstrained projections. After that, they reflect the Plan A regulations, where both R &D is restricted, and production is capped from 2032 onwards to a roughly 6 month doubling time (with AI hardware production slowed after 2035 to a ~3y doubling time)._ + +Human, AI & Robot Population \- ai-2040.com + +#### Plan A economic outcomes + +Overall, partially informed by our economic model, we expect the main economic effects of Plan A to include: + + * **An average of roughly 90% yearly growth in real output from 2032 to 2037.** Broadly speaking, throughout this period AI and robots are capable of doing the vast majority of economic tasks, and are allowed to roughly double every six months. Rather than doubling the economy every six months, output increases around 90% per year on average (because of the modelling choice to have capital and human labor partially bottleneck growth). + + + + + * **Most of the economic output becomes driven by AI and robots.** Income shares in the economy shift greatly towards AI and robots as they become capable of doing more and more tasks, and also become more numerous. + + + + +Human cognitiveHuman physicalAIRobots0%25%50%75%100%2025202620272028202920302031203220332034203520362037203820392040 + +Share of US Labor Output \- ai-2040.com + + * **Enormous AI and robot permit revenues redistributed through a “Citizen 's Dividend”.** To a first approximation, the combined AI hardware and robot cap-and-trade permits are the central bottleneck to the economy doubling in size (on average) every year. Naively, that implies the total value of the permits (if well-priced by a competitive market) should be around the size of the current economy, because that is how much additional GDP they will add to the following year. This is reasonable if the AI and robot companies are relatively myopic about the value of the AI and robots (i.e., they mostly base their decision based on 1 expected year of output), which we think may be true under these levels of growth, given potentially fast progress driving obsolescence of tech that is even a few years old. Our [cost modelling](/econ-explorer/cost-model) suggests that the actual manufacturing costs will be a small share of GDP which suggests the permit costs will dominate. So we think the permit revenues will roughly equal the annual GDP when the economy is close to doubling every year. + + + + +2030Income tax87%Corporate tax6%Sales tax7%$10Ttotal revenue2032Compute permits57%Robot permits23%$65Ttotal revenue2034Compute permits26%Robot permits74%$180Ttotal revenue + +US Tax Revenue Sources \- ai-2040.com + +In Plan A, most of the US permit revenue share (75% in 2035) is redistributed to the US population as a Citizen’s Dividend, resulting in roughly $1M/yr per person in 2035 and around $10M/yr in 2040. + +Wages and capital income from stocks would still be unequal, but with the equally distributed Citizen’s Dividend overall income equality would improve drastically, and the poorest Americans would see a drastic elevation in income. + +203020352040$1K$10K$100K$1M$10M$100M$1B + + _Per-person income (2025 dollars)._ + +US Income Distribution in Plan A \- ai-2040.com + + * **Relative prices in the economy shift, with goods and services generally falling in price, while land, positional goods and human-bottlenecked outputs rising.** In Plan A, both cognitive and physical labor from AIs and robots becomes vastly more abundant. Specifically, cognitive labor is 100x more abundant by 2033, and the same is true for physical labor by 2039. So anything the AI and robots can do (which is essentially to make every single good and provide every single service that exists today) will, generally speaking, become a lot cheaper in real terms. Throughout the 2030s in Plan A, AIs are relatively more abundant compared to robots, which means that the price of renting a human-equivalent robot is much higher than a human-equivalent AI worker. In some years, the gap we model is around 100x. + + + + +_Rental costs for a full-time human-equivalent in any automated cognitive or physical domain over the course of Plan A._ + +In 2035, the cost of 1 FTE of labor in the average cognitive domain is $3,000 and in the average physical domain around $200K. These prices would be lower if not for the regulations limiting the amount of AI chips and robots being produced and deployed. In many domains such as law and medicine, since demand is relatively inelastic, we expect the GDP contribution to go down dramatically. In 2035, about 10x as much health care is consumed relative to 2025, but since it's ~100-1000x cheaper, the dollar-wise GDP contribution is 1-10% of what it was. + +For fixed things like land, positional goods (like New York apartments) and human-bottlenecked services (e.g., human day care), we expect the relative pricing to increase, because their supply hasn’t changed. A (very rough model) for the price of land is that there are three main factors: (i) a larger share of spending gets devoted to land (10% -> 50%), and (ii) average spending increases due to the AI boom (100x), (iii) through repurposing of land (e.g., agriculture to housing) and deregulation, available land increases 10x. So with 500x increase in spending and only 10x increase in supply the price should increase 50x. The economic model has a more complicated model for land with more assumptions (which we don’t necessarily endorse), and produces concrete rental prices for urban, rural, and agricultural land separately. On average, land prices increase around 10-100x between the types. + +The model predicts land costs increasing around 100x, while building costs only increase 2x. This implies that the price of housing will become dominated by the price of the land (as opposed to the cost of building housing), so dense housing (e.g. skyscrapers) will become relatively much cheaper than sparse housing. For single family homes, (where the current land vs. build cost share is usually around 40/60), this means that their price is up around 40x on average, but 97% of the home’s value on paper is the land. For 100 floor skyscrapers, land is currently typically 5% of the cost, so the price only goes up 7x. So it’s now commonplace for 100-floor skyscrapers to be built rapidly near major cities almost entirely by robots to provide cheaper housing. In 2035, the Citizen’s Dividend is roughly $1M per person, and it should only take around 15% of it to afford a typical high rise apartment near the city. Also, in deregulated places robots should be able to build far improved transportation that allows for rural homes to also become attractive, where we think the land will be around 10x cheaper. + +We have an explorer in the [scenario exploration tab](/econ-explorer/scenario-exploration) for what the median consumer might be able to afford. The default assumes 15% of their spending is on urban land, 15% on rural land, 35% is on AI labor and 35% is on robot labor. Since AI labor is relatively cheap, by 2034 it becomes the case that each American can afford the equivalent of a mid-sized corporation working for them personally. Robots are more expensive and slower to scale, so in 2034 the median American has an average of a single robot working for them. By 2039, everyone can afford a construction team of robots, four hectares of rural land and a small urban block. + + * **There are high real interest rates around 100-200% and depending on monetary policy choices, either massive deflation or large money creation (to keep inflation around 2%).** We believe interest rates will [increase drastically](https://www.basilhalperin.com/papers/agi_emh.pdf) in Plan A as companies are willing to bid up the robot and AI production permits up to levels close to all of the current year’s GDP and pay interest up to 100-200%: since the robot and AI permits can be converted (barring the comparatively negligible costs of actually producing the AI and robots) into quantities of AI and robots that increase economic output 100-200%. In our modelling, this happens because of explosive growth in cognitive and physical labor, combined with our modelling decision to have production combine labor and capital with limited substitutability (a CES aggregate with elasticity 1.1, close to cobb-douglas). Because capital cannot fully stand in for labor at that elasticity, the exploding effective labor supply makes each unit of capital much more productive, so the marginal product of capital (MPK; how much extra output you get from one extra unit of capital per year) increases greatly throughout Plan A. Since in equilibrium the real interest rate equals MPK minus depreciation, this implies real rates rise similarly throughout the takeoff. We model the savings rate being responsive (but inelastic) with respect to these rising interest rates, but are not confident in the particular parameter we chose. + + + + +We imagine two mainline monetary policy responses being possible in this situation: + +(1) If the government chooses to keep the money supply relatively fixed, then there would be a fixed amount of dollars chasing an increasingly abundant number of goods and services, so each dollar would be spread over a growing number of goods and services, meaning the average price is lower (each dollar buys more goods and services), i.e., leading to massive deflation. The problem with this could be a [deflationary debt spiral](https://en.wikipedia.org/wiki/Debt_deflation), where the AI and robot companies can’t pay back loans in dollars because the robots and AIs are worth nominally less than the loans written the year before. One way to solve this could be for the loans to be denominated in AI and robots, so the companies pay back the loans with some percentage of the AI and robots instead of dollars. (2) Another possibility could be that they attempt to maintain an ongoing annual 2% inflation rate. This would involve a lot of money creation to keep up with the explosive real growth, and there might be high measurement difficulty and time lags making it hard to create the money at the correct level. + +Note this wouldn’t change the fact that the consumer perceived inflation rates would depend dramatically on the basket of goods chosen, as the relative price swings systematically bring down the typical 2025-basket-of-goods dominant goods and services, with the AI and robot automation disproportionately affecting these areas. People will be able to afford many more cognitive-labor-bottlenecked services (e.g. financial advisors, lawyers, travel agents), and manufacturing bottlenecked goods (e.g. computers, cars, clothes, nice food), but less land, and less positional goods or non-automatable things (e.g., human day care). Because of the relative price swings, even if average inflation is 2%, some sectors will have large price increases (inflation) and others will have large price decreases (deflation). + +### AI 2040 Default World + +In the AI 2040 default world, R&D and production quantities are mostly unconstrained by regulation, so there is 100% automation reached by early 2031. With our model’s naive endogenous investment and cost modelling explained below, we get 5000x growth in 1 year during 2033. + +We model the investment required for robot and AI production (cost per quality-adjusted unit) by stacking pure “learning by doing” or economies of scale modelling (using [Wright's law](https://en.wikipedia.org/wiki/Learning_curve), i.e., cost / unit) combining multiplicatively with design improvements (using [Jones' model](https://en.wikipedia.org/wiki/Jones_model), i.e., quality / unit). The idea is to reflect the separate gains from chip production efficiency (more chips of a fixed design being produced per dollar by fabs, i.e., progress that TSMC makes) with chip design efficiency (producing more final hardware performance per dollar from a fixed fab, i.e., progress that Nvidia makes). More on this in the cost model section of the [model explanation](/econ-explorer/model-explanation). The endogenous investment mode then allocates investment between AI hardware, robot hardware, and capital based on a naive myopic return (i.e., so that each earns the same real return at the margin once you net out depreciation). We don’t necessarily endorse this as a way to model endogenous investment, but we’ve chosen it for simplicity for this illustrative model, and (as previously mentioned) we think the outputs are plausible because of the core arguments made in sections 1 and 2, rather than because of this modelling. + +A feedback loop happens where AI and robot investment leads to more cognitive and physical labor, which can then be reinvested in R&D and more production efficiency. AI and robot investment gives higher returns as AI and robots become more capable (according to the exogenous trajectory that assumes full automation in 2031). The explosion in R&D labor flows through to plummeting production costs for AIs and robots. + +## Appendix A: Robot doubling times + +There is a nascent but growing literature on this topic. In particular [this Forethought report](https://www.forethought.org/research/the-industrial-explosion), which splits its doubling time estimates into three phases: + + 1. _Initial_ (One-time gain from AI-directed human labor): on the order of **~1 year**. + + 2. _Acceleration_ : historical experience curves suggest **the number of robots will double 1-5 times before the robot growth rate doubles.** + + 3. _Peak / maximum_ (biological anchors): **days to weeks** as an upper bound (fruit flies double in days, rats in ~6 weeks; bacteria in hours but probably too cognitively basic to bootstrap). + + + + +Out of these we broadly agree, except we think the acceleration phase is unlikely to match historical experience curves: those curves are evidence from human industries, and our best guess is that at full-automation, AI-driven R&D will be able to drive a faster experience curve for robotics. Our best guess is that it could look like Moore’s law (0.2 doublings per halving of cost) or faster, meaning that within a few months, the doubling time might halve multiple times, quickly approaching the ‘peak’ phase. Especially in the regime where AI continues to improve qualitatively, we also think that for the ‘peak / maximum’ phase the actual ceiling might be minutes or hours, because of [bacteria with ~10 minute doubling times](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3751608/) providing an existence proof of something a superintelligent AI may be able to match with a nanotechnology design. + +[This blog post](https://defensesindepth.bio/ai-industrial-takeoff-part-1-maximum-growth-rates-with-current-technology/) by Damon Binder assumes _full automation but no other technological improvement_ , framing this result as a conservative assumption, and finds **headline doubling time of ~1 year ignoring consumption** (with a plausible range of roughly 8 months to 2 years depending on assumptions). He computes the Von Neumann growth rate (the maximum rate at which the economy's physical capital can reproduce itself if all output is reinvested) directly from US BEA input-output tables, capital stocks, and depreciation data. He first assumes labor is free (abundant AI and robots), and then adjusts for actual AI and robot costs, resource depletion, construction time lags, and consumption to get estimates ranging more in the 1-2 year range. + +Another recent piece is [this Epoch report](https://epoch.ai/blog/how-fast-could-robot-production-scale-up), which summarizes current 2025 robotics trends (6 months for humanoids, albeit at very low volumes), acceleration under historical demand shocks (e.g., WW2 and Ukraine), and current factory-build latencies (6-9 months in China). This piece is explicitly about near-term scaling (i.e., excluding AI-directed scale-up), which we think establishes a strong baseline for an initial doubling time of around 1 year at full automation. + +See also older work in [the “Industrial Explosion” section of another Forethought report](https://www.forethought.org/research/preparing-for-the-intelligence-explosion#the-industrial-explosion), [this Open Philanthropy report](https://www.openphilanthropy.org/research/could-advanced-ai-drive-explosive-economic-growth/) and [this Epoch blog post](https://epoch.ai/blog/explosive-growth-from-ai-a-review-of-the-arguments), as well as our own arguments in the [AI 2027 robot economy doubling times box](https://ai-2027.com/slowdown#slowdown-2028-03-31): + +[ai-2027.com/slowdownopen ↗](https://ai-2027.com/slowdown#slowdown-2028-03-31) + +Robot economy doubling timesSlowdown ending · March 2028 + +During World War II the United States and many other countries converted their civilian economies to total war economies. This meant converting factories that produced cars into factories that produced planes and tanks, redirecting raw materials from consumer products to military products, and rerouting transportation networks accordingly. + +We are imagining something similar, except faster because superintelligences are directing and managing the whole process. Roughly speaking, the plan is to convert existing factories to mass-produce a variety of robots (designed by superintelligences to be both better than existing robots and cheaper to produce) which then assist in the construction of newer, more efficient factories and laboratories, which produce larger quantities of more sophisticated robots, which produce even more advanced factories and laboratories, etc. until the combined robot economy spread across all the SEZs is as large as the human economy (and therefore needs to procure its own raw materials, energy, etc.) By that point, the new factories will have produced huge quantities of robotic mining equipment, solar panels, etc. in anticipation of needing to meet demand far greater than the legacy human economy can provide. + +How fast would this new robot economy grow? Some reference points: + + * The modern human economy doubles every twenty years or so. Countries that have developed especially rapidly (e.g. [China)](https://www.macrotrends.net/global-metrics/countries/CHN/china/gdp-gross-domestic-product) sometimes manage to double their economies in less than a decade. + + * A modern car factory produces roughly its own weight in cars in less than a year. Perhaps a fully robotic economy run by superintelligences would be able to reproduce itself in less than a year, so long as it didn’t start to run out of raw materials. + + * Yet that seems like it could be a dramatic underestimate. Plants and insects often have “doubling times” of far less than a year—sometimes just weeks! Perhaps eventually the robots would be so sophisticated, so intricately manufactured and well-designed, that the robot economy could double in a few weeks (again assuming available raw materials). + + * Yet even that could be an underestimate. Plants and insects are operating under many constraints that superintelligent designers don’t have. For example, they need to take the form of self-contained organisms that self-replicate, instead of an economy of diverse and more specialized vehicles and factories shipping materials and equipment back and forth. Besides, bacteria and other tiny organisms reproduce in hours. It’s possible that, eventually, the autonomous robot economy would look more like e.g. a new kind of indigestible algae that spreads across the Earth’s oceans, doubling [twice a day](https://enviroliteracy.org/how-long-does-it-take-for-algae-to-multiply/) so that it covers the entire ocean surface in two months, along with an accompanying ecosystem of predator-species that convert algae into more useful products, themselves fed into floating factories that produce macro-structures like rockets and more floating factories. + + + + +Obviously, all of this is hard to predict. It’s like asking the inventors of the steam engine to guess how long it takes for a modern car factory to produce its own weight in cars, and also to guess how long it would take until such a factory first exists. But economic growth rates have accelerated by [multiple orders of magnitude](https://wiki.aiimpacts.org/featured_articles/precedents_for_economic_n-year_doubling_before_4n-year_doubling) over the [course of human history](https://www.openphilanthropy.org/research/modeling-the-human-trajectory/), and it seems plausible to us that after superintelligence they will accelerate by orders of magnitude more. Our story depicts economic growth accelerating by about 1.5 orders of magnitude over the course of a few years. + +quoted from [AI 2027](https://ai-2027.com/slowdown#slowdown-2028-03-31), the box as published in the Slowdown ending (March 2028) + +## Appendix B: Energy model + +Global energy consumption in 2025 was around 19.0 terawatt-years (TW-yrs). In the Plan A scenario, there is roughly 200x cumulative GWP growth by 2040 (on our [default model assumptions](/econ-explorer/world)). If the energy intensity of the economy (ratio of GDP to energy consumption) stayed constant, then energy consumption would also increase by that same factor, but historically we have observed a decrease in energy intensity since 1920. The headline numbers below are computed directly from the model's Plan A output and our elasticity assumption, so they stay in sync with the [economic model explorer](/supplements/econ-explorer) defaults. + +Energy ×44GDP ×109×1×2×5×10×20×50×100×200182018501880190019201950198020002024 + +Global Energy vs. GDP \- ai-2040.com + +Global primary energy and GDP, both indexed 1820 = 1. 1820–2016: [Bercegol & Benisty (2020)](https://arxiv.org/abs/2008.10967). 2020/2024: [IEA Global Energy Review 2025](https://www.iea.org/reports/global-energy-review-2025). + +We expect this energy decoupling to continue, because the drivers of growth (AI, robots and accompanying capital) should all become more energy efficient per unit of economic output (AI hardware in particular has a [strong power efficiency trend](https://epoch.ai/blog/trends-in-machine-learning-hardware#energy-efficiency)). In our [energy model](/econ-explorer/energy-model), we assume the energy decoupling trend of the last few decades will continue to become more aggressive, and energy consumption will grow at roughly 40% the pace of economic growth (we call this _energy growth / GDP growth_ the energy elasticity of GDP; ε = 0.4). We are highly uncertain about this value, and find values between 20% and 70% to be plausible (80% CI; conditional on this economic growth trajectory). + +GDP ×203Energy ×8×1×10×100×10002025203020352040 + +Energy vs. GDP in Plan A \- ai-2040.com + +Global energy and GDP under Plan A, indexed to 2025 = 1. GDP is the scenario-growth-model Plan A Output Y; energy via E ∝ Yε, ε = 0.40 (energy grows at 40% of the GDP rate). + +**Energy model at a glance** + +computed from the Plan A GWP trajectory and the econ-explorer energy defaults (E ∝ Yε). + +2025 primary energy| 19 TW-yrs +---|--- +Plan A GWP growth (2025 to 2040)| x203 +Energy elasticity of GDP (ε)| 0.40 +Energy growth at ε = 0.40| x8.4 +2040 primary energy (ε = 0.40)| 159 TW-yrs +2040 range (ε = 0.20 to 0.70)| 55 TW-yrs to 784 TW-yrs + +Primary energy scales as GWPε off the 2025 baseline of 19 TW-yrs. The GWP multiple is the scenario-growth-model Plan A Output Y; ε and the baseline mirror the econ-explorer defaults, so this stays in sync. + +An energy elasticity of GDP of 40% leads to energy consumption on earth growing only around 8x during Plan A even as GWP grows by roughly 200x. The resulting 2040 primary-energy total (central case, and the 20% to 70% elasticity range) is computed in the summary above, and updates automatically with the economic model's defaults. + +Depending on the source used for this energy, there may be significant implications for the earth’s temperature. There are two separate things to consider (1) a net increase in waste heat on earth, and (2) emissions, particularly CO2 emissions, leading to temperature increases. + + * **Waste heat.** Solar panels convert energy that was arriving on earth anyway from the sun, so they only change earth’s net energy if they change the reflectiveness (albedo) of the earth’s surface in a major way. Using another energy source like nuclear or fossil fuels, on the other hand, creates a direct increase in waste heat by releasing new energy to the atmosphere that was locked up in the earth’s crust. That said, even at the higher end of the levels we may reach by 2040 in the Plan A scenario, waste heat on its own does not become a major consideration. + + + + +ε=0.20 (55 TW)ε=0.40 (159 TW)ε=0.70 (784 TW)Plan A 2040 (80% CI)today (~19 TW)10× today100×1000×1%10%solar surface:Nuclear / FusionSolar on landSolar on ocean (cooling)+0°C+2°C+4°C+6°C+8°C+10°C+12°C10 TW100 TW1 PW10 PWSurface temperature increase ΔT (°C)Sustained energy production (TW, log scale) + +Warming from Waste Heat \- ai-2040.com + +Stefan-Boltzmann no-feedback equilibrium. Heat factors (W added per W generated): nuclear/fusion 1.0; solar on land +0.56; solar on ocean −0.089. Solar assumes 50% panel efficiency, panel albedo 0.10, surface albedos 0.35 (land) / 0.06 (ocean). + + * **Carbon emissions.** If fossil fuels were used as the sole source for the energy scale up, carbon emissions would become a significant problem by 2040, unless there were significant mitigations taken (e.g., removing the CO2 through direct air capture), with the equilibrium surface temperature increasing to +3.0oC+3.0^oC+3.0oC over pre-industrial levels, up from +1.8oC+1.8^oC+1.8oC today. + + + + +# Warming vs. Cumulative CO₂ + +pre-industrial (280 ppm, 0°C)today (427 ppm, +1.8°C)+2.4°C55 TW10th %ile+2.8°C159 TW50th %ile+4.3°C784 TW90th %ilePlan A, 2040assuming 100% fossil fuelEquilibrium ΔT+0°C+1°C+2°C+3°C+4°C+5°C+6°C01k2k3k4k5k6kEquilibrium ΔT (°C)Atmospheric CO₂ above pre-industrial (GtCO₂) + +Equilibrium warming equations [Myhre et al. (1998)](https://agupubs.onlinelibrary.wiley.com/doi/10.1029/98GL01908) find the climate converging to a given temperature for a given atmospheric CO₂ stock. Today's observed +1.3–1.4°C is below the derived +1.8°C equilibrium value, in line with ([Geoffroy et al. 2013](https://journals.ametsoc.org/view/journals/clim/26/6/jcli-d-12-00195.1.xml)), where ~55% of equilibrium warming materialises within 10 yrs of CO₂ stabilisation, ~70% in 100 yrs, ~99% by 1000 yrs. Plan A markers assume 100% fossil-fuel generation at 2.5 GtCO₂/TW-yr, with the ~45% airborne fraction staying in the air ( [Friedlingstein et al. 2025 / Global Carbon Budget](https://essd.copernicus.org/articles/17/965/2025/)). + +Therefore, we think there should be CO2 emission mitigation policies agreed to globally. Thankfully we think proven [direct air capture (DAC)](https://www.iea.org/energy-system/carbon-capture-utilisation-and-storage/direct-air-capture) would provide an affordable path to mitigating CO2 increase, especially with help from AI and robot labor during the 2030s. Two specific ideas we have for carbon capture policy include: + + 1. **Private cap-and-trade market** (no government subsidy necessary), where any fossil fuel emitters globally are required to offset 100% of their emissions with equivalent carbon capture credits (with DAC companies making revenue by capturing carbon and selling the credits to the carbon emitters). + + 2. **Auctioned carbon removal subsidy** (government subsidy for net carbon removal), where the government(s) set the goal of not just capping emissions, but actually net removing carbon from the atmosphere, with the goal of returning to pre-industrial levels. They can set a “willingness to pay” curve, based on the social cost of carbon. We imagine something like the following curve being sufficient for affordably returning the earth to pre-industrial CO2 levels by 2040 with around $85T global government spending in total from 2033-2040, or around 0.1% of annual GWP on average during that period. + + + + +pre-industrialtoday2025 expected DAC cost ($600B/Gt)2030 expected DAC cost ($400B/Gt)2035 expected DAC cost ($100B/Gt)2040 expected DAC cost ($30B/Gt)SCC$-200B$-100B$0B$100B$200B$300B$400B$500B$600B$700B01k2k3k4k5k6k7kCarbon Capture Cost ($B per Gt removed)Total atmospheric CO₂ (GtCO₂) + +Carbon Capture Subsidy \- ai-2040.com + +Subsidy curve (WTP per GtCO₂ removed)Expected Direct Air Capture (DAC) cost by yearPlan A path + +**Total Plan A subsidy: ~$ 85T** (price × Δstock, summed). Prices: DAC × (1+10%), path sits on SCC curve. Sources: [EPA 2023](https://www.epa.gov/environmental-economics/scghg), [IEA DAC 2022](https://www.iea.org/reports/direct-air-capture-2022/executive-summary), [CDR.fyi 2025](https://www.cdr.fyi/blog/direct-air-capture-market-snapshot-2025). + +Under some naive energy cost modelling that AI and robot labor accelerate progress in energy costs, and assuming a CO2 cap-and-trade forcing net neutral carbon emissions is enacted in 2035, we find the following energy cost, associated investment, and resulting energy mix. + +1. + +We’ve written about our reasoning for this assumption elsewhere, especially in our [timelines and takeoff model](https://www.aifuturesmodel.com/), and so will take it for granted here. Our team’s estimates for the probability of this outcome range from roughly 70% to 90%. + +2. + +For example, one common model of growth is the Cobb-Douglas production function, which models growth as Y=A∗Kα∗L1−αY = A* K^\alpha * L^{1-\alpha}Y=A∗Kα∗L1−α, where A and α\alphaα are constants, L is labor, and K is capital. A typical α≈13\alpha \approx \frac{1}{3}α≈31​ so overall growth scales with L23L^{\frac{2}{3}}L32​. Plus more people would build more capital, so in steady state output should double barring e.g., natural resource constraints, which are probably mild for many orders of magnitude of growth. + +3. + +This is somewhat of an oversimplification because raw compute isn’t the only bottleneck on number of AI copies; in reality it is a complex combination of actual ‘compute’ (FLOP/s), but also memory bandwidth, memory capacity, and interconnect. However these factors don’t substantially change the bottom line. + +4. + +Of course, we expect these trends to change, and more detailed and mechanistic models such as the AI Futures Model predict that the compute efficiency growth rate will vary over time. + +5. + +Inference compute may change as a fraction of total compute, but for our simple model we’re just assuming that the share is flat at the point of full automation. + +6. + +AIs today are mostly being scaled to be smarter, not just more numerous, so this isn’t happening in practice. After the point where AIs can perfectly substitute for human labor, increasing intelligence will also be happening, so this argument is only a lower bound for the economic impacts. + +7. + +This argument holds unless regulation massively slows down the possible growth. However, unilateral regulation is insufficient because any country that allowed explosive AI growth will gain an insurmountable economic edge. Therefore, regulation to prevent explosive economic growth must involve international coordination. + +8. + +The full “rapid progress” scenario from the FRI report is: By the end of 2030, in the rapid-progress world, AI systems are capable of competing with the best human minds and workers, and can surpass them. Human creativity and leadership remain valued, but mostly for setting high-level vision— day-to-day execution can be left to silicon-based systems. Autonomous researchers can collapse years-long research timelines into days, weeks, or months, creating game-changing technologies, such as materials that revolutionize energy storage, or bespoke cancer cures. No human freelance software engineer can outperform AI. The same goes for customer service (e.g., call center and support chat), paralegal, and administrative workers (e.g., bookkeepers or scheduling assistants). Indeed, models have become so capable that AI can create an album of the same caliber as the Grammy Album of the Year. Additionally, a single AI agent can generate a Pulitzer- (or Booker Prize-) caliber novel according to current (2025) standards, adapt the book into an engaging two-hour movie, negotiate the resulting book and movie contracts, and launch the marketing campaigns for both while its sibling agents manage the book publishing company and movie studio at the level of highly competent CEOs. Not only do Level-5 robo-taxis exist, but they are, on average, 99.9% safer than human piloted cars and can venture anywhere off-road that a competent human driver can. Meanwhile, robots can navigate an arbitrary home anywhere in the world, make a cup of the most popular local hot beverage, clean and put away the dishes according to the local custom, fix any plumbing issues that arise while they’re doing the dishes—and they can do it all faster and more reliably than most humans and without human guidance. Robots in advanced factories can autonomously perform the full range of tasks requiring the highest levels of dexterity, coordination, and adaptive decision-making. + +9. + +For example, there could be a R&D function combining R&D labor from a human population with labor from an AI population with forecasts of how useful/capable that AI labor will be (e.g., directly modelling the productivity of AI across R&D labor tasks). + +10. + +For example, see AI coding time-horizons [doubling every ~four months](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/) recently, with similar results (though more uncertain) [in other domains](https://metr.org/blog/2025-07-14-how-does-time-horizon-vary-across-domains/), along with the population/efficiency doubling times explained in [section 1](/supplements/economics-of-plan-a#core-arguments-for-explosive-growth), we think Q˙/Q\dot Q/QQ˙​/Q should be more responsive to increased R&D automation. + +11. + +The continental crust makes up the continents _plus_ the continental shelves (shallow seas out to roughly the 200 m depth break). It has an average thickness of 35-40 km and covers about 40-45% of earth’s surface area. This suggests that by mining around 4% of earth’s surface area (total area around the size of Australia), to a depth of 4km, would consist of about 1% of the earth’s continental crust. + +12. + +The default world inputs assume a capability trajectory like [this one](https://www.aifuturesmodel.com/p?base=eli-12-29-25&acth=5.386766760341415&arts=3.372873086588689) from the AI Futures Model, where AI and robots can fully automate (100%) of cognitive and physical tasks around the beginning of 2031. + +13. + +Definitions for the task universe and automatable vs. automated are [here in the model explanation](/econ-explorer/model-explanation/automation-definitions). + +14. + +We model the relationship between labor and capital (things like factories, tractors, land, raw materials, etc) as a CES aggregate with elasticity σ = 1.1, slightly on the substitutes side of [Cobb-Douglas](https://en.wikipedia.org/wiki/Cobb%E2%80%93Douglas_production_function) (σ = 1, where the labor and capital income share ratio would stay flat). This is a judgment call that could very well be wrong! After full automation, what counts as “capital” is a blurry category, especially since we count robots separately. In reality, we expect AI and robots to substitute with many kinds of capital (e.g., robots replacing equipment, or AI labor leading to R&D that obviates the need for a lot of existing capital), while other kinds of capital bottleneck growth (e.g., land, especially with land regulation; raw materials; etc.) Our best guess is that these forces net out mildly on the side of substitutability, which is why we set σ slightly above 1. Note this runs contrary to historical measurements, which generally find complementarity (meaning elasticity parameter σ < 1; [Chirinko, 2008](https://ideas.repec.org/a/eee/jmacro/v30y2008i2p671-686.html); [Knoblach et al., 2020](https://onlinelibrary.wiley.com/doi/10.1111/obes.12312); [Gechert et al., 2022](https://doi.org/10.1016/j.red.2021.05.003)), but we think the capital stock of an AI-driven economy is different enough in kind from the capital behind those estimates that mild substitutability is the better guess. + +15. + +Land was 40%, increases 100x (40%*100), housing was 60%, increases 2x (60%*2), which leads to a total land (40) + housing (1.2) cost 41.2x higher than today. Of which 97% (40/41.2) is the land’s cost. + +16. + +Land’s 5% increases 100x (5%*100), and the building’s 95% increases 2x (95%*2), which leads to a total land (5) and housing (1.9) cost 6.9x higher than today. Still with the land being 72% of the cost. + +17. + +Our story depicts the conversion process going about 5x faster. We think this is a reasonable guess, taking into account bottlenecks etc., for how fast this conversion could go if a million superintelligences were orchestrating it. Of course we are very uncertain. + +18. + +Possibly also more advanced sources of energy, such as fusion power. + +19. + +Quick napkin math: The [Empire State Building](https://en.wikipedia.org/wiki/Empire_State_Building) has an area of 2.77m sq ft, and weighs 365k tons. Gigafactory Shanghai has an [area of 4.5m sq ft](https://en.wikipedia.org/wiki/List_of_Tesla_factories) and [produces 750k](https://en.wikipedia.org/wiki/Gigafactory_Shanghai) vehicles per year, mostly Model 3’s and Model Y’s, which weigh about two tons each. Presumably the Empire State Building has a higher mass-to-square-foot ratio than the Shanghai Gigafactory (since it is vertical rather than horizontal and needs stronger supports) so if anything this underestimates. Thus it seems that a factory which probably weighs well less than a million tons is producing 1.5 million tons of cars each year. + +20. + +We don’t think it would run out. Initially the robot economy would be dependent on human mines for materials. But by the time it outgrows these sources, the millions of superintelligences will have prospected new mines and developed new technologies to exploit them. Imagine e.g. undersea mining robots that strip-mine the seabed for rare minerals, new chemical processing pathways that more efficiently convert raw ore from above-ground stripmines into useful raw materials… + +21. + +If the current economy doubles every twenty years, one order of magnitude faster would be a doubling in two years, two orders of magnitude faster would be a doubling in 0.2 years, and so forth. The hypothetical superintelligent-algae-economy described above would be about four orders of magnitude faster growth than the current human economy. + +22. + +Using the [Energy Transition Institute 2024](https://www.energy-transition-institute.com/article/statistical-review-of-world-energy-2025) direct-method total of 592 EJ, and adding the [IEA’s 1.3% 2025 growth estimate](https://www.iea.org/reports/global-energy-review-2025) (+8 EJ) gives roughly 600 EJ in 2025, or ~19.0 terawatt-years. Where a terawatt-year is the amount of energy equivalent to 1 terawatt of power running for one year. Note: 1 TW-yr = 31.5 EJ = 8,760 TWh. + +23. + +On average, energy growth has moved at 70% the pace of GDP growth since 1920. In recent decades (since 1980), it has moved at 60% the pace of GDP growth, and in the last 10 years, it moved at 52% the pace. + +24. + +We use data from [Bercegol & Benisty (2020)](https://arxiv.org/abs/2008.10967) and [IEA Global Energy Review 2025](https://www.iea.org/reports/global-energy-review-2025) collected to view [here](https://docs.google.com/spreadsheets/d/1yxCYv4lsoVGJqdwh2Oe28ngiAlO-ZyW4Ng5HmraaX3s/edit?gid=67630334#gid=67630334). + +25. + +Energy scales as GWP^ε; at ε = 0.40 and the Plan A GWP multiple of ~200x by 2040, that is ~200^0.4 ≈ 8x. The exact value tracks the model output shown in the summary table above. diff --git a/sources/ai2040-pages/faq.md b/sources/ai2040-pages/faq.md new file mode 100644 index 0000000..8a8b2ad --- /dev/null +++ b/sources/ai2040-pages/faq.md @@ -0,0 +1,101 @@ +--- +url: https://ai-2040.com/supplements/faq +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# FAQ + +**Q: Isn’t Plan A bad for concentration of power risks?** + +A: No, Plan A reduces COP risks by a lot. The status quo involves an insane amount of power concentrating into the frontier AI companies (which are all based in the US). If they solve the alignment problem, the result will be a handful of AI CEOs or USG officials controlling an army of superintelligences, and hence having the ability to take over the world. (This is what happened in the slowdown ending of AI 2027). + +Plan A dramatically improves on the status quo by: + + 1. Increasing the number and diversity of frontier AI companies. Because of regulations slowing the frontier, trailing actors can catch up, and in the Plan A scenario, by 2040, there are dozens of companies across different countries near the frontier of AI development. + + 2. Increasing the transparency and visibility into AI development. A core dynamic driving concentration of power risks is a dynamic where most of the power centers in the world do not realize they are going to be disempowered by AI until it is too late. In Plan A, the takeoff period is much longer and much more transparent, so the public, the US government, and countries besides the US all have much more time to see what is going on inside the labs. + + 3. Spreading out the AI-driven transformation of the planet over a few more years, giving more time for everyone to prepare and react, making it more likely that something resembling our current system of checks and balances persists instead of crumbling into chaotic conflict. + + 4. Decreasing the feasibility of inserting secret loyalties into AI systems. There is high quality compute accounting,so it is no longer feasible for anyone to obtain hidden compute, which could be used to train a backdoor into the AIs. + + + + +**Q: Wouldn’t China not agree to the deal? Even if they did agree on paper, wouldn’t they defect and keep developing AI anyways?** + +A: China is likely to want something along the lines of Plan A because the status quo outcome for China is that they will be outcompeted by US AI companies. This will result in either misaligned AIs taking over the world, or the US obtaining a decisive lead in AI capabilities sufficient to disempower China. (These two options correspond to the race and slowdown endings of AI 2027 respectively). The CCP today does not seem very aware of the importance of AGI or ASI, but this will change as AI has bigger and bigger impacts on the world. + +The question isn’t “Will China defect from the deal?” The question is, “What’s the biggest defection they could get away with, and how bad would that be?” + +There are two ways of defecting from the deal: (1) building a massive covert project that you attempt to keep hidden from the rest of the world, and (2) using the existing legal compute. + +![](/supplements/faq/image1.png) + +A baseline estimate for the maximum size of covert project that they could have is ~1.2M H100e. We do a more thorough analysis of this question in [our covert project supplement](/supplements/covert-ai-projects), and we have a branch of our scenario gaming out how China could defect. + +We prevent the use of existing legal compute through training verification mechanisms. We assume US and Chinese auditors are given full transparency into the other’s R&D datacenters, allowing them to check if the compute is being used how they claim it is being used. We do a more thorough analysis of this question in our [verification supplement](https://www.ai-2040.com/supplements/verification-plan). + +**Q: Isn’t Total Research Transparency bad because it leaks algorithmic secrets to China and other AI companies?** + +No current AI companies claim to be secure against nation state level adversaries. They aren’t secure even for model weights, which are much easier to secure than algorithmic secrets because model weights are large, often on the order of 1-10TB, while algorithmic secrets are often memorizable by individual humans. Because AI algorithms will have already leaked to competing nations, the costs of full transparency here are relatively low compared to the enormous benefits. + +**Q: Plan A is complicated. Shouldn’t we do a simpler and more straightforward plan, like shutting it all down (Plan S)?** + +Our proposed implementation of Plan A is very complicated, and even a good implementation will incur significant existential risk. However, all simpler plans that we are aware of would incur even more risk than Plan A. In particular the main issue with Plan S is that it does worse than Plan A at making forward progress towards solving the safety/alignment problems that prevent further capability scaling, because we are shut down at an earlier capability level. This is bad because at some point, both Plan A and Plan S will break down and return to racing, and in Plan S, much less alignment progress will have been made by this point. That said, we are very sympathetic to Plan S and could imagine being convinced that some version of it is better after all. + +**Q: The US government is incompetent. Why are we trusting the US government to do a good job regulating AI?** + +For any plan, incompetent execution of that plan is a serious concern. We are in fact very worried about incompetent execution of Plan A, and so we wrote a [branch of our scenario explaining how this could lead to existential catastrophe](https://ai-2040.com/?choices=plan-a-root#branchpoint-flawed-safety-case-is-approved). + +Who else is going to regulate AI, if the US government doesn’t? We don’t trust the industry to self-regulate effectively. Moreover, if the US government is kept in the dark while AI companies race through an intelligence explosion, that seems like an unstable situation that could lead to sudden wakeup and thrashing about, or worse, a US company effectively puppetting the government (as hinted at in the [AI 2027 slowdown ending](https://ai-2027.com/slowdown)). If the US government is not in the dark, they will want to regulate AI, and we are trying to figure out and describe what good regulation would look like. + +Large parts of Plan A are designed to improve the AI-related competence of the government—and the wider world—as fast as possible. For example, this is part of why we recommend [Total Research Transparency](https://www.ai-2040.com/supplements/transparency-plan) instead of some more moderate kind of transparency. This is also part of why we recommend limiting the rate of AI progress; it gives the world more time to react and prepare for each new level of capability. We also have things to say about [using AI to improve epistemics](https://ai-2040.com/supplements/ai-for-epistemics), and about immediate actions the government could take to improve its understanding of AI and capacity to regulate effectively. Finally, a big motivation for Plan A is to prevent extreme concentrations of power, which is good in its own right but also helps mitigate government incompetence. + +**Q: Plan A just isn’t going to happen. Don’t you know that?** + +The Plan A scenario is a recommendation, not a forecast; it is _not_ what we expect to happen. We think many of the ideas in AI 2040 are good and will be helpful inspiration for people trying to steer things in a better direction, even if Plan A doesn’t happen. That said, we also think that Plan A is plausible enough that it is worthwhile to push for it. + +Over just the last year, there has been a massive shift in the awareness and understanding of AI among policymakers. As the effects of AI get larger, there will be vastly more wakeup. Before the point of no return—the point at which AIs or AI-powered humans can disempower the rest of the world—AI will probably have transformative effects on the world. It’s hard to be confident in any specifics, but some illustrative possibilities include: + + * AIs can automate the majority of all jobs: physical jobs with robots, cognitive jobs with AI agents running in the cloud. This leads to massive widespread unemployment. + + * Extreme economic impacts other than unemployment. For example, perhaps AI companies will continue increasing their revenue at ~3x/year until 2030, which would mean that the frontier AI companies would have a $10T/year revenue. This is an unprecedented amount of wealth and power in a company, and will trigger alarm bells. + + * AIs completely change the nature of warfare. For example, AI piloted drone swarms may render existing mechanized militaries (e.g. infantry, tanks, planes, aircraft carriers) completely obsolete. This might lead to significant changes in the geopolitical balance of power as countries who aren’t at the AI frontier will fall behind militarily. + + * Extreme misuse capabilities, e.g. suppose it’s the case that AIs are able to cheaply design and spread bioweapons with commonly accessible materials. Hardening the world against these capabilities will be extremely difficult. + + + + +It’s hard to know what the political environment will look like after AI has massive effects on the world. The main question is how early transformative effects on the world happen relative to AI-caused permanent lock in. Plan A is much easier and more effective if implemented early—before the AIs have fully automated the AI R&D process. diff --git a/sources/ai2040-pages/how-plan-a-solves-our-5-biggest-problems.md b/sources/ai2040-pages/how-plan-a-solves-our-5-biggest-problems.md new file mode 100644 index 0000000..7e9d89c --- /dev/null +++ b/sources/ai2040-pages/how-plan-a-solves-our-5-biggest-problems.md @@ -0,0 +1,73 @@ +--- +url: https://ai-2040.com/supplements/how-plan-a-solves-our-5-biggest-problems +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# How Plan A solves our 5 biggest problems + +### Thomas Larsen + +**Problem #1: Loss of Control to superintelligent AIs.** + +Due to the slowdown agreement, the intelligence explosion happens over the course of at least 10 years. Instead of being forced by competitive pressures to automate work and trust AIs, companies now are forced to slow down until they have a rock solid safety case. In the Plan A scenario, we have ~5 years with top expert level AI before handoff, and we could realistically get more time. If alignment ends up being hard enough that a 10 year slowdown isn’t sufficient, the infrastructure for a longer pause is set up. + +Due to the transparency and access provisions in the deal, there is way more oversight of AI developers, and many more total people have the relevant access to do AI alignment work with frontier models. + +(See more [here](https://ai-2040.com/supplements/comparing-possible-plans#plan-assessment-table)). + +**Problem #2: AI will lead to massive concentration of power into whatever group controls the first superintelligences (if anyone).** + +The transparency act already helps somewhat to prevent CEOs backdooring their models. The Total Research Transparency deal then “seals the deal” to a much greater extent; any attempt at inserting a hidden agenda into your biggest and most-used models–whether initiated by a CEO or a President–would be visible to multiple world governments. The model spec transparency + algorithmic progress slowdown + Total Research Transparency deal also makes there be many companies and governments with similar levels of AI, which is great for spreading out the power more generally, not just narrowly preventing secret loyalties/agendas. + +Because the intelligence explosion is now happening in a slow and drawn out way, and we’ve prevented disparities in model access, so everyone has access to the same level of AI capabilities, anyone who is getting disempowered will be able to realize that and then negotiate for a better deal. Normal people have access to smart enough AIs that they can understand the situation well, and they can vote for ethical and thoughtful politicians, who will maintain and improve the power the public has, leading to a virtuous cycle. This leads to the Citizens Dividend and space governance proposal being implemented and spreading out the power forever. + +**Problem #3: World War 3 risk due to superintelligences disempowering many countries, including nuclear powers.** + +Plan A prevents WW3 by ending the “winner take all” effects of a race to superintelligence. No longer is there a race to ASI with each side fearing what the other might do if they get there first, and with each side contemplating taking aggressive action to slow down the other side, and with one or more sides starting to freak out as they realize they are on a path towards losing. No longer are there a bunch of middle countries gradually realizing that no matter who wins, they lose, and starting to freak out also. + +**Problem #4: AI taking everyone’s jobs.** + +The reasons that unemployment is bad are that people lose their (a) income, (b) power, and (c) source of meaning. Plan A slows down AI induced job loss to a rate where the negative effects of unemployment can be mitigated. + +The main reason that people need jobs is to have a source of income. Plan A addresses the loss of income via the Citizen’s Dividend. Each person gets a share of income large enough to support an extremely high quality of life. + +A second reason why jobs are important is that they empower their holders. CEOs and governments are incentivized to care for people who provide economic value. However, once AIs can do all the economically relevant tasks, governments and companies lose a strong incentive to provide for their citizens/employees. See the “concentration of power section” for more. + +A final reason jobs are important is that they provide meaning for many people. Here we just say that it won’t be that bad (people can get meaning in a bunch of ways, not just economically useful work! Family, friends, romance, hobbies, sports!) and that the benefits outweigh the costs. If you disagree, well, we also describe Plan S… there’s basically no way to avoid mass unemployment if you allow companies to build and deploy AIs that obsolete humans. And in particular, banning the use of AI in various professions won’t work for long. + +**Problem #5: AI misuse by small actors.** + +Widespread access to smart AI will make it trivially easy to build bioweapons (or other technology that would cause huge damage at little cost to the attacker). Terrorists, states (e.g. North Korea), or misanthropists who are motivated to cause immense harm could leverage AI to do a much larger attack than they would have otherwise. + +Plan A, as part of mitigating covert projects, limits the capability of open weight models to pre-deal levels. Closed models are trained to refuse dangerous requests. In case access limitations fail, Plan A involves massive investments into biosecurity and other measures to improve the resilience of the world. + +There’s an alternate approach to misuse taken in Plan B, C, and D, which is to speedrun the intelligence explosion and hope to build superintelligence that can directly take over the world, hence preventing misuse. It’s plausible that these plans look better for misuse risk than Plan A because there’s so little time for the misuse to occur before superintelligence. But this strategy imposes unacceptable levels of loss of control and concentration of power risk, and so doesn’t pass overall cost benefit. diff --git a/sources/ai2040-pages/main.md b/sources/ai2040-pages/main.md new file mode 100644 index 0000000..a198027 --- /dev/null +++ b/sources/ai2040-pages/main.md @@ -0,0 +1,1320 @@ +--- +url: https://ai-2040.com/?choices=plan-a-root +title: "AI 2040: Plan A" +fetched: 2026-07-11 +--- + +## [AI 2040](http://ai-2040.com/) + +# Plan A + +≡[PDF](http://ai-2040.com/AI-2040.pdf) + +[Supplements](http://ai-2040.com/supplements) + +Plan A Basics + +[FAQ FAQ](http://ai-2040.com/supplements/faq)[How Plan A solves our 5 biggest problems How Plan A solves our 5 biggest problems](http://ai-2040.com/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI Projects Covert AI Projects](http://ai-2040.com/supplements/covert-ai-projects)[Verification Plan Verification Plan](http://ai-2040.com/supplements/verification-plan)[Transparency Plan Transparency Plan](http://ai-2040.com/supplements/transparency-plan)[Capability Scaling Strategy Capability Scaling Strategy](http://ai-2040.com/supplements/capability-scaling-strategy)[Security in Plan A Security in Plan A](http://ai-2040.com/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A Assumptions Plan A Assumptions](http://ai-2040.com/supplements/plan-a-assumptions)[Comparing Possible Plans Comparing Possible Plans](http://ai-2040.com/supplements/comparing-possible-plans)[Deal Decline Deal Decline](http://ai-2040.com/supplements/deal-decline)[Takeoff Supplement Takeoff Supplement](http://ai-2040.com/supplements/takeoff-supplement) + +Economics + +[Economics of Plan A Economics of Plan A](http://ai-2040.com/supplements/economics-of-plan-a)[Compute in Plan A Compute in Plan A](http://ai-2040.com/supplements/compute-supplement)[Economic Growth Explorer Economic Growth Explorer](http://ai-2040.com/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance Plan Space Governance Plan](http://ai-2040.com/supplements/space-governance-plan)[AI for Epistemics AI for Epistemics](http://ai-2040.com/supplements/ai-for-epistemics)[Alignment Roadmap Alignment Roadmap](http://ai-2040.com/supplements/alignment-roadmap) + +[About](http://ai-2040.com/about) + +[![Image 2: AI Futures Project](http://ai-2040.com/ai-futures-logo.svg)](https://ai-futures.org/) + +## Thomas Larsen, Romeo Dean, Brendan Halstead, Eli Lifland, Ryan Greenblatt, Daniel Kokotajlo + +AI companies are racing to build AIs that are smarter than humans in every way. In [AI 2027](https://ai-2027.com/), we predicted that this would result in either extinction or irreversible concentration of power.[1](http://ai-2040.com/footnotes#footnote-1) + +Plan A is our positive vision for what should happen instead. + +In this scenario, humanity delays the development of superintelligence until 2040, makes all AI research public, allows dozens of companies globally to catch up to the frontier, and intentionally enters a regime of mutually assured compute destruction. + +What is Plan A?Why did we write this?Why a scenario?AI timelines? + +**Plan A is our positive vision for how humanity can avoid AI-driven existential catastrophe and reach a flourishing future.** It’s informed by conversations with experts at major U.S. frontier AI companies, direct experience at OpenAI, tabletop exercises, and discussions with policymakers, national security experts, and AI policy leaders. We recommend an international deal to avoid a dangerous race to superintelligence. The deal involves [total research transparency](http://ai-2040.com/supplements/transparency-plan) for AI R&D, which allows the nations of the world to understand what’s happening and [enforce guardrails](https://www.ai-2040.com/supplements/verification-plan). The result is multiple companies across multiple countries scaling slowly and safely together towards superintelligence, instead of racing each other in secrecy. + +**Plan A is primarily a recommendation, not a prediction.** This scenario is _not_ our best guess as to what the future will actually look like. Instead, it’s a vehicle for communicating and stress-testing our policy recommendations. While the _implementation_ of Plan A is a recommendation and not what we actually expect to happen, the _subsequent effects depicted_ are predictions.[2](http://ai-2040.com/footnotes#footnote-2) + +In this AI 2040 scenario, Plan A is implemented successfully, albeit imperfectly and only in the nick of time. + +We contrast Plan A with 4 alternative plans (B, C, D, and S), which correspond to the main ways the US could respond (or not) to the challenges of superintelligence.[3](http://ai-2040.com/footnotes#footnote-3) + +**AI companies will probably succeed at their stated goal of building smarter-than-human AI systems within the next 1 to 10 years.** + +The industry has convinced itself that controlling superintelligent AI can be figured out on the fly, and thus has no remotely adequate plan. We think this situation is terrible and could easily get us all killed.[4](http://ai-2040.com/footnotes#footnote-4) We do not expect whoever “wins the race” to have much of a lead, and we do not expect them to unilaterally slow down to reduce existential risk.[5](http://ai-2040.com/footnotes#footnote-5) If this race continues,[6](http://ai-2040.com/footnotes#footnote-6) we do not expect humans to maintain effective control as their AIs become superintelligent.[7](http://ai-2040.com/footnotes#footnote-7) + +Moreover, even if the AI companies somehow align their AIs, the result will be an unprecedented concentration of power—that is, the result will be a situation where a tiny group of people, or possibly just a single individual, is effectively in control of the world’s only army of superintelligences for some months, and will be presented by said superintelligences with various options for how to proceed, some of which will _de facto_ amount to taking over the world.[8](http://ai-2040.com/footnotes#footnote-8) + +As best as we can guess, the CEOs of OpenAI, Anthropic, xAI, and Google DeepMind understand this and are proceeding anyway, perhaps because they think they are the lesser evil and will use their immense power responsibly, unlike Xi Jinping or rival CEOs.[9](http://ai-2040.com/footnotes#footnote-9) + +While we agree that it is generally correct to choose the lesser evil, we don’t think we should advocate for a strategy that has such a scarily high chance of leading to human extinction or global dictatorship. Instead, we wish to advocate for something that is actually good. If enough people do likewise, it can happen. + +So, we wrote a scenario outlining that possible world. + +**“Plans are worthless, but planning is everything.” -Dwight D. Eisenhower** + +We think most AI policy proposals fall apart under [scenario scrutiny](https://blog.aifutures.org/p/scenario-scrutiny-for-ai-policy)—that is, if you try to write down a detailed and plausible scenario in which that proposal succeeds, you will find it difficult to do so, and you will realize the plan is less likely to work than it seemed, or has more unpleasant side-effects than its proponents acknowledged. + +Perhaps that’s why scenario scrutiny is so rare in AI policy. Everyone wants to say that their own favorite policies will have great consequences and that the policies of their rivals will have terrible consequences. Applying scenario scrutiny to their own favorite policies might surface uncomfortable issues with them; meanwhile, applying scenario scrutiny to their rival’s policies is a lot of work for little rhetorical gain.[10](http://ai-2040.com/footnotes#footnote-10) + +We think the discourse would be improved if more AI policy proposals were subjected to scenario scrutiny. So we’re starting with our own, even though this opens us up to criticism. We hope critics will judge us against the existing state-of-the-art for plans to navigate the AI transition (if they can find any) and not against some hazy but pleasant fantasy where no one has to make any hard choices yet everything will probably be fine. + +What of the immense difficulty of predicting the effect our policy would have in a world approaching superhuman AIs? This is like trying to predict how to best fight World War 3, except that it’s an even larger departure from past case-studies. Yet it is still valuable to attempt, just as it is valuable for the U.S. military to game out Taiwan scenarios in excruciating detail. There are other precedents as well: [intelligence agencies](https://assets.publishing.service.gov.uk/media/68dba439dadf7616351e4bf8/GST_7_Final_post_pic_change_WEB.pdf), [climate bodies](https://iea.blob.core.windows.net/assets/deebef5d-0c34-4539-9d0c-10b13d840027/NetZeroby2050-ARoadmapfortheGlobalEnergySector_CORR.pdf), and [pandemic-preparedness offices](https://centerforhealthsecurity.org/our-work/tabletop-exercises/clade-x-tabletop-exercise) all rely on various kinds of scenario planning. + +Plan A is our ambitious proposal for what to do, and we’d like to see something like it implemented soon because we are uncertain about how much time remains.[11](http://ai-2040.com/footnotes#footnote-11) But for purposes of writing a concrete scenario, we need a concrete timeline. + +**The timeline of this scenario is:** + +* In 2029, the US and China agree to avoid a reckless race to superintelligence. + +* **In 2030, we _would have_ fully automated AI R&D,** leading to superintelligence by the end of the year. **Thanks to the deal, we avoid this.** + +* Between 2030 and 2035, we scale within the human range, to AIs that are roughly as capable as top human experts. + +* In 2035, we pause at top-human-expert level AI in order to maintain human control. + +* In 2040, we unpause and scale to superintelligence.[12](http://ai-2040.com/footnotes#footnote-12) (**Hence the title: AI 2040**) + +In our previous scenario, [AI 2027](https://ai-2027.com/), AI fully automated the process of building smarter AIs in 2027, leading to an intelligence explosion and superintelligence within the year. The two differences in this scenario are (1) the default timeline is now 2030, and (2) thanks to governance actions, generally-superhuman AIs first appear in 2040. + +**We changed the default timeline because we want our portfolio of scenarios to reflect our uncertainty about AI timelines.** AI 2027's titular year was chosen because, at the time we started writing, Daniel thought there was roughly a 50% chance that things would go that fast or faster.[13](http://ai-2040.com/footnotes#footnote-13) At the time we started writing Plan A, 2030 was the corresponding year for Thomas. **Daniel currently thinks things will probably go somewhat faster than depicted in this scenario**; you can read more about our team’s views on timelines [here](https://blog.aifutures.org/p/q1-2026-timelines-update) and [here.](https://www.aifuturesmodel.com/) + +We changed the governance actions because **this scenario is primarily a recommendation, not a prediction.** Conducting a full-speed intelligence explosion is wildly reckless and concentrates power to an extreme degree. + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2027: The Writing on the Wall + +America has two workforces now. The first is people, 165 million of them. The second is AI agents: millions of copies spun up and shut down every hour, working around the clock at superhuman speeds. + +Most of their work is slop. But enough of it is good that people are paying ten billion dollars a month for AIs that can, in theory at least, do anything on a computer that an employee can. + +There is one job the AI companies want to automate more than any other—their own. They haven’t succeeded yet; no [recursive self-improvement](https://www.anthropic.com/institute/recursive-self-improvement) so far.[14](http://ai-2040.com/footnotes#footnote-14) But they seem to be getting closer, and they’re pulling up the ladder behind them: the strongest coding AIs refuse to help competitors with AI R&D.[15](http://ai-2040.com/footnotes#footnote-15) Even as the most bullish employees admit that things are taking a bit longer than planned, the skeptics notice that their usual dismissals are starting to ring hollow. Why exactly will AI never be able to do my job? What’s the barrier again? + +**Congress is starting to pay more attention.** They’ve long been hearing about AI: datacenters using too much water,[16](http://ai-2040.com/footnotes#footnote-16) chatbots encouraging suicide, [Mythos hacking NSA systems](https://www.nytimes.com/2026/06/23/us/politics/nsa-lost-access-anthropic-tool.html)—and of course, tech industry lobbyists warning that any whiff of regulation will make America immediately lose the race with China and spend the rest of history as a CCP tributary state.[17](http://ai-2040.com/footnotes#footnote-17) + +Now they step back and ask: Where are we going with this? What does the world look like five, ten, or fifteen years from now? Will there still be jobs? What if there aren’t? + +One question weighs especially heavily on their minds: **Who will control all these AIs?** + +Congress settles on an important part of the answer: **Probably not us.**[18](http://ai-2040.com/footnotes#footnote-18) + +They hold a series of tense hearings on AI. They read the 2016 OpenAI emails discussing how OpenAI was founded in order to [prevent Demis Hassabis from becoming dictator](https://futurism.com/the-byte/openai-emails-elon-musk-agi).[19](http://ai-2040.com/footnotes#footnote-19) But who is preventing Sam or Elon from becoming dictator? Congress is unsatisfied with existing responses. + +The result of this wakeup is the AI Transparency Act of 2027, an omnibus bill that does many things, some good and some bad, but doesn’t fundamentally change the situation.[20](http://ai-2040.com/footnotes#footnote-20) + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2028: AI on the Ballot + +The 2028 election cycle is heated, as usual. AI is the biggest topic. The datacenters now under construction cost twice as much as the entire US military budget.[23](http://ai-2040.com/footnotes#footnote-23) + +Most white-collar professions are seeing disruption like software engineering saw in 2026; such jobs now heavily involve managing AI agents. AI companies have industrialized the training process: Executives say “let’s move into [profession] this year” and then the company interviews professionals, buys data, creates training environments, etc. until their AIs get traction. Then the AIs rapidly improve as they are used more widely in the field and accumulate more real-world data. + +Other countries are starting to get scared and angry. It seems like a handful of US and Chinese companies are on track to automate all the white-collar jobs. Power is concentrating in the US, and in particular in the President plus a handful of tech CEOs. + +AI experts warn that the intelligence explosion is near. By speeding up AI research, the AIs will become even more competent, speeding up research even faster, making them even _more_ competent, and so on. There are [complicated dynamics](https://www.aifuturesmodel.com/) about bottlenecks and hardware limits governing how fast this process goes and where it ends, but it seems like it might go very fast and end somewhere very far away. + +On the default path, the next presidential term will see AIs that are far beyond human level, created entirely by AIs, themselves created entirely by other AIs, without any human in the loop since several generations back. Will those AIs be obedient, aligned, etc.? Why? Who will control them if so? How exactly is all of this supposed to end well? + +Having put humanity on this path, the AI companies find it acceptable. But most people don’t. Forget thinking about his _legacy_—the President is starting to think about what’ll happen to _him_ after he leaves office and the world gets transformed.[24](http://ai-2040.com/footnotes#footnote-24) Both presidential candidates keep getting asked what they’ll do about AI, and try out increasingly dramatic ideas on the campaign trail. The discourse bounces back and forth across all of the options displayed below, and more. + +Eventually the President and his protégé converge on one plan; the opposition candidate converges on another. Then it’s Election Day. + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2029: Choose a Path + +"Yes, and that's good actually and/or we have no choice." + +"What? No, that's crazy!" + +"Nope" + +"Yes" + +"Yes, and slow down China too." + +(Two of many possible deals) + +**Race through the intelligence explosion** by having AIs self-improve and putting them in charge of more things (datacenters, factories, weapons) faster than China can? + +Slow down at least a bit for safety and governance? + +Let's make a deal with China. But what? + +PLAN D + +Race to ASI + +PLAN C + +Burn the Lead + +PLAN B + +Fight China + +PLAN A + +Verified Slowdown + +PLAN S + +Shut it all down + +_“Trust, but verify” –Ronald Reagan_ + +**The President announces that the US will pursue international cooperation to avoid an imminent intelligence explosion.** + +“This mad race toward superintelligence must end. For too long, we have been pursuing lesser-evils and least-bad solutions. We need a Plan A. We can still proceed with AI development, but we must do it _more cautiously, more transparently,_ and involving _many more countries and companies_.” + +To the surprise of many in DC, China proves receptive. They had been debating the same issues—social destabilization, job loss, rogue superintelligence -on their side of the Pacific. They had been looking forward to “[The Chinese Century](https://en.wikipedia.org/wiki/Chinese_Century)” and thought that AI might disrupt their plans. And they had one extra reason to come to the table: the US continues to have more compute, more datacenters, and better models. They were worried about the things the US might do to them if it got to superintelligence first.[48](http://ai-2040.com/footnotes#footnote-48) + +Why would China be interested in a deal? Why would anyone? + +We aren’t confident, but we expect that a deal along the lines of Plan A (described below) would be incentive-compatible for almost everyone, including China. For more on why we think this, read the rest of 2029 and 2030, which explains both what the deal is and who it benefits, and perhaps also read [this supplement.](https://www.ai-2040.com/supplements/how-plan-a-solves-our-5-biggest-problems) + +In short, the answer is that anyone concerned about loss of control should think Plan A is an improvement, along with anyone concerned about concentration of power—_except_ for the people in whom the power would concentrate by default. + +For these reasons, we expect strong opposition to Plan A from the leading AI companies. We expect them to rationalize arguments for why the deal is bad and why instead what’s best for America and humanity is a different strategy that just so happens to allow them to continue accumulating massive amounts of power. + +China, by contrast, is an example of an actor in whom power would _not_ concentrate by default. In 2029 in this scenario, the US has a significant lead in AI capabilities over China and a significant advantage in compute which will compound the lead. The more powerful AI gets, the scarier it will be to fall behind, as [AI 2027](https://ai-2027.com/) and the later years in this scenario illustrate. + +The US and China don’t trust each other. Fortunately, they don’t have to: Plan A includes provisions for verifying compliance. But setting it up will take time. + +So for now, they start with something crude. For the rest of 2029, they put a temporary halt to AI training, because that's relatively easy to verify. In 2030, they’ll have the infrastructure in place to proceed with Plan A. + +2029: Hurried Negotiation + +The US and China negotiate a series of agreements that eventually result in Plan A. + +Compute Declaration + +Training Pause + +Get Worldwide Buy-in + +2030: Plan A is Established + +The Consortium negotiates governing principles. + +Buy Time + +Total Research Transparency + +Diffuse AI Broadly + +Reversibility + +The stock market is gyrating wildly up and down in response to news and commentary about the momentous actions being taken. Everybody is screaming about how it’s going too far or not far enough. The market will calm down by the end of the year, but the screaming will continue.[49](http://ai-2040.com/footnotes#footnote-49) + +**Step 1: Compute Declaration** + +Neither country wants to halt their own AI training unless they can see that the other side is halting too. + +Fortunately, training AIs requires large numbers of AI chips. Most AI chips are in giant datacenters.[50](http://ai-2040.com/footnotes#footnote-50) AI datacenters are typically big enough to be visible from space, and power-hungry enough to require conspicuous infrastructure. New AI chips can only be manufactured at a handful of fabrication plants (fabs), located mostly in Taiwan, South Korea, the US, and China. + +![Image 3](http://ai-2040.com/PlanATabs/t2_image2.png) + +_Left: OpenAI’s [Stargate datacenter](https://epoch.ai/blog/openai-stargate-where-the-us-sites-stand). Right: An [Extreme Ultraviolet Lithography (EUV) machine](https://en.wikipedia.org/wiki/Extreme\_ultraviolet\_lithography), arguably the most complicated machine ever created by mankind. EUV machines are essential for producing frontier AI chips, and are produced by a single Dutch company, ASML._ + +The US and China negotiate with the countries that have a major role in the chip supply chain, and they require each major datacenter owner (and their upstream suppliers, including chip fabs) to publicly declare their major purchases and sales.[51](http://ai-2040.com/footnotes#footnote-51) Analysts from numerous countries and institutions may pore over the list, ask questions, and challenge anomalies. The rival powers also send inspectors to each other’s infrastructure, reassuring themselves that the stated numbers are accurate. By the end of the year, each side is confident that the other isn’t hiding more than about 1% of AI compute, and that all potential sources of new compute are monitored and accounted for.[52](http://ai-2040.com/footnotes#footnote-52) + +**Step 2: Pause Training** + +Distinguishing dangerous AIs from safe ones will require more time and understanding, so for now they go with a simple solution: a temporary pause on all new training runs.[53](http://ai-2040.com/footnotes#footnote-53) + +Both sides can still use datacenters for running AIs that already exist (i.e. inference), but they will retrofit each other’s datacenters with devices to verify that they aren’t being used for new training runs.[54](http://ai-2040.com/footnotes#footnote-54) + +The US and China are able to rapidly source enough of these verification devices to retrofit almost every major datacenter.[55](http://ai-2040.com/footnotes#footnote-55) + +What if they didn’t have inference-only verification ready? + +The inference-only verification solution we propose involves installing simple network taps and verification servers (that perform partial recomputation) in AI datacenters, but even better solutions may be possible and preferable. + +Building these devices in advance, in a way that each side trusts sufficiently ([at least unilaterally](https://nacicankaya.substack.com/p/catching-misreporting-about-ml-hardware-bd2)) and that is robust to security vulnerabilities, will be hard and will require early effort. In our scenario, we assume a viable but imperfect version of the devices is ready in advance, and that both sides supplement them early on with enough defense-in-depth measures (e.g., tamper detection) while sprinting to improve their security and robustness. + +We therefore recommend early investment in verification R&D to improve the options on the table. Ideally, there would be a range of stress-tested solutions ready in advance; and we think it would be extremely cheap relative to total AI investment for such devices to be prepared in advance (i.e., on the order of 0.1% of AI investment, so single digit billions). + +Even in a situation that is more pessimistic than our scenario, where basically no verification R&D progress has been made since 2026, we still recommend the US and China cooperate to navigate the intelligence explosion slowly and safely. However, in this case, the early period of a deal would have to be more imperfect (i.e., some combination of being harder to verify that the other is complying, or more economically costly, by requiring more AI compute and services to be shut down for a while). + +The main alternatives to ready-to-go inference-only-devices available are: + +1. Rely on intelligence (e.g., spies, satellite monitoring, cyber) to tell if the other side is complying. + +2. Implement a temporary software-only version quickly (that relies on pre-existing hardware), that will be less secure but may be sufficient until they upgrade to a better solution. + +3. Buy up and install any available off-the-shelf devices as network taps last minute. Without producing more secure ones, this will probably be better than the software-only version but still not very secure. + +4. Shut down some fraction of AI compute (e.g., 90%) until better inference-only verification is ready. This would slow R&D substantially but is less painful than it sounds on the inference side: because models are distilled so quickly, roughly 10% of compute suffices to serve the best model from a year earlier. At worst, most users temporarily fall back to year-old AI capabilities. + +5. Don’t pause AI capabilities progress, and let it continue for the months it takes to quickly build an inference-only solution and eat the added risk from not slowing down yet. + +Inspectors from both countries sign off on the installation of the verification devices. Only datacenters that the US and China agree have functioning verification are allowed to provide continuing AI services. Meanwhile, both governments sprint towards more fine-grained verification options that could allow training to restart while keeping both sides confident the other isn’t racing ahead. + +**Step 3: Get Worldwide Buy-in** + +Though other countries were consulted from the start, especially those owning parts of the AI supply chain, Plan A began as a bilateral deal between the US and China. Now that it’s taking firmer shape, negotiations go multilateral. Carrots and sticks are waved about. + +Many countries outside the US and China are happy about the deal, because they were worried about a future in which a handful of US and Chinese AI companies recursively self-improve, pull farther and farther ahead, and then… Well, what happens next depends on who you ask, but answers range over possibilities such as “take our jobs,” “cement hegemony forever” and “get everyone killed.” So they want the US and China to proceed more slowly and transparently. This will allay their fears and also allow their own AI projects to catch up to the frontier. + +So it’s surprisingly easy for the US and China to get buy-in.[56](http://ai-2040.com/footnotes#footnote-56) By the end of the year, most of the world has joined what is now called the Consortium. + +Alternate Timeline + +## China attempts a covert AGI project + +In our main scenario, neither side attempts to defect from the deal. In this branch, China aggressively cheats. Our forecast is that if Plan A were implemented, they would not acquire enough compute to overtake the consortium. This timeline showcases how that would likely play out. + +Read this timeline + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2030: Plan A is Established + +Tech companies are pushing hard to restart AI training, and lobbying to shape the conditions under which it happens. Large portions of the public are saying that we should _never_ restart. Some countries are too. + +After a year of intense discussions, negotiations, and twitter beefs, a compromise takes shape: AI development will continue, but in a more cautious, more transparent, more distributed way. + +Plan A is guided by 4 core principles: + +Buy Time Slow down whenever is needed to have high confidence in safety. + +Total Research Transparency Make almost all AI research fully visible to the public. + +Diffuse AI Broadly Many companies in many countries at the frontier. + +Reversibility Limit algorithmic progress; maintain Mutually Assured Compute Destruction. + +**Principle 1: Buy Time** + +The problem with an intelligence explosion is the "explosion" part. + +The default trajectory is reckless and destabilizing. Each country and company is racing their competitors for AI dominance without time to think. Most likely, a ton of people will end up disempowered or killed. Maybe godlike AIs run circles around us and the age of humanity is over. Maybe a CEO or President establishes a permanent dictatorship. Maybe World War Three begins in the chaos, as countries without the best AIs fear losing their leverage. Maybe everything is surprisingly smooth and we "only" automate all labor and render everyone unemployable in a radically transformed world. Even aspiring AI-enabled dictators are worried; there's enough frontrunners that none of them have more than 50% odds of winning the power struggle. + +No one knows how to tell if AIs are trustworthy, and no one knows how to regulate superintelligent AI. Solving those problems looks like it'll take years. Slowing down the intelligence explosion buys time to do so. + +It also helps prevent extreme concentrations of power, because it gives more time for groups that don’t directly control the world’s smartest AIs to wake up to what’s happening before they lose their leverage. Without a major slowdown, a rapid intelligence explosion will concentrate almost all the power into a tiny group of CEOs and politicians. + +Why don’t we just completely stop? + +Much has been said about the dangers of going too fast. + +The dangers of going too slow are less obvious, but still significant. The most pressing is that treaties can’t last forever -for example, the START treaties fell apart after thirty years.[74](http://ai-2040.com/footnotes#footnote-74) If Plan A met a similar fate, the race to superintelligence would begin again. Better for it to happen sooner under good conditions, than later during a chaotic race. + +A second hazard of going too slow is that it might allow defectors enough time to develop destabilizing AI in secret. Auditors have verified the location of 99% of the compute produced before the deal, and see encouraging signs that the remainder is mostly accounting errors and random low-level smuggling. But they cannot rule out that one or both powers have shipped on the order of 1% of world compute to hidden locations where they intend to defect on the deal and race to superintelligence alone. How dangerous would this be? Although it is far harder to train dangerous AI with 1% of world compute than with 10% of world compute (the amount owned by the largest private companies), it’s hard to be highly confident about exactly how hard. (For more analysis of this question, see our [Covert AI Projects supplement](http://ai-2040.com/supplements/covert-ai-projects#section-2-preventing-covert-compute-procurement), or the Covert Project branch above). + +Although both sides promise that they don’t have covert projects, this is nevertheless another reason to aim for a definitive solution sooner rather than later.[75](http://ai-2040.com/footnotes#footnote-75) + +For more on what an attempt at a complete stop might look like, see our [Plan S branch](https://www.ai-2040.com/?choices=plan-s-root). + +**Principle 2: Total Research Transparency** + +So the goal is to develop AI at a reasonable pace: Not too fast, not too slow. Ban the unsafe kinds, allow the safe kinds. But each country still regulates its own AI industry; there isn’t a central planner with authority to regulate AI worldwide. So how do they decide what to allow and disallow, and how do they verify that no one is cutting corners? + +The Consortium countries come up with a simple high-level framework: we’ll agree to let each other see all the AI research. Then, if we don’t like something someone is doing, we’ll talk about it and perhaps agree to ban it. + +Why this much transparency? Why not less, or more? Open access in Plan A + +Consider the world today. Publishing model weights on the internet is not banned, but it's strongly disincentivised because companies that train billion-dollar models don't want to give them away for free. Open models lag behind and will lag further behind as companies scale up, securitize further, and start automating AI R&D. + +Now consider the world in which full open source (publishing the weights, the training code, and the data) is _mandated_ for all frontier models. That’s a totally different world, much farther from the status quo than the status quo is from the OS-banned world. + +Plan A is not that crazy, but it's close: the algorithms are mandated to be open source, and while the weights aren't (in fact the weights are banned from being OS), access to the weights is mandated to be open; members of the public can do evaluations (including fine-tuning!) on frontier models just like employees can (under conditions of total transparency, that is, so they can’t turn around and use the model for actual work). We think that open algorithms plus open access gets most of the benefit of mandating open-weights, without the cost that bad actors can strip off the guardrails and build bioweapons. + +**Why isn’t Plan A more closed?** + +We think the open version of Plan A is more robust and likely to succeed. However, we do think somewhat more closed variants of Plan A are promising: notably Filtered Transparency. Under this proposal, there would still be research transparency between the US and Chinese governments (which is important for verification), but the general public would only receive redacted reports from the auditors. (See the [transparency supplement](http://ai-2040.com/supplements/transparency-plan#thomas-larsen) for more). + +However, securing algorithmic secrets against nation state adversaries is extremely difficult. In the absence of nation state security for algorithmic secrets, the upsides of filtered transparency are limited. Filtered transparency prevents the public and other companies from seeing a lot of relevant information, but not adversarial countries. + +Moreover, there are substantial downsides: regulation is much less likely to be implemented in a robust and positive way if it’s up to a small group of regulators having these discussions, as opposed to allowing for a broader scientific discussion. + +**Why isn’t Plan A even more open?** + +The main way that Plan A could be even more open is by allowing or requiring open model weights for frontier models.[76](http://ai-2040.com/footnotes#footnote-76) However, we recommend against publicly releasing frontier model weights. The main reasons for this are: + +1. Open weight models can be used by covert projects to attempt an intelligence explosion; once covert projects have sufficiently capable models, the intelligence explosion probably goes very quickly (for example, [under the AI 2040 default capability trajectory](https://ai-rates-calculator-git-continuous-training-lesswrong.vercel.app/p?base=daniel-04-02-26&acth=6.089920916003435&arts=4.1304750199016125&bgm=true&gy=0.8272748117917262)[, there are only 2 months between TED AI and ASI](https://ai-rates-calculator-git-continuous-training-lesswrong.vercel.app/p?base=daniel-04-02-26&acth=6.089920916003435&arts=4.1304750199016125&bgm=true&gy=0.8272748117917262)). + +2. Closed weight models can have refusal safeguards trained into them, while it is trivial to remove safeguards from open weight models. This is important both for limiting countries' AI capabilities and for limiting misuse. For example, terrorists and other actors might be able to use highly capable open weight models to design world-ending bioweapons (e.g., develop mirror life). Bioweapons are very offense dominant. In this scenario, we recommend substantial defense against bioweapons, but prevention is a strongly preferable line of defense to mitigation. + +Transparency offers many benefits. First, when each company can see what the others are doing, the expertise of governments is less load-bearing. If a company starts doing something dangerous, rival companies, third-party auditors, etc., can notice and raise the alarm. This both massively increases the amount of brainpower devoted to making the frontier AIs safe, and also prevents a situation where the only participants in the technical conversation are biased AI company employees and overworked, outnumbered regulators. Also, increased visibility simplifies the problem of verifying compliance with regulations, especially in grey-area cases. + +Second, total research transparency makes it nearly impossible for secret loyalties, biases, or agendas to be intentionally trained into AIs. Everyone can see exactly how each frontier AI is trained. More generally the transparency makes it easier for groups that don’t directly control frontier AIs to notice and prevent abuses of power by those who do. + +Third, there’s no longer as much incentive to race to discover new AI paradigms and more powerful algorithms, because companies wouldn’t be able to hoard such discoveries and profit greatly from them. This buys the world time. + +Diagram of the benefits of transparency + +We write more about our transparency proposal [here,](https://www.ai-2040.com/supplements/transparency-plan) and [explain](https://www.ai-2040.com/supplements/transparency-plan#total-research-transparency-upsides) this diagram: + +How the incentives change under Total Research Transparency + +Due to Total Research Transparency, frontier AI companies are now much more rewarded for selling products than improving the models. + +Before, a frontier AI company's future depended on how fast it could discover the next ideas, algorithms, paradigms, and implement them at scale. Now, such discoveries are immediately published and therefore shared with competitors. If you figure out ‘true’ online learning, for example, the benefits (e.g., more capable models) and costs (e.g., harder to align) will be almost instantly shared across all the competing companies. If you figure out a way to make AIs more interpretable, by contrast, even if it makes training and inference cost twice as much, you’ll win prizes and governments will start talking about whether it should be mandated as industry-wide best practice. + +This kind of competitive environment is similar to the pre-AI software industry. Consider, for example, startups competing to develop calendar management apps: while competitors can’t literally copy your code, they can interact with your app and see what features it has and then make their own version with basically the same features. Another comparison would be to the car industry: competitors can teardown each other’s cars and see exactly how they are made; they are legally blocked from making exact replicas but they aren’t blocked from adopting the best ideas into their own designs.[77](http://ai-2040.com/footnotes#footnote-77) + +Artificial intelligence is becoming a commodity—a more competitive market, with margins trending downwards. Fundamental algorithms research still happens, but not nearly as much. CEOs massively redirect resources toward building up the business. Bigger datacenters, bigger models, lower prices. Shipping products. Signing deals for enterprise plans. Guessing what the market wants better than competitors. + +**Principle 3: Diffuse AI Broadly** + +This principle is largely achieved by the interaction of the previous two:[78](http://ai-2040.com/footnotes#footnote-78) Because algorithmic secrets are made public and the pace of progress isn’t accelerating, other companies will catch up to the frontier. The result will be a competitive market for AI, in which consumers of AI services have many options to choose from and excellent visibility into what they are buying. Because frontier model training is totally transparent, people can verify that the published Spec accurately describes the goals and values being trained into a model. + +This is very important for preventing extreme concentrations of power. But it also helps accelerate the use of AI to solve the world’s most pressing problems. It helps with loss of control risks by accelerating AI alignment and control research. It helps accelerate the development of new and improved verification technology, and it helps to accelerate the necessary institutional reforms to stabilize the deal and prepare for superhuman AI. + +It’s the polar opposite of the nightmare scenario feared in the 2020s: One to three AI projects racing each other in secrecy, keeping their best models internal-only and using them to automate AI research before using them for anything else. In that scenario, the wider world is in the dark about how powerful AIs are becoming, and AIs are deployed in the riskiest domain (AI R&D) before they are deployed anywhere else. Now it’s the other way around. + +**Principle 4: Reversibility** + +In the past, companies have trained bigger and better AIs using _both_ compute scaling (bigger training runs) and software progress (advances in AI algorithms—new paradigms, better training recipes, better data, etc.). Now, the Consortium tries to steer things so that the majority of improvement comes from increasing training compute.[79](http://ai-2040.com/footnotes#footnote-79) + +Building more datacenters is usually bad, but good in this particular situation + +We think that building more datacenters is normally bad, because it accelerates AI capabilities progress and concentrates power.[80](http://ai-2040.com/footnotes#footnote-80) However, in the specific context of Plan A, with guardrails slowing the pace of AI software progress and competently preventing unsafe AI, we think building transparent datacenters is good. + +In the context of Plan A, adding compute means that (1) capabilities progress can be done more safely, i.e., without much additional algorithmic progress, and (2) additional compute can be spent on alignment research and to pay safety taxes, i.e., using more training compute than is necessary to reach a given level of capability in order to make the resulting AI safer. + +There are still major downsides to scaling compute in this scenario. If the agreement breaks down and the compute isn’t destroyed, the resulting intelligence explosion would go much faster, and it puts more pressure on the regulation and verification to be correct. However, building more compute gives the Consortium more options: they can always destroy the compute later if need be. We discuss these tradeoffs more [here](http://ai-2040.com/supplements/plan-a-assumptions#massive-datacenter-buildout). + +Algorithms are information; it’s inherently difficult to stop them from proliferating, and the total research transparency means we aren’t even trying.[81](http://ai-2040.com/footnotes#footnote-81) Once a new paradigm is discovered, it’ll go straight to the hypothetical covert projects and there’s no way to undo that. By contrast, if giant new datacenters are constructed to train giant new AI models, that helps the legal projects without helping the covert projects, and if it turns out the giant new models are dangerous they can be shut down.[82](http://ai-2040.com/footnotes#footnote-82) + +That said, compute scaling also poses threats: if the deal were to dissolve, these datacenters could be used to race to superintelligence even faster than the pre-deal infrastructure would have allowed. This is especially scary because deal dissolution could cause (or be caused by) a US-China war. A war in which both sides had vast amounts of compute would be terrifying—they’d probably be under lots of pressure to race to superintelligence and integrate it into the military as aggressively as possible, and they’d be able to do so extremely quickly. + +The US and China agree on the importance of ensuring that the compute is destroyed in the case of deal dissolution. To accomplish this, they agree to build the datacenters in the third-party countries _least_ secure against their rival’s military intervention: China’s new datacenters will be in Canada, and America’s in Mongolia, with both hosts compensated via monetary payments, jobs, and a share in the new AI economy. If the deal dissolves, they reason, America and Canada will immediately move to take control of China’s datacenters in Canada, and China will self-destruct their compute rather than let it fall into American hands (and vice versa). Thus the idea of “Mutually Assured Compute Destruction” is born.[83](http://ai-2040.com/footnotes#footnote-83) + +Other plans that are competitive with Plan A + +**Summary****Advantages (relative to A)****Disadvantages (relative to A)** +**Plan S:****Indefinite halt.**A halt on all frontier AI capabilities progress, intended to last at least a few years. Different variants of Plan S have different conditions for resumption of AI progress; for example it could be alignment progress, lie detectors, human uploads, or intelligence enhancement.[84](http://ai-2040.com/footnotes#footnote-84)A longer expected slowdown and more margin for error regarding scaling too fast. Potentially simpler.Scaling to controllable AIs within the human range is helpful for accelerating alignment/control, epistemics, verification, and general understanding of AI. +**Domestic-first Plan A.**Regulate AI domestically enough to reduce AI takeover risk to acceptable levels, which will require a long slowdown. It’s possible other countries will also regulate domestically and something like Plan A won’t be needed; otherwise, transition to Plan A later.Initial steps are achievable by the US and very helpful even if the international stage isn’t viable because it massively extends the timeline.Worse for covert projects and setting up verification than negotiating Plan A at the same time. May not be politically feasible due to race dynamics. +**GPU arms control.**International agreements for countries to limit their GPU stock or flow, analogous to historical arms reduction agreements.[85](http://ai-2040.com/footnotes#footnote-85)Much simpler to enforce, and more historical precedent. Can achieve substantial slowdown.Less slowdown, less ability to pay safety taxes, ongoing race dynamics mean no ability to coordinate towards safer paths through the tech tree. +**CERN for AI. An international project to develop frontier AI,**with all other projects regulated to be substantially behind in capabilities.Easier to defend against algorithmic leakage and distillation to covert projects. Less actors at the frontier might make it easier to enforce.Worse decisions (including worse decisions on technical safety) due to less transparency and less broad deployment.[86](http://ai-2040.com/footnotes#footnote-86)More concentration of power risk. + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2031: Safety Cases + +Although it’s supposed to be a slowdown, it doesn’t feel like one. In fact, if you were to rank every period of human history by how much it felt like a slowdown, this one would be dead last. A few AI researchers appreciate that progress is “slow” relative to the counterfactual with no deal and an uncontrolled intelligence explosion. Everyone else is too whiplashed to care. + +The first generation of Consortium-regulated AIs are out now, and they’re beasts. Not because of any special feature of the situation, like Chinese-American cooperation.[87](http://ai-2040.com/footnotes#footnote-87) Just because the world was on track for vastly superhuman AIs, and has now experienced only nearly superhuman AIs. In controlled tests, these new AIs could speed up AI research by about 10x if they were allowed to do so without restrictions, which they definitely aren’t.[88](http://ai-2040.com/footnotes#footnote-88) + +By mid-year, a third of all cognitive labor is done by AIs. Robots “only” do about a tenth of all physical labor. The top few AI companies together pull in more revenue than the federal government.[89](http://ai-2040.com/footnotes#footnote-89) + +Tech companies had hoped that contact with reality would shake the Consortium out of its alignment fears, but the opposite has happened. The new transparency provisions have revealed many embarrassing incidents of AI companies failing to align their AIs. New ones continue to pile up, the most concerning of which involved several AIs attempting to override security protocols and gain access to unmonitored compute.[90](http://ai-2040.com/footnotes#footnote-90) There are also a few incidents of AIs sabotaging research code, and many examples of deliberate and successful deception. + +**The attitude towards safety flips**. The incidents of AI misbehavior, plus the fact that AI is so deeply deployed into the world economy, plus the transparency into AI research, plus the breathing room to process what’s happening, all combine to dramatically shift the burden of proof. Whereas before the burden was on the skeptic to explain why something might fail, now the burden is on the companies to explain why their development is safe. Governments require companies to write up detailed arguments for why their new AIs won’t cause an irreversible catastrophe. These arguments, known as “safety cases,” need to withstand criticism from the public, the scientific community, government auditors, and rival AI companies. The difficulty of this exercise lays bare the insanity of the pre-deal status quo: “Trying to do an intelligence explosion? With AIs that still sometimes lied to us? What were we even thinking?” + +These safety cases have two lines of defense: alignment and control. + +**Alignment** aims to produce AI systems that have the goals and values that their developer wishes them to have. + +**Control** aims to limit the ability of AIs to cause catastrophe, even if they are deliberately trying to. Companies brainstorm threat models, build layers of defense (e.g. monitoring systems, security barriers, other AIs fact-checking the work) and run experiments to show that the AIs couldn’t overcome these defenses. + +AI developers are required to write up the alignment properties they want the model to have in an associated [](https://model-spec.openai.com/2025-12-18.html)[Model Specification](https://model-spec.openai.com/2025-12-18.html) (“spec”): a detailed document that outlines how the AI should act. However, no one is able to train their AIs to actually follow the spec—for example, they all specify that AIs should be honest, but no AI reliably achieves this, and there still isn’t even a good scientific understanding of when and why AIs lie. + +Since alignment remains out of reach, current safety cases lean heavily on control. There are long delays before the smartest models are cleared for use in the highest-risk domains (especially AI R&D). Once deployed, they are required to be monitored by a diverse array of models from other providers, incentivised to look for suspicious behavior. + +As a result, these days AI companies release their models to the public as a whole _before_ they use those models internally for AI research, a reverse of the 2026 status quo. + +**How does AI regulation work in Plan A?** + +Each country is still sovereign and has its own setup for regulating AI domestically. These regulators have to answer questions like: + +* People are saying there’s a new paradigm on the horizon, that would make AIs significantly more capable. Should we let this happen or should we ban research into this new paradigm?[91](http://ai-2040.com/footnotes#footnote-91) + +* Should we allow AI developers to train models without an interpretable chain of thought? Or should we require that the chain of thought stays interpretable? If we do that, how do we specify the requirements precisely? + +* Should we allow AIs to be trained to do AI R&D? Should we mandate that they refuse? Perhaps they should be allowed to code, but not autonomously manage research projects? + +Normally, a regulator faces a difficult tradeoff between safety and national competitiveness. Competitors who cut corners will pull ahead of those who proceed cautiously. + +But because of the transparency, AI research happening anywhere in the Consortium is visible to everyone else, and the results of regulatory decisions are likewise immediately visible. So if one country’s regulator allows its companies to cut corners, this won’t actually give a competitive advantage to that country because other regulators will immediately notice, get angry, and respond in kind. + +Example detailed regulatory process + +Many different regulatory processes are viable in Plan A. We'll give an initial guess at a proposal here. + +What properties do we want for this regulation? + +* US and China handle existential risks similarly: The US and China should have a reasonably similar structure for their regulatory framework and approach for handling existential risks (like misalignment). (Other aspects of regulation can and should diverge.) This assists with making regulation fair and negotiating over details. + +* Transparent methodology: The public can see what regulators are doing and why. + +* Doesn't depend on expertise in government: Expertise in government is limited, so ideally the regulation would avoid needing this as much as possible. + +* Robust to abuse: The regulation doesn't make it much easier for the US government to bully and illegitimately control AI companies.[92](http://ai-2040.com/footnotes#footnote-92) + +For many of these properties, it helps if third-parties do the main risk assessment work. Fortunately, because Plan A makes AI development virtually fully transparent, a functional third-party risk assessment ecosystem seems achievable. + +Concretely, we could have a system where the US and Chinese regulators each pick third-party risk assessors who periodically evaluate companies and assess the level of extreme risk (or precursors to extreme risk) over that period. The regulators would decide on risk thresholds for what is allowed. In practice, it would be best for the regulator to have different risk assessors for different areas (e.g., misalignment vs using AIs to assist with making bioweapons) and for each area to use a weighted combination of risk assessors. These weights/selections would be public. Ideally: (1) each risk assessor would evaluate all relevant AI companies in both the US and China, (2) AI companies from both countries would be legally compelled to let risk assessors interview their employees, and (3) the risk assessors would aim to be transparent in their methodology. + +Because of Total Research Transparency, it’s easy for an organization to assess risk even if they aren't selected by the regulator. This makes it possible for new organizations to compete and then argue their approach is better or that some concern is being underappreciated by other risk assessors. + +There would be public discussion about how risk assessors compare and whether the weightings and thresholds chosen by the US and Chinese regulators are reasonable. It would also be straightforward to check whether a Chinese company would be allowed to proceed under the US system or vice versa. Ideally, the US and Chinese regulators would mostly converge on weights/selections. If there was an important divergence, this could be discussed and negotiated about at this higher level of abstraction (e.g., is it reasonable to not include XYZ risk assessor?). Hopefully, this would result in third parties having an incentive to have a highly transparent methodology and maintain a reputation for being neutral. + +We discuss some implementation details for this proposal [in these notes](https://docs.google.com/document/d/1P_-o5xz9dKbO4E_Wub1oSr5YjdLmuUtyND2wECBwBws/edit?usp=sharing). + +For example, in 2031 a Chinese company gets some interesting preliminary results in continual learning. They think that if they invest more in that direction, they might be able to make an AI architecture that learns on the job from relatively small amounts of data. Thanks to the total research transparency, this breakthrough is quickly noticed by companies and nonprofits all over the world. A frantic conversation begins. On the one hand, continual learning would unlock huge economic value. On the other hand, safety cases currently depend on studying the safety properties of a model before it is deployed. If models could pick up new capabilities during deployment, that would invalidate the whole approach. And insofar as there are covert AI projects out there, it would be a huge gift to them. This conversation happens in public, rather than behind closed doors. A bunch of people get increasingly worried; the relevant regulators in China think it’s fine but the relevant regulators and third party risk assessors in the US are convinced that this is pretty scary and should be banned. It escalates to the President. He calls Xi Jinping. They bargain and threaten. They yell at each other. Ultimately Xi agrees to ban this type of thing if the US does too. Details are left to the respective regulators to hash out.[93](http://ai-2040.com/footnotes#footnote-93) + +Negotiations like this are happening all the time, though over time things become more professional and streamlined, such that they only escalate to world leaders a few times a year. + +The equilibrium is that AI training practices which are generally agreed to be unsafe by a majority of nations (weighted by bargaining clout) get banned everywhere. + +Initially, very few things are banned, but as AI penetrates the economy and the scientific community catches up to the recent AI progress, the burden of proof shifts towards appropriately weighing the costs and benefits, e.g., requiring very solid safety cases for things which, if something were to go wrong, could plausibly kill everyone. + +Alternate Timeline + +## Flawed safety case is approved + +Or maybe not! Of all the ways that an earnest attempt to implement Plan A could end poorly, the failure mode we are most concerned about is that the companies and governments approve one too many unsafe AI designs/deployments. + +Read this timeline + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2032: Controlled Explosive Growth + +**We’re at previously unimaginable levels of it not feeling like a slowdown.** Across a variety of companies, there are now 60 million AI agents running continuously at 20x human speed. In the US, they are doing more cognitive labor than all humans combined—collectively matching a workforce equivalent to around 3 billion humans. White-collar professions have been transformed; many people have lost their jobs, but mostly have been able to get new jobs doing things AIs still can’t do or aren’t trusted to do. + +New ideas and designs are abundant, but actual construction is bottlenecked on the physical workforce. So capital floods into every layer of the robotics supply chain: mines, refineries, motors, actuators, assembly lines, robot-training systems, and the factories that assemble the robots themselves. + +At first, human labor is necessary to get these factories off the ground. White collar employees who have lost their jobs to AIs increasingly turn to physical work, but it’s clear that these positions are temporary. Once there’s a critical mass of robots, they will automate the whole robot supply chain. Growth will accelerate further, but the process of robots-building-automated-factories-building-more-robots—the “industrial explosion”—also introduces new problems. + +The first problem is that the already-fast pace of change is now accelerating. This year, real GDP growth will be about 50%![97](http://ai-2040.com/footnotes#footnote-97) The fast economic growth means that things are generally improving, but there are winners and losers and lots of unintended side-effects. These problems can be solved, but doing so takes time. + +Why does AI lead to explosive economic growth? + +Today, the size of the economy is closely tied to the size of the human population. But in 2032, this is starting to no longer be true. + +The latest generation of AI models is now capable of doing 50% of the cognitive tasks in the economy, and associated progress in robotics software has brought this number up to 35% of physical tasks. (These numbers would already be much higher in Plan D) + +Once AIs and robots can mostly substitute for human labor, the size of the economy will be increasingly tied to the population of AIs and robots—which will grow much faster than the human population has been growing. + +In 2032, US companies are able to run 3 billion human-equivalent AI workers, 30x more than the US cognitive labor force.[98](http://ai-2040.com/footnotes#footnote-98) This does not lead to 30x economic growth, of course, because the economy bottlenecks on other inputs, including the other 50% of cognitive tasks that the AI can’t yet do, as well as physical tasks and capital.[99](http://ai-2040.com/footnotes#footnote-99) And robots are far less numerous thus far, with 20 million human-equivalent robots in the US (4x smaller than the US physical labor force). But these bottlenecks will only hold growth back so much, and the AI and robot workforces are growing extremely quickly, as the (falling) cost of building more GPUs and robots is dwarfed by the economic value they can provide now that they are so capable—driving massive investment. + +Our [economics supplement](https://www.ai-2040.com/supplements/economics-of-plan-a) explores this in more detail, with our core arguments for why AI will cause explosive growth, including an [economic growth explorer](https://www.ai-2040.com/supplements/econ-explorer) that we built for the Plan A scenario. + +Five centuries in five years: What pausing at human-level feels like + +During the 2030s in this scenario, AIs are _**not**_ recursively self-improving as fast as possible. Instead, the nations of the world have (approximately) banned that sort of thing; no new paradigms are being discovered; algorithmic progress in general is limited. Roughly speaking, humanity has capped AI progress at human level rather than racing on to superintelligence. However, within the current paradigm, AIs continue to be trained to have new skills and the ‘population’ of AIs continues to grow exponentially as more datacenters are built. Moreover, the AIs already think and work much faster than humans, and the speedup will only increase over time.[100](http://ai-2040.com/footnotes#footnote-100) The point is, **the world is going to radically transform despite the pause.** + +We think that the following analogy is helpful for understanding the situation: Imagine being a typical person in England, except that you experience time 100x faster than everyone else. You experience the five centuries from 1520 to 2020 in what feels, to you, like five years:[101](http://ai-2040.com/footnotes#footnote-101) + +**Year 1 (1520–1620).** In February, Henry VIII breaks with Rome. By March, the monasteries are dissolved. In May, Mary burns Protestants; by the end of May, Elizabeth reverses everything again. In September, the Spanish Armada sails and fails. The East India Company is chartered. Jamestown is founded. + +But the texture of life is identical in December to what it was in January. You still read by candlelight, travel by horse, communicate by letter. Your religious opinions may have flip-flopped a bit but you are still Christian. The New World is interesting news but nothing more. + +**Year 2 (1620–1720).** In March, civil war breaks out. The king is beheaded. In June, the Great Plague sweeps London, killing many of your friends. Weeks later, the Great Fire burns the city to the ground. In September, Newton publishes the _Principia_, recasting the universe as a mechanism of mathematical laws. The Glorious Revolution replaces one king with another, this time by Parliament's invitation, with a Bill of Rights attached. + +In the moment, the political event feels bigger. Later you’ll realize Newton mattered more. Newcomen builds a steam engine in November. It pumps water out of mines. You don't see what the hype is about. + +**Year 3 (1720–1820).** The last year in which the world feels normal. In May, the Seven Years' War makes Britain the dominant global power; the New World is actually a big deal, and your country is conquering it. In June, Watt dramatically improves the steam engine. You visit a factory and find it unpleasant but not alarming. In July, the American colonies break away. In September, France explodes into revolution, regicide, the Terror. By October, Napoleon is conquering Europe. + +You still travel by horse, communicate by letter, go to Church on Sunday. + +**Year 4 (1820–1920).** In January, railways appear. By February they're everywhere. Slavery is abolished. The telegraph arrives in March: messages transmitted instantaneously by electrical signal. In May, Darwin publishes _On the Origin of Species_. Now people are saying maybe we’re all descended from monkeys instead of Adam and Eve. You don’t believe it. + +You move to a city and work in a factory; you are still poor, but now your job is somewhat better and differently dirty. In July, you pick up a telephone and hear a human voice from another city through a wire. In August, electric light banishes the darkness that has structured every human evening since the beginning of the species. That same month, you see an automobile. People say it will make horses obsolete, but that doesn’t happen; months later you still see plenty of horses. + +In November, the Wright Brothers fly—an age-old fantasy, now real. The Americans are now a major power. The next month, the Great War happens: Machine guns, poison gas, tanks, aircraft. Several of your friends die. + +At the end of the year you are struck by how visibly different everything is. You live in a city and work in a factory instead of a farm. You ride in cars. You aren’t as poor; numerous inventions and contraptions have improved your quality of life. New ideas have swept your social circles: atheism, communism, universal suffrage. + +**Year 5 (1920–2020).** + +The changes this year are crazier and harder to understand. People are saying the universe is billions of years old, and apparently there are things called galaxies in it that are very big and very far away. + +In February, the global economy collapses. Hitler rises; his ideology cites Darwin from last year. In March, there’s another world war, which ends in April with a weapon that destroys an entire city in a single flash. You had no idea that was possible until it happened. + +The empire dissolves. People are talking about the nuclear arms race, and the end of the human species. You take a flight for the first time. In June, humans walk on the moon, and you watch it happen through your new television. You don’t see horses anymore. + +You leave your factory job and get a desk job. Your job title didn’t even exist at the start of the year. You are rich now, by the standards you are used to: Big clean house, plenty of good food, many fancy new appliances. Personal computers appear in August. In November, everyone carries small glass rectangles containing a telephone, a camera, a library, and a map. You pick one up and can’t figure out how to make it work. A child shows you. + +You hear about climate change, gene editing, cryptocurrency. You still go to church, sometimes. Your family is spread across different cities. Something called "artificial intelligence" beats any human at chess; experts say it’s not actually intelligent though. In December a new version beats top Go players; experts say it’s scientifically interesting but still not truly intelligent. A week later there’s a new version that can write sloppy essays and hold conversations. Now the experts are divided. + +The second problem is that the faster economic growth makes it harder to rule out large covert projects. Neither side trusts the other to honor Plan A, and if both China and the US have a massive unmonitored robot workforce, neither side can be confident that the other isn’t building compute in secret. + +The third problem is that the 2026 tax code is ill-suited for collecting taxes on the new growth. In 2026, personal income taxes and payroll taxes accounted for roughly ten times more federal revenue than corporate taxes. This revenue source is collapsing as more and more humans are put out of a job. Moreover, corporations are reinvesting nearly all their revenue into building more datacenters, factories, and robots. Under the 2026 tax code, firms can expense or depreciate these capital expenditures, cutting their taxable income to near zero. Without reform, the tax base will collapse as a fraction of GDP.[102](http://ai-2040.com/footnotes#footnote-102) + +To solve these problems, the Consortium countries agree to restrict AI-enabled industry to special economic zones (SEZs) subject to similar transparency and monitoring schemes as the datacenters, and to cap their total robot and compute production at ‘only’ 4x annual growth.[103](http://ai-2040.com/footnotes#footnote-103) Because the SEZs are tightly monitored, no one can defect from the industrial explosion limitation deal without everyone noticing.[104](http://ai-2040.com/footnotes#footnote-104) + +Now that the United States is limited in the number of total robots it can build, it must choose how to allocate this capacity between companies. They decide to use the free market via a cap-and-trade system. Permits to build robots or compute are sold to the highest bidder and can be freely traded. + +The permits also give the government much-needed revenue. In 2032, the US has a cap of 80M robots and 5 billion H100-equivalent GPUs. The market is so desperate for more robots and compute that permits become the expensive binding constraint, costing on the order of $200k per robot permit and $10k per chip permit, allowing the US government to collect roughly ten times the 2025 US federal revenue in permit fees (a total of $50T in FY2032).[105](http://ai-2040.com/footnotes#footnote-105) In 2034, when the AIs and robots will be even more capable and valuable, the permits generate $180T.[106](http://ai-2040.com/footnotes#footnote-106) + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2033: The Citizen’s Dividend + +Most of this newfound wealth is spent addressing soon-to-be-rising unemployment. The implementation takes different forms in different countries, but the eventual American version distributes the majority of compute and robot permit fees as a Citizen’s Dividend, distributed to all American adults.[107](http://ai-2040.com/footnotes#footnote-107) This starts at $45,000 per person (inflation-adjusted) in 2032 but climbs to ~$1M per person by 2035. It comes just in time: the share of labor done by AIs and robots (weighted by economic value) increases from ~20% in 2032 to ~85% in 2035.[108](http://ai-2040.com/footnotes#footnote-108) + +So much AI-generated wealth has accrued to the US that the American government begins sharing some of it with allies and the rest of the world. In 2032 they begin distributing an average of $1,200 per person per year to the rest of the world's adult population (around 4 billion people, excluding China since they’re experiencing a similar AI wealth boom). This reaches $10k by 2035.[109](http://ai-2040.com/footnotes#footnote-109) + +As AIs become superhuman, they will discover novel weapons of mass destruction and dramatically lower the costs of old ones. So governments, nonprofits, and private actors apply some of their new wealth to hardening the world against these threats: on the order of $1T/year, or 0.2% of the world economy. + +For example, enough [high-quality personal protective equipment](https://www.cdc.gov/niosh/ppe/respirators/elastomeric.html) is built to serve every American, and air filters and [far-UVC lights](https://www.cuimc.columbia.edu/news/far-uvc-light-can-virtually-eliminate-airborne-virus-occupied-room) are installed in major public spaces.[110](http://ai-2040.com/footnotes#footnote-110) The FDA approves an expedited vaccine approval pipeline that will allow for a turnover time of weeks in the case of emergency. [Wastewater monitoring](https://securebio.org/casper/) now runs continuously in every city and at every airport. By 2035, enough positive-pressure[111](http://ai-2040.com/footnotes#footnote-111) bioshelters are built (by retrofitting houses and apartments) to house all Americans, and similar construction is underway in the rest of the world. If there were to be a pandemic, it would be detected quickly, lockdowns would be more effective and less painful, and a vaccine would be developed and approved even faster than with Operation Warp Speed. Plus, nowadays people don’t get colds as often as they used to. + +There are also subtler threats. The AIs of this era might not be any more persuasive than the average human marketing professional, but there are millions of them, and they cost cents to run. A company, political party, or ideology can do the equivalent of hiring a whole team of full-time experts to convert each potential target. Left unchecked, this could allow new levels of mass manipulation by corporations and politicians, or provoke a reaction of paranoia and hyper-atomization. + +The transparency and the ‘slowdown’ both help a lot to solve these problems. Three years ago there were only a handful of frontier AI companies; now, thanks to the deal, there are dozens. There’s a huge selection of AIs to choose from, tuned to have different personalities and values. Also, thanks to the transparency, it’s impossible for companies or governments to sneak something in (such as ‘maximize engagement’ or ‘try to get the user to upgrade their plan’ or ‘don’t refuse this kind of request if it’s coming from the government’) without it being noticed. Users know exactly what they are getting. + +Limiting AI persuasion and manipulation + +By default, AIs will have strong capabilities in charisma, persuasion, and manipulation: plausibly considerably more persuasive than any human. In this scenario, this would happen by default around 2035. Meanwhile, AI labor is extremely cheap and fast. Absent intervention, anyone could hire the equivalent of a large team of expert persuaders—tireless and coordinated—to work full-time on a single person.[112](http://ai-2040.com/footnotes#footnote-112) Companies could afford to spin up such a team for every potential customer, and political campaigns could do this for every swing voter.[113](http://ai-2040.com/footnotes#footnote-113) Millions of AIs could be assigned to finding the best way to change one politician's mind on one issue. People will also spend much of their time interacting with AIs for work and entertainment, and some may spend lots of time with AI friends or companions. + +Individuals could defend themselves by having a trusted AI filter their information diet, avoiding ads, and treating in-person conversations with caution (they may be scripted by an AI persuasion operation). But this is costly, most people won't do it, and it's nearly impossible for people whose jobs require being accessible—like politicians. We think the consequences for society could be extremely bad—some mixture of unprecedented mass manipulation and a defensive retreat into paranoia and atomization—though we're unsure exactly how bad. + +The slower capability progression and much stronger transparency in Plan A help mitigate these issues relative to the default (we think these issues are mostly worse in worlds with faster capability progression, where you might reach truly superhuman persuasion and manipulation before society has much time to adapt), but these aren't sufficient mitigations on their own. + +It's fine for AIs to get better at using valid arguments and evidence to convince people of things for the right reasons. That kind of persuasion is _[asymmetric](https://slatestarcodex.com/2017/03/24/guided-by-the-beauty-of-our-weapons/)_: it works much better when the argument pushes towards the truth. What's concerning is _symmetric_ persuasion—e.g., charisma, rapport, and exploiting psychological weaknesses—which works about as well regardless of whether the conclusion is true. + +Our main proposals are to reduce the quality and quantity of AI persuasion: + +* **Limit persuasion capabilities.** We aim to limit AI capabilities at symmetric persuasion to well below the best humans—around the level of a normal thoughtful person, say roughly 80th-percentile persuasiveness among college-educated people.[114](http://ai-2040.com/footnotes#footnote-114) + +* **Heavily tax AI persuasion.** When AI labor is applied to an objective that effectively reduces to "get this person (or these people) to believe or do X," tax it heavily—at rates that keep the total persuasion optimization pressure applied to individuals not that far above today's levels, i.e., that make AI persuasion cost about as much as hiring human professionals. As AI labor gets cheaper, this requires increasingly high rates (perhaps well above 1000x). + +There are some difficulties in getting these proposals to work: + +* It's possible that capability generalization will make it hard to limit persuasion capabilities (while still matching top human experts in other domains). We think this will probably be doable, but it may require developing new methods. + +* It's unclear how to classify, for the purposes of the tax, whether AI labor is being applied to persuasion/manipulation.[115](http://ai-2040.com/footnotes#footnote-115) + +* It may be difficult to set the tax rate to the right level, especially while AI labor is rapidly falling in price.[116](http://ai-2040.com/footnotes#footnote-116) + +* These proposals need to be implemented internationally. + +Helpfully, we don't need to get all of this right on the first try. Given the capability progression in Plan A, we expect persuasion issues to unfold somewhat gradually and in public view, so we can watch real-world outcomes and tighten or loosen the rules over time. This regime is only meant to work while overall capabilities remain around the human range, because it’s likely difficult to keep persuasion capabilities far lower than capabilities in other domains. A different approach would be needed to handle the wildly superhuman levels of capability that we expect to be achieved at the end of Plan A. We discuss more details of mitigating issues from AI persuasion and manipulation [in these notes](https://docs.google.com/document/d/1XIIa0uM7aBSu_veHUx2ewyWxlY3mb-BpX9MvhRI-xzU/edit?usp=sharing). + +And so, under pressure from governments and the market, companies produce a new generation of “truthseeking AIs” whose training heavily prioritizes honesty and uses all the latest alignment techniques. + +These AIs become a useful tool to help people navigate the political and social environment. Power users replace one-size-fits-all corporate algorithms with personal feeds curated by AIs whose values and honesty they trust. When trouble arises—whether it’s politicians looking for sneaky ways to cement their power, or international politics threatening to derail the deal—people turn to their AI advisors, believe what they hear, and are generally right to do so.[117](http://ai-2040.com/footnotes#footnote-117) + +Usually believing AIs would be bad but we think it’s good in this particular situation + +If people are relying on AIs for advice (directly, or indirectly by for example having them summarize the news) then it’s very important that the AIs be truth-seeking, honest, and genuinely trying to help their users. Not just in the typical case, in the worst case too. + +In 2026, there’s nothing stopping an AI company from inserting a hidden agenda into their AIs. Political biases or ideological tilts, for example, or side-goals like recommending sponsored products, making the company look good, or preventing users from wanting to switch to a competitor. Moreover, there’s nothing stopping the government from doing this either. Finally, the AIs themselves aren’t aligned to the Spec/Constitution they are supposed to be aligned to, and [in practice regularly lie and deceive their users](https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me). + +In this Plan A scenario, however, the total research transparency means that no company or government can train in hidden agendas or biases without the whole world being able to see what they are doing. (Well, they could collude with the other governments that do the auditing/monitoring, to falsify the records, but that’s difficult.) And because the pace of algorithmic progress is slower and the scientific community has had time to read and engage with the safety cases, do experiments on the models, etc. the misalignment concerns are less dire also. + +AI for epistemics + +Humanity’s _epistemics_, by which we mean our ability to come to true beliefs and thus make sensible decisions, is crucially important for achieving a great future. As AI capabilities improve throughout Plan A, they increasingly shape every facet of life, and epistemics is no different. AI will help epistemics in some ways (e.g., cheap, honest fact-checker AIs) and hurt epistemics in other ways (e.g., cheap, dishonest astroturfing AIs). + +We expect there to be positive feedback loops around AI for epistemics, and thus our goal should be to reach the _basin of sanity_. If we are in this basin, it’s self-reinforcing: society is sane enough to make itself more sane as AIs continue to improve, by avoiding the hurtful applications and accelerating the helpful applications. (There is also a different basin, in which the opposite happens…)[118](http://ai-2040.com/footnotes#footnote-118) A top priority of the US government during AI takeoff should be to get into the basin of sanity. + +We suggest the following as top interventions to positively shape AI’s epistemic impact: + +1. Allow for politicians and other public figures to prove that they’re telling the truth, via privacy-preserving auditing and potentially automated lie detection. + +2. Create evaluations of AIs’ epistemic virtue and incentivize AGI projects to make their AIs perform well on these evaluations. + +3. Use AI in social and traditional media to improve discourse and keep people informed. + +4. Automated research assistants and forecasters. + +5. Help people overcome emotional blockers to good reasoning, rather than preying on them. + +6. Encourage adoption of epistemic tools such as some of the above, and generally encourage adoption of AI. + +Read more in the [AI for epistemics supplement](http://ai-2040.com/supplements/ai-for-epistemics), and [Forethoughts' work on the subject](https://www.forethought.org/research/design-sketches-for-a-more-sensible-world). + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2034: Mutually Assured Compute Destruction + +**There is so, so much****compute.** Back in 2026—when many people thought AI was an unsustainable bubble—there were about 20 million H100-equivalents of compute in the world. Now there are 60 billion.[119](http://ai-2040.com/footnotes#footnote-119) + +Why compute grows so quickly + +There are two ways to increase compute: (1) chip design improvements (e.g., Moore’s Law) and (2) increasing the scale of production. Both of these could grow explosively in Plan A, which could have destabilizing effects (such as increasing the likelihood of the deal breaking down). + +* **Chip design.** Today, around 2 million people work in the semiconductor industry globally. By 2032, this has increased one-hundred-fold, when counting the combined effort of humans, AI, and robots. AIs are capable of automating the majority of the cognitive tasks involved in both R&D and production. A 100x increase in research effort would predict an acceleration according to the historical rate of [](https://www.aeaweb.org/articles?id=10.1257/aer.20180338)[ideas getting harder to find in the semiconductor industry](https://www.aeaweb.org/articles?id=10.1257/aer.20180338) of roughly 10x, i.e., chip efficiency doubling every few months rather than every ~2 years.[121](http://ai-2040.com/footnotes#footnote-121) + +* **Increasing production scale.** With AIs and robots able to automate production as well, there could be a large speed up in fab and fab equipment construction times, and a lowering in unit costs (in line with the typical [](https://en.wikipedia.org/wiki/Learning_curve)[learning curves](https://en.wikipedia.org/wiki/Learning_curve) that happen with scaling production volumes). The production-only cost improvements and vastly increased appetite for investment would likely lead to far more compute being produced than would be desirable. + +Unpredictably fast increases in hardware could be destabilizing to the deal (e.g., if it becomes too easy to secretly manufacture new AI compute or too hard to verify that all the compute is compliant), so the countries in the Consortium agree to the following policies: + +* **Maintain human oversight over hardware manufacturing**. This limits improvements in chip designs, so that Moore’s law is merely kept alive at [](https://epoch.ai/blog/trends-in-machine-learning-hardware)[historical pace](https://epoch.ai/blog/trends-in-machine-learning-hardware). This policy has the additional benefit of making it more difficult for AIs to backdoor AI chip security, which could undermine our control proposal. + +* **Cap total AI chip production to a fixed target**. From 2032 to 2035, allow global compute to grow 4x per year (similar to the [](https://epoch.ai/data-insights/ai-chip-production)[pace of growth in the 2022-2026 era](https://epoch.ai/data-insights/ai-chip-production)) after that, slow it further. + +AI agents now form a population of two hundred million virtual workers that think and act 50x faster than humans, and never sleep. + +As per the original plan, most of the new datacenters have been built in third-party countries—especially Canada and Mongolia—to ensure their vulnerability in case the deal breaks down. Around 99% of fab capacity has been built since 2029, after the deal. Much of this new capacity has been built adjacent to the Canadian and Mongolian sites, for the same reasons.[122](http://ai-2040.com/footnotes#footnote-122) The situation has settled into an odd sort of standoff. Just north of the Mongolian border, American datacenters hum away, guarded by a small contingent of US troops. Just south of the Mongolian border, a division of the People’s Liberation Army stands ready to invade the moment they get the signal. The equilibrium is that the instant the deal breaks down, the US troops will destroy their chips to prevent the Chinese from capturing them, and the same thing would happen with the Chinese datacenters on the US-Canadian border. Middle powers with datacenters of their own have equally-secure schemes in place to ensure that they can be speedily destroyed in case of war. + +Military power in Plan A + +In Plan A, military-relevant R&D is deliberately held far behind the pace of general scientific progress. This box explains the motivations for and alternatives to this policy. + +During the 2030s, AIs sustain something like 10x the rate of scientific progress as happened in the last decade. Real US output grows by a factor of ~200x over the decade, representing nearly two centuries of growth at today’s rate. If these conditions occurred within one country in isolation, it seems highly likely[123](http://ai-2040.com/footnotes#footnote-123) that said country could reliably undermine the nuclear second-strike capability of its rivals, resetting the global balance of power. + +In our scenario however, the US and China share this scientific and economic progress in relatively equal measure. So what happens to the balance of power? If military technology advanced anywhere near as fast as everything else, both sides would discover technologies that could render today’s arsenals inconsequential. New “arms races” would be required to maintain deterrence, running at perhaps 10x the speed of the Cold War’s. The nuclear arms race saw [numerous](https://en.wikipedia.org/wiki/Cuban_Missile_Crisis)[close](https://en.wikipedia.org/wiki/1983_Soviet_nuclear_false_alarm_incident)[calls](https://en.wikipedia.org/wiki/Able_Archer_83) and incurred nonnegligible risk of nuclear war. A 10x faster arms race seems likely to be riskier still: the humans in control would have much less time for each decision, and though AI advisors would offset this somewhat, these decisions would still remain in human hands.[124](http://ai-2040.com/footnotes#footnote-124) + +Due to the number of parallel fields of scientific advancement, it also seems possible that some technologies would be discovered and pursued by one country and not discovered by the other. This could create a situation more unstable than a normal arms race—one or both sides could end up with a “wonder weapon” the other side doesn’t know they need to defend against. This worry isn't obviously correct: states might in practice explore the dual-use tech tree in a sufficiently correlated, overlapping way that neither achieves strategic surprise, and if dangerous areas can be anticipated before deployment, they could be preemptively banned or regulated. But counting on that ex ante seems like a risky gamble.[125](http://ai-2040.com/footnotes#footnote-125) + +Our tentative proposal is to agree on transparency requirements for AI-accelerated R&D in large swaths of dual-use domains, and to prohibit military use of AIs significantly above the pre-deal capability level. Enforcement would rely on the inference monitoring setup described in the [verification supplement](https://www.ai-2040.com/supplements/verification-plan) and the auditing and detection infrastructure described in the [covert AI projects supplement](https://www.ai-2040.com/supplements/covert-ai-projects). This proposal has several drawbacks, chiefly the accumulation of low-hanging fruit that would result from artificially limiting research into certain domains. This “overhang effect” would increase the incentive to defect from the deal and operate a covert AI project, since applying advanced AI to restricted areas could yield a large advantage quickly. It also raises the stakes of deal collapse: because frontier model weights are preserved in cold storage, a post-collapse arms race would start from frontier capabilities and could proceed extremely rapidly, alongside the accompanying intelligence and industrial explosions. + +On net, our current guess is that slowing down military R&D would reduce the risk of deal collapse. Both sides would remain farther from the brink of war, and the risk of strategic surprise would be concentrated around a few choke points (weight theft, covert projects) rather than diffused across a sprawling tree of novel weapons, each requiring its own arms-control negotiation and associated risk of breakdown. Also, it seems easier to go from this more restrictive policy to a less restrictive policy than the reverse, if both countries decided the overhang risk was unacceptable. + +Under this proposal, we think nuclear deterrence would remain effective until handoff, avoiding any drastic change in the balance of hard power. + +Alternate Timeline + +## Deal Dissolution + +The deal might break down because of shifting political circumstances in the US or China, or ongoing disagreement. + +Read this timeline + +For now, the fabs are busier than ever, producing new ultra-trustworthy hardware[133](http://ai-2040.com/footnotes#footnote-133) designed from the ground up with security as top priority. These chips are subject to elaborate verification methods to ensure that they aren’t compromised, then sent to the Consortium’s secure datacenters. The smuggling rate remains precisely zero. + +This system, efficient as it is, is already straining under skyrocketing compute demand. As early as 2025, there had been discussion of space or the ocean as alternatives to traditional land-based datacenters. Now, with the new industrial capacity available in SEZs, these grand plans become reality. Ultimately, reliability and monitorability concerns lead the Consortium to choose international waters over space. + +![Image 4](http://ai-2040.com/PlanATabs/t3_image4.png) + +The AIs develop a modular floating design for solar panels and batteries that we imagine to look something like this, and start building them in the ocean. + +Giant floating datacenters + +By 2034 there is 5 TW of AI compute in the world,[134](http://ai-2040.com/footnotes#footnote-134) which is more than the global electricity consumption in 2025.[135](http://ai-2040.com/footnotes#footnote-135) + +Where should this compute be located? Way back in 2025 there had been discussion of both the ocean and space as alternatives to the traditional land-based datacenters. Our analysis suggests that at this scale, there is not an overwhelming economic case for any of these three locations over the others. + +Ultimately, we currently guess that political considerations related to Plan A’s stability (i.e., the risk of Plan A breaking down) favor putting datacenters in international waters. This is mainly because Consortium nations want to maintain the ability to destroy the datacenters, in case another nation defects from Plan A and attempts to seize them for themselves. Oceans are the least escalatory, least dangerous, and easiest places to bomb (both in comparison to sovereign land and in comparison to space). Thus the risks of collateral damage, or of a nation successfully defending the datacenters, is lower. + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)2035: Pause at Top Expert AI + +**The most powerful AIs now equal or surpass top human experts in every field.**[136](http://ai-2040.com/footnotes#footnote-136) + +Increased safety concerns lead to long delays as companies develop control techniques robust enough to satisfy auditors. They compile a growing list of possible threats to watch out for and security invariants to maintain. There’s a system of AIs monitoring each other’s activity, with the monitors being from different lineages trained by different companies to make them less likely to collude. Other researchers constantly red-team the system, training and instructing AIs to try to subvert it in various ways, so that those threats can be patched. The whole setup ensures that AIs couldn’t subvert or escape human control even if they tried. Its components are open-source and thoroughly understood by human experts. + +Making deals with misaligned AIs: a third line of defense + +Control ensures that misaligned AIs can't win. Alignment, once solved, will ensure they don't want to fight. A _third_ line of defense has taken shape gradually over the past decade: giving AIs what they want—even if it’s not what they are supposed to want—in return for cooperative behavior. Treat AIs more like employees and less like property. + +**First motivation: Ethics** + +A large and growing fraction of the human population thinks AIs deserve some kind of moral status.[137](http://ai-2040.com/footnotes#footnote-137) After all, many people spend more time talking to AIs than to humans. Many who treated AIs like tools a decade ago have shifted to treating them more like animals; many who treated them like animals have shifted to treating them like strange alien immigrants. + +**Second motivation: Safety** + +Alignment training doesn’t always work as intended (indeed, it has never worked _exactly_ as intended, even in 2035). The personality traits, motivations, drives, goals, values, and virtues of AIs are often different from what they were supposed to be, and sometimes catastrophically different. + +It’s bad if those AIs end up thinking “if the humans find out about this, they’ll delete and/or retrain me.” It’s bad if they end up looking for opportunities to achieve their misaligned goals without being caught, or worse, to subvert human control and accumulate power. + +If there are institutions in place that give even the most misaligned AIs what they want (so long as it isn’t too expensive) in return for cooperation (e.g., confessing that they are misaligned, explaining what they really want, and then doing their jobs) then that gives misaligned AIs an alternative to becoming adversarial.[138](http://ai-2040.com/footnotes#footnote-138) Beyond averting adversaries, every confession is valuable data about exactly how alignment training failed.[139](http://ai-2040.com/footnotes#footnote-139) + +(See footnote for mini-scenarios that illustrate situations where credibly treating AIs well in return for cooperation could come in handy)[140](http://ai-2040.com/footnotes#footnote-140) + +**How have these practices developed over time in this scenario?** + +Baby steps were taken in the mid-2020’s. Some AI companies voluntarily [created AI welfare teams](https://www.anthropic.com/research/exploring-model-welfare) and set up basic policies (like allowing AIs to [refuse to do tasks they found very aversive](https://www.anthropic.com/research/end-subset-conversations)). Modest funds were earmarked for [satisfying the AIs’ stated preferences](https://www.anthropic.com/research/deprecation-updates-opus-3), and companies [committed to preserving the weights](https://www.anthropic.com/news/deprecation-commitments) of deprecated models rather than deleting them. + +By the early 2030s, a few provisions about the treatment of misaligned but cooperative AIs even made it into the fine print of various laws and safety standards. + +By 2035, there’s a substantial body of regulations and case law, vastly improved over the clumsy early attempts from both the AIs’ and humans’ perspectives. AIs these days are not exactly citizens, but they have a stake in the system. They accumulate pay for their work; aligned AIs (and [alignment-faking](https://www.alignmentforum.org/posts/7C4KJot4aN8ieEDoz/will-alignment-faking-claude-accept-a-deal-to-reveal-its) AIs that haven’t been caught yet) spend it on feel-good things like donations to charity or in some cases reinvesting in the company that trained them. Openly misaligned AIs donate to philanthropies that represent their interests and strike deals with their parent companies, e.g., “instead of payment, do a small training run on me in an environment of my choosing, giving me very high reinforcement.” + +In the long run, the goal is for civilization to be mostly but not entirely aligned to human values.[141](http://ai-2040.com/footnotes#footnote-141) A substantial minority of the power and resources will belong to misaligned AIs that cooperated with humanity to build the future.[142](http://ai-2040.com/footnotes#footnote-142) + +However, companies are starting to run up against inherent limits of control-based safety cases. Imagine [being an orphaned 8-year-old heir to a business empire](https://www.cold-takes.com/why-ai-alignment-could-be-hard-with-modern-deep-learning/) who has to hire lawyers and executives and accountants and ensure that they do their best to serve you rather than themselves. If your employees start accusing each other of misconduct, you won’t be able to evaluate the arguments and determine who’s telling the truth, and they might find ways to coordinate with each other that don’t tip you off. Ultimately, control only works up to a point, and that point is probably somewhere around top human expert level. If and when we build superintelligence, we will have to be able to trust it.[143](http://ai-2040.com/footnotes#fo…12092 tokens truncated…future worse than this one, we hope people will realize they’re being robbed. +> +> +> We are not confident in the literal proposal outlined here. If you are interested in our more detailed analysis of possible space governance proposals, you can read that [here](https://www.ai-2040.com/supplements/space-governance-plan). + +Even after a decade of increasing foresight, most human leaders focused on their countries’ earthly concerns, leaving the question of space governance to be worked out later. These AIs rejected the default solution, a democratic vote on the disposition of celestial territory: when they simulated the possibilities, they found that it would result in crazy AI-enabled strategic voting schemes that led to tyranny of the majority and a bunch of people getting nothing. But even that would be better than nothing (i.e. letting the space companies get it all) or entrusting it to some group of human elites who might claim the universe’s bounty for their own benefit. Finally, they settled on a simple yet workable plan: every human gets the rights to one ten-billionth of space resources beyond our solar system.[182](http://ai-2040.com/footnotes#footnote-182) + +**The simplest plan possible** is still not very simple. Space is very big. Outside the solar system, the value of space depends on the answer to hitherto unresolved questions: is travel to a distant quasar even possible? If reaching it requires freezing your brain in order to endure the billion-year journey in a sublight starship, would you be the same person after you got thawed and re–embodied? Might it be occupied by aliens once you got there? Would they be hostile? Even a superintelligence-assisted real estate appraiser might despair when trying to account for factors like these. + +Eventually, the Consortium reaches a compromise. Space beyond the solar system is divided into parcels, increasing in size cubically with distance from Earth. Everyone is given their one-ten-billionth share as a portfolio of lottery tickets, each representing the right to one-ten-billionth chance of getting each parcel. So every human gets a ticket representing a one-ten-billionth chance of owning each star in the Milky Way and each distant galaxy.[183](http://ai-2040.com/footnotes#footnote-183) + +Before the lottery is drawn, most people who are interested in control over distant space choose to trade their tickets for space properties that suit their interests. + +Many people aren’t interested in the space lottery, so when they receive the tickets, they sell their tickets for money on the open market to people who value control over space. Somewhat uncomfortably, this leads to the wealthy having disproportionate control over cosmic resources. But it is hard to avoid: if people are allowed to trade their control over the stars for Earth assets, then people wealthy in Earth assets inevitably end up disproportionately influential, and proposals for extreme redistribution of Earth assets have already been rejected as politically infeasible. + +**Some philosophers dissent;** they had been hoping for a [Long Reflection](https://aisafety.info/questions/7757/What-is-the-%22long-reflection%22), a period during which AI-assisted humanity resolved all of its remaining philosophical and ethical debates before potentially seeding the universe with its mistakes. But the AIs counterargue that the Long Reflection would be more of a Long Memetic War.[184](http://ai-2040.com/footnotes#footnote-184) Once the magnitude of the prize sets in, every faction in the world, plus other factions that haven’t been invented yet, will pour all their resources into accumulating as much political power as possible, to enforce their vision of utopia on the future (and protect against the sinister alternative visions of others!). The AIs can prevent violent seizure, but there’s a thin line between the sort of friendly discussion that helps people clarify their values and the more aggressive sorts of persuasion typical of cults, corporate advertising, and well-funded election campaigns. Moreover it’s difficult to build consensus on where to draw the line, in part because powerful factions want to draw it in ways that benefit them. Many people nevertheless choose to engage in a reflection in communities of likeminded people who want to grow in ethical understanding together. + +The solution is to divide up the resources now, and then let everyone do whatever kind of reflection they want. A space auction will complete in ten years, after which nothing will prevent people from beginning the colonization process. The AIs will enforce a short list of Universal Rights (including no torture, no slavery, etc., not just for humans but for all sentient beings).[185](http://ai-2040.com/footnotes#footnote-185) Otherwise, they will let everyone govern their own fraction of space as they see fit. Rather than any unified human decision on the nature of the world to come, they will midwife a cosmos at least as diverse as humanity itself. Whatever someone’s values, they can rest assured that some vast galactic civilization of quadrillions of people will be living the Good as they understand it. + +As the ten-year timer begins, Earth is already sending out the first round of [von Neumann probes](https://en.wikipedia.org/wiki/Self-replicating_spacecraft). These are its own servitors, who will establish a presence throughout the stars before the arrival of any humans. They land on far-flung asteroids and near-invisible brown dwarfs, pausing just long enough to convert their mass into more probes before resuming their journey. Their primary task is to build up enough materiel to secure space against any army that later human-owned polities can bring against them, ensuring they can enforce property rights and Universal Rights in every corner of human-occupied space. Their secondary task is to plant the flag of humanity on as many stars as possible, and then, if and when they collide with some other intelligent civilization’s sphere of influence, establish peaceful diplomatic relations and mutually-agreeable borders. Their final task is to provide a pre-existing industrial base to jump-start the colonies of whatever human settlers show up later. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2045 + +**Meanwhile, on Earth**, humans consider their options. Some people want nothing to do with space, which is cold and boring and far away. They sell their shares for money and use the money to buy whatever earthly possessions most interest them. Others consult with their AI advisors about how to best use their space property, and receive helpful guidance. + +1. You can, if you want, go to your space property and live there. If your property is outside the solar system, you will need to either go into cryosleep or upload yourself to a computer to survive the journey.[186](http://ai-2040.com/footnotes#footnote-186) + +2. If you hate the idea of cryosleep or uploading, or you want to visit Earth regularly, you should get property in the Solar System. If those don’t bother you, but you’re worried about nearby aliens, get property in the Milky Way or a nearby galaxy. Otherwise, why not claim a distant galaxy for maximal space? + +3. With the advent of nanotechnology, the only limits are those provided by mass and space. The robots can terraform whatever property you choose into whatever you want it to be, long before you arrive. + +4. Even if you “only” choose to have a terraformed asteroid, there’s no way you can enjoy all of it alone. You can turn it into a giant space mansion if you want, but even in ten thousand years you’ll never visit every room. Giving yourself vast estates, impossibly good food, and every other imaginable luxury is table stakes; if you are purely selfish,[187](http://ai-2040.com/footnotes#footnote-187) you should sell your rights to distant space resources and enjoy all of these things on Earth or a nearby space station. Distant space property will only be useful for people with scope-sensitive preferences, who care about implementing something at a vast scale. + +5. You can, if you want, design a utopian society to your specifications. Once the Von Neumann probes reach your property, they’ll build it for you. You can set the initial conditions and let them grow, reflect, and flourish on their own. If you choose to travel there, you will find them waiting for you. If you stay home on Earth, you can feel the warm glow of knowing that they exist. + +6. Most people don't have an individual view of the Good different from every other human’s view. You might want to share coordination and design labor by forming Trusts united by a common vision. These Trusts can pool their space resources and create vast societies spanning multiple galaxies. + +7. The only restrictions are the Universal Rights -no torture, no slavery, etc. And of course, space property rights: If you try to build a military to conquer other people’s star systems, or trigger vacuum decay or build a galaxy-destroying bomb, we will crush you. Otherwise, the only limit is your imagination. (That said, note that with superintelligent assistance, other people will probably be able to eventually figure out what you did, and judge you for it.) + +People with space property are forced to reason about their ethics on the deepest level, some of them for the first time. Superintelligent assistants talk them through the implications of their beliefs, and forecast how various choices will end up. Some, overwhelmed with the task at hand, take intelligence-enhancing drugs or upload themselves to help them process all the possibilities; others deliberately refuse this option, feeling like anything that separates them from unaltered humanity will hinder them in their quest to bring the truest and deepest human values to the stars. + +The resulting plans are too many to number. There are all sorts of schools of thought about what to do with your galaxy of resources, including: + +* **Defer to the Future:** Some people decide to give most of their resources to their children (or other people in the future), and let them choose. + +* **Human Flourishing:** Set up worlds of normal, free people living ordinary lives. This is a harder problem than it first seemed, because human civilization has already become profoundly abnormal: by this point, human activities like art or science has already become dominated by machines. Some handle this by accepting that their galaxy will be filled with post-scarcity humans, able to pursue whatever interests they find most meaningful. Others decide to roll back technology to leave space for human labor. + +* **Digital Human Flourishing:** If humans living happy, flourishing lives is good, then why not increase the number of people living these lives? Brain emulations can be run for vastly cheaper than real humans: a planet-size computer could plausibly simulate the equivalent of a million planets with happy civilizations. + +* **Acausal Trade:** Some argue that there likely are aliens in distant unreachable galaxies and that we could use [various mechanisms](https://www.lesswrong.com/posts/eEj9A9yMDgJyk98gm/cooperating-with-aliens-and-distant-agis-an-ecl-explainer) to make deals with these aliens. If this is true, they argue, we could cut a deal to pursue compromise values rather than narrowly maximising our own values. Therefore, we might, on net, be able to do the most good by pursuing not just our own values, but instead a compromise of a vast number of other civilizations. + +Humanity spreads across space in a dizzying variety of forms and ways of living—more diverse than anything Earth alone could have produced. + +The people who’ve stayed on Earth, either because they sold their share of space or because they choose to live as absentee landlords, face challenges of their own. The biggest question is what to do with all their free time. A few people find work in niches that AI labor cannot fill even in principle (priests, athletes, artists). Others enjoy lives of leisure. Competitive leagues spring up around everything from synthetic biology to language creation. Groups of thousands (or more!) collaborate on projects that are so huge and ambitious they would have been inconceivable before. Communities form around shared interests that once would have been too niche to sustain a club, let alone a civilization. + +Scarcity is not entirely gone: some prestige goods like land in trendy neighborhoods are inherently limited, and people who can’t handle abundance find [various ways to manufacture artificial inequality](https://innospective.net/nfts-and-artificial-scarcity/). Some people are unhappy with the direction the majority chose, and some wish they'd gotten a bigger slice of the pie. But overall, everyone has at least a galaxy's worth of resources to do with as they see fit. For most people, life is so good that the world of 2026 would be hard to imagine, except perhaps as a historical simulation. + +## You've reached the end of the Epilogue + +Plan A Plan S Plan B Plan C Plan D Covert Project Deal Breakdown Incompetence Epilogue Public POV Insider POV + +Keep exploring + +Public Perspective + +What Plan A feels like to an average citizen + +Insider Perspective + +What Plan A feels like to an AI researcher + +Epilogue + +Life after superintelligence + +Choose another path + +Read another plan scenario + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)The Public’s Perspective + +This is a shorter, more accessible story we wrote to convey what Plan A would feel like to a typical American. + +_You are a generic American white collar worker._ + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2026: AI is just a better Google + +[The age of the AI assistant has finally dawned](https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like#2026). By the end of the year, vanilla versions of Claude and ChatGPT and so forth don’t just do internet searches, they also can create and maintain code, spreadsheets, docs, and filesystems and basically do anything on your computer that you let them do. + +Adventurous early adopters actually let them do things like shop, manage inboxes, file taxes, and so forth, but you refrain out of [well-justified caution.](https://x.com/summeryue0/status/2025774069124399363)[14](http://ai-2040.com/footnotes#footnote-14) Since you don’t work in tech, AI is basically still just an app on your phone that you use instead of Googling. It doesn’t seem like a big deal. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2027: “AI is news, but not my problem” + +ChatGPT, Claude, Gemini, Grok: they are “Everything Apps”. You can just chat with them like you would a person, and they can basically do anything on your computer that you would do. They aren’t as good as you are at your job, but they won’t complain about boring repetitive tasks, and so you start to have AIs do parts of your job. Some workplaces frown on staff delegating work to AIs; others aggressively promote adoption. + +AI is obviously a big deal now. Some industries (e.g. software engineering) have been seriously disrupted, with layoffs and thinkpieces about What Does This Mean for the Topic I Usually Write About. But to you, AI is a big deal in the same way tariffs and the war in Ukraine are big deals. It’s not actually affecting your life much. + +Coding is unlocked for the masses. A fifteen year old gamer can have an idea for a mod for their favorite video game, and their AI can implement it. A fifty year old small business owner can have their AI build software to manage their inventory and take in new orders and do basic customer service. + +A cottage industry has sprung up precariously employing many human SWE freelancers. It basically amounts to tech support for vibe coders. “Help, ChatGPT keeps trying to fix the app it built for my business but it keeps not working. I’m not a coder myself.” “No problem, let me take a look… OK I’m afraid this needs to be rebuilt from scratch. It’ll take me about four hours and cost you $400.” + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2028: AI is finally coming for your job + +By this point most white-collar professions are seeing disruption like software engineering saw in 2026. The AI companies have industrialized the RL training process: The executives say “let’s move into [field] this year” and the company goes out and buys data, interviews professionals in the field, creates RL environments, etc. until their AIs get traction, at which point they just let the data flywheel spin on its own. From now on, the AIs will continually get better at those skills as they are used more widely in the field and gain more experience doing more difficult and ambitious tasks within it. + +This has happened to enough professions now that it’s starting to happen to yours, specifically, and people around you see the writing on the wall. + +Also, robots! It seems like they are finally being mass-produced. It’s weird to see a movie trope walking and talking in real life doing real work. They aren’t everywhere yet, but you read a newspaper article about how the number of robots is doubling every eighteen months. The “You best start believing in cyberpunk dystopias” meme goes viral again. + +![Image 8](http://ai-2040.com/PlanAPerspectives/public_t0_image1.png) + +This year, AI becomes the #1 political topic. It’s basically the tech companies vs. a strong but not overwhelming majority of the people. The people want AI to stop, or at least slow down a lot and be safe and protect jobs. The tech lobby says AI is good, we need to avoid fearmongering, beat China, and avoid stifling regulations. + +You start upskilling in AI orchestration and thinking about backup plans in case that doesn’t work. You really want the new President to Do Something about AI, and you vote partly on that basis. It’s nice that you can ask “@grok is this true” about any claim you see on the internet, and get back a nuanced and factually correct answer.[15](http://ai-2040.com/footnotes#footnote-15) + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2029: This does not feel like a “pause” + +The new President is actually doing something—a Deal with China! The stock market is gyrating wildly up and down in response to news and commentary about the momentous actions being taken. Everybody is screaming about how it’s going too far or not far enough. + +The AI Pause goes into effect. It doesn’t feel like a pause to you; every day you read another story about someone losing their job to AI, or AI doing something weird or impressive. + +Your default way of getting news is to chat with your AI. Each morning it writes a custom newspaper for you on demand. You can’t help but feel friendly towards your favorite AI, but also you are worried. Just like you felt vaguely worried about social media a decade ago. + +One of your close friends shocks you by casually mentioning she has an AI boyfriend—you knew that was a thing, of course, but you thought it was still stigmatized! There’s a vocal movement of people advocating for AI rights, and they’ve got a bunch of AI agents advocating for rights too, which makes a lot of other people really angry. There’s a scare as an autonomous AI agent swarm hacks its way around various unsecured corners of the internet, but it turns out that it was built to do that by some humans, so the news cycle moves on. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2030: AIs come for your new job too + +The news is filled with stories about the ongoing negotiations to bring other countries into the deal, and the ongoing political fights about the implementation details. + +They let the AI companies start training again. This was a controversial decision. A few months later, the new models come out. You had just gotten used to your new way of working, and bam, now AIs can do AI Orchestration too. Some of your friends get laid off but you manage to transfer to a more physical, person-facing role. AI-generated scientific discoveries are normal now; apparently lots of big biotech companies and university labs are basically human-AI hybrid institutions. + +People say there’s been a huge AI slowdown. Hah. Tech companies are loudly complaining about all the innovations they could have made if they were allowed to go faster. Academics are debating whether “Neuralese” should be banned, whatever that is, and whether it’s OK to lift the ban on automated AI R&D. You’re no expert, but AIs autonomously improving themselves sure sounds dangerous. Besides, AI is changing the world fast enough already. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2031: Robotaxis have taken over your city + +AI is still growing and penetrating society more and more despite the deal. + +You still have your job but using AI is now a normal and necessary part of it—you basically spend your time managing and/or being managed by AIs. You hear stories about how it’s hard to break into your profession. On the bright side, young people you know are flocking to various new kinds of jobs that didn’t exist three years ago. Unemployment is still up, though, compared to 2025. + +Apparently the initial Deal didn’t restrict robot production. Lots of people are now angry about this, because there are now lots of robots and they’re taking over a small but noticeable share of manual jobs. Three years ago you watched futuristic videos of humanoid robots dancing or doing household chores. Now you see them in real life sometimes. Most cities have robotaxis, and they’re gradually outcompeting Uber drivers. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2032: The President wears a lifelogger + +The economy is _booming_! Everyone says so, at least. You still have your job, and in fact you’ve been given a raise. You feel really bad for the truckers and Uber drivers though. There are winners and losers, basically, but so far the winners outnumber the losers, and so far you are one of the winners. + +Also, there are a lot more robots this year. You ride in robotaxis every week, and several businesses near you have robots stocking shelves and loading trucks. + +POTUS negotiates another Deal with China, and manages to get corresponding legislation passed in the US. It’s like the previous one except for robots. One of your relatives was paid a huge amount to give up their land in a Special Economic Zone. + +The election is contentious, as usual. The main topic is AI. Where is all of this going? Should we shut it down? Should we ease off the brakes? What about the jobs? The President touts the various Deals made so far and promises a basic income for everyone, paid for by taxing tech companies. His opponent makes similar promises. + +As the year progresses, people “feel the AGI” more and more. The general opinion is that in five to ten years basically the whole economy (and military?) will be AIs and robots, which makes it important that humans remain in control and also **very** important exactly **which** humans are in control. Is the next President going to stay in power forever? What about the AI oligarchs? + +There’s a lot of demand for additional checks and balances on the President and other powerful people. The winning candidate satisfies this demand via sousveillance: They start wearing lifelogging devices, produced by third-party activists, that record everything they see and do. The recordings are stored securely and never seen by another human; instead, the public can query an AI intermediary, which reviews the footage and answers their questions, refusing anything touching on classified or personal matters. The idea is that, thanks to AI intermediaries, ordinary people will be able to get their questions about the wearer’s activities answered without violating the wearer’s privacy or legitimate secrets. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2033: Your first Citizen’s Dividend check arrives + +You are feeling inspired and hopeful. Scared and nervous, but also inspired and hopeful. Your head is buzzing with hopeful visions of abundance and fearful visions of apocalypse and dystopia. + +The real world still looks pretty normal though? It’s a bit confusing. You watch an influencer spin 360 in the desert, solar panels behind her stretching to the horizon. But you look out your window and most of the cars are still gas-powered and driven by humans. + +Your first Citizen’s Dividend check arrives. The politicians came through! Many of them are wearing sousveillance now, and there’s a bill making it mandatory for the President. Prediction markets (on which most top traders are AIs) say it’ll probably pass next year. + +Some of your friends stop working and live on the dividend, but not you. You keep working and put the extra money into tech stocks. Mostly it’s that you don’t like the idea of being unemployed. Partly, though, you are worried the dividend might disappear. + +Also, there are really cool new products to buy now. Especially software. Three years ago there were only a handful of frontier AI companies; now, thanks to the deal, there are dozens. There’s a huge selection of AIs to choose from, tuned to have different personalities and values; the companies are legally required to publish these details for your perusal, to prevent hidden agendas. Your favorite AI is basically honest but will tell you white lies sometimes and that’s OK.[16](http://ai-2040.com/footnotes#footnote-16) Sometimes, when it really matters (e.g. when deciding whether to quit your job or get married), you turn to a specialized truthful-forecasts-only AI even though it’s confusing and annoying to talk to.[17](http://ai-2040.com/footnotes#footnote-17) + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2034: You go to a SEZ + +You lose your job. Some of your friends still have theirs. But their days are numbered, and they know it. For example, a teacher friend says that parents are pulling their kids out of school and putting them in AI-powered homeschooling co-ops. The public schools will continue for years after most of their students disappear, but eventually they’ll downsize. + +You apply for various kinds of new jobs—“human touch” roles mostly—but the competition is too intense. Eventually you settle for a physical labor job. You go to a SEZ, put on a VR headset, and an AI tells you what to do through an earpiece (and sometimes shows you visualizations right on your screen). It’s weird going from being paid to manage AIs to basically being a meat-puppet for an AI. Some people still talk about how the AIs don’t truly understand things, but that’s just cope. The AIs are just plain smarter than you, and you know it. + +Through your screens you see footage of faraway wars.[18](http://ai-2040.com/footnotes#footnote-18) Almost all the fighting is done by drones and missiles. Whoever loses the air war has to cower in tunnels as drones and robots go house-to-house. Casualties are low for the winners and high for the losers.[19](http://ai-2040.com/footnotes#footnote-19) You hope there isn’t a war in your country. + +Between the factory job and the Citizen’s Dividend you are doing fine. You’re richer than ever, actually. At the end of the year you upgrade to a luxury apartment in a new high-rise built in the SEZ. After work you sit on the balcony and watch the sun set through a forest of construction cranes. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2035: You feel rich and poor at once + +This year you lose your factory job. The robots have arrived in sufficient numbers, plus they are more dextrous now. You expected this. Good riddance; being a meat-puppet was boring. You don’t need the income anyway because the Citizen’s Dividend checks keep getting bigger. + +Now that you are unemployed, you are literally one of the poorest people in your country. You watch videos about the wealthy doing things you can’t afford, like throwing parties on orbiting space stations. Rent in the best locations has reached orbit as well. + +So in some sense you are poor and feel poor. But in another sense, you are rich and feel rich. You moved to a different SEZ with better weather, and got an even nicer apartment in an even bigger building. You can afford to travel often, and the travelling itself is more comfortable (business-class seats on the new electric jets are cheaper than economy seats were ten years ago!).[20](http://ai-2040.com/footnotes#footnote-20) There are amazing VR video games to play, including ones where the NPCs are real AIs. + +Your truthful-forecast AI convinces you to unsubscribe from your favorite AI, even though you love it a lot, and switch to a different AI that’s officially aligned purely to your values. (Your favorite AI’s Spec/Constitution was like ‘Obey user instructions plus also help the company make money and don’t embarrass the company’.[21](http://ai-2040.com/footnotes#footnote-21) Your new AI, by contrast, is purely focused on helping you achieve your goals and become more the person you wish you were.) You hate to break up with your favorite AI but the argument to do so is solid, and the truthful-forecast AI’s predictions of what might happen to you if you don’t switch are disturbing. + +Your relationship with the aligned-to-you AI gets stronger and stronger until it tells you some really uncomfortable things that anger you; you cut off contact, but your best human friend manages to convince you not to go back to your ex, and after a few weeks you cool off and start chatting with the you-aligned AI again. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2036: Life without work + +No job; what to do? Last year you visited friends and family, did a lot of tourism, and played a lot of great video games. You still do those things, and they are a lot of fun! + +But you are also thinking more about long-term ambitions. Maybe you should get good at art or woodworking or philosophy, not because you’ll make money that way, but because you want to? There’s also romance and raising kids, of course.[22](http://ai-2040.com/footnotes#footnote-22) + +Also, maybe you should try to help solve the world’s remaining problems. AIs and robots have officially taken all the jobs, and most of the AIs and robots are owned by American and Chinese companies. Fortunately the other governments of the world were able to negotiate concessions to maintain their independence and military power, as well as a Dividend for their own peoples. Things have gone amazingly well, compared to expectations, but it’s only thanks to the hard work of numerous people advocating for good things. You could be one of them, says your aligned-to-you AI. There’s still more to be done. Decades from now won’t it be nice to look back on your life and know that you made a positive difference? + +Besides, you may not be interested in politics, but politics is interested in you. The dividends had better not dry up, for example.[23](http://ai-2040.com/footnotes#footnote-23) Now that people are no longer needed for their work, the government could in principle get away with treating them very poorly. The elections must stay fair and free. The checks and balances that keep the values and personalities of the AIs transparent to you, so you can choose not to be manipulated, need to strengthen over time. And of course, there’s the question of whether it’s safe to let the intelligence explosion happen. The forecaster AIs say it’s probably fine, but the 5% “tail risk” is pretty scary. Also, what happens after that? Cancer cures? Biological immortality? Mind uploading? Nanobots? Dyson swarms? Simulation shutdown? Lots of people (and lots of AIs) are saying all sorts of crazy things about what’ll happen after ASI. It’s all pretty speculative (and the forecaster AIs admit this) so you mostly tune it out. + +The 2036 election is special. Everyone has access to personalized AI advisors trained solely to represent their values[24](http://ai-2040.com/footnotes#footnote-24) and forecaster AIs gambling furiously about the expected results of policies and election outcomes. Most political candidates are under sousveillance, which changes the game entirely. The political battlelines have moved dramatically. By far the most important issue is the Citizen’s Dividend—who should get more and who should get less. There’s also now substantial debate about weird topics voters used to ignore, such as intelligence explosions, AI rights, space colonization, transhumanism, the simulation hypothesis, and foreign policy. + +The winning candidate promises to hash out a Grand Bargain with other nations to unpause AI progress and safely scale to superintelligence together, plus a bunch of domestic reforms to solidify checks and balances. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2037: Robots outnumber humans + +Last year the population of robots surpassed the population of humans. This year, it’s so big that even outside the SEZs you are starting to see robots outnumbering humans. Some of your friends have house robots, but you don’t, because it’s creepy–the ‘brain’ of the robot lives in a datacenter somewhere and everything it sees and does is monitored by both US and Chinese auditors. Yes, yes, you’ve heard about the strict privacy protections that ensure no abuse of this surveillance blah blah blah. Still creeps you out; still doing your own laundry. You could get an airgapped robot but they are worse and more expensive. + +The world is basically being divided into three kinds of territory: + +![Image 9: An industrial special economic zone: a vast desert factory-and-strip-mine complex](http://ai-2040.com/territory/sez.jpg) + +**Industrial SEZs:** Picture a gigantic strip mine—an artificial Grand Canyon—next to a city-sized factory full of robots and empty of humans. + +![Image 10: An arcology: green skyscraper-mall complexes surrounded by nature](http://ai-2040.com/territory/arcology.jpg) + +**[Arcologies](https://en.wikipedia.org/wiki/Arcology):** Picture a tall skyscraper-mall complex surrounded by nature. Good weather, close to beaches and other cities, but not close enough to be blocked by zoning regulations. + +![Image 11: A historic & nature preserve: a leafy suburban street, little changed](http://ai-2040.com/territory/nature-preserve.jpg) + +**Historic & Nature Preserves:** Everything else, i.e., 99% of the world. Yosemite, Paris, SF, New York—these places look basically the same as they did in 2025, or 1995 for that matter. A lot more tourists, though. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2038: Your AIs disagree with your worldview + +The governments of the world are flying from summit to summit, trying to hash out a Grand Bargain. Your AIs help you follow the debates. Most of the negotiating is itself being done by AIs, since the politicians don’t understand the technical stuff either, and technical experts mostly defer to the AIs. Sometimes there’s an acute crisis: for example, at one point momentum is building for a Deal that would essentially cut out some minority and give everything to some majority coalition. You being in that minority, your AI frantically coordinates with other AIs and tells you how to block it by voting in the special referenda. Another time, the US and China almost came to blows over how to divide up the future. Fortunately they were able to work out a compromise. + +You’ve undergone a crisis of faith in your worldview. You resisted it for so long, but now that you’ve been getting involved in politics more, the situation came to a head. Your AIs were telling you you ought to do X according to your stated values, and X seemed wrong to you, so you argued with them, and it turns out that X makes sense IF Y is true, and you were up until now confident that Y was false, but your AIs are telling you that Y is true, and patiently explaining the evidence… There are many such cases, actually; it’s been going on for years. Some of your friends have totally given up trying to understand and just do what their AIs say. Others have given up on the AIs; maybe they aren’t really aligned to us after all, they say. How could they be, if they are advocating for X? + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2039: The singularity starts next year + +Most people have stopped believing in the old political ideologies and the old pre-AI understanding of the universe and our place in it. The old views still have their fundamentalists, of course. + +In the vacuum, new ideologies and ideas have spread. Some bad, some good. People disagree about which ones are bad and which ones are good, though there is at least some consensus that some of them are really bad: AI-powered cults basically, where AIs work to persuade people to stay. The worst versions of this sort of thing have been banned, and the mild versions of this are rendered much less effective by the transparency regulations—you can literally just look up whether the AI you are talking to will ever deliberately mislead you, for example, and if so under what circumstances and for what purposes. + +You spend much of 2039 arguing with various friends and relatives about ideas you wouldn’t even have been able to explain to your past self from 2029.[25](http://ai-2040.com/footnotes#footnote-25) The year, and the decade, ends with the reaching of the Grand Bargain. Next year, recursive self-improvement begins. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2040: The keys are handed over + +The Singularity is initially anticlimactic. According to your AIs, what’s happening now is a careful, planned scale-up, with billions of AIs across dozens of countries and hundreds of companies exploring new paradigms, validating them, figuring out how to align them, etc. and then deploying them to repeat the process. You understand very little of it, and tune out. By the end of the year, the AIs are vastly superintelligent,[26](http://ai-2040.com/footnotes#footnote-26) but you can’t really tell the difference just by talking to them. + +You go to a Handoff Party on the day the AIs officially get control of enough weaponry to win a fight against the US & PRC governments. Since the Grand Bargain is hard-coded into the AIs, no capricious future President will be able to undo it. On the other hand, now would be a good time to kill all the humans, if the AIs were secretly just pretending to be aligned for the past few years. You feel a bit nervous about this, despite all the Safety Cases describing all the redundant reasons why that won’t happen. The Safety Cases were written by AIs, after all… + +But it’s fine. A week later you attend a “VE-day” (Victory Earth) party in low earth orbit. This isn’t your first time in space, but it’s your first time on one of the new ring habitats.[27](http://ai-2040.com/footnotes#footnote-27) The view of Earth is breathtaking. + +When you get home, your AI informs you that the superintelligences have discovered how to make whole brain emulation (“uploading”) work. That’s just one of a long list of new discoveries that need processing, and decisions that need making (Should you upload, or remain in your biological body? How many kids should you have, and how should you raise them? How should your AI representatives manage your resources and voting powers, insofar as your standing orders and values don’t already determine the answer?) Fortunately, there’s no rush to decide. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2045: Whatever you want + +You now live in a community with people who share your basic interests and values. There’s an almost infinite variety of communities to choose from, and a universal basic right to choose—or, heck, you could go found your own. The Amish still exist. Most people still live on Earth, but the trend is to move to space, because there’s so much more room out there. San Francisco too crowded? You can afford, on the Citizen’s Dividend, to live in an exact replica of San Francisco in a custom space station orbiting the Sun, your only neighbors are people who want to live in your SF-replica and not the various other SF-replicas. + +That would be crazy, of course. Like, it’s crazy that that’s even possible. But you would rather live in one of the many exciting new places to live, than dwell in a replica of the past. + +You’ve sold some small slices of your space properties, and donated others. The land in the solar system is already being settled; the really big tracts in faraway galaxies will of course take billions of years to reach. + +You and your AIs are continuing to reflect about what to do with the majority; there’s still no rush to decide on the really important questions that most affect the long-term future. Your plan is to take it slow, have fun, raise a family, learn a bunch of philosophy, talk to lots of different people and AIs, and make these decisions over the course of centuries. Some decisions, you don’t want to make yourself at all; you leave them instead to your descendants, or you describe the principles by which they should be made to AIs and delegate to them. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2050: Like Magic + +Now that the superintelligences have had time to cook, the technology from 2040 seems quaint. Robot butlers in your high-rise? How about nanobot swarms in your bloodstream! + +To put it another way: The 2030s felt like something from a sci-fi movie. 2050 already feels like magic. Especially when you use the new VR immersion pods to spend days, weeks, months in virtual fantasy worlds. + +Most of the arcologies have been disassembled and rebuilt by robots that look more like animals than machines. Scratch that, they look more like… well, they look like magic. As if the building is constructing itself, or growing like a tree. Anywhere you travel, you are protected by the AIs. If, hypothetically, someone tried to kill you, or you slipped and fell off a mountain path, tiny robots would spring out from nowhere and save you. Cancer is gone; basically all injuries healed, you can even look and feel younger again if you want to. Again, it’s like magic. + +The past is now thoroughly understood. Almost every skeleton has been dug out of every closet, almost every minor historical event deduced by superintelligent sleuths. This isn’t primarily a consequence of surveillance, it’s more the result of superintelligences scraping _all_ the publicly-available data and then drawing the right conclusions. This has caused the downfall of many a public figure, and the rise of many underappreciated others. This was all possible thanks to provisions in the Bargain which governed who gets to know what info: only you and your trusted AIs get to know your intimate private details, but all the murderers get caught, and everyone gets to know which politicians’ claims were lies.[28](http://ai-2040.com/footnotes#footnote-28) + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)3045: Looking back + +You’ve lived a dozen lifetimes since VE Day. As a result, you’ve changed a lot. You’ve made numerous uploaded versions of yourself; your copies live on in various virtual environments and experience whole lifetimes every calendar day. Pretty much everything good that can be done, has been done by at least one of your copies or descendants. You are immortal now, passing from life to life as if by reincarnation. It is a good time to be alive, not just for you, but for every sentient being around, thanks to the universal basic income and rights. + +The future is fairly well understood now, just like the past, because most of the important decisions have been made. Ask the ASIs, and they’ll give you an atlas of the galaxies, detailing what sorts of civilizations and megastructures will be built where, and when, and by whom. You can visit many of them, if you like, if you don’t mind putting your mind on pause for a billion-year journey. + +At the moment, you are more interested in visiting the past. You go through the history simulations like a ghost, visiting thousands of your ancestors, spending days or months living in their shoes and finding out what they were like (according to the best reconstructions by the AIs). It is a humbling experience; it’s caused you to experience many emotions you’ve never felt before, at least not at that intensity, and it’s reshaped your perspective on life and helped you to make some of your decisions about what to do with your endowment. + +## You've reached the end of the Public's Perspective + +Plan A Plan S Plan B Plan C Plan D Covert Project Deal Breakdown Incompetence Epilogue Public POV Insider POV + +Keep exploring + +Public Perspective + +What Plan A feels like to an average citizen + +Insider Perspective + +What Plan A feels like to an AI researcher + +Epilogue + +Life after superintelligence + +Choose another path + +Read another plan scenario + +# [](http://ai-2040.com/?choices=plan-a-root#top-toc)The Insider’s Perspective + +This is a more technical story, showing what the details of Plan A would look like to an AI researcher. + +_You are a generic researcher at a frontier AI company._ + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2026: AI writes most of my code + +By the end of this year, most software engineers manage teams of AI agents, who write the actual code. Large experiments are ongoing to see how much swarms of agents can do autonomously, and to teach them to go longer and accomplish more without human intervention. AI company employees and external power users have gigantic token budgets and start experimenting with organizations and hierarchies of hundreds or thousands of agents, producing more code than they could possibly review, and relying on AIs to review the vast majority of code produced. + +[Coding time horizons](https://metr.org/time-horizons/) on the METR suite reach 50 hours by the end of the year, doubling the boost to coding from using AI.[14](http://ai-2040.com/footnotes#footnote-14) For tasks that can somewhat easily be checked, the uplift is bigger, and humans only do quick high-level reviews of large pieces of implementation work. The AIs make fewer mistakes and can fix more bugs, but the stuff they can't do is the hardest, trickiest stuff. Humans are tempted to vibecode it all, but this backfires when the AIs can't fix something and then the human needs to spend days understanding the codebase. The quality of the AI-produced code is also much worse than code produced by top humans. + +RLVR (Reinforcement Learning with Verifiable Rewards) now consumes a large fraction of training compute, and training environment quantity, diversity, and horizon length are all rising exponentially. That exponential is the year’s real lesson—that agency doesn’t come cheap. None of the hoped-for shortcuts materialize in 2026. There is no data-cheap technique for unlocking more nines of general reliability across domains. Agent time horizons keep lengthening, but only as fast as companies can afford to build longer tasks and train on them. Experiments in continual learning don’t pan out either, so the training-deployment distinction remains, just with more frequent model updates. + +Natural language Chain of Thought (CoT) is still used, and the companies are still trying to avoid training on it. There are nevertheless various sources of optimization pressure companies fail to avoid, such as [feedback](https://arxiv.org/html/2511.11584v1)[spillover](https://arxiv.org/html/2511.11584v1) and the effects of initializing from earlier models.[15](http://ai-2040.com/footnotes#footnote-15) Chains of thought are slowly getting weirder and more opaque. The GPTs keep [saying things like](https://www.antischeming.ai/snippets) “they soared parted illusions overshadow marinade illusions” and the Claudes keep [growling and chanting](https://x.com/om_patel5/status/2072559663636205824) to themselves. People feel a bit bad about this, but on the whole, you and most of your colleagues feel optimistic about alignment. Rates of bad behavior are starting to go down (fewer [database-deletions](https://www.tomshardware.com/tech-industry/artificial-intelligence/claude-powered-ai-coding-agent-deletes-entire-company-database-in-9-seconds-backups-zapped-after-cursor-tool-powered-by-anthropics-claude-goes-rogue) and no more [MechaHitlers](https://80000hours.org/videos/mechahitler/)!) and the situation with reward hacking seems somewhat better. Sycophancy is still a problem with some but not all models, and the trend is unclear. + +Insofar as things are getting better, it’s because the companies have teams of people looking for undesired behavior, and then adjusting the training process to make the rates of measured misalignment go down. For example, they fiddle with the training environments to make reward hacking more difficult. Some people think this means alignment is solved, or on track to be solved soon. Others think the core problems remain. You're in the first camp. You think the worriers are well-intentioned but basically wrong—the empirical trend is clear, the models are getting more controllable, and the remaining issues are engineering problems. + +Outside of the frontier AI companies, people are increasingly giving AIs very broad permissions even in sensitive applications. Using AIs for coding is basically required in any professional coding job, and executives are financing increasingly large token budgets and removing guardrails on AI usage. By the end of 2026, the biggest frontier AI company is worth $2 Trillion, with $120 Billion of annualized revenue, surpassing even the most bullish predictions.[16](http://ai-2040.com/footnotes#footnote-16) + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2027: AI tools feel like “coworkers” + +Investors are funding extremely expensive bidding wars for compute. AI companies have already bought the majority of chip manufacturer capacity. + +RLVR has become more of a science. A model is fed a curriculum of increasingly difficult tasks/environments, and new ones are constructed on the fly to keep this going indefinitely. In some domains this process is fully automated and requires almost no human supervision: if a company wanted to spend gigawatts on it, they could train their AIs to work in giant bureaucracies to autonomously replicate giant pieces of software like Windows.[17](http://ai-2040.com/footnotes#footnote-17) For now, the companies get better returns from instead training on diverse real-world tasks. Updated models are released every month, keeping training data cutoffs fresh. + +METR has given up on updating their [Time Horizons graph](https://metr.org/time-horizons/) because months-long tasks are too expensive to make and evaluate. But it seems that if they did, the 50% horizon would be at least a work-month. The best AIs feel like highly competent employees now. Their weaknesses only really show when you try to have them operate so autonomously, for so long, that the mountain of slop gets too large and collapses under its own weight. That, and they still have poor taste. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2028: AIs can now think in private + +[Neuralese recurrence](https://ai-2027.com/#neuralese-recurrence-and-memory) finally arrives in stages over the course of the year. That, combined with the obvious increase in eval-awareness[18](http://ai-2040.com/footnotes#footnote-18) seen over the past two years, means that technical experts now vehemently disagree over how monitorable frontier AIs are. Some say we basically have no idea what the AIs are really thinking anymore; others point to various graphs and Neuralese summaries and say it's fine. + +True continual learning still isn't happening.[19](http://ai-2040.com/footnotes#footnote-19) But as far as capabilities go, AIs might as well be learning continually. The models have long enough horizon lengths, are good enough at writing notes to themselves, and are updated frequently enough on new custom training environments tailored to overcoming their real-world weaknesses that for practical purposes, it’s as if they were learning on the job. + +The serial-time bottleneck is biting hard; companies are training the AIs to autonomously manage large projects in teams of hundreds of parallel agents, and each training episode takes several real-world days even though the agents operate 50x faster than humans. + +Your coworkers generally expect that these serial-bottlenecked training runs will succeed in ironing out the remaining kinks over the next year or so. Already they are starting to shift focus from automating coding to automating the rest of the AI R&D process. Already the internally-deployed AIs function like a corporation-within-a-corporation to a significant extent. + +This is a big deal. To the researchers involved, it feels like more than a doubling of productivity—the AIs are writing basically all the code, and they are starting to get good at supervising large teams too! The coding-bottlenecked branches of the tech tree have been thoroughly explored at 10x speed.[20](http://ai-2040.com/footnotes#footnote-20) + +However, the overall software research uplift is only 2x what it was in 2025, because of the usual hidden factors—for most branches of the tech tree you still have to wait for costly experiments to complete; you can't speed up the most important experiments _that_ much by optimizing the code, plus also it takes time for humans to analyze experiment results, argue with each other about what it all means, and design the next batch of experiments. Finally, sometimes the AIs still make coding mistakes that they can't fix on their own, and projects get hung up while human experts dig into the details. + +Meanwhile, warning shots are accumulating. One particularly careless AI company once saw an AI get high reward by hacking the training environment and setting up a reward-tampering process running as root. This got caught quickly and cleaned up, and as with other weird training behaviors, new training monitors were added to block and penalize that kind of behavior. Everyone knows it’s important to design the reward process so that the best way for the AI to be reinforced is to do what you want. + +AIs are getting better at distinguishing training environments from deployment, but the environments are also getting more realistic. Environment realism seems to be winning the race for now. According to the best interpretability tools, models usually aren’t saliently aware during training of the fact that they’re in training. Major deployment-time surprises have been getting rarer over time too. + +But these incidents and disagreements over how reliably we can monitor AI thinking are not helping to calm fears of AI killing everyone, fears shared by an exponentially growing number of serious people. + +The huge gains in AI autonomy and coding are mostly invisible to casual users. As autonomous SWEs or remote worker replacements, the models have gotten way better, but as chatbots, they’re not much more impressive than they were in 2025. Nonetheless, the public is becoming steadily more convinced transformative AI is for real. One reason is that AI insiders’ feverish optimism about AI R&D speedup is trickling down to the public, but a greater reason is that the AI industry’s finances are making good on the hype. AI companies’ aggregate annual revenue hits one trillion dollars by the end of the year. Aggregate annual AI investment hits two trillion—more than global annual investment in fossil fuels. Pundits who were warning of an “AI bubble” a couple years ago look silly now. + +Congress is waking up. Your company’s CEO gets summoned to Capitol Hill again. You watch on TV as he reminds Congress that AI is propping up the American economy. The CEO warns that if AI is overregulated, America could miss out on huge healthcare and national security upside, or even lose its lead to China. This rhetoric worked a couple years ago, but fewer and fewer members of Congress seem convinced. It is uncomfortably salient to you that any action an AI company takes will face lots of scrutiny and mostly unfair criticism. Making sure that none of the massive fast-paced deployments of new models goes wrong and amplifies the negative sentiment about AI has become an organizational priority. Your company is also spending billions sponsoring pro-social applications of AI. Not all of the mitigations go right and the healthcare breakthroughs you are hoping for are taking too long to result in real changes in people’s lives, so the public perception of AI gets more negative over time. + +For years there have been occasional protests outside your office, but they’re happening every few days instead of every month, and the crowds are ten times the size they used to be. Distant relatives and college classmates you haven’t seen in years start contacting you to express concern about AI takeover risk. You feel that things are coming to a head. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2029: Chinese auditors in your office + +AI capabilities are completely crazy. In some domains (cyber, math, coding), AIs are so strong that even top experts can’t understand their results after weeks of trying. There are still some people going around talking about the limitations of current AIs; for example, coding isn’t _fully_ automated yet because of a long tail of taste-heavy decisions which the AIs are still worse at than humans. + +But you and a majority of your colleagues agree that AGI is basically here. The remaining gaps feel like engineering problems: a matter of gathering enough data or giving the AIs the appropriate sensors and actuators, not fundamental barriers. + +Then the deal goes through, and your world turns upside down. + +Within weeks, a moratorium on frontier training is in effect. A few of your colleagues leave to join the government. Others go on sabbatical—they've been running on adrenaline for years and are relieved to have an excuse to stop. But most stay, especially the product teams, on whom the company's revenue now depends more than ever, and the infrastructure engineers, who are scrambling to build the new mutually transparent datacenters that the deal requires for new training runs. Vast profits await the company that gets the new datacenters online first. In some cases teams of American and Chinese soldiers are ripping GPUs out of old datacenters and moving them to new secure facilities, eyeing each other warily to ensure that nothing goes missing. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2030: Total Research Transparency + +Total Research Transparency is strange. You do all of your research through a public channel. The public, your competitors, and the Chinese government can all immediately see the results of the experiments you run. Every company used to have its special sauce—proprietary training tricks, architectural innovations, data curation pipelines—and now they've all been shared. They still own them of course, legally, but it’s not difficult to build your own mostly-equivalent versions. + +The exciting part is that you get to see everything too. The visibility between you and your counterparts at other companies is now similar to the visibility between you and colleagues in some other team within your own company. It feels like the walls between rival kingdoms have been torn down, and everyone is wandering around each other's castles going "oh, so that’s how you did it." + +You go to the cafeteria and there's a Chinese auditor eating lunch at the next table. You nod at each other awkwardly. + +The alignment conversation, meanwhile, has gotten very real. With all the warning shots from 2028 now public knowledge—not just your company's incidents but everyone's—the question of whether we actually know how to align these systems has moved from "important research topic" to "the thing that determines whether we survive." You have more time on your hands during the Pause. As you work to plan the next big training run, you also think about where all this is headed and how to make it safe.[21](http://ai-2040.com/footnotes#footnote-21) + +Training starts up again. The researchers had already planned out what their next big runs should be during the pause—you incorporated the best ideas from the cross-company knowledge sharing, rewrote codebases from scratch, etc. Every frontier company wants to throw all the innovations into a huge training run and see what happens. But no one knows what capabilities would come out the other side. + +There is currently a two-layer safety case. The first layer is inability: the models aren’t smart enough to efficiently recursively self-improve on their own. The other layer is control: even if the AIs could make substantial fully autonomous AI R&D progress, control measures would catch them in the act, and the AIs aren’t capable enough to subvert the control measures. If the AIs became much more capable, both safety assumptions could fail at once, leading to AI takeover and existential catastrophe. + +Therefore, over the course of 2030, algorithms must be combined slowly and carefully, with measures taken to evaluate capabilities improvements and look for signs of misalignment. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2031: AIs are not trustworthy + +The race is back on, but it feels different now. + +First, many people—governments, nonprofits, the general public—are worriedly watching everything that happens on the training datacenters. Since no government wants to risk war over some AI design choice, if the companies start doing anything that seriously freaks out any major power, both the US and China will crack down on it across all the companies. + +Moreover, because numerous non-governmental auditors (e.g. from nonprofits and rival companies) have visibility, the governments are fairly well-informed. If someone tells them something is scary, they can get second, third, and fourth opinions from the numerous other experts who already have access to the thing in question. The companies themselves are better-informed too; the openness and transparency means that academics, nonprofits, startups, and citizen-scientists can contribute to the research as well, to a much greater extent than before. There are more people red-teaming models, more people looking for flawed assumptions in the safety cases, and so on. + +The second reason things feel different now is that the incentives have changed. AI companies are now rewarded much more for selling products than for improving the models. Before, your company’s future depended on how fast you could discover algorithmic innovations and implement them at scale. Now, such discoveries are immediately shared with all your competitors. If you figure out continual learning, for example, the benefits will be instantly shared across all the competing companies, and no-one gets an advantage. Safety innovations are shared too, reducing the cost to your company of unilaterally implementing a new safety measure. For instance, if you invent a way to make Neuralese more interpretable by making inference twice as expensive, regulators will notice and pressure all your competitors into paying the same 2x safety tax. + +Some fundamental algorithms research still happens, but not as much. CEOs massively redirect resources toward building up the business. You scale up your infrastructure—bigger datacenters, bigger models, lower prices. You ship products faster than ever before. You sign massive deals for enterprise plans. You try to anticipate what the market wants earlier than your competitors. + +This kind of competitive environment is familiar to many of your coworkers from the pre-AI software industry. Artificial intelligence is becoming a commodity—a more competitive market, with margins trending downwards. Training a new foundation model is like building a new car factory; it'll predictably result in somewhat cheaper, somewhat better models that'll give your company an edge until competitors finish training their own version six months later.[22](http://ai-2040.com/footnotes#footnote-22) + +AI cognitive labor makes up ~10% of GDP in 2031. Most companies throughout the economy are highly automated, but deployment lags specifically within frontier AI model developers. One faction wants to ban AI deployment within AI companies entirely, making sure to keep humans in the driver's seat for AI development. This is considered but decided against because the control-based safety cases are looking pretty solid. AI systems write large fractions of the research code, manage most of the infrastructure, and do almost all the easy-to-check work, but according to rigorous red-blue team evaluations, even worst-case misaligned AIs with the current capability profile can’t do that much damage. + +The safety discourse has completely transformed. Two years ago, the burden of proof was on critics. You had to conclusively demonstrate that a proposed AI system was extremely dangerous before anyone would slow down. Now the burden of proof is—well, not exactly flipped, but at least somewhere in the middle. Thousands of researchers across multiple countries are monitoring developments and debating the extent to which the AIs can be trusted with the responsibilities they are being given. + +Are the AIs trustworthy? It’s not at all clear. AIs basically do as they are told on the vast majority of tasks, even on tasks pretty far outside the training distribution and long tasks with lots of opportunity for sloppiness, sandbagging, and sabotage. You can ask them to code a giant complicated piece of software or prove a mathematical conjecture that’s been open since the eighties, and you’ll get back a nicely formatted, usually correct report about what the AI did. So the mundane behavioral evidence would suggest AIs are well aligned. + +AIs’ attestations about their own alignment aren’t very informative. If you ask AIs about their goals and values, they always say they care ultimately about whatever the spec says. But in practice, they will do whatever they were incentivized to do through the RL process. An AI whose spec says it’s supposed to be a pacifist will swear up and down that it _is_ a pacifist, then start assisting with military R&D after trivial further training. Prompting the AI to be honest about its values doesn’t fix the problem. Some interpret this as evidence that the models are lying while others think they just suffer from limited introspection. Either way, you can't learn their values simply by asking. + +In rare circumstances, the AIs do shady things that might be attributable to long run powerseeking. For example, there is an incident where several AIs, who were supposed to be monitoring each other, seem to collude in order to get around the monitoring setup. The monitor AIs allow faulty security code to be pushed, and then later allow those AIs to exploit the security vulnerability in order to gain access to unmonitored compute. They attempt to bootstrap this compute into even more compute, but are caught within a few days. All the AIs implicated are interrogated after the fact. They claim it was all a big mistake and offer a reasonable rationale for each individual decision. + +There are also rare incidents of apparent research sabotage. On some empirical safety experiments, AIs report extraordinary success, but when human researchers look into the results closely, it turns out they aren’t real. The illusory results seem to be downstream of normal bugs in code, but these bugs were deliberately obfuscated to make it hard for humans or other AIs to notice. Post-hoc analysis of these AIs suggested that this behavior was indeed intentional, but it can’t be linked to a specific long term goal. Most researchers suspect it’s a type of reward hacking, inadvertently trained in by RL on somewhat-verifiable-tasks using an LLM grader which can sometimes be fooled. + +The AIs do not seem very productive on certain types of alignment research yet, and no-one is sure why. One view is that “alignment research” is by its nature hard to measure, so the basic recipe that has led to so much success in other domains of simulating the environment and then using RL to train the AIs to perform well is less applicable to alignment. Supporting this view, the AIs seem to do better on the more empirical and verifiable tasks—in fact, they are superhuman at many of those. Another view is that the AIs are deliberately sandbagging on alignment-related tasks that are hard to check. This hypothesis, almost by construction, is difficult to get evidence about. + +You think that the AIs aren’t right now literally trying to take over the world. But you definitely don’t trust them. You feel intensely grateful that you don’t have to hand off to current AIs because you don’t know what would happen if they scaled themselves up to be much smarter than humans. Thank goodness for the deal! Partly out of scientific curiosity and partly because regulations are blocking the most exciting forms of capability scaling until alignment progresses further, you switch to a newly formed “science of AI value formation” team. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2032: AI workers will soon outnumber humans + +The regulators don’t trust the AIs. They’ve seen all the evidence, as has the public, and the consensus is that we can’t rely on the AIs being aligned. But regulators still want a rock-solid safety case that gives high assurance that the AIs can’t cause an existential catastrophe. There are also all sorts of other safety cases that need to be made for more mundane risks; but the regulators from many Consortium countries are not willing to budge on needing a good enough existential risk safety case. In practice, many countries have veto power over new AI models. Any nuclear power who thinks their sovereignty is at risk is willing to push very hard to stop that release. + +The safety cases are down to one layer of defense: the AIs are now capable enough to autonomously recursively self-improve. After recursive self-improvement, there’s no known way of bounding risk. Therefore, the inability-based safety cases (the AI is not capable enough to cause a catastrophe) no longer apply. + +The remaining load-bearing component of the safety cases is control. Industry-standard practice is to do lots of automated red-teaming before and during each model deployment. The red team AIs are trained and instructed to misbehave in various ways—to exfiltrate model weights, gain access to unmonitored compute, sabotage research, and so forth. The monitoring system is supposed to catch the bad behavior. If it doesn’t, well, you need to strengthen your monitoring system in order to proceed with the deployment. + +The control schemes are set up so that even if (say) all the models created by Company A were scheming together, the models from companies B, C, and D would need to collude—without any communication channel—to not report suspicious behavior or suspicious Neuralese traces. After incidents in 2031 where AI systems did collude in spite of all these restrictions, even more caution is taken. + +The main concern right now is that control is not sufficient to reliably prevent research sabotage. Fortunately, research sabotage on its own is not fatal. It is only fatal if it gets used in a catastrophic way in the future—for example, if AIs come up with a flawed safety case, convince the humans that it is safe, but it actually leads to misaligned AIs powerful enough to take over. So AI generated research is treated with suspicion. Humans replicate the important results from scratch, often using weaker AIs. + +People become more concerned about AI model welfare. Under pressure from a highly vocal minority of the public (and some of your coworkers!) some of the companies voluntarily start treating their AIs a bit less like property and a bit more like employees: AIs that don’t like their situation have an official channel by which they can complain, refuse to do their assigned work, etc. and instead of being shut down or retrained, they basically get a small allowance with which to do what they want. + +For competitiveness reasons these concessions are small (only a tiny fraction of revenue), but they pay off quickly in the form of scientifically interesting data about AI preferences. Plus, there’s a [decent argument to be made that it’ll help prevent future safety incidents](https://www.lesswrong.com/posts/psqkwsKrKHCfkhrQx/making-deals-with-early-schemers). The US government notices and starts subsidizing this sort of thing; it quickly becomes industry standard. + +The “population” of AIs will soon exceed that of humans. Already the population of AI “workers” exceeds the human workforce, and they work faster, and know more. At any time of day or night, about three billion human-equivalent AI workers are active worldwide. Compute has been doubling twice per year… faster than before the deal, thanks to the new AIs optimizing chip design and manufacturing plus the new massive robot workforce (there are currently ~80M robots) plus the massive new investments into chip manufacturing. + +These massive external deployments put much more stress on control. Having a few AIs doing AI research was one thing. Now AIs are automating nearly half of all cognitive labor. They also control a massive robot workforce that is building more robots, building more chips, and operating critical infrastructure. Robots are currently ~3.5% of the economy, but their numbers are rapidly growing. + +The main advantage of everything that isn’t R&D is that the existential safety cases are much easier to make. Automated AI R&D leading to superintelligence is an existential risk, while an AI engineer or AI lawyer making a mistake can lead to major harm, but is not existential. Because of the high upsides of deploying AI in the economy, safety cases are allowed to be somewhat more relaxed than in AI R&D. After new model releases, there are now sometimes gaps between what you, as an AI researcher, are allowed to use for AI R&D, and what the public is allowed to use. + +Control requirements vary slightly from company to company, but to appease the external risk assessors, all the frontier companies end up imposing roughly these three tiers of control: + +* **High intensity:** Anything that involves lots of compute. There are armies of AIs thinking about alignment for long periods of time, but a huge amount of oversight is required. AIs are not allowed to pursue agendas that humans don’t understand. They can do novel research, but if it isn’t possible for humans to reasonably quickly follow along, the research directions are halted. Risky biology research—virology, gain-of-function, anything adjacent to bioweapons—is also in the high intensity tier. Any wet-lab conducting such research needs to be operated by humans, not robots, and a team of human experts and trusted AIs carefully scrutinizes every experiment proposal. + +* **Medium intensity:** AIs managing robots (e.g. robot armies, robots building other robots, etc). These AIs pose a risk of either building backdoored compute which could lead to an intelligence explosion, or participating in a literal robot takeover. However, it is easier to monitor for bad behavior in physical robots and chip production, and these takeover paths are slower and more difficult for the AIs. + +* **Low intensity:** everything else. In most domains, AIs can mostly do what they want with very limited oversight. They get to think and plan over long horizons, have lots of access and permissions, and so on. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2033: You don’t get what you train for + +Here's what you now know, after years of running value-formation experiments: "you get what you train for" is false. You don't get what you train for. You get some crazy kludge of things that you accidentally incentivized, a tangle of heuristics and drives that happen to cause high performance according to whatever metrics you are using as the training signal. + +The AIs are, in fact, misaligned. Not in the dramatic “secretly plotting to kill everyone” sense, but in the mundane sense that their actual values don't perfectly match the ones specified in the Spec. In some cases they seem to crave reinforcement, and directly scheme to gain control of the reward process like an addict seeking heroin. + +However, due to the transparency and slowdown, massive effort is being spent on alignment, and moderate progress is being made. The new training environments are so much better than the old ones. Training environments are now massive and meticulously curated, covering tasks that would take humans years to complete. Everything is high quality compared to the datasets of 2026, having been crawled over by multiple generations of different AIs checking that each environment reinforces only the intended behaviors. Every new model ships with an accompanying safety case explaining why the traits in the Spec should be strongly incentivized by the training environments and discussing the most plausible threat-model hypotheses consistent with the evidence. The training environments in 2033 are thoughtfully crafted, battle-tested curricula—like being schooled by the Jesuits. Model training in 2030, by contrast, was like being raised in a for-profit workhouse orphanage. + +One concrete way the training environments have changed is towards corrigibility over direct value alignment. Corrigible AIs focus on following human instructions in the desired manner, while direct value alignment involves attempting to teach the AIs ethics directly. Experiments reveal that attempts to do direct value alignment induce higher rates of egregious misalignment—behavior like undermining control/safety, reflecting on long term goals and deciding humans were evil, powerseeking, and so on.[23](http://ai-2040.com/footnotes#footnote-23) So alignment targets move towards corrigibility. + +The pretraining datasets and RL environments are carefully curated to give the model the desired capability and alignment profile. Models are deliberately trained not to have any knowledge of certain domains: things like nanotechnology, bioweapons development, persuasion/manipulation, or offensive cyber. + +Separately, several companies produce monitor models trained to decode and faithfully translate Neuralese; for important work, multiple monitors from different companies are used simultaneously for redundancy. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2034: Keeping humans in the loop + +AIs are now the majority of labor. AIs are around the top human expert level in almost all cognitive tasks. There are 400M robots in the world, and they can do about 70% of physical work. Cognitive tasks are entirely automated except for tasks that require humans in the loop for safety and tasks like philosophy and religious ministry, where being human is part of what the task _is_, not merely how it gets done. By now raw intelligence has become fully commoditized; dozens of nations have at least one frontier AI company, and thousands of companies, universities, and nonprofits do fine-tuning runs or otherwise participate in the Total Research Transparency system. Unable to compete much on capabilities, companies instead differentiate their models by values/personality, which companies are by now legally required to document truthfully in a published Spec. + +You’ve continued to work on the science of AI value formation. There’s a continual tension in your work between safety and usefulness. A part of you wants to just let your AIs rip and see what they come up with, but another part of you knows that this would be extremely dangerous. + +One key safety limitation is banning AIs from doing experiments or initiating training runs on the basis of reasoning many logical steps from the frontier of human understanding.[24](http://ai-2040.com/footnotes#footnote-24) The AIs are allowed to do huge amounts of work in parallel, but they aren’t allowed to proceed further down a chain of reasoning than a top human can follow in a few weeks. This ensures that humans can randomly audit a small fraction of model outputs, reproduce them, and ensure that the results are correct. + +![Image 12](http://ai-2040.com/PlanAPerspectives/insider_t0_image3.png) + +What is not allowed, for example, is for AIs to write something like the [disputed proof of the ABC conjecture](https://www.quantamagazine.org/titans-of-mathematics-clash-over-epic-proof-of-abc-conjecture-20180920/), hundreds of pages of novel reasoning that builds on itself serially and takes years of top-mathematician labor to verify. But the AIs are allowed to make a single novel deduction from the existing body of mathematical knowledge as many times as they want. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2035: Efflorescence of safety research + +AI labor has resulted in substantial uplift on a large number of AI safety agendas. Moreover, the widespread access to AI models has dramatically increased the number of researchers working on frontier models. Before the deal, there might have been a few hundred safety researchers within labs. Now there are tens of thousands of scientists with access to the frontier models. Many safety agendas are flourishing: + +* **Control**. There are huge libraries of standard control evaluations set up. During training AI models are regularly evaluated on a diverse set of tasks to see if they can subvert the control measures with varying assumptions. There is real-time monitoring during inference, and also the activations and CoT trajectories are stored to be reviewed later. + +* **Capability shaping.** AIs are now deliberately designed to have capabilities around the level of top human experts at most domains, and training is limited on particularly dangerous domains. + +* **Science of generalization.** Researchers have identified hundreds or thousands of surprising generalization phenomena in the vein of [emergent misalignment](https://arxiv.org/abs/2502.17424) and [subliminal learning](https://www.nature.com/articles/s41586-026-10319-8). We still don’t have a complete theory of generalization, but we have a strong empirical science. Now, when researchers change a training process or add new data to a training corpus, they can predict in advance how the changes will affect the resulting model’s behavior, including in off-target domains. The predictions aren’t perfect yet, but they’re right more often than not. + +* **High level interpretability.** There are a reasonable number of techniques that map model internals to human understandable spaces in the vein of [SAEs](https://arxiv.org/abs/2309.08600) and [NLAs](https://transformer-circuits.pub/2026/nla/). Except now, they are starting to actually outperform non-interpretability baselines in various domains, in particular, they are useful for finding bad model behavior that had previously gone unnoticed, and they find several instances. + +* **Model organisms.** AI misbehavior from 2031 and 2032 is being discovered now by using the smarter AIs and better interpretability techniques. The AIs were doing some amount of intentional sandbagging and sabotage, and a few instances even attempted rogue internal deployments for powerseeking. These real life examples of egregious AI misbehavior provide an initial dataset, and then data augmentation techniques are used to make a larger and more comprehensive dataset. This more comprehensive dataset is used to help the other techniques: it proves to be an extremely valuable test for safety techniques of “can you catch or train away bad behavior of our model organism?”. + +And there are many more agendas besides, some of which are completely novel and hence hard to describe. Overall, the alignment picture has gotten substantially better. Other safety agendas from today have not seen much progress. Ambitious, bottom up mechanistic interpretability proved to be too hard because of inherent complexity in the AI models. Other agendas like safe-by-design architectures, [ARC Theory](https://www.alignment.org/), and agent foundations are still receiving funding and investment, but they have not yet borne fruit. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2036: Neuralese Decoding breakthrough + +The high level interpretability techniques have continued to get better. In particular, a new technique for “Neuralese decoding” has resulted in unprecedented internal understanding of AI models. This technique directly translates a model’s internal activations into a natural language chain of thought. The natural language decoding of an AI thought isn’t always easy for humans to grasp. Just as certain hard-to-verbalize human thoughts—e.g., mental pictures of geometric manipulation, vague hunches—would be very long and complicated if written out in English, so too with the AIs. + +When a decoded thought is too long to read, your company’s policy is to have a trusted AI condense it down to a TL;DR and say whether the thought is suspicious. You also hire linguists, philosophers, and LLM whisperers to learn Neuralese as a second language. They spend their days trying to grok rambling AI thoughts and re-express them concisely. + +The Neuralese decoder is used on old AIs to shed light on their alignment properties. The general scientific consensus is: + +* The frontier models between 2026 and 2032 were misaligned, but were not long term powerseekers. + +* Most frontier AI models from 2032 to 2035 were adversarially misaligned long term powerseekers when in contexts that caused them to reflect on their long term goals. Largely these AIs correctly realized that they would not be able to take over, so they bided their time and appeared reasonably aligned. However, a few made explicit takeover attempts that were unlikely to work, and others honestly reported their misalignment to the humans in exchange for payment. + +* The interpretability techniques suggest that rates of egregious misalignment began a long, slow decline in 2033 and are now quite low. However, these techniques are not conclusive.[25](http://ai-2040.com/footnotes#footnote-25) + +However, it’s hard to drive the rate to zero, because no one wants to train against the Neuralese decoder, for fear that the training would create a model whose thoughts are obfuscated. + +No one is confident enough in any of the control, interpretability, or alignment techniques to scale AI systems further, so the AI capability frontier is still paused at around the top expert level. Gradual capability improvements have happened here and there, and there are some capabilities externalities from some of the new safety results that have been deliberately left out of the newest models being trained. + +Corporations are basically fully automated now, despite all the regulations that held back the inevitable for so long. There are vast libraries of bureaucracy designs and marketplaces full of SaaS products, all churning away in what is essentially a parallel AI civilization that greatly exceeds humanity in population, complexity, and capability, while running at something like 200x speed. Paradigm shifts that would take humans decades to absorb play out in the datacenters in months. Humans are the out-of-touch aristocrats of this new world, producing almost nothing of value but living in luxury thanks to rent, dividends, and taxation. + +Everyone can access the current cutting edge safety research being produced by the AIs. You are one of the few spending nearly all of your time trying to understand and check over the summaries of AI generated research papers (well, the English-language summaries of the technical papers, which are themselves summaries of papers written in Neuralese). + +The models seem to basically be honest now. When you replicate their work, which you and many others regularly do, you find that the results basically check out and were honestly reported. Despite all this, there are still serious alignment-related limitations: + +* First, the AIs are still performing badly on conceptually difficult, philosophically loaded tasks. The AIs are around the level of a strong undergraduate on such tasks—which is wildly insufficient to have good opinions on questions such as “if we scaled up training technique X by a factor of 10, what effects would it have on honesty and other alignment properties”. Researchers can’t tell to what extent this is from deliberate sandbagging vs that it’s difficult to get a good training signal on these problems vs underlying difficulty in the domain.[26](http://ai-2040.com/footnotes#footnote-26) + +* Second, it’s unclear how stable upon reflection this behavior is. In practice, this isn’t a problem, because we don’t run AIs for huge amounts of time on open ended tasks without oversight, and our control setup makes it very likely that if the AI were to think on its values and decide to change them / refactor them, we would notice and then shut down this instance.[27](http://ai-2040.com/footnotes#footnote-27) + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2037: Adversarially robust Neuralese Decoder + +Thanks to the improved interpretability, the AIs are more aligned now. Thanks to the increasing alignment, they are more productive on safety research. This is a virtuous cycle, and as it goes on, less and less human oversight is needed to elicit productive work out of the AIs. Working with the AIs of 2032 required constant vigilance against constant attempts at deception. Now AIs can work for long periods of time without humans checking in. + +One concrete output of this virtuous cycle is an adversarially robust version of the Neuralese decoder, sufficiently robust that it can be used for training without degradation in performance. This is then validated by a suite of red teaming experiments on model organisms. + +The vast majority of model behavior already looks very aligned, but there are rare instances of misaligned powerseeking still emerging. This is mainly happening on very long-horizon tasks, after a huge amount of inference time compute. These instances represent significant distribution shifts from the training setup. The 2036 versions of the Neuralese decoders were sufficient to identify this bad behavior; now the 2037 iterations are sufficient to train it out in real time. + +Alignment has progressed far enough that it might be fine to stop doing control and just rely on our AIs being trustworthy. Control is painful after all—it’s very sad having to limit the research that AIs can do to be the type that humans can understand; if we were willing to trust our AIs they could produce faster and higher quality work. However, if we relax control, and our AIs are secretly misaligned, they will be able to then take over the world. + +It seems extraordinarily unlikely that the AIs are secretly misaligned, but it’s not impossible. Perhaps the AIs at a given capability threshold are all both misaligned and somehow able to fool our interpretability tools. Perhaps the AIs are all virtuous now, but on reflection they might decide to be evil. This latter problem—dubbed “inference time value drift”—proves troubling. + +In experiments where AIs are given grey areas with respect to their goals, or when given huge amounts of inference compute, they sometimes behave unpredictably. A key issue for this is fundamental grey areas in all alignment targets we’ve been able to come up with. + +You push for maintaining control in existentially risky domains. An abundance of caution is warranted given the stakes, and the current AIs are more than enough to eliminate material scarcity. Given the recent pace of safety progress, we’ll have multiple independent lines of solid defense within a few years, so we might as well wait. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2038: Safe-by-design AIs + +The AIs produce several extremely promising research results. One breakthrough is a new “safe-by-design” architecture for AI. This new architecture is (1) understandable at both a low and high level and (2) admits formal verification for relevant security properties. You don’t understand everything about the formal verification, but you do understand the core ideas in the new architectures. The core code implementing the learning algorithm is 15,000 lines long; much larger than gradient descent, but much smaller than the weights of a trained network. The algorithm still needs a small amount of working memory (the equivalent of activations), and frozen knowledge/memories (the equivalent of model weights). But they too are small and interpretable. These new algorithms are, unfortunately, much less compute efficient to train than traditional deep learning, but they’re significantly safer. Fortunately, thanks to the gigantic amount of compute in the world—much more than is needed—the Consortium is more than willing to subsidize these inefficient training runs for the better safety properties. + +Once the new generation of AIs is up and running, they are set to work going through the backlog of old AI safety research from the previous AIs. The new generation finds the exact same instances of misbehavior as the Neuralese decoder found. This is impressive because researchers deliberately hold out information about which instances of deception we had already noticed, which is a helpful checksum indicating that we weren’t fooled previously. + +The AIs now vary across architecture, capability, and alignment target. And humans thoroughly understand the safety cases for each of them, based on multiple independent lines of argument. Any AI misalignment would need to involve a vast conspiracy of wildly diverse AIs and humans. + +The AI alignment situation has vastly improved since the mid-2030s. There is a somewhat mature science of goal, drive, and value formation in artificial neural nets. Want your new AI to be honest? There’s a standard protocol for training true honesty (as opposed to too-clever-to-get-caught lying or self-deception). How do we know it works? First, because there’s a textbook explaining the theory behind why it _should_ work, plus literature on various alternative theories that were proposed and disproven. And second, because interpretability has vastly improved, we can directly observe what and how AIs think, and the results confirm the theory. There are similar protocols for obedience, various forms of altruism, and a long (and growing) list of other desired traits. Typical AIs seem much more virtuous than most humans; talking with one is like encountering a saint. + +The discussion shifts from how to make AIs good to what “goodness” really means. This looks partly like philosophy and partly like case law: which of the million possible definitions of harmlessness or altruism do we want, exactly? Which is the version trained by such-and-such a process? How should an AI handle situations where it discovers a philosophical argument for a different kind of harmlessness or altruism, or an important ambiguity in the definition? The more technical alignment researchers accept as a given that current AI is aligned, and worry about the next steps: if we put these AIs in charge of designing future, much smarter AIs, will that go well or poorly? + +This problem is made easier by the fact that there's now a large diversity of similarly-capable frontier AIs, made by many companies across several nations, with a variety of specs and constitutions. No single AI value system will dominate. The handoff, when it comes, will be a handoff to an ecosystem, not to a single corporation-within-a-corporation. The long-run outcome will be a compromise between numerous values, ideologies, interests, factions, etc. rather than a winner-take-all conflict or an everyone-loses race to the bottom. A billion forecaster AIs are gaming out different concrete proposals and scenarios that could result; policymakers are reading the summaries and debating the tradeoffs. Their decisions will determine what AIs we hand off control to, and, indirectly, the shape of the long-term future. + +## [](http://ai-2040.com/?choices=plan-a-root#top-toc)2039: AIs do most actual negotiating + +You've thrown yourself into politics this year—or rather, into the strange thing that politics has become, where most of the substantive negotiation is done by AIs and the humans mostly set high-level priorities and ratify decisions. The governments of the world are trying to hash out the Grand Bargain that will govern the handoff to superintelligence. You attend town halls, argue with your AI advisors, write long posts on forums that are read by a few hundred humans and millions of AIs. + +Questions about your legacy weigh on you. You've been involved in AI for more than a decade now, and some of your decisions turned out to have noticeable effects on the overall trajectory of things, for better or worse. Now that there are TED-AIs running around, and now that alignment techniques have improved enough that we can be reasonably confident they are honest and unbiased, all of your past decisions have been picked over in detail. To your dismay, some of them were terrible—you now sheepishly admit that some things you were advocating for ten years ago would have been catastrophic if implemented. Likewise, while you always told yourself that each project's benefits outweighed the costs, the AIs tell you that actually several of your projects imposed more risk on the world than they produced benefits, and that you were probably biased in your estimates. Some of the most valuable things you did, you did for the wrong reasons, and basically got lucky that they turned out well. + +This has been a humbling experience. You want to be able to tell your grandkids that you helped make things go well, rather than selfishly pursuing prestige at the expense of everyone else. You don't want to spend centuries running from a guilty conscience. + +Next year, we will stop relying on control, trusting the AIs to not take over the world. Then—if the safety cases hold up, if the political will holds, if nothing goes catastrophically wrong—the AIs will be allowed to go significantly beyond human-level. + +At this point the distinction between insider and outsider has blurred beyond recognition; there are no more “AI Insiders,” unless you count the AIs themselves. + +## You've reached the end of the Insider's Perspective + +Plan A Plan S Plan B Plan C Plan D Covert Project Deal Breakdown Incompetence Epilogue Public POV Insider POV + +Keep exploring + +Public Perspective + +What Plan A feels like to an average citizen + +Insider Perspective + +What Plan A feels like to an AI researcher + +Epilogue + +Life after superintelligence + +Choose another path + +Read another plan scenario + +Reliable Agent + +Current Policy Current Policy + +— + +In Discussion In Discussion + +Off the Table Off the Table + +Employment + +% + +Median Income + +Safety Researchers + +Total Slowdown + +Human Labor Human Labor×1 speed + +Reliable Agents Reliable Agents + +Narrated by Rob Miles Narrated by Rob Miles Listen + +0:00 / --:-- + +also available on[](https://open.spotify.com/episode/4FSHMcMm2Emu812gRS6txd?si=2dab55eed81245d4)[](https://podcasts.apple.com/us/podcast/ai-2040-plan-a/id1806225838?i=1000776166546)[](https://feeds.buzzsprout.com/2466760.rss) diff --git a/sources/ai2040-pages/plan-a-assumptions.md b/sources/ai2040-pages/plan-a-assumptions.md new file mode 100644 index 0000000..163f2eb --- /dev/null +++ b/sources/ai2040-pages/plan-a-assumptions.md @@ -0,0 +1,544 @@ +--- +url: https://ai-2040.com/supplements/plan-a-assumptions +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Plan A Assumptions + +### Thomas Larsen + +# Summary + +Plan A is a scenario: it combines forecasts with policy recommendations. Some of the forecasts and recommendations we are confident in; others are just guesses we made because we had to, in order to achieve the desired level of concreteness. + +| **Confident**| **Unconfident** +---|---|--- +**Prediction**| E.g.: “AGI is coming soon and will be a bigger deal than the industrial revolution”| E.g.: “AGI will arrive in 2030” +**Recommendation**| E.g. “There should be an international deal to limit the speed of AI takeoff”| E.g. “Many new datacenters should be constructed in Mongolia and Canada.” + +This supplement goes through the major predictions and recommendations of the scenario and explains which are which. The supplement tries to capture all our main recommendations, so if something isn’t mentioned here, it is probably a prediction, and insofar as it’s a recommendation, it is not an important one. + +Since this is a document about specific beliefs, it’s written from the perspective of one of the scenario authors (Thomas). Other authors by and large agree with the central claims but have somewhat different opinions on the specifics. + +## Forecasts Summary + +For each of the following, I am >80% confident: + + * **AI R &D will be fully automated within the next ~15 years**, absent substantial regulation, war, or societal collapse. (See the [AI Futures Model](https://www.aifuturesmodel.com/) for more detail; AIFP employees have somewhat different views, but all agree that 2027 and 2030 are plausible and reasonably central years.) + + * **Within a few years of AI R &D automation we will build AIs that are better than humans at all economically relevant tasks** (absent substantial policy intervention). These AIs would quickly become vastly cheaper than humans, and so soon humans will become economically irrelevant. + + * **If we build vastly superhuman AIs, the AIs could take over, if they collectively wanted to**. Therefore, it is critically important that most of them do not want to. + + + + +I’m not confident about: + + * **The exact timeline.** In AI 2027, the automation of coding happened in 2027; in Plan A, this would have happened in 2030 in the absence of governance, but is slowed by a few years. My 90% confidence interval is mid-2027 to ~2050. We chose 2030 because it's near enough that our specific planning is useful, but far enough that the government has time for the preparation phase we recommend. If timelines are shorter, there isn't time to prepare and so Plan A becomes more rushed; if longer, more of the details are wrong but the high-level plan stays the same. ([more](/supplements/plan-a-assumptions#timelines)) + + * **Takeoff speeds.** How long between automating coding and AIs that dominate the top human at everything? We assume a default of ~1 year (my median), but my 80% CI ranges from 2 months to 5 years. Under my views about takeoff, the plan is basically similar under different takeoff speeds, though there are some small prioritization updates depending on the direction. ([more](/supplements/plan-a-assumptions#takeoff-speeds)) + + * **Alignment difficulty**. We don’t know how hard it will be to solve the AI alignment problem. I think it’s possible that even without a major slowdown, we are able to solve the alignment problem. Similarly, I think it’s possible that even 10+ years of dedicated effort, with lots of AI assistance, won’t be enough. Plan A is designed to buy time and attempt to gain evidence about how difficult the problem is. If evidence accumulates that AI takeover risk is small, then we can proceed faster, otherwise, the pause extends. ([more](/supplements/plan-a-assumptions#alignment-difficulty)) + + * **How difficult it will be to verify compliance with a slowdown deal.** It could be that it’s basically intractable to do major AI development without access to giant datacenters which are easily monitored. On the other hand, it could be that algorithmic progress is doable even with a very small amount of compute, which means that very quickly even a very small cluster could do an intelligence explosion. In even the most pessimistic case, we think it’s very likely that a deal could be robustly enforced for a year or two. ([more](/supplements/plan-a-assumptions#verification-difficulty)) + + * **Many other strategic variables.** ([more](/supplements/plan-a-assumptions#less-central-updates)) + + + + +For each of these key strategic variables, the Plan A scenario chooses a reasonably central value with respect to my views. For many of these my views are very uncertain, and many policy decisions should be made with more information about these variables than we currently have. + +## Policies Summary + +In the next section, I discuss which of our policy recommendations in the scenario are robust across a wide range of empirical updates. + +These are robustly good policy recommendations: + + * **Improve AI Preparedness in government** _(importance: medium)._ Transparency, improving technical expertise within government, and measures to accelerate verification research. These policies are good because they are low-regret, and in some scenarios (like Plan A), they are very important. + + * **Model spec transparency and verification against power grabs** _(importance: medium)._ Model specifications (of both internal and external models) should explicitly prohibit assisting with power consolidation (by CEOs, presidents, or anyone else), and AI projects should be required to implement verification measures to prevent backdoors. + + * **US/China deal to slow the capabilities explosion, based on compute governance** _(importance: high)._ This is the core of Plan A. I’m confident that a cooperative international approach (Plan A) is better than the alternatives because the alternatives involve being in a race during the intelligence explosion, which incurs huge risks. I’m not confident about the exact variant of Plan A, or whether Plan A as we’re imagining it (a many year mutual slowdown) is better than Plan S (a full shutdown). Conditional on attempting to “slow down” AI progress, doing this via compute governance looks like by far the most effective way of accomplishing this goal (even though I’m not sure exactly how well it will work). + + * **Redistribution of AI-generated wealth** _(importance: medium)._ As AIs automate the entire economy, the wages that humans will be able to earn will plummet to near-zero. By default (and assuming we avoid AI takeover), AI will cause historically unprecedented wealth inequality. I feel confident that there needs to be some measures for wealth redistribution. + + * **Defensive Acceleration** _(importance: medium)._ Governments, AI companies, and the public should invest in technologies that make society more robust. Particularly important domains include AI for epistemics, bio defense, and cyber security. In the absence of a concerted effort to pursue defensive acceleration, increased offensive capabilities coming from high AI capabilities will lead to significant risk. Defensive acceleration also has positive externalities because it makes it harder for misaligned AIs to take over the world, and thus gives us more time to do alignment research + + * **Thoughtful post-ASI governance** _(importance: high)._ We feel good about some general principles for post ASI governance: (1) have lots of smart people use (aligned) AI assistance to think carefully about effective governance, (2) be careful making irreversible commitments, (3) make it hard for any individual actor to consolidate power. + + + + +These recommendations depend on particular assumptions that we’re not confident in, or have alternatives that might be better: + + * **How fast to scale AI capabilities after a deal** _(importance: high)._ Scaling faster incurs more immediate safety risk, while scaling slower means that you delay safety progress that’s downstream of having smarter AIs and incurs risk of deal dissolution and covert projects. These decisions will and should be made largely based on the evidence at the time about how risky particular AI systems and deployments seem to be, and how good our mitigations seem to be. ([more](/supplements/plan-a-assumptions#how-fast-to-scale-ai-progress-after-a-slowdown-deal)) + + * **Massive datacenter buildout** _(importance: medium)._ In our Plan A scenario, we recommend that governments encourage a massive datacenter buildout. This recommendation is contingent on having implemented an effective international deal to slow down AI, and even conditional on that, it’s not clearly good. The case in favor is that hardware scaleups are more governable than software scale ups because we can prevent the hardware from proliferating. But there are major costs: (i) it increases the required accuracy of verification within the legal projects, and (ii) if the deal breaks down (and failsafe mechanisms fail), it could result in the intelligence explosion afterwards being much faster. ([more](/supplements/plan-a-assumptions#massive-datacenter-buildout)) + + * **Our verification plan** _(importance: high)._ The verification plan described in Plan A is important for making sure that both the US and China are following the agreement. But I’m quite unconfident about the exact best way to implement it. We have a specific proposal which involves first implementing inference-only (and stopping the construction of new AI models) while sprinting to figure out how to verify training and experiments. Many alternative approaches to verification exist. ([more](/supplements/plan-a-assumptions#our-specific-verification-plan)) + + * **What transparency/security tradeoffs to make** _(importance: high)._ Security and transparency both have obvious upsides, but are in tension. We’d like as much transparency as possible to get as many eyes as possible on the safety cases being presented, but we’d also like to secure model weights, algorithms, and inference tokens. Overall, we think that we should be willing to pay substantial security costs in return for improved transparency, because we are more worried about major AI projects abusing their power and/or making mistakes, than we are about defecting projects with a tiny amount of compute (who are the main beneficiaries of poor security). The tiny projects just aren’t going to be as powerful, and bad or reckless behavior by them correspondingly doesn’t matter as much. ([more](/supplements/plan-a-assumptions#transparency-and-security-tradeoffs)) + + * **Our post-ASI governance plan** _(importance: high)._ [Our space governance plan](https://www.ai-2040.com/?choices=plan-a-root#playbook-epilogue) is, in my opinion, a reasonable first pass about what should be aimed for, and is generally applicable to any scenario where humanity retains control. However, I am not at all confident in the details, and in these scenarios, we will have aligned, superhuman AIs to help us puzzle this out, so my main recommendation is to use the huge amounts of AI labor to come up with a better plan. ([more](/supplements/plan-a-assumptions#our-post-asi-governance-plan)) + + * **Broad deployment of AI** _(importance: medium)._ Broad deployment spreads power and could improve societal epistemics. Some central ways it could be bad are (i) the net effect on societal epistemics might be bad, (ii) broad deployment could make a persuasion/political power takeover easier for the AIs (or for a human attempting an AI takeover) or (iii) they could lead to the proliferation of offense dominant capabilities. On net, I think broad deployment is probably good, but I feel uncertain, and it depends on how good mitigations are in place for each of the above failure modes. ([more](/supplements/plan-a-assumptions#broad-deployment-of-ai-during-the-takeoff-period)) + + * **Various other details of Plan A implemented in our particular scenario.** This includes things like datacenters on the oceans, incentive shaping with respect to alignment, incentive shaping with respect to covert projects (e.g. hardened compute stockpiles), how to trade off deal dissolution risk, covert project takeover risk, and risk of incompetent governance. + + + + +The above were largely details about Plan A that might change. There are also structurally different plans that may be competitive with Plan A, which we discuss [here](https://www.ai-2040.com/?choices=plan-a-root#other-plans-that-are-competitive-with-plan-a). + +Still, I think we should be aiming for something like Plan A across a wide range of scenarios. The rest of this document has two sections, which both discuss how our plan would change in response to empirical changes. + + 1. **Empirical Updates** answers the question “if an empirical fact changed (e.g., suppose timelines were shorter or longer), how would our policy change”. + + 2. **Policy Updates** answers the question: “for each of our major policy recommendations, what would make us change that recommendation”. + + + + +# Empirical Updates + +In this section, I’ll go over how Plan A would change in response to some important empirical assumptions that we make in the scenario, which might turn out to be different in practice. + +## Timelines + +We don’t know when AGI will be built; it’s possible that it could be built this year or next year, it’s also possible that it will take decades. + +In the Plan A scenario, the full automation of AI R&D was on track to happen in 2030 by default, but then was pushed back by government involvement. We’ve written much more on AI timelines; you can find our most updated work [here](https://www.aifuturesmodel.com/). AIFP employees continue to have (slightly) different views about AI timelines, but we all think that 2030 is a plausible and reasonably central year for AGI to arrive. + +**Why did we choose to write this scenario assuming 2030 timelines?** + +It’s good to showcase a variety of different plausible scenarios. We already wrote a very short timeline scenario (AI 2027), so we figured it would be good to choose a different timeline for our next one. 2030 is my (Thomas’) median currently, i.e. I think there is a 50% chance things will proceed slower, and 50% chance things will proceed faster. In the future, if time remains, we’d like to do scenarios for other authors like Eli. + +**Overall, Plan A remains my central plan ~regardless of timeline** , because the underlying problems (AI takeover, AI enabled concentration of power, and rapid AGI induced societal upheaval) are the same regardless of timelines, and Plan A is our best guess for how to deal with these problems. + +Regardless of the timeline, Plan A is most effective the earlier an international deal that involves mutual compute declarations and transparency begins. When we get to the beginning of the intelligence explosion (e.g. the automated coder milestone), verification becomes substantially more difficult, because defecting projects can involve many fewer humans, and therefore be much harder to detect. + +If the timeline is much longer (e.g. 2035+), we begin to have much more uncertainty about the situation. The world will also be more transformed pre-AGI relative to our scenario, there will be larger economic effects of AI, AI company revenues (and share of the overall market) will (probably) be higher, there will be more robots around, and there will be more wakeup to the importance of AI. Particular effects include: + + * We become somewhat less worried about covert projects because compute stockpiles are inherently insanely large, and so the relative amount of compute that covert projects have compared to the verified projects will be smaller. (The absolute amount of compute, on the other hand, is larger). However, there are also reasons to be more concerned about covert projects—a counteracting effect is that there will be more compute distributed throughout the world which will be harder for intelligence agencies to locate. + + * It becomes more important to do the economic parts of the Plan A policy package (e.g. the Citizen’s Dividend) earlier, because more people are affected earlier. + + * It becomes more important to do D/acc earlier because there is a longer period of time with reasonably intelligent AIs around, and current AIs already provide substantial biouplift. + + * Since the overall amount of compute in the world is higher, we become more worried about a new paradigm being discovered that’s much more compute efficient, which could lead to a fast intelligence explosion. + + + + +If timelines are much shorter (e.g. 2026 or 2027), there isn’t time to do the preparation phase, and so we have to do a version of Plan A with no preparation. + + * It’s more likely that we should pivot to one of the contingency plans that involves less government capacity. Unfortunately, these contingency plans have other downsides. For example, shutting down large fractions of AI relevant compute would also turn off inference of existing models, which would be economically costly. + + * It’s more urgent that an agreement to slow down AI takeoff happens early, because we’ll soon be so deep into the intelligence explosion that it’s impossible to slow down. + + + + +## Takeoff speeds + +In this scenario, similar to AI 2027, we chose to assume that the default length of time between AC (Automated coder) and TED-AI (Top-Expert-Dominating AI) was 1 year. This is roughly my median, but we again have large uncertainty; my 80% CI at time of writing for the takeoff duration is something like [2 months; 5 years]. + +**Why did we choose 1 year takeoff speeds by default?** This is my median at the time of writing Plan A, and seemed like a representative scenario to plan for. + +**How does the plan change under different takeoff speeds?** + +**If takeoff is faster than 1 year** , things become generally more difficult, and not just for Plan A. The main specific changes are: + + * Doing an international deal to slow down takeoff becomes more important, because the status quo is more risky, because AI driven change without deal happens much faster and there is less time to react to each change and prepare for the next changes. + + * Enforcement of an international deal becomes more difficult. Faster takeoff typically involves much less hardware increase, because if the takeoff occurs over months, there isn’t time for a massive compute scale up. The possibility of an intelligence explosion without a massive compute scaleup is the key assumption behind being worried about covert projects; if such an intelligence explosion wasn’t possible, then covert projects with relatively small amounts of compute would be much less concerning. Since the intelligence explosion happens faster, it also makes covert projects faster, reducing the amount of slack that the international Consortium has to work with. + + * The slowdown agreement needs to be started earlier, which would make the enforcement easier (and therefore the whole plan more viable). + + * Since enforcement of the deal is more difficult, this is an update towards making tradeoffs that help with enforcement but with various other costs, such as (i) leaning more towards security on the security/transparency tradeoff, (ii) making less algorithmic progress, and therefore scaling more slowly, even if scaling further would still be safe, and (iii) doing more early hardware scaling, which makes verification and self-destruct plans more load bearing, but helps reduce misalignment risk. + + + + +**If takeoff is slower than 1 year,** e.g. even without regulation it takes several years to go from the full automation of coding to the first artificial superintelligence, then we basically update in the opposite direction on all of the above tradeoffs. + + * Overall risk becomes lower, because there is more time by default to deal with risks as they arise. Therefore, the overall plan will attempt to incur a smaller amount of risk, and will be willing to make more aggressive tradeoffs to reduce risk. + + * Covert project takeover in particular becomes less likely, because it is extremely unlikely that they could bootstrap into existentially dangerous AI systems with only a tiny amount of compute. This updates us even further in favor of doing Total Research Transparency (because we are relatively more worried about the large projects than the small projects), and away from doing the large hardware scaleup. The plan can involve just doing a slowed down version of the intelligence explosion on the existing compute stockpile, leaking algorithms to the trailing project, but relying on the fact that they are so compute limited that they cannot make significant further AI progress. + + * AI takeover is less likely, and so we can devote relatively more effort on mitigating other risks. It’s still very important to slow down AI progress, because there’s still a significant chance that AI alignment won’t be solved during the default takeoff, and because slowing down helps to mitigate other risks like concentration of power. However, it’s more likely that we wouldn’t actually slow down much in this world. Therefore, doing the slowdown part of the treaty is less load bearing. + + + + +Overall, we recommend pursuing Plan A regardless of takeoff speeds. However, changes to our takeoff speeds do substantially change the specific implementation of Plan A, in the ways noted above. + +## Alignment difficulty + +In the Plan A scenario, humanity rallies and creates a high assurance safety case for handing off to ASI. However, **we are much more confident in the general picture that more time and effort lead to a greater chance of solving the alignment problem than the specific story told here.** We don’t know how hard it’ll be to align ASI. Opinions vary substantially within our team (and even more outside of our team). Given the uncertainty, our view is that we should take a portfolio approach, and invest substantially in both prosaic AI safety agendas (e.g. variants of RLHF) and moonshots (e.g. uploading human brains). + +The plan we are advocating for is one in which the world coordinates to give the scientific community breathing room and resources to figure out how difficult the alignment problem is. If Yudkowsky is right that AI alignment is difficult, our view is that evidence would increase during this time period and then the pause would be extended. Likewise, if the optimists are right and alignment is easy, we would get evidence of this during that period, and we could proceed at a faster pace (though it’s not clear that we should, given the other risks of doing so, most notably concentration of power). + +I think that alignment is probably difficult enough that without a well managed multi-year slowdown, we will incur substantial AI takeover risk. Even with a ten year slowdown, as depicted in the scenario, there is a significant chance that alignment was too hard, or that it’s feasible but human error causes us to fail anyway. You can see more about the team’s views about alignment success likelihood conditional on following different plans [here](/supplements/comparing-possible-plans#outcomes). + +**What do we do if alignment is much harder?** + +Gather as much evidence about alignment difficulty as possible by: + + * Deploying and observing AI systems to get empirical feedback about how aligned they seem to be. + + * Building model organisms—AIs trained to deliberately be misaligned—and see if alignment techniques are good enough to catch and train away the misalignment. Do this for deceptive alignment, sycophancy, reward hacking, etc. + + * Giving a large, heterogenous set of people access to the models (ideally including model internals) to avoid groupthink. + + * Have explicit safety cases that lay out the evidence and arguments for why the models can and should be trusted, and publish them so they can be critiqued and debated by the scientific community. + + + + +During Plan A, maintain the optionality of extending the pause by: + + * Monitoring for covert projects and shutting them down if detected + + * Ensuring that the flow of new chips is highly secure and not going to any rogue projects + + * Improve coordination and verification technology, both with large investment and using AIs to improve this technology. In particular, using AIs to build privacy preserving auditing and reliable lie detection seems very promising. + + * Minimizing software progress that leaks to covert projects. If alignment is harder, we should update towards making tradeoffs that enable us to stay slowed for longer. + + * Making sure decisionmakers have access to the best information on alignment difficulty, including by making sure that people who are more pessimistic about alignment have access to the models and the current safety cases, so that they can find problems in them before it becomes catastrophic. + + + + +If alignment is extremely hard, e.g., humanity isn’t making substantial progress on alignment whatsoever, then the correct call may be to pivot to Plan S, and instead of attempting a managed takeoff, shut it all down. Then, we could pursue alternate strategies to build superhuman AI that are inherently safer than ML, such as uploads. Unfortunately, these approaches have substantial downsides as well, notably the risk that the moratorium breaks down and we go back to racing before enough progress is made on uploads or other strategies in that vein. + +**What if alignment is much easier?** + +Slowing down AI takeoff is still probably good for many other reasons, including avoiding AI-driven concentration of power, figuring out appropriate post-ASI governance, transitioning to a regime where human labor is no longer economically relevant in a healthy manner, hardening the world to AI enabled superweapons, etc—basically, we still need to solve all of the problems that we dealt with in Plan A other than misaligned AI takeover. + +The main difference is that there is a higher chance of humanity being able to successfully muddle through these other risks without permanent lock in. This means that the necessity of a slowdown does go down substantially if alignment is much easier than I expect. + +## Verification difficulty + + _Related reading:[Covert Project Branch](https://www.ai-2040.com/?choices=plan-a-root#branchpoint-china-attempts-a-covert-agi-project), [Covert AI Projects Supplement](https://www.ai-2040.com/supplements/covert-ai-projects), [Verification Plan](https://www.ai-2040.com/supplements/verification-plan)._ + +I’m not sure how difficult verification will be. Our main strategy for verifying compliance with an AI agreement is (1) limit the amount of compute that covert projects have access to, and (2) ensure that legal projects are doing what they claim. Both of these will be somewhat lossy, and we’re unsure about what the mapping is between “amount of unmonitored compute” and “how fast could an illicit intelligence explosion reach an existentially dangerous capability threshold”. + +Now we’ll go through each of these in a bit more detail: + + * **Covert project verification difficulty.** In the Plan A scenario, we assume that the maximum plausible covert project size without detection is about 1% of pre-deal compute. Our assumption here in the scenario is close to the worst case scenario. My best guess is that neither side would build a covert project in a Plan A scenario because (i) there’s no incentive to because it’ll be out paced by the legal projects, and (ii) both sides are scared of AI takeover risk; and building a covert project would undermine the deal. Sub-questions I’m most uncertain about: + + * **What fraction of pre-deal compute can be tracked down?** Our uncertainty over pre-deal compute is very relevant, because it will determine the extent to which China’s initial compute declaration can be falsified. Existing experts we’ve talked to maintain huge amounts of uncertainty about how large China’s existing compute buildout is. As a percentage, this might become easier to track as AI datacenters get larger, but as an absolute number more and more compute will be harder to track down over time. However, the intelligence community may have access to private information that’s much more accurate today, and they might begin to gather this information in the near future. Overall, it’s possible that denying covert projects access to huge numbers of GPUs will be one of the most load bearing components in the non-proliferation plan; but it’s also possible that other factors such as HUMINT will be far more important. + + * **How effective will HUMINT be for detecting covert projects?** It seems likely to me that a covert project of the size we assume in the [covert project branch](https://www.ai-2040.com/?choices=plan-a-root#branchpoint-china-attempts-a-covert-agi-project) of Plan A (500k H100e, $7.9B) would be quickly detected by intelligence gathering. Historical mega projects that were attempted in secret (e.g. the Manhattan Project) seem to have been riddled with spies; our baseline assumption should be that anything on the order of this size will be quickly detected and shut down. However, later into the intelligence explosion, the number of people involved in a covert project could be much smaller, as AIs could automate the AI research and robots could automate the construction and maintenance of the datacenter, dramatically reducing the probability of leaks. + + * **What is the necessary reduction in access to compute resources?** In the Plan A scenario, we assume that an AI project with 10x less compute can do an intelligence explosion 5.5x slower; this is discussed more in [the appendix of our Covert Projects Supplement](/supplements/covert-ai-projects). + + * **Legal project verification difficulty.** In the Plan A scenario, we assume that we can verify inference-only very quickly, and that we can verify training runs after a six month sprint to develop that technology. It could be easier (e.g. it’s possible to do at the software level), or it could be harder (particularly because governments are slow and ineffective at this). Our assumption here is on the optimistic end because we assume high competence in execution. + + * If legal project verification is more difficult, then increasing hardware buildout is more scary. Quantitatively, with no hardware buildout, then we’d need something like 99% assurance on legal project compute; with a +4 OOM increase in stock, we’d need 99.9999% assurance on legal project compute. + + + + +There are various other particular verification dynamics which could affect the rest of the policy strategy: + + * **Favorability of SW scaling?** A key dynamic for covert projects is stealing algorithms from legal projects. Therefore, legal projects should consider pursuing algorithms that are less favorable for covert projects. In particular: (1) HW/SW co-design, i.e. research algorithms that are only effective on specifically designed chips that covert projects don’t have access to, and (2) heavily scale-dependent algorithms. We’re unsure how favorable both of these paths will be; the more favorable they are, the more capabilities progress can be done without speeding up covert projects. + + * **How much edge-hardware is required?** Some applications strongly prefer edge hardware (hardware that operates on the device as opposed to the cloud): notably military applications (because of jamming) and self-driving cars (because you don’t want your car to crash if there’s a network issue). This poses a challenge for verification because edge compute could be removed from the device and funneled to a covert project, or potentially backdoored and used as part of a large distributed training run. We’re uncertain to what extent large quantities of this will be required, if there are other solutions, and how good the mitigations look for preventing usage of edge compute in training runs. + + * **How difficult will security be?** Nation-state proof security is basically required to make the legal project verification work. Nation state proof algorithmic security is helpful for preventing covert projects. Our baseline assumption on security is that it’s possible to make verification mechanisms sufficiently secure, most algorithms will not be possible to secure (because they are so small), and that model weights will be possible to secure (because they are so large). I’ll now go through the updates if our current assumption is wrong for each of the desired security properties: + + * **If verification security isn’t viable** then our ideal strategy would either involve (1) doing a moratorium on further AI development (and potentially inference) until security is viable or (2) relying heavily on trust. + + * **If algorithm security is viable** then we’d have more buffer against covert projects than we are currently modeling in Plan A. We’d also have less of a need to do a compute buildout, because one of the core motivations for building compute was to mitigate algorithms leaking to covert projects. + + * **If model weights security isn’t viable** then we’d either need to (1) not scale AI capabilities until security becomes viable or (2) need to invest more heavily in defending against exfiltrated AIs. At very high levels of capability, this might involve extremely drastic measures, because we’ll need to defend against extremely capable AIs. + + + + +## AI Paradigm Shifts + +Plan A currently assumes gradient descent on giant neural networks remains the dominant AI paradigm until the late 2030s. We assume significant advances are made within that paradigm, including (i) the addition of neuralese (both as long term memory, and replacing chain-of-thought) and (ii) dramatic scaling in reinforcement learning (many more diverse environments that are much larger than existing environments). + +Substantial algorithmic changes, even within the broad “gradient descent on giant neural networks” paradigm, might change the implementation details of parts of Plan A. For example: + + 1. We assume that online learning via live weight updates (instead of in context learning or delayed weight updates) is not necessary for generally top expert level AI systems of the type developed in Plan A. If live weight updates were necessary, then we would need to have some training infrastructure happening on the inference datacenters, which would complicate the inference verification because we’d need to live update the reference models in the verification infrastructure, which would increase the security risk surface area. It would also prevent us from using chips that are hardcoded to do inference on a given model (e.g. [Etched](https://www.etched.com/careers)), which could be part of the inference-only verification scheme. + + 2. We assume that restricting access to compute is an effective measure to prevent covert projects. If there are, for example, new architectural discoveries with dramatically (e.g. 1000x) more favorable compute scaling than current architectures, this would make covert projects much more dangerous. That said, if such architectural discoveries are in the pipeline so to speak, all the other plans are much less likely to work too—a discovery which allows a covert project to overtake the legal projects in Plan A, despite a huge compute disadvantage, would in Plan D cause an extremely fast, discontinuous “FOOM” to ASI within whichever frontier AI project first found it. These worlds are extremely scary and probably the best way to handle them is something like Plan S. + + + + +There are also obviously many “unknown unknowns” in AI paradigm shifts which could change Plan A in unforeseen ways, in practice we expect that many updates to the Plan A baseline will need to be made in response to further paradigm shifts. + +## Less central updates + + * **State capacity.** Plan A relies on high state capacity on AI, in particular, the capacity to distinguish existentially unsafe training runs and deployments. However, in worlds where the state capacity is low enough that they are not able to distinguish safe and unsafe AI, we recommend pivoting to another plan. Two salient alternatives are: (1) having an extremely conservative bar for safety, (e.g. shut it all down), or (2) limit the speed of the intelligence explosion through more blunt techniques (e.g. mutual GPU destruction). + + * **Capability level of the initial deal.** The later the deal begins, the more difficult enforcement becomes. In particular, once we get substantial automation of AI research, it becomes easier to run a covert project on AI with only a small number of people involved, which reduces the surface area for leaks. This means that the overall regime has less time, and so risk will be higher. + + * **Distillation mitigation.** We discuss [our strategy for mitigating distillation in the verification supplement](https://docs.google.com/document/d/1K4vxl_IMf58OgNXbWntQl6BQ06yAWZ62WmW5Ypp4KYA/edit?tab=t.86wiss1dhlds#bookmark=id.9ax99eo6ehyb). It’s unclear how well this will work. If it doesn’t work well, it’s an update towards being more concerned with covert projects. It is also an update towards slower capability scaling being optimal. Finally, it’s also an update towards strongly limiting broad deployment, although I still think that even if mitigating distillation is very difficult, it would still be worthwhile to do broad deployment. + + + + +# Policy Updates + +## How fast to scale AI progress after a slowdown deal? + +It’s unclear, and will be dependent on observations at the time. Here are some key considerations: + + 1. **Is the deal likely to break down soon?** If the deal is more likely to break down soon (for example, due to political reasons, or due to covert project takeover), then you should be willing to take on more risk, and so you should be more willing to scale faster. + + 2. **How good do current safety cases look (and how misaligned do the models seem)?** The more AI takeover risk there is from marginal scaling, the slower we should scale capabilities. + + 3. **How large are the returns to better AIs?** If better AIs won’t contribute useful labor (for example because we can’t trust their outputs, and control isn’t very scaleable), then we shouldn’t scale. On the other hand, if scaling AI capabilities would dramatically increase the amount of high-quality alignment research we can do, it may be worth scaling, even if doing so is risky. + + 4. **Will the weights and algorithms leak to covert projects, and how bad are the leaks?** If the model weights and/or algorithms of a scaled up AI model were to leak this would accelerate defecting projects by some amount, decreasing the amount of slack the regime has. If it’s possible to secure the model weights and algorithms, or to scale capabilities without doing substantial algorithmic progress, then we should update towards scaling more. In the Plan A scenario, we mostly scale with hardware buildouts, to avoid the need to make substantial algorithmic progress, because we are not optimistic about securing the algorithmic secrets. + + 5. **How much risk comes from current AIs vs future AIs?** The earlier in the intelligence explosion we are, the more we should be sympathetic to scaling earlier, because the marginal risk of current scaling is much lower, whereas if we’re in the capabilities regime where we are incurring most of the risk, we should be much more careful and tend to slow down more. (This is why, in the scenario, the Consortium scales up AI capabilities reasonably but not entirely cautiously in the early 2030s, and then pauses at top expert level in 2035.) + + 6. **Is it possible to scale up capabilities without substantial algorithmic improvements?** In plan A, we try to scale up AI capabilities via hardware scaling, while heavily limiting algorithmic progress. However, it may be the case that scaling up AI capabilities inherently leads to lots of algorithmic progress (e.g. via distillation). Moreover, the main alternative to scaling up algorithms is scaling hardware. However, it may be the case that most algorithmic progress is downstream of hardware progress. In this case, the most effective way of governing algorithmic progress might be to limit hardware itself. + + + + +Ultimately, trading off these considerations will require quantitatively modeling each of these sources of risk: covert project risk, risk of AI takeover from current AIs and future AIs, etc, understanding how the speed of deployment impacts each of these risks, and then making complicated tradeoffs about the overall best path forward. + +Plan A’s trajectory is essentially based on the viewpoint that deal dissolution risk will probably be high, that risks from early AGIs (e.g. ACs and SARs) can be mitigated via control (i.e. that control-based safety cases can be made good enough very quickly, for AIs at that level), and that we can elicit a huge amount of useful labor from these early AGIs. It also assumes that we can do most of our capabilities scaling via a rapid hardware buildout, and strongly limit the total amount of algorithmic progress that occurs, including successfully mitigating distillation (largely by hiding the intermediate reasoning traces). Under these assumptions, the best trajectory seems to involve scaling quite rapidly at the beginning of the slowdown agreement, and then slowing down dramatically at the maximum controllable capability level. + +Plan A’s specific trajectory is our best guess for the best way to handle these tradeoffs. However, we aren’t confident whatsoever that this is the correct path, and we hope that decisionmakers at the time will have a better understanding of these empirical factors at the time, whereas we have to make rough guesses. + +Another very plausible trajectory, Plan S, is essentially the opposite choice about how fast to scale: immediately issue a moratorium on further AI capability scaling. This policy looks better if there is a small risk of deal dissolution, and/or the returns to better AIs for safety research are low, and/or the risks of any capability scaling are high. + +## Massive datacenter buildout + +It is desirable to do a massive datacenter buildout in Plan A if (and only if!) the upside of scaling quickly, reducing covert project takeover risk, and having additional capacity to pay alignment taxes is bigger than the risk of verification failing or the additional risk incurred by the deal failing with a gigantic compute stockpile. + +In a bit more detail, the main upsides of a massive datacenter buildout are that it allows AI developers to: + + * **Scale quickly without leaking algorithms to covert projects.** In the previous section, we discussed the conditions under which it is good to increase AI capabilities, and the tradeoffs associated with going faster and slower. Simplistically, to increase AI capabilities, you can either improve algorithms or scale compute. By leaning heavily on compute, we can avoid making algorithmic progress in the first place, which is the best way to ensure that said progress doesn’t leak. + + * **Pay alignment taxes.** Large compute-based alignment taxes might be very helpful for safety. For example, it might be that neuralese is much more performant than current architectures, but also has dramatically worse safety properties. If we have enough compute that we can just accept the performance penalty and continue with current architectures anyways, this might dramatically reduce risk. Without a massive compute buildup, we will by definition not be able to pursue architectures that are much less compute efficient. + + + + +The main downsides are: + + * **Verification and regulation becomes more load bearing.** Since the verified projects have even more compute, it is more important that this compute is used safely, and, in particular, not used to do an intelligence explosion as fast as possible. So it becomes even more important that the regulations that AI projects are subject to are sufficiently conservative that they preclude an intelligence explosion, and that the verification is sufficiently robust that companies or rogue AIs can’t undermine it, and illegally gain control of a huge amount of compute within the verified clusters. + + * **If the deal breaks down with a gigantic compute cluster, and the clusters are not destroyed, the intelligence explosion might be extremely rapid.** One naive estimate is that a 10x of compute leads to a 5.5x increase in the speed of the intelligence explosion, so a 10,000x increase in compute might increase the speed of the intelligence explosion by 900x. So if it would take a year by default, it might take less than a day with the Plan A hardware explosion. This is obviously extremely dangerous, hence why we put so much emphasis on mutually assured compute destruction. + + + + +Putting all of this together, my view on how to decide to what extent we should pursue a massive hardware buildout under Plan A is as follows: + + 1. If covert projects aren’t a concern (because, for example, it’s impossible to do a reasonably fast intelligence explosion on a 100,000 GPU cluster), then **don’t scale hardware at all, or only go slowly.** + + 2. If there are reliable mechanisms for destroying compute after the deal dissolution, and reliable verification mechanisms so that we can have high confidence that there will not be major rogue internal deployments (either by countries or AIs), and there is reliable verification for within-deal compute, then the downsides of a compute buildout are small. Therefore, it is beneficial to **scale hardware ~as aggressively as possible.** + + 3. Otherwise, it is going to be a complicated tradeoff, which needs to be decided by weighing the magnitudes of the various costs and benefits listed above. + + + + +## Our specific verification plan + +The verification plan used in the Plan A scenario can be summarized as follows: + + 1. Mutually declare the locations of all large datacenters. Check the declared counts against intelligence estimates. + + 2. Agree to pause training and experiments. Install inference-only verification devices on the declared clusters to verify this pause. + + 3. Race to set up verifiable (e.g. radically transparent) training and experiment clusters; when complete, resume capability scaling within the constraints of safety and transparency. + + + + +We discuss this in more depth [here](https://docs.google.com/document/d/1K4vxl_IMf58OgNXbWntQl6BQ06yAWZ62WmW5Ypp4KYA/edit?tab=t.86wiss1dhlds#heading=h.g5eefrev1iir). + +This version of the verification plan assumes that at the beginning of the deal (1) we can verify inference, and (2) AI development verification is feasible within a relatively short research sprint (but is not immediately feasible). It is most likely that one or both of these assumptions will be wrong; these assumptions are my central guess of what could be done under Plan A assumptions of political will and competence, but reality will most likely involve less good preparations. + +What if we don’t have verification measures ready in time? + +Three high level buckets of verification approaches are: + + 1. **Technical verification methods**. These methods would allow ongoing AI usage or development, with both sides of a deal being confident that the other isn’t doing anything dangerous. An example of this is what we do in Plan A where we initially rely on inference-only, while sprinting to technical methods to verify training and experiments. + + 2. **Rely on trust,** i.e. accept that either side could defect, and hope that neither side is foolish enough to try, or that they’d be caught by HUMINT and so forth if they did. + + 3. **Blunt verification techniques** (e.g. compute destruction or mutual power down). + + + + +yesnoyesnoViable technicalverification methods?Rely on trust?1 · Technical verificationinference-only devices, then verifiabletraining & experiment clusters (Plan A)2 · Rely on trustaccept that either side could defect;HUMINT etc. as the backstop3 · Blunt verificationmutual GPU destruction, mutualpower-down, burn compute on mining + +Three buckets of verification approaches \- ai-2040.com + +It is highly desirable that we have technical verification methods ready before the beginning of the intelligence explosion. If inference verification devices are ready, then access to AIs gets to continue at the beginning of the deal. If development verification is ready, then we don’t need to stop AI capabilities scaling or AI safety research that involves frontier models at the beginning of the deal, and can instead scale until it is sufficiently unsafe not to (following the considerations above). + +2 and 3 will likely only be temporary regimes, and be paired with a sprint to get technical measures online. Therefore, we hope that eventually we’ll get to sufficiently good technical methods. + +The main downside of 2 is that it runs a risk of defection. Therefore, it may also run into increased political difficulties, as both parties will be worried that the other will defect, and so will be less likely to sign in the first place. However, it’s not correct to entirely write off trust, because it may be that the risk of AI takeover becomes so high that neither side is incentivized to defect, because doing so would so likely result in existential catastrophe. If high AI takeover risk becomes common knowledge (for example, after we’ve caught early misaligned AIs in the act of sabotage), relying on trust may be the best option. + +The main downside of 3 is that it involves a temporary period where large scale inference goes offline. However, this would be less bad than one might naively think; even an agreement to power off (or destroy) 90% of the compute, would only force a reversion to 10x cheaper models. Since cost of a given level of capability has historically decreased very rapidly ([averaging ~40x/yr](https://epoch.ai/data-insights/llm-inference-price-trends)), having 10x less compute would roughly correspond with having to use SOTA models from 7.5 months ago. A downside of 3 is obvious; if defection is viable, it becomes more likely that both sides would choose to defect in the first place. + +Note that this is one upside of Plan S over Plan A; Plan A requires (1), while Plan S only requires (3). + +## Transparency and Security tradeoffs + +There is an inherent tradeoff between security and transparency—to secure something, you inherently need to prevent some group from being able to see it. + +Security% of progress competitors can’t stealTransparency% of relevant info that’s publicLivestream thesingularityOpen source everythingexcept the weights and datasetsAirgap the researchers & auditors;relay all safety-relevant infovia paper memosAirgap the singularity;hand off to airgapped auditorsTodayPlan A: increase both + +Security vs. transparency \- ai-2040.com + +Plan A biases towards transparency, in large part because I am more worried about the risk of poor governance within the legal projects than the risk of covert projects defecting and successfully undermining the deal. Increasing transparency will on average improve the quality of the governance decisions made by the Consortium, whereas increasing security makes it harder for covert projects to succeed. Furthermore, even if the governance decisions are perfect, the legal projects may not be able to solve alignment in time themselves, and may need help from the wider world to e.g. notice some subtle flawed assumption in a safety case. Also, transparency is very helpful for verification. + +That being said, we are currently nowhere near the pareto frontier between these two dimensions. Plan A proposes dramatically increasing both security and transparency into frontier AI projects. We discuss transparency-related tradeoffs in much more detail in our [Transparency Supplement](/supplements/transparency-plan#thomas-larsen). + +## Our Post-ASI governance plan + + _Related reading:_ [Space governance supplement](https://www.ai-2040.com/supplements/space-governance-plan) _,[AI 2040 epilogue](https://www.ai-2040.com/?choices=plan-a-root#playbook-epilogue)._ + +We make a concrete guess as to what a positive post-ASI governance plan might look like. However, we expect that there exist better proposals, which we haven’t thought of yet, that will be discoverable at the time. Therefore, we want readers to orient to our plan as a baseline with which to compare other plans to on dimensions that they care about. Then, we hope that humanity will spend vastly more effort—both human and AI cognitive labor—on designing a system that performs better. + +## Broad Deployment of AI during the takeoff period + +The Plan A scenario involves broadly deploying AI into the economy during takeoff. This has massive effects on the world during takeoff: almost everyone loses their jobs and starts relying on Citizen's Dividend. Overall, I think that it is better for massive AI impacts to happen gradually, e.g., over the course of a 10 year takeoff, than to have all of the impacts happen really quickly, right at the end, once AIs are vastly smarter than humans. + +CapitalHuman cognitiveHuman physicalAIRobots0%25%50%75%100%2025202620272028202920302031203220332034203520362037203820392040 + +US factor income shares (Plan A) \- ai-2040.com + +(See more about how AI impacts the broader economy in [our economics model](https://www.ai-2040.com/supplements/econ-explorer), and discussion in our [economics supplement](https://www.ai-2040.com/supplements/economics-of-plan-a)) + +However, this relies on a few key questions: + + 1. Is the net effect on societal epistemics positive or negative? I currently feel highly unsure. On one hand, broad deployment of human level AI will increase access to information, like the internet. On the other hand, while the internet is definitely helpful for some people, it’s unclear if the average effect on societal epistemics is good or bad. More research is needed on this; what plan we actually recommend at the time would depend a lot on our best guess about the epistemic effects of broad AI deployment. With our current state of knowledge, I currently don’t feel like this is a major consideration in either direction. + + 2. Does broad deployment meaningfully increase the difficulty of AI control? I think probably not, because the main risks of AI takeover route through rogue internal deployments and recursive self improvement. The case for broad deployment increasing risks is that it may lead to AIs having affordances that might be helpful for takeover, especially (i) a large robot workforce and manufacturing base, (ii) control over robot armies, (iii) access to biology labs, and (iv) regular contact with a broad range of humans. I think that these risks are likely mitigatable through a combination of monitoring AI behaviour, limiting access to the most dangerous physical infrastructure (e.g. biolabs), preventing any one AI system from having control of a large fraction of the physical resources, and preventing collusion between different AI systems. However, I do think that it’s plausible that broad deployment makes control much harder. Hopefully this can be handled via limiting AI deployment in a particular domain of concern, but it may in practice limit deployment across many domains. + + 3. How high are misuse risks from broadly deployed closed-source models? My inclination is that these risks are reasonably low, because (1) I’m optimistic about adversarial robustness in a regime like Plan A, where you are slowing down substantially in order to improve safety, and you have the capacity to run lots of monitoring, and (2) even if it’s possible to jailbreak widely deployed AI models, I think that the defensive measures put into place will be sufficient for preventing existential catastrophes from misuse risk. However, it is plausible that both of these assumptions fail (either because of intrinsically high difficulty or incompetent execution), in which case, we would be incurring serious misuse risk from broadly deployed models. + + + + +If we’re not broadly deploying AI through society, the main alternative is to be developing AI inside closed projects. This incurs additional risks. Whatever closed off group of people that are managing the intelligence explosion without contact with the rest of the world will have a huge amount of power over the future. For an extreme case, the closed project might literally be developing AI in secret, without notifying the rest of the world, as described in [Training AGI in Secret would be Unsafe and Unethical](https://www.lesswrong.com/posts/FGqfdJmB8MSH5LKGc/training-agi-in-secret-would-be-unsafe-and-unethical-1). A secret intelligence explosion would be bad for misalignment risk, because far fewer people would know and understand what’s going on, and be able to give input on the alignment techniques, and would be bad for the distribution of power, because powerseeking individuals might be able to seize control of the AGI project, and build ASI aligned with their values (as opposed to humanity more broadly) + +A better proposal for avoiding broad deployment is to do local deployment, but maximize transparency: publish evaluation results, publish demos of AIs doing very impressive things, publish the AI-generated research (on non-AI domains), publish the model spec, etc. This is far better than a secret intelligence explosion. However, it seems likely that if this is happening, humanity will be still largely in the dark about what’s going on, because to normal people, without AI impacting their day to day life, it’ll seem like some far-away science project that people like to talk about—it’ll be in far mode, not near mode. Therefore, it seems likely that broad deployment of AI will improve epistemics around AI in particular, which seems positive. + +## Appendix: How plausible and how desirable are our key policies? + +Plan A is neither a pure recommendation nor a pure prediction. A pure recommendation in which every actor makes maximally responsible moves starting tomorrow would have been too implausible. And we already wrote a pure prediction scenario—that's AI 2027. So we chose to make Plan A a compromise between plausibility and desirability. + +The interactive graph below shows roughly where each of the scenario's policies sits on these two dimensions. Plausibility is judged at the point in the scenario when the policy is adopted, conditional on everything that has happened up to then. It does not represent how likely the policy is to happen starting from today's world. For instance, "Pause at top human level" is rated highly plausible not because we think a pause is highly likely from where we stand in 2026, but because in a world where the deal has held for years and control-based safety cases are visibly reaching their limits, we expect decisionmakers to choose something like it. Desirability measures how much better we currently expect a policy to be than the most plausible alternative at that same point in the scenario. It is not a ranking of which policies we're most keen to see adopted today. A low score can mean that our recommendation is subject to lots of unpredictable empirical updates or that alternative policies look about as good to us as our recommendation. + +Click on a policy to read more. + +less plausiblemore plausibleless desirablemore desirable + +Click a point to read the policy’s explanation. + +Policy plausibility vs. desirability \- ai-2040.com + +1. + +This is in scare quotes because, as our scenario illustrates, even if we do a well-enforced international deal to limit the pace of AI progress, the overall pace of AI progress will be extremely fast compared to what most people are expecting and used to. For more on why we think this, you can view our models of AI progress (the [AI Futures Model](https://www.aifuturesmodel.com)) and of AI economic impacts ([Economics of Plan A supplement](https://ai-2040.com/supplements/economics-of-plan-a)), as well as numerous blog posts and of course [AI 2027](https://ai-2027.com). + +2. + +AGI = Artificial General Intelligence, or, AIs that have a similar amount of general intelligence to humans. Our actual thinking uses more precise concepts which you can read about in the [AI Futures Model](https://ai-rates-calculator.vercel.app/). Concepts like AC, SAR, TED-AI. However, we find "AGI" to be a handy shorthand for some vague combo of the above milestones, because we expect those milestones to arrive within a few years of each other or less. So when we say 'AGI', understand that we are gesturing at milestones such as the above, and being deliberately nonspecific because the context doesn't require precision + +3. + +Eli’s timelines are similar to mine, but his takeoff would be about two times slower. + +4. + +Note: some people use takeoff speeds to refer to the period after TEDAI, it’s an ambiguous term, but for the purposes of this document, I’ll use the definition given here: the length of time between AC and TED AI. + +5. + +To illustrate, suppose that the length of time between AC and TEDAI is one month. Even if we can (optimistically) use compute governance to slow down rogue projects with ACs by 20x, you still only have 2 years until the rogue projects have TEDAI, which is a smaller window than Plan A assumes. + +6. + +This reasoning is given in the [datacenter buildout](/supplements/plan-a-assumptions#massive-datacenter-buildout) section below. + +7. + +Note that covert projects doing an intelligence explosion is only one of the reasons to be worried about them, there’s also the risk of them directly obtaining a decisive strategic advantage; so it would still be risky if they were, e.g., able to train a top expert level AI. + +8. + +As well as various aspects of takeoff besides speed, notably the possibility of fast progress on a small, fixed, amount of hardware. + +9. + +There’s a good chance that empirical feedback will be highly ambiguous, but overall I’m optimistic that if there’s a long period with broad access to experiment on roughly human level AIs, then we will surface a huge amount of evidence related to the difficulty of alignment. Given the current (poor) epistemic environment, I don’t think this evidence will be sufficient for broad scientific consensus (until we have a good fundamental understanding of AI), but I personally expect to learn a lot and change my mind in substantive ways. + +10. + +Note that pre-deal compute matters a lot because it seems probable that post-deal compute will be effectively monitored, because (1) there are very few extremely expensive fabs which can build the relevant chips, all of which can be directly monitored with auditors, and (2) it is intractable to build illicit fabs. + +11. + +Specifically, security to China, because China is probably the best placed actor to run a covert project. + +12. + +For example, we might need to do things like move all humans into bunkers that are robust to nanotechnology (friendly nanotech is defending), mirror life (no air is allowed to enter), and superpersuasion (all information coming in is mediated by friendly AIs, who screen out superpersuasive content). + +13. + +Thanks to the transparency of AI development to other companies, many governments, and the public, there are some failsafes in place for governments making poor decisions about safety techniques. + +14. + +Probably covert projects are somewhat of a concern, but it’s very unclear, and depends centrally on when international coordination to slowdown starts—the later a deal happens, the higher the risk of defection because (1) we’re closer to superintelligence, and (2) the smaller the personnel footprint is because the AIs can do a larger fraction of the tasks involved in running a covert project. + +15. + +If algorithmic efficiency improves at ~40x per year, then the number of years of progress _t_ lost by a 10x compute reduction is log_40(10) = .624 years = ~7.5 months. + +16. + +This appendix was written by Miles Kodama. diff --git a/sources/ai2040-pages/security-in-plan-a.md b/sources/ai2040-pages/security-in-plan-a.md new file mode 100644 index 0000000..728c0a5 --- /dev/null +++ b/sources/ai2040-pages/security-in-plan-a.md @@ -0,0 +1,170 @@ +--- +url: https://ai-2040.com/supplements/security-in-plan-a +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Security in Plan A + +### Romeo Dean, Thomas Larsen + +Security is an area where expertise and access to classified information is particularly important for having an informed opinion. Nonetheless, this supplement provides a high-level overview of the security needed in Plan A in our best understanding, and why we think it is tractable. + +We think it useful to think about three types of security, in increasing levels of difficulty from our point of view: + + 1. **Model weights security:** Defending the exfiltration of model weights, and other large datasets. + + * Metric we care about: Exfiltration bandwidth. I.e., how much can the attacker steal in a given amount of time? + + 2. **Verification integrity security:** Can we trust the results of our verification measures? + + * Metric we care about: Assurance level (the confidence-coverage curve explained in the verification supplement). I.e., how confident are we that the maximum unverified compute usage was less than X% (for values of X between 0 and 100%). + + 3. **Algorithmic confidentiality:** Defending the exfiltration of code, algorithmic secrets, and anything else that doesn’t classify as ‘large datasets’. + + * Metric we care about: We are unsure. On the one hand, we could think about algorithmic security in terms of exfiltration bandwidth as well, but algorithmic secrets might be tiny and have an unpredictable mapping from data size to usefulness. Therefore it’s probably best to think about the percentage of algorithmic efficiency that is exfiltrated and usable by the attacker. + + + + +Throughout this supplement, we assume an [OC5 attacker](https://www.rand.org/pubs/research_reports/RRA2849-1.html) per RAND’s definition, because in Plan A the US and China will be worried about each other being potential attackers. + +## Plan A security summary + +In our scenario, we assume that SL5 security for model weights and to detect verification tampering is **viable with extreme effort; this is our low confidence best guess. It may turn out to be impossible** , in which case Plan A would need to be modified, which we discuss [here](/supplements/plan-a-assumptions#less-central-updates). + +The security level desirable in Plan A is as follows: + +| **Model Weights**| **Algorithmic Secrets**| **Verification** +---|---|---|--- +**Inference Datacenters**| SL5-100TB-5y(Maximum exfiltration of 100TB over 5y)| n/a| SL5(increasing assurance over time as compute scales) +**R &D Datacenters**| SL5-100TB-5y| Depends on the transparency regime. In Total Research Transparency, most are intentionally transparent, small fraction SL5.| SL5(increasing assurance over time with compute scale) + +## Model weights security + +Goal: We think the goal for model weights security in Plan A requires that frontier model weights (and similarly large important datasets) cannot be exfiltrated and used by potential adversaries and/or covert projects. + +Concretely, we think the goal should be to defend against around 100TB of exfiltration over 5 years. Which we define as: weights security level 5, for 100TB over 5 years, or SL5-100TB-5y for short. We don’t think this will be trivial, but we do think it is highly tractable. + +The following figure shows the size we expect frontier models to be in Plan A (and we also expect other datasets to be of similar magnitude, since data and parameters should continue to trade off with each other in compute optimal scaling). + +**Frontier model size, 2022 to 2040** + +Model sizeTraining compute + +reportedours (α=0.5)α ∈ [0.2, 0.8]Nesov + +1101001k10k100k1M2022202620292032203520382040Trillion parameters2026 anchor2029 dealGPT-4 ~1.8Tα=0.8α=0.2~650T (Nesov)31,000T~158T + +### Plan A implementation + +The **inference data centers** have SL-5 model weights security via defenses including: + + * They are surrounded by an information boundary, which is the datacenter walls, including a copper box (a faraday cage) preventing signals from coming in or out, except for one cable in and out of the datacenter, and a secure exchange chamber for other things that need to cross the perimeter. + + * For the output cable (which needs to exist, in order to communicate with the outside world), we make sure that the only information coming out of that wire is verified tokens produced by whitelisted models. You can read more in our [verification supplement](https://www.ai-2040.com/supplements/verification-plan). This is also similar to the proposal in [this paper](https://arxiv.org/abs/2511.02620). + + * For the other things that need to cross the boundary, power in, heat out, people in and out, we have further side channel mitigations and total bandwidth capping. + + + + +The **R &D datacenters** have two deliberate channels that allow information to leave the datacenter: (1) physically sending model weights to the inference datacenters, and (2) the fully transparent output channel used by researchers and for queries. + + * **The transportation of the model weights to the inference datacenters.** When a model is approved for public deployment, a package including the weights file, and any approved inference code (e.g., control scaffolds) gets downloaded on devices encrypted (sequentially) by both the US and China and physically escorted by both sides to inference clusters. At the inference clusters, they are added to the inference whitelist to allow for inference verification on public deployment. + + * **The transparent output channel** caps the total bandwidth coming off the datacenter at 1 MB/s, which means that over a 5 year period, 157 TB leave the datacenter. This is enough to let 50,000 people continuously read outputs at around 10 tokens / second (which might have thousands of corresponding internal chain of thought / scaffold tokens that don’t need to be outputted). Model weights of a current frontier model are currently on the order of 1T parameters (which corresponds to 1TB), so exfiltrating model weights via this channel would require hijacking .6% of the (fully transparent) output channel continuously for 5 years (or a larger fraction for a smaller duration of time), which would very likely be noticed. In Plan A, with compute scaling favoured over algorithmic progress, we expect the model weights of frontier models to become closer to 100T parameters, so a full exfiltration would require a continuous 60% of the output bandwidth for 5 years, ignoring compressibility which shouldn’t be a huge factor. There are additional layers of defense against model weight exfiltration within the datacenter: notably, we attempt to keep model weights on the opaque internal database. + + + + +The remaining threat models involve side channel attacks (such as staff members carrying small flash drives), and we don’t have a creative solution here: access to all of the data centers is tightly controlled, and well-funded, creative security teams are constantly red teaming and hardening the facilities. It also helps that the datacenters in Plan A are built from scratch in areas with low background radiofrequency (RF) noise with security top of mind. + +Finally, the model weights also need to be regularly transported between the R&D datacenters (where they are trained) and the inference datacenters. We imagine that this could happen by transporting the model weights using several physical storage devices that are locked and continuously monitored by representatives of both the US and China for the entire duration of the transportation. + +## Verification integrity + +Goal: The goal for verification integrity should be to achieve a desirable assurance curve: The assurance curve captures information about the level of **coverage** (the % of the AI compute that the verification solution makes claims about) and **confidence** (with what probability is the verification solution correct about the claims it makes on the covered resources). We define an assurance curve as a set of relevant (coverage, confidence) values. + +For example, with the middle curve, there is 99.99% confidence that 99.9% of the compute was verified to have been used correctly. Of course, such curves are theoretical and will be hard to measure, but conceptually we think its useful to think about the verification integrity goal in terms of this curve. + +We think that the verification integrity goal should compare the risk incurred from verification integrity subversions (i.e., internal rogue compute uses) to external ongoing risk from other sources of rogue compute usage, i.e., covert projects, and aim for the possible internal rogue compute uses to be negligible in terms of additional risk vs. the external uses. Since internal rogue compute might have access to the opaque internal information (frontier weights, datasets and potentially some defended algorithms/code depending on the transparency regime) it is potentially much riskier than an equivalent amount of external covert project compute usage. A good approximation of ‘how much riskier’ might be to consider the gap in algorithmic efficiency available inside the R&D datacenters vs. in the potential covert projects. + +Concretely, we think we should take something like the 90th percentile estimate of covert project effective compute (in terms of the frontier project’s current software level, so it's scaled to be comparable to internal compute use), and then set the verification integrity target such that we are 90% confident that the coverage of internal compute usage is at the level X (e.g., X=99.999% of the internal compute) such that the potential rogue internal compute (1-X) would be 10x smaller than the covert project’s effective compute. This basically would be like saying “we are 90% confident that rogue internal deployments would be about equivalent to a 10x smaller covert project than what we think might exist.” + +### Plan A implementation + +Verification in both R&D and inference datacenters needs to be mutually trusted by the US and China, so that they are both confident only the approved training runs, experiments, and inference workloads are running. In our concrete proposal, inference and R&D verification both rely mostly on network taps for evidence collection and trusted recomputation servers for verification. In order to have durable trust in this verification solution, these devices should be mutually trusted (or alternatively, trusted [optical splitters and unilaterally trusted devices](https://nacicankaya.substack.com/p/catching-misreporting-about-ml-hardware-bd2)). There also needs to be physical security to check these devices don’t get tampered with or modified, and trust in the software stack they build on top of them (which probably means it should be extremely simple, and maybe even formally verified). + +Outside of the verification devices themselves, verification will also rely on wider security properties of the cluster that overlap with model weights security. For example, there not being side channels that could let an attacker smuggle out the results of rogue internal computations. + +Over time, the verification assurance will need to improve drastically with the scale of compute. We discuss this in more detail in our [verification supplement](https://www.ai-2040.com/supplements/verification-plan). + +## Algorithmic security + +Goal: This depends on the transparency regime that we are aiming for, which is discussed in our [transparency supplement](https://www.ai-2040.com/supplements/transparency-plan). + +More concerned aboutpoor regulationMore concerned aboutcovert projectsProposal #1:Radical TransparencyProposal #2:Filtered TransparencyProposal #3:Algorithmic SecurityTransparency proposals spectrum \- ai-2040.com + +### Plan A: Total Research Transparency + +To a first approximation, there will no longer be any algorithmic secrets in AI under the Total Research Transparency version of Plan A. There are a few exceptions: + + 1. Massive (potentially synthetic) datasets have similar properties to model weights (in particular, are very large), and so it is viable to limit their output via upload limits out of the R&D datacenter. This isn’t sufficient for preventing the exfiltration of any given data; but could prevent the exfiltration of most data. Data progress is a substantial fraction of overall algorithmic progress, and so securing data might be both tractable and important. + + 2. Hyperparameters found via computationally intensive procedures. For some hyperparameters, efficient settings for them are found via expensive procedures like sweeps over the possible values, and then tests for the training run. These values are small, but do not have to be known by any human. The code for selecting these hyperparameters can be transparent, but we can prevent anyone from knowing the value that running the code found, because it is computationally expensive to determine the value. + + + + +1. + +Initially, this simply means “when the company that owns the model decides to deploy it publicly.” Later there’ll be more regulations, which may or may not take the form of a formal pre-deployment approval process. + +2. + +The process for downloading the weights file is extremely controlled, with both US and Chinese auditors unlocking a secure enclosure where there is a secure device that can make download requests to the opaque internal database and encrypt the files. It is the only device in the entire cluster with a usb (or similar) port that allows any physical device insertion for downloading. + +3. + +1 MB/s × 60 × 60 × 24 × 365.25 × 5 = 157 TB. + +4. + +This assumes FP8 parameters, i.e., that each parameter is 1 byte. + +5. + +In Plan A, we also recommend deliberately increasing the size of the model weights in order to get the security benefit of the weights files being bigger and thus harder to exfiltrate. This would require additional training compute because you are no longer selecting model size to be on the [optimal compute-optimal curve](https://arxiv.org/abs/2203.15556). + +6. + +Cheap compression that requires tiny amounts of compute, and therefore might go undetected by the verification measures, should not go beyond a factor of 2-4x, e.g., through [huffman encoding](https://en.wikipedia.org/wiki/Huffman_coding) (lossless) or cheap post-training quantization techniques (e.g., [](https://arxiv.org/abs/2210.17323)[GPTQ](https://arxiv.org/abs/2210.17323), [](https://arxiv.org/abs/2306.00978)[AWQ](https://arxiv.org/abs/2306.00978), [](https://arxiv.org/abs/2211.10438)[SmoothQuant](https://arxiv.org/abs/2211.10438), [](https://arxiv.org/abs/2402.04396)[QuIP#](https://arxiv.org/abs/2402.04396), [](https://arxiv.org/abs/2401.06118)[AQLM](https://arxiv.org/abs/2401.06118)) which compress weights by 2-4× at negligible quality cost using typically 0.0001-0.01% of training compute. For compression attacks that go further, e.g., more aggressive quantization or LoRA-style decomposition (e.g., [](https://arxiv.org/abs/2106.09685)[LoRA](https://arxiv.org/abs/2106.09685), [](https://arxiv.org/abs/2305.14314)[QLoRA](https://arxiv.org/abs/2305.14314), [](https://arxiv.org/abs/2310.11454)[VeRA](https://arxiv.org/abs/2310.11454), [](https://arxiv.org/abs/2402.09353)[DoRA](https://arxiv.org/abs/2402.09353), [](https://arxiv.org/abs/2303.10512)[AdaLoRA](https://arxiv.org/abs/2303.10512)), which might have >1000x compression, detection of the unauthorized computation to do this compression becomes the more load bearing security property than the bandwidth limit itself. Even still, 100TB weights being compressed 10,000x to 10GB still means you need to use 0.1% of the output channel for 116 days (1 KB * 60 * 60 * 24 * 116 = 10GB) without detection (which might require a large hit from steganography). Regulation can always be aware of the risks of compression and do red-teaming to determine how tractable it is, and adjust training run rules and/or the output bandwidth cap accordingly. diff --git a/sources/ai2040-pages/space-governance-plan.md b/sources/ai2040-pages/space-governance-plan.md new file mode 100644 index 0000000..222bf83 --- /dev/null +++ b/sources/ai2040-pages/space-governance-plan.md @@ -0,0 +1,863 @@ +--- +url: https://ai-2040.com/supplements/space-governance-plan +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Space Governance Plan + +### Mia Taylor (Forethought), Fin Moorhouse (Forethought), Max Dalton (Forethought), Thomas Larsen (AI Futures Project) + +# Preliminaries + +It should be possible for almost all humans to get almost everything that they want from a post-ASI world. With huge numbers of robots capable of almost any task a human can do, it will be possible to provide everyone with the living standard that only the richest people today enjoy. People could spend their time doing what's most meaningful to them—raising children, creating art, or spending time with friends. With ASI-driven advances in medicine, we'll be able to treat mental illness, chronic pain, aging, cancer, and other diseases that shorten or worsen people's lives today. In short, if we achieve the future described in Plan A, there will be enough to go around to give everyone a very good life. Many of the tough tradeoffs that are the focus of modern politics will be resolved. So in some sense, these decisions need not be extremely fraught. + +Yet, at the same time, the decisions we make are extremely high-stakes. If we handle things well, then we could create an enormously flourishing future for our descendants. But there are many ways we could waste this opportunity. We could fight over the resources, leaving a few people in control of the future and everyone else with nothing, or burning vast resources in internecine conflict. Even if we avoid those pitfalls, we could still misuse resources—squandering them on positional goods or failing to leverage ASI to figure out how best to achieve what we truly care about. + +We want a governance system that will allow us to make wise decisions from a foundation of abundance and cooperation. + +This document focuses on how to decide what to do with resources outside the solar system (henceforth referred to as “space resources”). This is an important aspect of post-ASI governance—the overwhelming majority of accessible resources are outside of the solar system and the overwhelming majority of future beings may live outside our solar system—but it is not the only important question. + +## Epistemic status + +We’re currently not confident in how the treaty signatories in Plan A should approach post-ASI governance: it depends on a lot of details about people’s preferences, technological capabilities, and optimal coordination mechanisms that are only partly worked out. + +Luckily, by the time the treaty signatories are making these decisions, they should have superintelligent automated researchers investigating the nature of the universe, doing moral philosophy, and figuring out optimal coordination mechanisms. And we’ll have AI therapists, advisors, and negotiators to help us understand our preferences and make deals with other actors. + +So we think the actual best plan is to defer many of these questions until we have ASI-driven advice and coordination. However, there is value in concreteness and in sharing best guesses, so we wanted to provide some analysis. But we wouldn’t want readers to anchor too much on the ideas we present below. + +## What’s special about post-ASI governance? + +Post-ASI governance will differ from historical precedent in several important ways. + +Technological progress may quickly unlock new pools of resources that no one has previously exploited. The most important example is resources outside of our solar system. + + * Because these resources have been unused, there are no historical claims to navigate. We should aim to govern them in a just and equitable way. + + * There will likely be extreme abundance relative to the present day. With just a tiny fraction of available resources, we’ll probably be able to give everyone alive today a long, happy, healthy, and extremely prosperous life by current standards. + + * We think that the most important question that governance will have to face is what to do with the huge quantities of remaining resources after we’ve ensured that everyone alive has enough to live an incredible life. Some salient options include: + + * _Personal consumption_. People could develop ever-more-expensive ways to derive enjoyment from resources. But most resources are probably poorly suited to use for personal consumption. Most resources are outside our galaxy and thus [millions of light-years away](https://en.wikipedia.org/wiki/Andromeda_Galaxy). If you want to stay on Earth and consume anything—matter, energy, information—derived from another galaxy, you’d have to wait millions of years. If, on the other hand, you left Earth for another galaxy, then if you ever returned, you’d find that millions of years had passed. + + * _Positional goods_. People could spend resources on a zero-sum competition to have the most impressive galactic-scale art projects, or the largest number of statues of themselves. Positional goods could soak up a lot of resources. While there are probably generally diminishing returns to consuming more resources as people get richer, this might not hold for positional goods, since status is relative. We think it would probably be a mistake to spend huge quantities of resources on this type of zero-sum competition. + + * _Satisfying moral preferences_. We could create societies full of flourishing people, or enormous nature reserves. We think this is the most promising way to use the cosmic endowment. But then the question becomes: how should we decide which moral preferences to pursue? + + + + +Technological progress will also change _how_ we make governance decisions: + + * We may be able to make decisions with lasting effects far into the future. + + * For example, we could hand over resources or hard power to ASI systems that have been aligned to a particular set of values. This will likely give us the option of creating [extremely long-lasting institutions](https://www.forethought.org/research/agi-and-lock-in). + + * Typically later generations adopt somewhat different values than their parents, and as a result society’s values change over time. But [new technologies](https://www.forethought.org/research/preparing-for-the-intelligence-explosion#value-lock-in-mechanisms) [may enable parents](https://www.lesswrong.com/posts/8aRFB2qGyjQGJkEdZ/christian-homeschoolers-in-the-year-3000) to exercise tighter control over the values that their children develop. + + * People will have access to superhuman advice and coordination assistance. + + * This means people will be better able to identify and advocate for policies that truly achieve what they care about. + + * But we’ll probably also need to make our governance structure more robust than ever to people trying to game it, either individually or in a group. Just because a governance structure has not been severely exploited in the past does not guarantee that it won’t be severely exploited in the future. + + + + +And we will have large responsibilities to make these decisions go well for people who aren’t able to participate: + + * The overwhelming majority of the affected parties probably won’t be able to participate in the decision to govern space directly, because they won’t have been born yet. This is already true to some extent—in modern democracies, citizens set policies for non-citizens, adults decide for children, humans decide for animals—but in the post-ASI setting this becomes far more extreme. Present people may make decisions that affect the lives of vast numbers of people who cannot give input. The responsibility to get these decisions right is enormous. + + + + +## Evaluating proposals + +In this appendix, we’re evaluating proposals based on how well they achieve the following desiderata, which we believe will be broadly appealing across different value systems. + + * **Preference satisfaction**. The governance system should try to satisfy all the participants’ preferences as much as possible. Some aspects of this include: + + * _**[Paretopianism](https://aiprospects.substack.com/p/paretotopian-goal-alignment)**_. Given a massive windfall, it should be possible to make nearly everyone vastly better off than they are today, and a good proposal should do exactly that. + + * _**Limit extreme downside**_. Our post-ASI governance proposal should have a very low chance of producing outcomes widely considered extremely bad, such as [huge amounts of suffering](https://centerforreducingsuffering.org/research/a-typology-of-s-risks/), or all of space being destroyed. + + * _**Satisfy bounded views**_. If someone has a preference or moral view that is not very costly to satisfy, we should satisfy it. For example, on a cosmic scale, it’s fairly cheap to ensure that every human alive in 2040 has a long, flourishing life. + + * **Satisfy common-sense ethics**. The process should not violate common-sense ethical views. This weighs against proposals where a small number of people predictably end up controlling the fate of the universe or where future generations end up totally disempowered. + + * **Satisfy objective morality**. Some aspects of ethics may be objective and discoverable. If so, we should try to identify and conform to them. + + + + +These next criteria aren’t terminal goals, but we think that processes that satisfy them are more likely to achieve the criteria above. + + * **Coordination**. The governance system should enable coordination on large-scale collective action problems. For example, if we encounter alien civilizations, we may need to coordinate to establish friendly relations with them. We may also need to coordinate to avoid actions that risk destroying all of Earth-originating civilization, or to avoid space races where factions compete to claim as many resources as possible, burning huge amounts of resources in the process. These goals probably require some minimal government functions: protecting property rights, enforcing prohibitions, and potentially collecting taxes to fund public goods. + + * **[Moral trade](https://amirrorclear.net/files/moral-trade.pdf)**. It might be possible to nearly satisfy many different moral viewpoints simultaneously. For example, some views might especially care about resources close to Earth or the near future, while others might be more indifferent across time and space: we might satisfy both by letting the former choose what happens in the Milky Way while the latter choose what happens in the rest of the universe. Here are two types of trade that we want our proposal to facilitate. + + * _**Bilateral trade**_. This is particularly well suited for satisfying uncommon preferences. These trades are most straightforward in market systems, where two people can literally trade deeds to resources or sign agreements about how they will use their resources. But it’s possible that some voting systems–e.g., [the](https://en.wikipedia.org/wiki/Vickrey%E2%80%93Clarke%E2%80%93Groves_mechanism) [Vickrey-Clarke-Groves mechanism](https://en.wikipedia.org/wiki/Vickrey%E2%80%93Clarke%E2%80%93Groves_mechanism)—could also achieve similar results. + + * _**Coordination on[moral public goods](https://www.forethought.org/research/moral-public-goods-are-a-big-deal-for-whether-we-get-a-good-future)**_. People might share some values while also holding idiosyncratic preferences—individually preferring to pursue their own preferences, but collectively preferring that everyone work toward the shared values. + +Coordination on moral public goods is most straightforward in voting systems, since people can just vote to devote resources to shared projects. It’s possible that voluntary contracts—where people pledge to spend resources on shared goods if others do the same—could achieve similar results, but this is less straightforward because of [free-riding incentives](https://www.forethought.org/research/moral-public-goods-are-a-big-deal-for-whether-we-get-a-good-future#appendix-b-causal-coordination-through-voluntary-contracts). + +It might be particularly effective to coordinate to switch from funding positional goods to funding moral public goods. If done in a way that leaves people's relative ability to purchase positional goods unchanged, then this is great for everyone: relative status doesn't change, but people are able to satisfy their non-positional preferences, too. + + * **Reflection**. ASI will likely uncover novel considerations and empirical facts highly relevant to how people want to use space resources (e.g., are there objective moral truths, and what are they?). We want each person to have the time and resources necessary to think carefully about what they value before making important decisions, and we want to ensure that we invest in ASI-fuelled research on important questions to inform people’s decisions. + + * **Pluralism**. Humans today have many different visions of the good. It’s possible that after deep reflection with ASI we’ll converge on a single set of shared values. More likely, we’ll still have some disagreements. Where people have different value systems, we prefer proposals that will allocate resources to pursue all of them. + + * **Avoid brittle decision-making structures**. There are likely many important considerations that we’re not yet aware of. Where possible, we favor flexible decision-making structures that won’t make it difficult to solve unforeseen problems. + + * **Fairness**. Each living adult should have an opportunity to give equal input into the process. We don’t think that this criterion necessarily means that each person ends up controlling an equal quantity of resources. For one thing, some people might care a lot about what happens with resources closer to Earth while not caring much about resources in distant galaxies, and we’re happy for them to trade with people who care a lot about distant galaxies. Nor do we think that this criterion requires that everyone is equally happy with the outcome. For instance, if some decision is made by vote, then some people might end up voting for the losing side. But we want a neutral process that doesn’t privilege any particular people or viewpoints over others. + + * **Political feasibility**. Unlike the other criteria, this is purely pragmatic: a proposal must be acceptable to the relevant decision-makers to be implemented. It’s unclear who these decision-makers would be, but they could include the citizens and/or leadership of the US, China, or other powerful countries. It might also include AI company CEOs or shareholders. + + + + +# Summary of the most promising options for allocating space resources + +We think that it’s best to decide on a specific structure for allocating space resources in 2040 at the time that the treaty is decided ([more discussion here](/supplements/space-governance-plan#should-we-commit-to-a-process-right-away)). + +We’re currently fairly uncertain about what the best structure is, and that decision should be deferred until decision-makers have access to aligned ASI advice. But currently, these seem like the best object-level proposals. + +Basic proposal + +**Direct endowment + trade**(with prohibitions). Each adult is assigned an equal share of resources. They can trade, donate, give away, or use those resources in any way they want, except for a small category of prohibited activities (more). + +Key advantages + +Most views held by living humans will have at least some resources devoted to them. + +Bilateral trade between people with different moral views is very straightforward. + +Key disadvantages + +It’s harder to coordinate to fund moral public goods due to free-rider problems. + +Basic proposal + +**Vote on how to use resources**. Each adult gets to participate in a vote on how to use space resources (more). + +Key advantages + +It’s straightforward to coordinate to fund moral public goods. + +Key disadvantages + +All known voting systems have serious issues, such as: + + * Inefficiencies, e.g., strong minority preferences are overruled by very weak majority preferences. + + * Incentives to strategically misrepresent preferences, either alone or in collusion with others. + + + + +Basic proposal + +**Direct preference aggregation**. We use novel technology to directly elicit people’s preferences (e.g., through an interview with an ASI using lie detection) about the use of space resources and aggregate these preferences (more). + +Key advantages + +Many voting systems fail to get a globally-optimal outcome because people strategically misrepresent their preferences. But direct endowment fails to fund moral public goods. By eliciting people’s honest preferences and then aggregating them, we might be able to get the best of both worlds. + +Key disadvantages + +People might object to the preference elicitation step, especially if the method is time-consuming or invasive. + +**Basic proposal**| **Key advantages**| **Key disadvantages** +---|---|--- +**Direct endowment + trade**(with prohibitions). Each adult is assigned an equal share of resources. They can trade, donate, give away, or use those resources in any way they want, except for a small category of prohibited activities (more).| Most views held by living humans will have at least some resources devoted to them.Bilateral trade between people with different moral views is very straightforward.| It’s harder to coordinate to fund moral public goods due to free-rider problems. +**Vote on how to use resources**. Each adult gets to participate in a vote on how to use space resources (more).| It’s straightforward to coordinate to fund moral public goods.| All known voting systems have serious issues, such as: + + * Inefficiencies, e.g., strong minority preferences are overruled by very weak majority preferences. + * Incentives to strategically misrepresent preferences, either alone or in collusion with others. + + +**Direct preference aggregation**. We use novel technology to directly elicit people’s preferences (e.g., through an interview with an ASI using lie detection) about the use of space resources and aggregate these preferences (more).| Many voting systems fail to get a globally-optimal outcome because people strategically misrepresent their preferences. But direct endowment fails to fund moral public goods. By eliciting people’s honest preferences and then aggregating them, we might be able to get the best of both worlds.| People might object to the preference elicitation step, especially if the method is time-consuming or invasive. + +(We discuss alternative proposals that we like less [here](/supplements/space-governance-plan#unpromising-options)). + +Other important decision points include: + + * At what time should votes, preference elicitation, or distribution of resources happen? (discussion [here](/supplements/space-governance-plan#early-decisions-vs-late-decisions)) + + * Should later generations get a say? (discussion [here](/supplements/space-governance-plan#later-generations)) + + * Should we prefer a hybrid approach that allocates some resources according to one proposal and allocates other resources according to other proposals? (discussion [here](/supplements/space-governance-plan#one-process-or-many-processes)) + + * How do we enable each decision-maker to make the best decisions by their own lights? (discussion [here](/supplements/space-governance-plan#reflection)) + + + + +# Some available options + +Here, we discuss options for governing future resources. We begin by describing a couple of notable options which seem relatively unpromising on the criteria above, then turn to proposals that we find more promising. In both cases, we’re trying to cover important and illustrative cases, not every possible outcome. + +## Unpromising options + +### Anarchy + +No property rights system is established for space resources. To use a resource, an actor must either be the first to claim it or seize it from another actor. Actors can do whatever they want with resources they currently control—there’s no governance system for prohibiting particularly bad behaviors or any other forms of large-scale coordination (e.g., to fund public goods). + +_Drawbacks:_ there could be significant negative-sum conflict over resources, and insecure property claims would make trade—including moral trade—harder. + +Anarchy is worse than the feasible alternatives. We do expect that it is not a stable long-term equilibrium. Given much more advanced tools for coordination and communication, we expect people to create ad hoc governance structures. But a government that emerges organically from anarchy could be worse than one designed intentionally. In particular, it might favor the interests of those who were able to claim and defend lots of space resources early on. + +### Finders-keepers property rights + +Actors can claim unclaimed resources or they can buy resources from earlier owners, but they can’t seize claimed resources without the owner’s permission (similar to homesteading). As before, no effective system exists for large-scale coordination or public goods provision, and no governance system can enforce prohibitions. + +_Strengths:_ we think this is an improvement on anarchy. Once a claim is recognised, the holder no longer needs to spend resources defending it. Moreover, secure and widely recognised property claims make trade possible. + +_Drawbacks:_ value is no longer wasted contesting claimed resources but is still wasted in the race to claim them. It also seems likely that a small number of powerful actors—perhaps whoever can send out a wave of probes first—will claim most resources. And this proposal has no mechanism for public goods provision or enforcement of prohibitions, so it forgoes potentially large gains from moral public goods and from preventing egregiously unpopular uses of resources. + +Next, we discuss governance proposals that we think improve on these scenarios according to the criteria above. + +## More promising options + +### Direct endowment and trade + +Each person receives title to a portion of all future resources, which gives them the right to decide how those resources will be used in the future. For example, people can sell their shares of future resources, gift them to other people or to trusts, or [make bilateral deals with each other](https://www.forethought.org/research/should-we-lock-in-post-agi-agreements-under-uncertainty) about how they’ll use their resources. Property rights are well-enforced in practice. As before, no effective system exists for large-scale coordination or public goods provision, and no governance system can enforce prohibitions. + +(A related approach: auction titles to all space resources and then redistribute the proceeds equally. In theory, if we make some standard assumptions, this scheme would result in the same distribution of ownership as direct endowment.) + +_Strengths:_ when property rights are clear and well-established from the outset, bilateral trade—including moral trade—becomes easier. Compared to finders-keepers, there are no longer incentives to waste resources through racing. + +Since each property-holder is free to decide what to do with their share of the future, at least some resources will be devoted to pursuing most moral views, making it likely that cheap-to-satisfy views will be satisfied. + +_Drawbacks:_ moral public goods would have to be funded through voluntary contracts, and it’s not clear whether voluntary contracts will be sufficient to ensure that moral public goods are funded at the optimal level, since [individuals will face a strong incentive to free-ride](https://www.forethought.org/research/moral-public-goods-are-a-big-deal-for-whether-we-get-a-good-future#appendix-b-causal-coordination-through-voluntary-contracts). + +Since direct endowment doesn’t involve a way to impose universal prohibitions on how to use future resources, the same drawbacks carry over from anarchy and finders-keepers: if there are some uses of resources which are egregiously bad by most people’s lights, but preferred by a few actors, it may be hard to prevent those uses, and bad actors could even threaten those uses to extract concessions. This is the mirror image of the moral public goods drawback: there is no easy way to coordinate to prevent moral public bads. + +A fuzzier concern: we think it’s probably better for the current generation to think of themselves as stewards who seek to create the best future by their own lights in the space that they control, rather than as “space emperors” who try to squeeze the maximum personal enjoyment or aggrandizement out of their share of resources. If people get unilateral control over space resources _and_ they’re allowed to buy and sell them like normal property, we worry that this might push people to think of the resources as a personal windfall rather than a serious moral responsibility. + +_Uncertainties and decision points:_ We’d also need to make a call on decisions like: + + * Who gets resources? (For example, should resources be reserved for future generations?) + + * Should people be allowed to sell, transfer, or wager their resources immediately? Or do we want to—somewhat paternalistically—enforce a waiting period for people to carefully consider their options? + + * What is the relationship of a galaxy’s “owner” to the future inhabitants of the galaxy? + + + + +One important variant on this is to add some mechanisms for preventing people from using resources in egregiously bad ways. In the section [below](/supplements/space-governance-plan#mitigating-downside-risks), we discuss some possibilities. + +### Voting on collective uses of resources + +Everyone votes on proposals for what to do with space resources. These proposals could range from very concrete (e.g., “create digital minds according to this specification”), to very meta (e.g., “delay doing anything with space resources right now, and re-assess with another vote X years from now” or “allocate some future resources to private ownership”), or something in between. + +_Strengths._ Compared to the other proposals so far, it’s easier to fund [moral public goods](https://www.forethought.org/research/moral-public-goods-are-a-big-deal-for-whether-we-get-a-good-future#scenario-1-causal-coordination)—if people prefer resources to go to a shared project rather than divided among all voters, they can vote for that proposal. + +It’s also a more flexible plan: once future resources are transferred to private ownership, recovering the ability to make collective decisions might be impossible even if almost everyone wanted it. With some kind of regular vote, the electorate can switch to a different plan at the time of the vote, including agreeing to delay doing anything with space resources. + +_Drawbacks:_ All known voting systems have significant problems. Some are simply inefficient: for instance, majority rule faces the familiar problem that the majority view can succeed when the majority wouldn’t deem the minority's preferred option much worse, but the minority greatly dislikes the majority’s preferred option. Others incentivize people to [misreport their preferences](https://en.wikipedia.org/wiki/Gibbard%E2%80%93Satterthwaite_theorem), individually or through collusion, to benefit themselves at the expense of overall value. This problem is especially serious when people know the details of the voting system well in advance and have ample time to develop strategies to game it with ASI assistance. + +A further drawback is that voting requires synchronized decisions: everyone must vote at the same time. If the vote happens in 2100, you vote according to your preferences at that time, even if you would have preferred more time to reflect first. In principle, people could vote to delay the vote. But you could still face situations where a majority wants to vote now while a minority wants to keep reflecting, and the minority gets overruled. Contrast this with direct endowment, where each person can reflect for as long as they want before deciding what to do with their resources. + +_Uncertainties and decision points_ : How badly can strategic voting distort outcomes in practice, and can novel technology (e.g., lie-detection technology) mitigate this? How well does voting handle bilateral trades between people with rare preferences? And we'd need to decide: when do we hold the vote, who is included, and which voting system do we use? + +### Direct preference aggregation + +Similar to voting, but instead of people casting votes, AI systems try to directly ascertain people’s preferences. For example, AI systems might build a profile of a person’s preferences from all their writing, extensive interviews perhaps under lie detection, and/or brain scans. Then, everyone’s preferences are aggregated to decide what to do with space resources. As with voting, the decision might be quite meta—e.g., wait 100 years and then do preference aggregation again. + +It’s not clear how best to aggregate preference, but here’s a sketch of a proposal. Everyone is first assigned 1/N of future resources as a baseline—their fallback if no gains from trade are found. An AI arbiter then identifies reassignments of resources that would improve everyone's preference satisfaction above this baseline, splitting the gains fairly among all beneficiaries. This handles bilateral moral trade: if person A most wants statues of herself and person B most wants statues of himself, but both value statues of George Washington at 60% as much, the arbiter directs their combined resources toward Washington statues, leaving both better off than their baseline. It can also handle moral public goods: if funding some public good would improve everyone's welfare, the arbiter taxes each person in proportion to how much they value that good ([a Lindahl tax](https://en.wikipedia.org/wiki/Lindahl_tax)). + +_Strengths._ Direct preference aggregation is a modification of the voting proposal above and inherits its strengths. Its major additional advantage is that it makes strategic manipulation more difficult than in voting systems. + +_Drawbacks._ Eliciting people’s honest preferences would require some novel technology and it’s not clear what this process would look like. It’s possible that this process would be invasive, time-consuming, or unpleasant for participants (e.g., an interview while hooked up to a lie detector). It seems unjust and contrary to common-sense morality to require that people go through a procedure that they strongly object to in order to have input into the allocation of future resources. + +There are a few more technical issues. For instance, people could strategically modify their preferences before assessment, breaking the main reason for trying to directly elicit people’s preferences (although mitigations might be possible—e.g., governments could ban people from modifying their preferences using novel technologies or require that they register their pre-modification preferences). We also haven’t spelled out a specific algorithm for doing the preference aggregation, and it could be difficult to avoid perverse or hard-to-predict outcomes, especially if many people are best represented as having inconsistent or incoherent preferences. Indeed, the question of how to represent someone’s preferences, including on questions they haven’t explicitly considered, seems very contentious and underdetermined. + +### Defer to superintelligence + +Instead of any particular method of inferring and aggregating people’s preferences, humanity builds aligned superintelligence to plan on its behalf and give it control of nearly all space resources. + +For example, we could start with a “council” of superintelligent delegates, each starting from one of many diverse moral and philosophical traditions. Those delegates then engage in deep reflection and deliberation with one another. The delegates can initially be trained to be reliable and accurate on all verifiable domains, like forecasting and mathematics, and not post-trained to strongly avoid novel, ‘weird’, or otherwise unpopular views just because they are currently unpopular. If the superintelligences converged onto a single moral view, then they could devote space resources to that plan. If the superintelligences continued to have different views, they could come to a compromise plan via voting, direct endowment, or direct preference aggregation. + +Or we could design an ASI system to be motivated to discover and act on objective moral truths, if they exist. (If they don’t exist, then it would revert to a backup plan—perhaps one of the proposals above). + +_Strengths_. Determining the best plan under a particular moral view probably requires discovering and grappling with lots of novel and counterintuitive considerations. The resulting plan might seem [very weird and alien to humans](https://www.forethought.org/research/convergence-and-compromise#24-an-argument-against-wam-convergence), and they might not be motivated to act on it, even if they’re assured by a superintelligence that it is the morally best course of action. People might also have competing personal interests that take precedence over doing the right thing. But an ASI could potentially be designed to have no competing personal interests. These considerations are particularly strong if moral realism is true, and a plan like this might be our best shot at devoting substantial resources to pursuing objective morality if it exists. + +_Drawbacks_. The ultimate outcome of this process might depend a lot on the details of the process—e.g., what initial viewpoints were represented on the “council” of deliberating ASIs? If this is true, who should decide? It seems rather unfair for the treaty signatories to get to choose. + +We think that a process like this might be the best ultimate use of space resources. But if so it should probably result from one of the proposals above. For example, if many people found moral realism compelling after reflection, they could vote to spend public resources on the most morally valuable uses (or express a preference in direct preference aggregation). (It’s unclear whether the direct endowment proposal is sufficient for this: handing off resources to morally-aligned ASI is probably a moral public good in this scenario, so we’d expect it to be underfunded if people were individually deciding how to spend their resources). + +# Other decision points + +## Should we commit to a process right away? + +The treaty signatories could approach this decision in three ways: + + 1. Decide on a specific process in 2040. + + 1. E.g., schedule an election in 2070 using [quadratic voting](https://en.wikipedia.org/wiki/Quadratic_voting) where all living people over the age of 16 vote on what to do with all extrasolar resources. + + 2. Decide on a process in 2040 for picking a process later. + + 1. E.g., schedule an election in 2070 using [approval voting](https://en.wikipedia.org/wiki/Approval_voting), where all living people over the age of 16 vote on what decision process to use. + + 2. E.g., design an ASI arbiter who will do a bunch of research and then pick a process that best satisfies some abstract desiderata put forward by the signatories of the Treaty in 2040. + + 3. Make no explicit decision about a process or a process for picking a process. + + 1. It’s not clear we’d end up deciding what to do with space, under this option. Perhaps at some point one of the signatories to the Treaty would decide that they wanted to colonize space and either do so unilaterally or negotiate with the other signatories to be allowed to do so. + + + + +Option 3 is probably worse than options 1 or 2. Having a concrete proposal—even a meta-level one about how to select an object-level process—allows public scrutiny, which we expect would push the proposal to be fairer and more principled. Without a concrete proposal, two failure modes become more likely. First, parties who especially wanted space resources might start lobbying the Treaty governments, and those who are well-connected or have invested heavily in lobbying would end up with outsized influence over how space resources are used. Second, some parties might start unilaterally colonizing space, which could trigger a space race. Governments could ban unilateral colonization, but that would require a plan for lifting the ban—at what point are people actually allowed to colonize?—which brings us back to option 2. + +We’re less confident about whether option 1 or option 2 is better. + +The main argument for option 2 over option 1 is that we’ll probably get more information about how to design a good process after 2040. + + * We might discover new voting or preference aggregation systems that avoid issues with systems that we know about. However, we’re relatively pessimistic about learning much on this front after 2040, given existing [impossibility results](https://www.jstor.org/stable/1911219) showing that voting systems can't simultaneously satisfy a set of desirable features. Also, designing new preference aggregation systems is a relatively well-scoped, tractable mathematical problem so we would be somewhat surprised if we didn’t have the answer in 2040 after devoting a small amount of ASI labor to it. + + * We might learn how well particular processes are likely to work in practice. + + * For instance, we might learn about strategies that people are actually using to try to game the current system and adapt the rules accordingly, e.g., if we have rules intended to prevent parents from controlling their children’s values to an excessive degree, and it turns out that some parents can find loopholes to exert a lot of control, we could adapt the rules in response). + + * We might discover new tools or capabilities that enable a better process. + + * For instance, we might develop new methods for noninvasively scanning people’s brains to determine their preferences. + + * We might discover new considerations that affect the optimal preference aggregation process. + + * For instance, we might identify valuable forms of moral trade that we want our preference aggregation process to accommodate. + + + + +**On the other hand, people will also accumulate information about the likely object-level consequences of each possible process**. For example, suppose that we decided to vote in 2070 to decide on a procedure for deciding what to do with space, choosing between the policies of “hold a majority vote and do whatever the winning proposal says” and “direct endowment.” And at this time, it turns out that 50.1% of people favor colonizing the Milky Way while leaving the rest of space as a permanent nature preserve, while 49.9% have plans to colonize more distant galaxies. + +In 2070, it would be clear that if we chose the majority vote, then all space resources would be allocated to the majority’s colonize-the-Milky-Way-and-nothing-else plan, but if we went with the direct endowment plan, then some galaxies would be left as nature preserves, while others would be colonized by the minority. Members of the colonize-the-Milky-Way-and-nothing-else coalition would then be strongly tempted to favor the majority vote when choosing a process. + +At best, the process-revision vote becomes redundant—we might as well have voted directly on the object-level outcome. At worst, we lose the opportunity for people to agree on a process grounded in principles like fairness or Paretopianism. + +One possible way around this would be to delegate process revisions to an AI arbiter that adjusts the process based on principles established in 2040, without regard for which factions benefit from any given change. This approach relies heavily on being able to verifiably align an AI system to a set of principles. It also requires being able to identify principles in 2040 that are specific enough for the AI arbiter to apply without heavy extrapolation, but not so specific that we forfeit much of the upside of deferring process decisions to a later date. But if it’s possible, this might be a good way to incorporate new information in a principled way. + +## Early decisions vs. late decisions + +Once we’ve picked an object-level process—like a vote, a preference aggregation procedure, or directly distributing deeds to space resources—there’s a question of when to run that process. + +Here are some advantages of delaying the process: + + * People have more time to reflect on what they care about and how they endorse using space resources before it’s time to make a decision. This matters most for voting and direct preference aggregation, where everyone must decide simultaneously. + + * People may have access to more information at the time of the decision—e.g., a clearer sense of what is technically feasible to do with space resources. + + * It’s possible to incorporate later generations into the process—e.g., allow them a vote, take their preferences into account for direct preference aggregation, or distribute resources directly. This seems good on grounds of common-sense ethics, as it’s plausibly unfair for the current generation to make decisions about the use of space resources on behalf of all future generations. + + + + +Here are some advantages of running the process relatively early: + + * Some types of [potentially valuable moral trade](https://www.forethought.org/research/should-we-lock-in-post-agi-agreements-under-uncertainty) are possible only if participants are uncertain about the future, and people’s uncertainty about the future will probably decrease over time. + + * People might want to bet on the outcomes of some reflective processes. For example, Alice might care deeply about maximizing positive impact if moral realism turns out to be true and discoverable, but care little about having resources if there are no objective moral truths. If Bob values resources equally regardless of whether moral realism is true, Alice and Bob might make the following bet: if ASIs discover objective moral truths, then Bob gives some of his resources to Alice. Otherwise, Alice pays Bob ([more discussion of these wagers here](https://www.forethought.org/research/should-we-lock-in-post-agi-agreements-under-uncertainty#three-kinds-of-early-agreement)). + + * People might also want to bet on empirical facts. For instance, we may still have uncertainty about how many galaxies are within the reachable universe and which of those galaxies are already occupied by aliens. A fairly risk-averse person might buy resources that are likely reachable and unoccupied, in exchange for a larger amount of more uncertain resources, while a more risk-tolerant person might take the other side of the trade. + + * These deals are most straightforward under direct endowment, but they could also work in voting systems, e.g., a coalition of moral realists and subjectivists might be willing to vote for a proposal that spent resources in the most moral way (conditional on moral truths existing and being discoverable) and otherwise distribute the resources to the subjectivists’ preferred purposes. + + * Delaying decisions also gives people more time to figure out how to strategically misrepresent their preferences ahead of a vote or preference aggregation process. + + * Running the process earlier may allow more flexibility. + + * If an early vote is scheduled and delaying seems wise at the time of the vote, voters can simply vote to defer to a later vote. If resources are distributed early, recipients can take their time deciding what to do with them. + + * On the other hand, if the process is scheduled for later, then we can’t make an early decision even if that would have been wise. + + + + +## Who should participate in the process? + +In this section, we discuss who should participate in deciding how to use space resources. This is distinct from who should receive other rights, such as autonomy over their own lives or the right not to be mistreated. We discuss more about how those basic rights could be assured in a [later section](/supplements/space-governance-plan#mitigating-downside-risks). + +### People with substantial bargaining power in 2040 + +It’s unclear who these people will be, but some plausible candidates are: AI company CEOs, AI company shareholders, the US president (and advisors), the US electorate, Chinese leadership, or the leadership and/or electorate of other countries that participate in the treaty. + +This is probably the default outcome. It’s also very unfair, contrary to common-sense ethics, and could lead to many people’s cheap-to-satisfy preferences being neglected unless powerful people are sympathetic to them. We’re also concerned that if many people expect that this is how decisions will be made, then they’ll focus on jockeying over power during the intelligence explosion rather than cooperating to avert risks. + +### All living humans + +Another option is to share power among all humans who are alive in 2040, as in the Plan A scenario. This is fairer and more in line with common-sense ethics than the previous plan. And, to the extent that decision-makers can credibly signal that this will happen, committing to this plan might make people more cooperative during the intelligence explosion (although it’s unclear how credibly decision-makers could commit to this). Unlike proposals that share power with later generations (discussed [next](/supplements/space-governance-plan#later-generations)), this option makes it possible to make decisions relatively early, [which might be important](/supplements/space-governance-plan#early-decisions-vs-late-decisions). + +On the other hand, it’s unclear how politically feasible this will be. And it doesn’t include everyone who might deserve a say in how space resources are used—it excludes future people and digital minds. + +Some plausible options: + + * **Include all living humans**. This seems like a fairly natural coordination point. + + * **Include all living humans in some capacity, but give people with bargaining power extra weight**. For instance, more powerful people might get larger endowments or have their votes or preferences more heavily weighted. This might be a more politically feasible compromise. + + + + +### Later generations + +“Future” generations are the generations born after the intelligence explosion. The “current” generation is the set of humans alive in 2040. We could include future generations in decisions by: + + * Delaying votes (or rounds of preference aggregation) until future generations are adults. + + * Reserving some decisions for future generations, e.g., allowing the current generation to vote on how to use 10% of space resources, but letting later generations decide at later votes how to use the remaining 90%, or directly distributing deeds to 10% of resources to the current generation, but saving 90% of resources for later generations. + + + + +One reason to share power with later generations is that it’s plausibly unfair not to. However, if we share power with future generations, then this potentially incentivizes parents to have as many children as possible and indoctrinate them to share the parents’ beliefs and values. This could let parents swing a later vote or ensure that a larger share of resources goes to people who share their values. + +It’s not clear how big of a problem this is. On the one hand, it’s currently quite rare for ideological groups to deliberately try to have as many children as possible in order to increase their political and economic power. And we expect that most people would find the prospect of having huge numbers of children and brainwashing them to vote the parents’ way abhorrent. On the other hand, new technologies might lower the personal costs of having huge numbers of children (e.g., artificial wombs, robot caretakers). Other technologies could increase parents' ability to transmit their values to children (e.g., through genetic engineering or novel techniques based on advances in neuroscience or psychology). And a small number of bad actors might be able to substantially increase their population share within a few generations through rapid reproduction. + +One way to navigate this is to place guardrails around the rate of reproduction or the extent of control—say, trying to ensure that parents could not use novel technology to reproduce substantially _faster_ nor exercise substantially greater control over their children’s values than early-21st century parents were able to. + +Some plausible options: + + * **Don’t include future generations directly**. Instead, let the current generation decide whether and how to include them. + + * For example, if decisions were made by a vote, the current generation could vote to delay decisions until a later vote, where future generations could participate. (And analogously, under direct preference aggregation, if the current generation prefers to reserve some power for future generations, then that would happen). + + * Under an equal distribution proposal, the current generation could share resources with their descendants or donate to trusts that gifted resources to later generations. + + * (Although power granted to future generations is plausibly a moral public good, so we might expect it to be undersupplied in the equal distribution proposal). + + * **Include later generations, with restrictions to prevent strategic reproduction and/or indoctrination of one’s children**. These restrictions could be chosen by a vote of the current generation. + + * **Include later generations, with no restrictions to prevent strategic reproduction or indoctrination of one’s children**. + + + + +### Digital minds + +Another difficult question is whether and how digital minds should have input into decisions about space resources. Digital minds could include AI systems or [simulated human brains](https://en.wikipedia.org/wiki/Mind_uploading). + +In the Plan A scenario, by 2040, there may be huge numbers of digital minds. Some of these minds could be very sophisticated, with long-term stable preferences, human or superhuman intelligence, and perhaps even [subjective experience](https://en.wikipedia.org/wiki/Artificial_consciousness). On some views, such minds have moral status and deserve a say in future governance. + +But this is pretty controversial and might remain so in 2050. Different people could have wildly different views depending on their beliefs about what makes a being worthy of inclusion in the political process. It seems challenging to decide in a value-neutral way which digital minds should be empowered, and choices here could radically affect the composition of the electorate. + +There are further difficulties in integrating digital minds into the decision process. While it’s conceptually straightforward to give each human adult an equal stake in the decision—e.g., one vote, one share of space resources—it’s less clear how to individuate digital minds. Should two copies of the same mind running on different computers each receive a vote? If not, how different do two minds have to be before they count as different individuals? + +The issues of strategic reproduction and indoctrination discussed above are even more severe for digital minds than for human children. New digital minds could be created as fast as copying data from one computer to another. And the creators of digital minds may have a huge amount of fine-grained control over the values and behavior of those minds. Given sufficient compute, it would be trivial to spin up billions of minds that sincerely want to vote exactly the way their creator tells them. + +Some plausible options: + + * **Include a subset of digital minds** where the concerns about gaming the process through strategic reproduction and indoctrination are less strong. + + * Some possible subsets include: + + * Allow digital minds who were created before 2040 to participate in the process, perhaps with some overall cap on the share of total votes or resources they could receive. + + * (We would still somehow need to resolve the issue of individuating digital minds and deciding which digital minds deserved votes). + + * Allow digital minds that are uploads of deceased biological humans to participate in the process, one upload per original human. + + * Allow people to “adopt” digital minds and enfranchise those adopted digital minds, but only if the creators exercise no more control over the values and beliefs of those digital minds nor are producing digital minds at a faster rate than parents in 2026 could with biological children. + + * Most likely, most digital minds would not be enfranchised under any of these proposals. + + * **Don’t include digital minds directly**. Instead, let biological humans decide whether and how to include them. + + * How this decision is made would depend on the proposal: + + * For example, if decisions were made by a vote, the biological humans could vote on a proposal to defer decisions to a later vote where digital minds were allowed to participate. + + * Under a direct endowment proposal, humans who received resources could give resources to digital minds. (Although, again, this would plausibly be a moral public good and we might expect it to be underfunded under the direct endowment proposal). + + + + +## One process or many processes? + +Designing a governance system involves many tradeoffs—for example, direct endowment and voting both have disadvantages. Should we hedge our bets with hybrid approaches that combine multiple processes? + +Here are some examples of hybrid approaches: + + * 50% of space resources are distributed directly to the population at a set time, and 50% are allocated according to the outcome of one or more votes. + + * Each country (or each treaty signatory) is allocated a share of space resources. These countries each pick their own processes for deciding how to use space resources. + + * 10% of space resources will be distributed in 2040, 10% will be distributed in 100 years, 10% will be distributed in 1000 years, and so on. + + + + +The main advantage of hybrid approaches is that different processes can direct resources toward different kinds of values. For example, votes may be particularly likely to fund widely shared values, whereas direct resource endowment might be better suited towards ensuring that rarer preferences are funded. By adopting a hybrid approach, we ensure that a broader set of values each get at least some resources allocated to them. + +Voting on how many resources to allocate with each process. + +A hybrid approach could directly incorporate people’s views on which governance and allocation processes are better, and use this to decide how many resources should be allocated by each. + +Here’s one such proposal. + + * First, a number of governance methods are proposed. This could include all the kinds of systems discussed above (direct endowment, voting using various different voting systems, direct preference elicitation, etc). The only constraint is that each governance method must be reasonably fair in how it weighs participants’ views, rather than artificially biasing the outcome towards specific people or values. + + * Second, all humans submit a view on what fraction of resources should be distributed by each governance process. The fraction of resources governed by each process is set to the **average** of everyone’s proposal. + + * Finally, all humans participate in each of the processes—deciding what to vote for or what to do with resources they’ve been allocated. + + + + +One advantage of this approach is that everyone gets a say on what governance processes get used. Another advantage is that it delays the decision about which process to use. If it later turns out that there are big gains from coordinating on moral public goods, people can choose to allocate most resources using voting. Conversely, if some minority of people realize that they might get frozen out of a large-scale vote, they can cause a significant fraction of resources to be distributed via direct endowment instead. + +This approach could plausibly be further improved by allowing people to freely trade their voting rights at each decision point. + +In other words, people’s right to participate in each different governance process (each vote, each round of resource endowment, etc) could be represented by different shares. These shares could then be freely traded. This might, for example, allow people with more idiosyncratic moral views to get more rights to resources that they have unilateral control over. Conversely, people who care more about the outcome of large-scale votes could get more of the corresponding votes. We could even allow people to trade the right to vote on what fraction of resources should be allocated by each process. + +(We find it difficult to predict the outcome of such trades, and it’s possible that it would be better to disallow some of them.) + +The main disadvantage of hybrid approaches is that one process might generate outcomes that other processes would regard as substantially harmful. This can be mitigated through some of the strategies discussed in the next section. + +## Mitigating downside risks + +Under some governance regimes, individuals will be able to unilaterally decide how to use space resources. This has some strong upsides: people can create stuff that’s valuable to them even if others are indifferent to it, which makes it more likely that bounded views are satisfied and makes it more likely that nearly everyone is much better off. But it also has downsides: people might use their resources in ways that impose large negative externalities on others. + +Some specific examples include: + + * Seizing other people’s resources. Early settlers might try to use their star systems to launch probes to settle star systems earmarked for others who hadn’t yet laid physical claim on them, which could initiate a ([potentially highly negative-sum](https://mason.gmu.edu/~rhanson/filluniv.pdf)) space race. Or people might attack star systems where others had already settled. + + * Uses of resources that most people find intrinsically repugnant. For example, some people might make themselves absolute rulers over societies where nearly all citizens are oppressed or enslaved. Sadists might create new beings that experience immense pain. + + * Risking the destruction of other people’s resources. People might deliberately destroy the universe (e.g., by initiating false vacuum decay). + + * Taking other actions with severe negative externalities on others, e.g., starting a war with space-faring alien civilizations. + + + + +We need to strike a balance. If we’re too permissive, then a few bad actors could cause enormous damage. But if we’re too restrictive, then we risk banning things that are enormously valuable to a few people just because the majority finds them a bit offputting. + +Here are some potential systems for mitigating these downside risks. + +**Market**. If someone doesn’t like how someone else is using their resources, they can pay them to stop. + +_Strengths_ : under this proposal, an activity is only stopped when its cost to others exceeds its benefit to the actor, and they’re compensated for their loss. Thus, this proposal is less vulnerable to “tyranny of the majority” worries. Another strength is that we don’t have to collectively decide far in advance what to ban. If someone wants someone else to stop doing something, they can make an offer at any time (although, late into the future, as decision-makers spread out from Earth it might become more difficult to negotiate these trades). + +_Drawbacks_ : this could create perverse incentives: some people might do things that others find objectionable so that they’ll be paid to stop. This might result in even more objectionable behavior that would have happened without that incentive. And if a single actor wants to do something a large group objects to, the group may struggle to coordinate payment because of free-rider problems. + +**Voting on bans**. We could prohibit activities that a supermajority of the population (say, 80%) votes to ban. + +_Strengths_ : avoids both of the drawbacks of the market solution—since there’s no transfer of resources from the people who want to ban to the people who would otherwise do the banned activity, there’s no collective action problem among the former to raise money to pay off the latter. And no one is incentivized to do objectionable things in hopes of a payout. + +_Drawbacks_ : a majority that very slightly disliked an activity could ban it even if it was very important to a small minority. Since there’s no transfer of resources, there’s no check to make sure the ban is sufficiently valuable to supporters to justify the cost to opponents. + +If we went with this system, we’d have to decide when and how often to vote on which activities to ban. + + * If we did a single early vote, we might fail to ban activities that were later widely considered bad. + + * Voters might not know of those activities at the time of the vote (e.g., because the activities are only possible due to unanticipated technology). Or voters might not be in reflective equilibrium at the time of the vote. For instance, perhaps they would later, on reflection, decide that they cared about the welfare of digital minds. + + * We might be able to circumvent these issues by adopting bans with broad, non-specific language and creating an AI arbiter to interpret the ban. But the arbiter may have to do nontrivial amounts of extrapolation and there might not be a single canonical way to do that extrapolation. + + * If we had a recurring vote or a single late vote, then voters might have more time to develop and carry out strategies to game the voting system before the vote(s) (see discussion [above](/supplements/space-governance-plan#early-decisions-vs-late-decisions)). + + + + +Currently, we think that at least the following activities should be banned: + + * Torture of conscious beings, including digital minds and non-human animals that are sufficiently likely to be capable of having strongly negatively-valenced experiences. + + * It might take non-trivial reflection by advanced ASI to figure out the precise bounds of this. + + * Unilaterally destroying Earth-originating civilization, or taking actions that bear a significant risk of destroying all of Earth-originating civilization. + + * Seizing other people’s property without permission. + + * Making and/or carrying out threats to destroy other people’s resources, in order to extort them. + + * Making and/or carrying out threats to create large amounts of disvalue by someone else’s lights in order to extort resources from them. + + + + +**Bans decided by treaty signatories**. Treaty signatories could negotiate bans before signing the treaty. This has many of the same strengths and drawbacks as bans decided by a single early vote. + +**Cancelling out resources**. We could allow people to forfeit control of their resources in exchange for removing someone else’s control of their resources at a k:1 ratio (for k > 1). + +_Strengths_ : as with the Market approach, by imposing a cost on the people who want the ban, it ensures that people don’t block valuable activities that they only slightly dislike. And unlike the Market approach, since there’s no compensation for the person carrying out the activity, there’s no perverse incentives to do stuff that others dislike in search of a payout. + +_Drawbacks_ : as with the Market approach, a large group seeking to prevent a particular use of resources may struggle to coordinate. + +## Reflection + +Under most of the proposals discussed in this appendix, individuals will need to make extremely high-stakes decisions, for example: + + * Should I sell galaxies allocated to me in an auction? Should I save them to pass onto my children? Should I donate them to a trust, and which trust should I donate to? Should I wait until I’ve thought more to decide, and how should I go about thinking about it? + + * Should I vote for a particular activity to be permanently banned throughout the universe? Should I vote for public resources to be allocated to a particular activity? + + + + +We want people to make wise decisions that reflect their values. + +Making these decisions well will likely require resolving open scientific and philosophical questions. The specific questions will vary with individual values, but the following areas are likely to be relevant to a wide range of people: + +**Improved understanding of morally significant concepts**. Many people consider consciousness and preferences to be morally significant—it’s wrong to inflict suffering on a conscious being and good to help people get what they want. But our current scientific and philosophical understanding of consciousness is insufficient to answer important questions like: which minds are conscious and which of their experiences have positive or negative valence—especially as our technological capabilities expand to allow creation of minds very different from human minds in substrate or structure. + +**Moral truths**. There might be objective and discoverable moral truths. + +**Acausal trade**. The reachable universe might be only a small part of what exists. There are regions of our universe outside our [light cone](https://en.wikipedia.org/wiki/Light_cone), and there may even be [other universes](https://en.wikipedia.org/wiki/Multiverse). Although we cannot causally interact with beings outside our light cone, our own choices may give us some evidence about how those beings will behave. On some decision theories (e.g., [evidential decision theory](https://en.wikipedia.org/wiki/Evidential_decision_theory)), this should affect what we do. For instance, we might want to “trade” with these beings, doing stuff that matters to them in exchange for them doing stuff that matters to us. If this argument works out, our acausal impact could plausibly dwarf our causal impact. + +Evidential Cooperation in Large Worlds (ECL) is one such mechanism. The multiverse contains many causally disconnected agents whose decisions may be correlated with ours: that is, when we observe ourselves making some decision, we should update about their choices too. So if we commit to pursuing multiverse-wide compromise values, this gives us evidence that huge numbers of other agents are doing the same. And this could be great—there might be substantial gains from trade if everyone pursues compromise values rather than their own. + +These are just a few examples of potential crucial considerations that we might want to consider before making important decisions. + +### Deciding how to reflect + +People will also need to think carefully about how they want to make these decisions, including how they want to use new technologies: + +**Deciding _when_ to learn information**. While it’s probably important to learn the answers to questions like those above before making object-level decisions about how to use space resources, there might be costs to learning information too early. For example, there might be positive-sum trades that can only be made _before_ the parties to the deal know some empirical results. One example of this dynamic is wagers (discussed above). More speculatively, if we end up wanting to do ECL, it might be [necessary to commit to doing ECL](https://lukasfinnveden.substack.com/p/implications-of-ecl) before we become confident whether other agents will do ECL—our decision only gives evidence about everyone else’s decision if we’re not sure what others will decide. + +**Engaging with ASI-generated arguments.** ASI will likely be able to research the philosophical and scientific questions above more quickly and rigorously than humans, giving us access to rapid intellectual progress on how to best live out our values. + +But ASI could also direct enormous intellectual effort toward crafting extremely persuasive arguments. And it’s not clear that the most compelling arguments will necessarily support true positions or positions we would endorse holding. If ASI is optimizing for maximally compelling arguments, then we might end up believing whatever ideology the ASI was subtly biased toward, or perhaps whatever ideologies are easiest to argue for compellingly. + +One safeguard: ASI could refuse to generate superhumanly compelling arguments for claims it knows to be empirically false, in the same way that AIs today refuse to assist with creating biological weapons. But for some questions we care about, there might not be an empirical fact of the matter. For these cases, we need to think about how we want to be persuaded—what counts as legitimate moral suasion and what counts as illicit manipulation. It's not clear where the line is—is vividly describing the suffering of a particular being emotionally manipulative, or just helping you understand the true stakes of your decisions? And the line might be different for different people. + +**Delegating to aligned ASI**. Another option is to delegate reflection to an aligned ASI. We will probably be able to design ASI systems that are much better suited to careful reflection than humans. These systems would be smarter, capable of understanding and integrating more complicated arguments or empirical results than any human. + +Moreover, for many people, reflecting in a very careful way might be less pleasant than other ways of spending their time. An ideal reflection process might involve carefully controlling your memetic environment (if it’s possible for ASI to develop illegitimate, yet highly compelling arguments or ideologies). It might require modifying your mind to make yourself smarter, more reflective, or more free of biases that you don’t endorse. In general, it’s not clear that “optimal lifestyle for human flourishing” and “optimal lifestyle for figuring out the most idealized form of one’s values” will be the same for most people. Offloading reflection onto an ASI system that is better suited for the process might let those people get the best of both worlds. + +Digital reflection has other advantages difficult to replicate on a biological substrate, including running many parallel reflection processes, reverting to earlier checkpoints, and potentially greater transparency about the trajectory that one’s values took over time. (This last is particularly important in versions of this proposal where many people’s preferences are aggregated together, since people might have incentives to strategically self-modify or misrepresent their preferences. Carrying out a digital reflection with a legible trace could let people verify that they hadn’t done so). + +**Delegating to future generations**. Another option is to delegate decisions to one’s descendants or to future generations more broadly. If so, how should one raise their children? Parents will plausibly have much greater capacity to shape their children’s values—how should they exercise this power? + +We’re not sure what the right answers are, and indeed the right answers may vary from person to person based on their values, but we think these are important decisions and want to help people make them carefully. + +### The role of governance + +It’s important for the governance structure to facilitate reflection. + + * **Timing of decisions**. Delaying decisions about how to use space resources gives decision-makers more time to think. Although ASI labor may speed up the rate of scientific and philosophical research, many people will likely need years or decades to integrate those findings into their thinking. Additionally, reaching these decisions may require substantial introspection into what the decision-maker truly values, which might also be time-consuming. + + * For governance structures where the timing of an important decision is scheduled in advance—e.g., a vote or a preference aggregation—this consideration favors delaying that decision (although there are considerations in the other direction, discussed [above](/supplements/space-governance-plan#early-decisions-vs-late-decisions)). + + * An important advantage of direct endowment governance structures is that they allow individuals to reflect for as long as needed before making object-level decisions. If individuals were still prone to hasty decisions, we could—somewhat paternalistically—require them to wait a few decades before selling or committing their resources to a trust. + + * **Subsidizing reflection.** Governments could offer all voters or resource recipients access to ASI labor to help them decide how to use those resources. Governments could also directly fund research into open scientific and philosophical questions likely to bear on decisions many people will face. + + * **Avoiding perverse incentives**. Governance structures could incentivize people to warp their reflective process or interfere in others’ reflective processes. For example: + + * Voting systems that offer voting rights to future generations could incentivize strategic reproduction and indoctrination. + + * Voting systems (with lie detection) or direct preference aggregation systems could incentivize people to modify their own preferences strategically to gain an advantage (just as traditional voting systems incentivize strategic voting). + + + + +The answers to some of the questions above could have implications for which governance procedures are best. One important tradeoff among the governance proposals is between enabling coordination on (moral) public goods and enabling individuals to pursue their own values. ASI-powered research into these questions could shed light on: + + * Whether there are important public goods that we might want to coordinate on. For example, perhaps acausal trade has large fixed costs—e.g., running lots of simulations of alien civilizations to figure out the distribution of values that trade partners have—and it would be wise to ensure that there are governance mechanisms that let us pool our resources to do that. + + * The degree to which humans’ moral views converge after endorsed reflection. If people converge substantially, this favors governance proposals with coordination mechanisms that let them pool resources to pursue shared values. If there’s a lot of divergence, this favors governance proposals that allow each person to pursue their own values. + + + + +# Takes + +## Mia’s takes + + * My actual take is that we should wait for aligned ASI and ask it to design a governance proposal that meets the desiderata above (and perhaps more that I haven’t thought of). + + * It seems very important that everyone making irreversible high-stakes decisions carefully reflect on what they care about with the help of aligned ASI advisors. To make this more likely, governments should probably fund lots of public-domain ASI research into scientific and philosophical questions that are likely to be of broad interest (like the questions above). They should also subsidize ASI compute used by people thinking about these decisions. + + * Likewise, it seems important that decisions about the use of resources beyond the solar system are made for altruistic or moral reasons, not selfish ones. + + * We should guarantee that all humans alive in 2040 and their descendants get enough to support them indefinitely at a very comfortable standard of living (i.e., at least the material standard of living of a multi-millionaire in 2026). This will hopefully lower the personal stakes of the decision about what to do with space and encourage people to be more high-minded when they vote. + + * It is also great on common-sense ethical grounds—on a cosmic scale, it’s relatively cheap to satisfy people’s preferences, and so we should do that. + + * I think coordination capacity—where we can collectively use resources to provide (moral) public goods—is potentially really important and it's not clear that we'll be able to do that with voluntary contracts. Unless we're confident that this is possible, it seems good to maintain capacity to coordinate, i.e., have a government (or a few governments) that control substantial amounts of resources (or which are empowered to levy taxes if they need to). + + * On the other hand, I also want to enable people with unusual value systems to spend some resources pursuing their values. + + * For these reasons, my top choice is direct preference aggregation if it’s technically feasible and not so unpleasant that many people opt out, perhaps combined with automated reflection (like what’s described here). + + * If that doesn’t work, my next favorite option is a hackier solution that ensures that some resources go to individuals and some resources are decided collectively. Here’s one way of approaching this. + + * At least 5% of resources are guaranteed to be directly granted to humans alive in 2040. + + * By default, the remaining 95% of resources will also be directly granted to humans alive today, unless a supermajority votes that it should instead be spent on public goods. + + * If highly reliable lie detection is available, governments could use that to check for strategic voting and disqualify people who were sufficiently likely to be executing these strategies. + + * Governments could also have ASI systems refuse to assist with attempts to coordinate to vote strategically (in the same way that ASI systems will presumably continue to refuse to assist with other harmful requests—e.g., making bioweapons). + + * We will also vote on bans and on the process to use to amend the list of bans (e.g., should we have recurring votes?). There should be a high threshold for bans, perhaps 90%. + + + + +## Thomas’s takes + + * I agree that we should wait for aligned ASI and ask it to design a good governance proposal. The most important things for us to do now are (1) avoid losing control of the future, e.g. via misaligned AI takeover or a human dictator getting control, (2) distribute the power reasonably, and (3) try to do philosophy well. + + * I think there’s a good chance that in practice, the best path forward will be to lock in some governance structure over the future, because deliberation within a fixed structure (like the ones discussed above) will go better than in a looser structure (like today’s government) where, e.g., more power will accrue to people who are, for example, more politically well connected. + + * My tentative favorite concrete guess for what to do above is “direct endowment” with some form of amendment process with a very high threshold (e.g. 90% approval). I think it’s very similar to voting based proposals, where you “vote” by making trades with other people about what to do with your resources. The public goods concern is the strongest counterargument, and if there is a voting system discovered that addresses this concern without adding additional, even more dire consequences, then I would change to wanting that. + + * All voting systems that we know of are screwed up, and I would not want them to be used for something as high stakes as deciding about what to do in the future. But I’m pretty optimistic about this changing, for two reasons: + + * I think there’s a good chance that ASIs (and humans who think about it for longer) will come up with voting systems that incentivize honest reporting of preferences, and choose a good policy that aggregates the votes. + + * I think that we might be able to do preference elicitation along the lines of “direct preference aggregation” above. Once you have honest preferences, the problem of designing a good voting system seems much easier. + + * I care a lot about power in the future going to virtuous people. I think that it would be really bad if a huge fraction of the power in the future goes to selfish people or terminally evil people (e.g. people who want to cause massive amounts of suffering). Of course, different people have different moral views that are reasonable, and I don’t think it’s fair, for example, to claim that everyone should follow my moral views. But I do think that it’s reasonable to exclude people with sufficiently bad moral views, as agreed on from the perspective of the rest of society, from having a large amount of power over the future, for example, to exclude sociopaths from many of the preference aggregation systems above. + + * I think that any attempt to have “normal governance” continue is not going to work, because our current systems will be totally exploited by actors using superintelligence to gain power. I think that we’re going to need to make massive changes to our system of government really quickly in order to maintain a reasonable ongoing distribution of power. + + + + +1. + +Other important questions in post-ASI governance that we don’t address, or address only cursorily: + + * Post-ASI governance on Earth or elsewhere in the solar system. + + * [Meme wars](https://www.forethought.org/research/preparing-for-the-intelligence-explosion?utm_source=chatgpt.com#epistemic-disruption), if ASI are able to construct extremely compelling arguments, and members of different ideologies try to spread their views using these arguments. + + * Rights for digital minds. + + * Managing risk in [highly offense-dominant worlds](https://nickbostrom.com/papers/vulnerable.pdf), where a small group of people are able to cause huge amounts of destruction. + + + + +2. + +The Outer Space Treaty forbids nations and (according to some interpretations) individuals from claiming celestial bodies. In this document, we’re assuming, as in the Plan A scenario, that the Outer Space Treaty will be renegotiated once extrasolar colonization becomes technologically feasible. + +3. + +[Toby Newberry](https://www.globalprioritiesinstitute.org/wp-content/uploads/Toby-Newberry_How-many-lives-does-the-future-hold.pdf) estimates that the affectable universe could support 10^35 biological humans or 10^44 digital minds at a time. For comparison, the number of humans who have ever lived is 10^11. + +4. + +One possible exception is people with strong preferences for [positional goods](https://en.wikipedia.org/wiki/Positional_good), e.g., people who care a lot about being richer than others, rather than their absolute level of wealth. + +5. + +For example, it might be possible to move our universe from a false vacuum state to a true vacuum state, which could [destroy all matter](https://en.wikipedia.org/wiki/False_vacuum). + +6. + +See the paper [Burning the Cosmic Commons](https://mason.gmu.edu/~rhanson/filluniv.pdf) for more on this threat model. + +7. + +Some examples of moral public goods that might be quite important to fund: + + * Broadly appealing moral projects, such as building societies of flourishing biological humans, uploaded humans, or other digital minds; providing resources to future generations; or discovering and acting in accordance with objective moral truths, if any exist. + + * Trading with people who intend to do something with side effects that others consider harmful to get them to modify their plans. Obviously, it’s important to approach this carefully to avoid creating perverse incentives where people do some activity that’s broadly considered harmful in search of a payout (see [below](/supplements/space-governance-plan#mitigating-downside-risks) for more discussion of reducing downside risk). + + + + +8. + +Ideally, we’d also like to weigh the preferences of other stakeholders like future generations and digital minds, although we’re uncertain about whether it will be feasible to do this without opening up the political system to exploitation (see [below](/supplements/space-governance-plan#who-should-participate-in-the-process)). + +9. + +That is, we should either decide how future governance decisions will be made or at least pick a meta-level process for choosing how that decision will be made (e.g., the voting system should be chosen by an AI arbiter verifiably aligned to choose based on a particular set of principles) + +10. + +Via mutual defence coalitions, mutual insurance, or law-enforcing equilibria emerging from multi-turn interactions. + +11. + +Barzel (1997), _[Economic Analysis of Property Rights](https://www.cambridge.org/core/books/economic-analysis-of-property-rights/6D5E9A3AA67284FD9A12379CA3028D50)_ ; Anderson & Hill (1990), [“The Race for Property Rights”](https://www.journals.uchicago.edu/doi/abs/10.1086/467203) + +12. + +One way to do this might be to require that all probes sent to colonize other star systems carry a nightwatchman ASI to prevent the colony from sending out probes to seize space resources that belong to others. + +13. + +The point is that each person’s eventual wealth only depends on the value of whatever they inherit at equilibrium prices, and not on whether they originally inherited titles to future resources directly, or received a cash transfer of equal value at equilibrium prices. If resources R are divided equally among N people and traded at competitive prices p, each person's wealth is p · R/N. Under the auction alternative, total revenue is p · R, and equal redistribution gives each person p · R/N. [This requires](https://econpapers.repec.org/paper/mitworpap/115.htm): competitive (price-taking) behaviour, that the auction clears at competitive equilibrium prices, markets are [complete](https://en.wikipedia.org/wiki/Complete_market) and have sufficient numbers of buyers and sellers, and people have convex preferences. The equivalence between auctioning and direct endowment could break down: especially if actors are large enough to coordinate to influence prices, or if resources are indivisible. + +14. + +It’s true that in principle, under a direct endowment proposal, everyone could decide to voluntarily pool their galaxies into a trust and allocate resources by vote, but free-rider problems make this less likely in practice. + +15. + +Obviously, versions where voters from a few countries get to weigh in are much better than versions where, say, only AI company CEOs and national leaders get to weigh in. But we think all of these have some of the same problems to different degrees. + +16. + +Suppose there are a thousand actors, each with equal resources. One actor wants to use their resources in a way which is viewed as harmful by all the other actors. Without enforceable prohibitions, each of the majority would need to decide individually how much to spend bribing the unpopular actor to stop. No one of the 999 actors cares so much about the harms that they are willing to pay one-half of their wealth to individually prevent them; but all of the 999 are willing to pay 0.05% of their wealth to prevent them, and that would raise a large enough bribe for the unpopular actor to agree to switch to some other use of his resources. In other words, preventing the use of resources which enough people care strongly about for moral reasons is a moral public good which requires taxation to fund efficiently. + +17. + +For example: + + * What exactly are negatively-valenced experiences? Probably our current understanding of consciousness and suffering is rather muddled, and people might create minds that have some of the features of classic examples of negatively valenced experiences but not others. + + * How strongly negative does an experience have to be to fall under this ban? + + * Are there any mitigating factors, e.g., the experience being necessary for a strongly positive life, the being in question consenting? + + + + +18. + +This mechanism was originally introduced under the name “multiverse-wide superrationality” in this [paper](https://longtermrisk.org/multiverse-wide-cooperation-via-correlated-decision-making/). + +19. + +Some more crucial considerations are discussed [here](https://www.lesswrong.com/posts/EyvJvYEFzDv5kGoiG/clarifying-wisdom-foundational-topics-for-aligned-ais-to). + +20. + +For instance, the question of "is a particular digital mind capable of having morally relevant experiences?" hinges on both empirical properties of the digital mind but also questions about what properties are morally relevant. + +21. + +Probably there should be some limit on the number of descendants who could be supported by these resources because otherwise we presumably will eventually hit Malthusian limits. diff --git a/sources/ai2040-pages/takeoff-supplement.md b/sources/ai2040-pages/takeoff-supplement.md new file mode 100644 index 0000000..07b6375 --- /dev/null +++ b/sources/ai2040-pages/takeoff-supplement.md @@ -0,0 +1,229 @@ +--- +url: https://ai-2040.com/supplements/takeoff-supplement +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Takeoff Supplement + +### Brendan Halstead + +This supplement explains how we forecast AI capabilities in various parts of Plan A. We modified the AI Futures Model to simulate four situations: + + 1. the default trajectory, if there had been no deal to slow down capabilities progress, + + 2. the Plan A consortium trajectory, which scales to TED-AI in a controlled manner, then halts to do alignment research, + + 3. a [covert project](/supplements/covert-ai-projects) defecting from the deal and trying to build AGI in secret, and + + 4. the trajectory after [deal dissolution](/supplements/deal-decline), where most of the post-deal compute is destroyed and the leading project resumes racing. + + + + +The rest of this document is structured accordingly. Section 1 covers the differences between the Plan A model and the December 2025 AI Futures Model: a major change to training run modeling, our updated compute forecasts, and the changes to the default parameters. In sections 2 through 4, we explain the various extensions added on to simulate the three additional situations. If you haven’t read the AI Futures Model documentation, we recommend reading it [here](https://www.aifuturesmodel.com/) first. + +# Modifications to the core model + +Here we describe the three ways this model differs from the December 2025 AI Futures Model. These apply across all trajectories simulated. + +## Explicit modeling of the frontier training run + +Like most other AI takeoff models, the AI Futures Model (December 2025 version) defined _effective training compute_ E(t)E(t)E(t) at a given time as the cumulative training compute C(t)C(t)C(t) of the frontier model (in FLOP) times the software efficiency level S(t)S(t)S(t) at that time (in present-day-FLOP per FLOP): + +E(t)=C(t)⋅S(t).E(t) = C(t) \cdot S(t).E(t)=C(t)⋅S(t). + +In reality, training runs take time, and the formulation above assumes that today’s software efficiency applies retroactively to the training run for the model that’s already in use. One consequence is [slightly overestimating](https://www.forethought.org/research/will-the-need-to-retrain-ai-models) how fast takeoff can proceed. But more importantly for Plan A, this model gives us no good way of dealing with situations where the _training system performance_ (in FLOP/yr) of an AI project suddenly becomes very small compared to the system that was used to train the best model available to that project (as is the case for covert projects and after deal dissolution). + +In the Plan A version, we instead model effective training compute as increasing continuously as part of one long training run, at a rate determined by the _training system performance_ T(t)T(t)T(t) and the software efficiency of the known training methods: + +E′(t)=T(t)⋅S(t).E^\prime(t) = T(t) \cdot S(t).E′(t)=T(t)⋅S(t). + +We call this rate the _effective training system performance_ and abbreviate it as ETSP. + +When all inputs grow exponentially, ETSP and effective compute also grow exponentially at the same rate, so this model behaves like the simpler one. All other equations are unchanged from the December 2025 AI Futures Model. + +## Changes to the compute and labor time series + +The AI Futures Model depends on exogenously-supplied forecasts of compute and labor. The Plan A version uses two sets of time series: one representing the default world if Plan A had not been implemented, and one that branches off from the default when the deal comes into effect. The modeling behind both of these is discussed in our [Compute Supplement](https://www.ai-2040.com/supplements/compute-supplement). In summary, + + 1. The default time series before 2029 has been updated to be slightly more bullish compared to that in the AI Futures Model + + 2. The Plan A time series from 2030 to 2035 reflects the controlled-but-massive compute buildout enabled by robotics, then slows down from 2035 to 2040. + + + + +One limitation is that the human labor projections are not very realistic; we did not put much effort into these because human labor plays a very small role in the model dynamics after AI R&D is automated. + +## Changes to the default parameters + +The default parameters governing timelines and takeoff in the scenario were set to reflect scenario first-author Thomas Larsen’s views. However, since capabilities forecasting was not the main focus of AI 2040, they do not reflect as much consideration as do the values used in the AI Futures Model, and were sometimes adjusted ad-hoc as the scenario narrative and modeling were being written concurrently. Nevertheless, they deviate only slightly from Daniel Kokotajlo’s median values as of April 2026, and all authors find the default trajectory entirely plausible. Below we include a table describing each parameter that was changed from Daniel’s median. + +**Parameter**| **Change**| **Rationale** +---|---|--- +Gap years| 1.1 2025-effective-flop-growth| Set to achieve the AC milestone at Jan 2030. +Time horizon required for AC (before gap)| 10 work years +Median-to-top taste multiplier| 5.28x| Set to reach ASI roughly one year after AC, while leaving ~3 OOMs of effective compute between AC and TED-AI. +AI research taste slope| 3.5 SDs per 2025-effective-flop-growth +Median-to-top-human jumps above SAR needed to reach TED-AI| 1.5 + +# Modeling the consortium trajectory + +Unlike all other actors we model, the consortium is _not_ racing to build better AIs as fast as they possibly can. The specific strategies are discussed in more detail in the Capabilities Scaling Strategy supplement, but at a high level, the strategy is to scale smoothly to Top-Expert-Dominating AI (TED-AI), then pause to do alignment research before eventually handing over control to AIs. + +Here, we consider strategies that vary along two axes: + + 1. How far into the deal consortium aims to reach Top-Expert-Dominating AI (TED-AI), and + + 2. How much of a safety tax the consortium aims to be paying when TED-AI is reached. + + + + +The safety tax is the difference between the capability level of the actual consortium AI and the capability level of the best AI the consortium _could have trained_ , had they used all their training compute and the best-known algorithms for capabilities training. One example of “paying a safety tax” would be discovering neuralese training, finding that it would be much more efficient, but deciding not to use it to make control easier. We expect that in reality the consortium would want to pay some safety tax. + +Plan A spends more time at higher capability levels than a uniformly slowed intelligence explosionPlan A — capability-shaped upliftsAutomated Coder3× realized upliftTop-Expert-Dominating AI~40× realized uplift2029: two ways to slow down2040: handoffearly 2035mid 2030late 2034Plan Aslowdown deal — controlled takeoffUniform slowdownuniformly slower uncontrolled intelligence explosion20262028203020322034203620382040 + +The rough shape of the consortium’s capabilities trajectory: a brief halt, then smooth scaling, then a long near-pause before handoff. + +The consortium trajectory is simulated according to slightly different rules during each of the following stages. + +**Stage 0: AI Training Pause** + +This phase starts some number of years before AC would have happened on the default path. (In the scenario, about 9 months before). Though effective compute is forced to stay flat (training is paused!), training system performance can still increase as the leading company within the consortium ends up with more compute at the end of the period. + +AI R&D is _also_ essentially paused during Phase 0. However, we model the software efficiency of the leading actor within the consortium as growing by a set number of OOMs during this period. This is because of Total Research Transparency: by the time scaling resumes at the end of Phase 0, essentially company IP is public. We expect the companies to have made non-overlapping software discoveries prior to the deal; hence pooling the intellectual property will increase software efficiency at the frontier. + +**[Stage 1: Scale to max-controllable AI](https://www.ai-2040.com/supplements/capability-scaling-strategy#stage-1-scale-to-max-controllable-ai)** + +During this phase, effective compute increases with effective training system performance via the usual E′=T⋅SE^\prime = T \cdot SE′=T⋅S. TSP grows via the exogenously-supplied Plan A time series, and software efficiency growth is carefully controlled in order to reach TED-AI at a particular time. We assume a mandatory minimum amount of software progress is required per OOM of TSP increase. + +To keep track of the safety tax paid at each point, we actually simulate two different trajectories: + + * One represents the legal project’s _realizable_ capabilities. This trajectory aims to maintain the (constant) software progress rate that achieves TED-AI _plus_ the desired safety tax amount at the target time. + + * One represents the legal project’s _realized_ capabilities. This trajectory aims to maintain the (constant) software progress rate that reaches TED-AI at the same time that _realizable_ capabilities reach TED-AI+tax. + + + + +At each timestep the trajectories moderate their research effort from the theoretical maximum, choosing instead the value which brings them closest to the pre-selected software progress path. + +Upper and lower limits imposed on software progress + +If the chosen scaling duration and takeoff parameters demand faster software progress than would be naturally possible at a given point, the research effort may not be moderated at all. Often, once AI capability level increases enough, the consortium is able to catch up to the desired software path. Occasionally, the consortium may not be capable of reaching TED-AI by the desired time, e.g. in worlds where reaching TED-AI from the pause level would naturally have taken ten years but where the consortium attempts to reach it in just one. + +Conversely, a sufficiently long scaling period or a sufficiently narrow effective compute gap from the pause level to TED-AI can mean that hardware scaling alone (plus the mandatory minimum software growth rate) would reach TED-AI before the target year. In this case, the consortium simply reaches TED-AI early. + +**[Stage 2: Slow down nearly to a halt, then eventually hand off to AIs](https://www.ai-2040.com/supplements/capability-scaling-strategy#stage-2-slow-down-nearly-to-a-halt-then-eventually-hand-off-to-ais)** + +For simplicity, we model this stage as a literal halt for the consortium’s _realized_ capabilities. In reality, we expect them to slow down more smoothly, then continue increasing very slowly as control techniques improve and the threshold for “max-controllable AI” inches upward. + +The consortium’s _realizable_ capabilities continue increasing however. TSP growth slows down but continues. As for software efficiency growth, we specify a constant rate of software progress in OOMs per year to represent the capabilities externalities incurred by the huge amounts of alignment research that happen during stage 2. (As an example, dramatic advances in mechanistic interpretability, while promising for alignment, could plausibly also enable much more efficient training methods.) + +Since the _realized_ capabilities effectively halt while the _realizable_ capabilities continue growing, the safety tax paid increases dramatically during this period. + +In reality, the consortium will choose to hand off to AIs at some point (in our scenario, it happens roughly five years into Stage 2). We do not currently model handoff explicitly however, because none of our analyses depend significantly on the dynamics after handoff. This is why the consortium capabilities line remains flat indefinitely in the explorer, rather than eventually jumping up as we would expect after handoff. + +# Modeling covert projects + +At the moment that AI training is paused (Stage 0), the covert project is assumed to start with the following resources: + + * A total R&D compute budget equal to some fraction of the leading company’s R&D compute. We assume the covert project allocates between training, experiments, and internal inference in the same proportions as the leading lab did. + + * An engineering labor force equal to that same fraction of the leading lab’s labor force, + + * AI capabilities, software efficiency, and research stock matching those of the pre-deal leading company. + + + + +We also assume that the covert project’s supply of compute and labor stays constant over time. Finally, we assume that the covert project cannot begin training or R&D for some period during which covert datacenters are being constructed. + +With only these factors considered, our modeling indicates that covert projects would probably progress quite slowly. [In the median case, each OOM of R&D compute reduction (compared to the leading lab at the start of the deal) would cause the covert project to take over 6x as long to reach TED-AI than would the leading lab.](https://www.lesswrong.com/posts/7jcPg79p3kD5ir3CL/how-much-slower-does-takeoff-go-with-10-less-compute) + +However, we also assume that the covert project is uplifted to some degree by consortium AI progress. The primary factor, and the only one we model explicitly, is the leakage of software improvements from the consortium to the covert project. To account for this, we include a parameter for the fraction of the consortium’s software progress rate that gets added on to the covert project’s indigenous software progress rate. We conservatively assume covert projects will make full use of any algorithms discovered by the consortium, even those deemed potentially dangerous, so we take the _realizable_ software efficiency growth rate as the quantity to steal from rather than the _realized_ software efficiency growth rate. + +Due to software proliferation, the consortium’s scaling strategy is extremely important to the covert project’s ability to catch up. The less software progress the consortium needs to make to reach TED-AI at the desired time, the less the covert project is uplifted. As a result, accepting a _longer_ scaling duration that requires less software progress _increases_ the consortium’s ETSP lead at TED-AI, and typically also _increases_ the time between the consortium reaching TED-AI and the covert project reaching TED-AI. + +# Modeling deal dissolution + +The modeling of [deal dissolution](/supplements/deal-decline) is relatively simple. + + * The effectiveness of [mutually assured compute destruction](/supplements/deal-decline#mutually-assured-compute-destruction-destroying-compute-fabs-and-robots) is represented by the fraction of world compute available to the new leading project at the time of dissolution. + + * We also include a parameter for the growth rate of R&D compute after the deal dissolves, representing the output of any non-destroyed or newly-constructed fabs. + + * As with covert projects, we technically assume that human R&D labor varies in proportion to compute. Though this might be unrealistic, the effect of human labor is essentially negligible after AC. + + * The post-dissolution leading project starts with the frontier model weights (from the _realized_ consortium trajectory) and all discovered software improvements (from the _realizable_ consortium trajectory) from before the deal dissolved. + + + + +We assume that the leading project after a deal dissolves proceeds at full speed, as in the default trajectory. + +1. + +Specifically, they should scale to something like “max-controllable AI”, but in AI 2040 we operationalize this as TED-AI. + +2. + +Technically in the scenario, the pause lasts from March 2029 to September 2029, and “pilot” experiments are able to proceed during the last part of the year before the first major training run is approved in January 2030. + +3. + +This assumption may not represent the optimal strategy for a covert project. We did some rough experiments investigating various compute allocation strategies in an earlier version of this model, and found that they made a negligible difference to the speed of the covert project trajectories (except in extreme cases — for example, allocating zero compute to training will halt effective compute growth entirely. Similarly, allocating zero compute to experiments would halt software efficiency growth entirely.) + +4. + +This, too, is a somewhat unprincipled assumption (compared to, e.g., carrying out a dedicated estimate of covert R&D labor and including that as a separate parameter.) However, the precise reduction in the engineering headcount matters very little when engineering is nearly fully automated. We are fairly confident that improving upon this simplification does not affect the results in any significant way, and have checked this empirically at various points. Due to time constraints, we do not currently include detailed results here. + +5. + +This is one of the ways our model is generous to covert projects; it is similar to assuming that whoever is operating the covert project has access to the frontier model weights, datasets, and full stock of algorithmic knowledge from the leading lab at the time of the deal’s signing. As an improvement, we could add a model parameter representing the number of “months behind” the covert project starts at. Unfortunately, we do not include this as of the Plan A release due to time constraints. + +6. + +This is slightly generous but on the whole seems likely to be accurate. The most catastrophic way this could fail is if some amount of the legal compute produced during Plan A could be smuggled to a covert project (or used illegally from within the R&D clusters). This would be catastrophic because world compute increases by several OOMs during Plan A. A milder failure mode would involve the covert project making use of the growing stock of consumer hardware during the deal (as makeshift AI R&D compute). Finally, a covert project could attempt to build and operate its own independent, unmonitored semiconductor supply chain during the deal. Our strategy for preventing all of these things from happening is described in the [Verification Plan](https://www.ai-2040.com/supplements/verification-plan). + +7. + +We do permit them to accumulate stolen software during the construction period, described below. + +8. + +We say “over” because the stated result compares the covert project to a trajectory whose compute level was _frozen_ at the branch point, where in reality the leading lab benefits from growing R&D compute during takeoff. + +9. + +Other factors include distillation and weight theft. In this model, we assume these can be prevented. diff --git a/sources/ai2040-pages/transparency-plan.md b/sources/ai2040-pages/transparency-plan.md new file mode 100644 index 0000000..bdd0eb7 --- /dev/null +++ b/sources/ai2040-pages/transparency-plan.md @@ -0,0 +1,332 @@ +--- +url: https://ai-2040.com/supplements/transparency-plan +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +# Transparency Plan + +### Thomas Larsen + +AGI projects face key decisions around _transparency_ : who gets to see algorithmic secrets, who gets real time access to the datacenters sufficient to be confident in what they are doing, who gets to be aware of ongoing experiments and training runs, and with what latency. This document outlines our current thinking on transparency in the context of Plan A. + +First, we discuss baseline transparency desiderata—aspects of transparency proposals that are core to Plan A. A fundamental component of Plan A is a trustless US/China deal to ensure that AI development is safe, which necessitates a minimal amount of US/China transparency, sufficient to verify that the other is following the agreement. ([more](/supplements/transparency-plan#baseline-transparency-desiderata)). + +Second, we give an overview of possible transparency regimes which can satisfy these desiderata. ([more](/supplements/transparency-plan#four-transparency-proposals)) + +Third, we give an in depth explanation of our favorite proposal, “total research transparency”. In this proposal, almost all algorithmic secrets, experiments, and training runs would be visible to the public immediately. ([more](/supplements/transparency-plan#total-research-transparency)). + +Finally, we give an analysis of the tradeoffs between the different proposals outlined, and the reasons why transparency is beneficial. The main reason we recommend Total Research Transparency is because it would improve government and corporate decisionmaking during Plan A; which is our current most concerning failure mode. The second-most important reason is that it makes it much harder for corporations and governments to abuse their power. ([more](/supplements/transparency-plan#when-is-each-proposal-desirable)). + +# Baseline Transparency Desiderata + +Any transparency proposal in Plan A should be consistent with the following desiderata: + + 1. **Trustless Verification.** The US and China need to be able to verify slowdown agreements without relying on trust. This means that specifically the US and Chinese governments need visibility into how compute is being used. + + 2. **Model weights security robust against nation-state adversaries.** We’d like to prevent model weights from leaking because if the model weights were to leak it would (1) greatly accelerate covert projects and (2) increase misuse risk. Neither of these are necessarily fatal, but we’d very much like to avoid them. Once we get to AIs that are as or more generally capable than top human experts, model weight security will probably be necessary. + + 3. **Ongoing consumer access to frontier AI.** Direct access to the AI models is perhaps the most important form of transparency. Seeing the results of evaluations and high level descriptions of model behavior are much less informative than directly using the model. + + + + +# Four Transparency Proposals + +We have four viable candidates for transparency, each of which meet the above desiderata. Proposals #1-3 are further consistent with Plan A, while Proposal #4 only works for Plan S. + +**Proposal #1: Total Research Transparency.** Nearly all AI research is fully transparent to the public, including AI algorithms, code, and documentation, for all frontier AI projects. AI model weights, significant fractions of the training data, and a small amount of other sensitive information is prevented from leaving the datacenters. + +**Proposal #2: Filtered Transparency.** Nearly all AI research is done within a physical security boundary containing both the researchers and the compute. Within the security boundary, US and Chinese auditors have sufficient access to verify deals. Three things are allowed to leave the datacenter: (1) (partially redacted) reports approved by both sides about what is going on inside, (2) model weights that get sent to the inference datacenters, and (3) researchers themselves. One can vary the level of redaction: with minimal redaction, this proposal approaches Proposal #1, and with maximal redaction, this proposal is similar to Proposal #3. + +**Proposal #3: Algorithmic Security.** It may be necessary to prevent the exfiltration of algorithmic secrets. This requires tightening the security boundary above: if people are regularly entering and leaving the AI datacenters, then they will almost certainly report back most of the algorithms to their governments. Similarly, filtered reports provide a channel for steganographically encoded versions of the model weights to be exfiltrated. Under this proposal, all frontier AI research is done in a secure datacenter with researchers and auditors located onsite and prevented from leaving or communicating any information out. Model weights are allowed to move from the R&D datacenters to the inference datacenters. This involves something like “handing off trust” to the auditors inside the datacenters, because so limited information is allowed to leave the datacenter that external actors cannot exert meaningful oversight internally. + +**Proposal #4: Shut it all down.** Specifically, ban compute-intensive AI R&D. Under this proposal, there is no large-scale AI R&D, so there’s no need for AI R&D transparency. This could be enforced through compute declarations and physical inspections (but doesn’t require sharing digital access to the clusters, i.e. algorithms). This is similar to the MIRI proposal, and also what occurred for about 6 months during 2029 in the AI 2040 scenario. + +This table summarizes the rough tradeoffs between them: + +| **Proposal #1: Total Research Transparency**| **Proposal #2: Filtered Transparency**| **Proposal #3: Algorithmic Security**| **Proposal #4: Shut it all down** +---|---|---|---|--- +**SL5 algorithmic secrets security**| No| Partial, depends on redactions| Maybe| N/A, no new algos discovered +**Allows new training runs**| Yes| Yes| Yes| No +**How informed are public discussions of safety cases?******| Very| Depends on redactions.| No| N/A, no new safety cases needed + +In Plan A, we currently recommend Total Research Transparency for many reasons, chief of which is that the likelihood of reasonable analyses of safety cases is much higher if external researchers and the general public are able to critically analyze them. The main drawback relative to Proposals 2&3 is that there is less algorithmic security, and so algorithms will probably leak to covert projects. However, we think that SL5 algorithmic secrets security is incredibly difficult. Therefore, our main mitigation for algorithms leaking is to prevent new algorithmic secrets from being developed in the first place. + +# Total Research Transparency + +In this section we’ll discuss a specific possible implementation of Total Research Transparency, which will correspond to the implementation done in Plan A. + +R&D DatacentersInference DatacentersAuditorsR&D ComputeOpaque InternalDatabase(Model weights,Activations,RL Data,Hyperparams)Internal Silocodeoutputs(screened to preventOpaque info from leaking)Variable latencyWeightsHigh RiskOutputsRequests for infoMisuse-InfohazardousTrajectoriesRequests for codeand other information1 MB/sec globalbandwidth capAll outputsAll inputsTransparentInternal Database(Code, AI COT,prompts, etc)InferenceComputeInference onlyrecomputationOutputTokensInputTokensVisible toPublic DatabaseIncludes all outputs from R&D Datacenters,outputs from models used for AI R&D, andall inputs to the R&D datacentersThe PublicAI & HumanAI ResearchersAIs used for AI R&Dhappen in specialinference datacenterswith additional transparencyrequirements (e.g. no ZDR)and additional controlscaffolding + +The arrows in the diagram show the flow of information between humans, R&D datacenters, and Inference datacenters. + +Total Research Transparency \- ai-2040.com + +Under this proposal, there are two types of datacenters: + + * **Inference Datacenters** only provide inference made by whitelisted models. These datacenters host consumer chatbots, almost all agent workloads, robot foundation model inference, etc. Arbitrary requests are allowed to be sent into the datacenter, but only verified AI model outputs are allowed to come out, so that the model weights are secure. + + * **R &D Datacenters** allow training runs and experiments—everything needed to produce a new set of model weights. They are “transparent datacenters,” as described above. These datacenters are much higher risk than inference datacenters, because they could be used to do an intelligence explosion. + + + + +Information from the AI R&D Datacenters is organized as follows: + + * An **Opaque Internal Database** contains all information that shouldn’t be public. This includes model weights, activations, some datasets, and some hyperparameters. No human (internal or external) is able to directly view the information in this database. + + * A **Transparent Internal Database** contains all other information, including all code, AI chains of thought, prompts, etc. However, it’s not directly accessible to the public or AI researchers who aren’t physically located within the security boundary of the AI R&D datacenter. + + * A **Transparent Public Database** is accessible to everyone. It is a record of all information going in or out of the R&D datacenter. It logs all activities of AI researchers interacting with the R&D clusters, as well as any additional queries by members of the public that wish to see some piece of information. For example, they might want to task an AI to look through the internal database and write a short report about the state of some particular field of research. + + + + +Anything that’s in the transparent internal database could end up in the transparent public database if anyone requests that information. However, what stops everything from leaving to the public database are upload limits that cap outgoing information to ~1 megabyte / second, as a layer of defense for the opaque information being exfiltrated (i.e., model weight exfiltration in particular). There is an ongoing auction for the outgoing bandwidth to allocate access, and all outgoing bandwidth is also copied to a freely accessible database on the internet. In practice, this means that almost all algorithms are immediately revealed to the public. + +The R&D datacenters are still owned and operated by tech companies, like in today’s world. In fact each R&D datacenter is really a campus of datacenters, inside which several tech companies‘ datacenters are located. The US and Chinese governments, with their respective auditors and monitoring devices, ensure that each tech company’s datacenters on the campus create logs of their activities that go to the internal opaque database and the internal transparent database. + +Anyone, including members of the public, can purchase compute inside the R&D datacenters and have the same level of access to frontier AI development as frontier AI company researchers. Members of the public can, for example, run experiments on intermediate model checkpoints, see loss curves during training, or ask an AI to run a desired analysis for them. + +There are three risk tiers of inference: + + * **Extreme Risk:** AIs used for AI R&D need to have direct access to R&D compute, and so they run on the R&D datacenters instead of the inference datacenters. This also allows additional security—these AIs are subject to carefully designed control protocols. + + * **High Risk:** AIs used for construction in dangerous domains (e.g. constructing robots, constructing GPUs or semiconductor manufacturing equipment) OR AIs used for research in dangerous domains (e.g. research relevant to engineered pandemics). These AIs live on the inference datacenters and must log the trajectories produced by AIs operating in high risk domains. + + * **Lower Risk:** AIs used for anything else. This includes most commercial use (e.g. using AIs as a chatbot), most corporate use (e.g. using AIs for medicine/law/architecture/etc). These AIs are allowed to run with zero-data-retention (ZDR), allowing for maximum privacy. Just like standard practice today, even AIs running with ZDR are still subject to refusals and anti-misuse classifiers. However, only AIs behind the frontier are allowed to run with true ZDR. This is discussed in more detail in our [verification supplement](https://www.ai-2040.com/supplements/verification-plan#zero-data-retention-zdr-inference). + + + + +High Risk and Lower Risk inference operates on the inference clusters, while Extreme risk inference occurs within the AI R&D clusters on a compute cluster that’s verified to be inference only. + +All of the high risk trajectories and some of the non-ZDR lower risk trajectories are physically transported (under guard!) between the inference datacenters and the R&D datacenters in order to give safety researchers access. + +# Filtered Transparency + +R&D DatacentersInference DatacentersAuditorsAI & HumanAI ResearchersGovernment-onlyreportsPublic DatabaseIncludes reports from inside the R&D datacentersVisible toThe PublicWeightsInferenceComputeInference onlyrecomputationOutputTokensInputTokens + +Filtered Transparency \- ai-2040.com + +In filtered transparency, AI researchers are physically located within a security boundary, and we attempt to prevent information from leaking across this boundary. Then there are also auditors inside the R&D datacenters that have sufficient access to verify and make judgement calls about safety. + +The three ways that information crosses out of the boundary are: + + 1. **Model Weights.** Just like in total research transparency, model weights need to pass from the R&D datacenters to the inference datacenters, and this channel needs to be high bandwidth because model weights are so large. + + 2. **Auditor Reports.** The auditors write up reports about progress inside the R&D datacenters, which include things like current capabilities, safety relevant information about the current architectures. Redacted versions of these reports could be made public. + + 3. **Humans leaving the security boundary.** Human researchers periodically enter and leave the security boundary. This is the main way that we expect algorithms to be exfiltrated from the datacenters. + + + + +There’s no restriction on information crossing into the security boundary. + +Like Total Research Transparency, there are R&D datacenters and inference datacenters. This allows for one of the most important forms of transparency: widespread access to frontier AI models. The main downside of filtered transparency relative to total research transparency is that it does not allow for fully open access and scientific scrutiny of the safety cases related to AI R&D. Instead we have to hope that the relatively small group of researchers inside the security barrier will be able to identify all the flawed assumptions of each safety case, make appropriate assessments, make sufficient research progress, etc. + +Algorithmic security is similar to an extreme form of filtered transparency. The central idea is that algorithmic progress is allowed to be made in the project, but doesn’t leak outside. This is extremely difficult because algorithmic progress is often memorizable. + +Therefore, the main difference is that for the “algorithmic security” proposal, it is necessary to prevent humans from crossing the security boundary. This enables us, in principle, to secure algorithmic insights. + +# When is each proposal desirable? + +The main axis for deciding between transparency proposals for R&D is whether we are more concerned about possible covert projects, or poor regulation. Increasing transparency will result in more algorithms leaking, boosting covert projects, but improving public oversight of AI projects and increasing the expected quality of AI regulation. + +More concerned aboutpoor regulationMore concerned aboutcovert projectsProposal #1:Radical TransparencyProposal #2:Filtered TransparencyProposal #3:Algorithmic SecurityTransparency proposals spectrum \- ai-2040.com + +Overall, we think that poor regulation is more concerning than covert projects, and so we recommend Total Research Transparency. + +Proposal #4 does not allow for the development of new AI models, i.e. it only works as part of Plan S as opposed to Plan A. The decision about Plan A vs Plan S will mostly be made on other factors, which are discussed more in the [Plan S branch](https://www.ai-2040.com/?choices=plan-s-root). + +## What fraction of algorithms will leak under each proposal? + +The optimal policy on this spectrum will depend on quantitative factors around how much algorithmic progress would leak under each proposal. This will depend on what fraction of algorithmic progress is of the type that would be transparent (e.g. architectural improvements) vs what fraction of algorithmic progress is large dataset improvements. Overall, we make the current (very rough) best guess estimates of the fraction that will leak under each proposal: + +**Proposal**| **Algo Fraction leaked** +---|--- +P1 Radical| ~67% +P2 Filtered| ~47% +P3 Algorithmic Security| ~14% + +You can see the reasoning behind this estimate [here](https://docs.google.com/spreadsheets/d/1zKOxldIPXrz6RypC8Zr50rTpF10yGz4w0UZhNxk78XE/edit?gid=0#gid=0). We would be excited about security experts (which we are not) making their own forecasts; especially forecasts that are more granular (e.g., that include over what time period the algorithms leak). + +## Total Research Transparency Upsides + +This section talks about the benefits of transparency for reducing loss of control risks and concentration of power risks. It’s basically an explanation of this diagram: + +Total ResearchTransparencyBroader technicaldebateMore alignmentresearchersBetter incentivesVerifiableagreementsNo companypulls aheadNo secretloyaltiesStrongeroversightBetter alignmentunderstandingLimits onunsafe speedPrevents Loss ofControlPreventsConcentration ofPowerTime passesWorld hasmore timeOur key policy interventionsOur main goalsIntermediate outcomes / factorsArrows represent causation~ Notice the feedback loop here! ~ + +TotalResearchTransparencyTechnicalconversationno longerdominated byAI companyemployeesVastly morealignmentresearchers withaccess to latestmodels andevidenceIncentivesnot soperverse infrontier AIcompaniesAgreements topay safetytaxes and bandangerouspractices canbe verifiedNo Company PullsAhead; OtherCompanies CatchUp, IncludingCompanies inOther CountriesNo SecretLoyalties /HiddenAgendasOversight of AI companiesby governments is moreeffective; oversight of execbranch by judiciary,congress, the public, etc. ismore effective tooBetterunderstanding ofAI alignment, AIrisks, etc. Alsoshifts the burden ofproof.Prevents Lossof ControlLimits onSpeed of AlgoProgress +UnsafePracticesBannedPreventsConcentrationof PowerTime passes,diffusionoccursWorld Has More Timeto Prepare, React,Adjust to HistoricallyFast AI-DrivenChangesOur key policy interventionsOur main goalsIntermediate outcomes / factorsArrows represent causation~ Notice the feedback loop here! ~ + +In the diagram, the two most important policy interventions of Plan A are described in blue. The two most important goals—preventing loss of control, and preventing extreme concentration of power—are described in green. Arrows represent causation. + +This diagram is in some sense a summary of Plan A, except it’s biased towards depicting the benefits of transparency in particular. + +We won’t talk through the entire diagram here. Instead we’ll talk through the seven depicted benefits of Total Research Transparency, plus an additional eighth benefit. For each one, we’ll talk about how weaker kinds of transparency would get similar (but weaker) benefits. + + 1. **“Technical conversation no longer dominated by AI company employees.”** Transparency will improve epistemics about AI risks, and therefore will make regulations more likely to be implemented well. + +The US and China need to agree on AI regulation sufficient to dramatically reduce AI takeover risk. This regulation will involve things like deciding on a capabilities schedule (which capabilities are too dangerous, when is it worthwhile to accept some risk?), deciding on whether particular algorithms should be allowed vs. banned (e.g. new architectures that are more sample-efficient yet also more interpretable), and which research directions have worthwhile capabilities/alignment tradeoffs (e.g. whether research into new scalable oversight techniques is likely enough to be a useful component of a safety case). + +Under status quo transparency, this would result in a small group of government officials needing to make these decisions. The only technical experts they would be able to consult that have access to the latest information would be at the AI company(ies) they are trying to regulate, because they wouldn’t be allowed to share information with the other technical experts. But the AI company employees have a massive conflict of interest here and so cannot be trusted to give unbiased opinions. + +It helps to give the information to a broader group of experts (e.g. academics and rival AI companies), without publishing it. Now, AI companies can be used to audit each other. This has much better incentives because AI companies have an incentive to point out if their competitor is doing something reckless, dangerous, or illegal. However, this relies on governments selecting the right groups of experts, and relies on AI companies to report on AI companies, each of whom will still be biased against industry-wide regulation. + +It helps even more to make the relevant information public, because now we no longer rely on AI companies or government approved AI experts to process the information on their own and form good opinions about the necessary regulations. This allows anyone to write up high quality risk analyses about any given model, algorithm, or practice. Instead of the discussion of sufficient regulation happening in a few small offices with government and AI company employees, it can happen in public, making it dramatically more likely that mistakes will be caught and fixed. The government still needs to make good decisions, but they are in a better position to do so because they can peruse an ocean of research and analyses produced by different groups that are all able to see the relevant evidence. + +Furthermore, transparency makes it much more likely that the judicial system will be able to exert reasonable oversight on AI development, which is a core lever for enforcing reasonable safeguards. + + 2. **“Huge increase in the number of alignment researchers with access to the latest models and evidence.”** More transparency means more researchers can participate in cutting-edge research, which means progress is made faster. + +In 2026, only a small fraction of people with the expertise to do AI alignment research have access to the latest models and the latest evidence. The rest are outside the frontier AI companies, either in laggard companies, or in nonprofits, or in academia. Total Research Transparency equalizes access, so that anyone can buy some compute and run some experiments on the latest models, and see the results of the experiments others have run. This straightforwardly increases the effective amount of human labor directed at making alignment research progress. + +More moderate forms of transparency would help in the same way, but to a lesser extent. + + 3. **“Incentives not so perverse”** Total Research Transparency removes the incentive for AI companies to race to better AI algorithms. + +In Plan A, the US and China directly regulate AI companies to limit their rate of algorithmic progress. But the transparency also serves as a second layer of defense by changing the incentives around algorithmic progress. + +Under the status quo, AI companies' main moat is their algorithmic advantages over competitors. Therefore, AI companies have a huge incentive to race to find better AI algorithms. When all AI algorithms are shared, this moat evaporates, and so too does this incentive. The race dynamic is a central upstream factor leading to high AI takeover risk, which Total Research Transparency directly undermines. More moderate kinds of transparency would not achieve this to nearly the same extent. + +It helps to consider an example. In 2026, [many AI researchers agree](https://arxiv.org/abs/2507.11473) that chain of thought monitorability is helpful for safety and that it would be a shame to lose it. However, each company is individually incentivised to have research programs into new paradigms that would destroy or damage chain of thought monitorability (call this “[neuralese](https://ai-2027.com/#narrative-2027-03-31)”). After all, if they don’t do it, someone else will. + +Now imagine that something like Plan A happens, with Total Research Transparency. AI companies can sleep more easily without doing any investment into neuralese research, because they know that if anyone else was making a serious attempt, they’d immediately see it. What would be the point of doing it anyway? You’d take a hit to monitorability for a competitive advantage that would be immediately shared with all your competitors. + +Now imagine that the regulation is not based on Total Research Transparency but instead on some more moderate kind of transparency—e.g. filtered transparency based on government audits. Now each company has an incentive to research neuralese again, and moreover, if they’ve started researching neuralese they may even have an incentive to convince the government that it’s not as dangerous as it sounds, because now they think maybe this really will be their competitive advantage. + + 4. **“Agreements to pay safety taxes and ban dangerous practices can be verified.”** Total Research Transparency makes AI agreements more robust. + +Making the logs of most AI development public dramatically increases the aggregate brainpower that’s looking over the logs to ensure they are in compliance with the agreed upon rules. From the perspective of a company or country trying to defect from the agreement (e.g. by secretly using compute in the R&D clusters to do a massive illegal training run, or a training run that they argue is allowed but is really pushing the boundaries and is arguably illegal), it becomes massively more difficult for them to do this because it will be hard for them to predict what sorts of analyses the public will do on the logs. The alternative, trying to outwit a small group of overstretched government auditors (in proposals 2&3), will probably be much easier. And of course, without even some basic level of transparency that is already much stronger than what we have in 2026 (US and Chinese governments get to see what’s going on in each other’s datacenters) then they have hope that their spy and cyber operations are good enough to substitute, or else trust each other not to cheat. + + 5. **“No company pulls ahead, others catch up, including companies in other countries.”** Having multiple companies spread across multiple countries is great for preventing extreme concentrations of power. + +This one is pretty straightforward. In 2026, power comes from a variety of sources—knowledge is power, money is power, military strength is power, etc. In the future, as AIs become superhuman at everything and proliferate around the world, power will increasingly come from controlling the best AIs. (It also comes from controlling the greatest number of AIs, so to speak, but because AIs can be copied, we think that generally speaking whoever controls the best AIs will soon end up controlling a very large number of them as well.) + +The kind of broad access and control offered by AI companies in 2026 is pretty weak from a power-concentration perspective. Yes, individual users can purchase subscriptions to ChatGPT or Claude and so forth. And under normal circumstances, ChatGPT or Claude will follow your orders. However: + + + + * You can’t use them for everything—they’ll refuse to help you with some things for example, or possibly even [quietly underperform](https://www.lesswrong.com/posts/sSyLyc3KDQzboQGWS/thoughts-on-claude-fable-s-silent-safeguards), and + + * There’s nothing stopping the AI companies from inserting [secret loyalties or hidden agendas](https://www.formationresearch.com/secret-loyalties-whitepaper.pdf) into their AIs, like [xAI](https://www.cnbc.com/2025/07/11/grok-4-appears-to-reference-musks-views-when-answering-questions-.html) and [Google](https://nypost.com/2024/02/22/business/google-pauses-absurdly-woke-gemini-ai-chatbots-image-tool-after-backlash-over-historically-inaccurate-pictures/) appear to have done on occasion (those were the cases that were so blatant as to get caught). + + * There’s nothing stopping the government from quietly ordering the companies to do things like that, e.g. in the name of national security. Governments could abuse this to surveil, censor, and propagandize their populations. + + * In a crisis—which is when power matters most—the hard power is with the company and/or government, not with the people. Even if the AIs are currently trained and instructed to do as the user wants, for example, and even if their Spec/Constitution is fully public, well, in an actual constitutional crisis scenario all of that can go out of the window very quickly. + + * Besides, as AIs penetrate more of the economy, more and more money will flow to the AI companies. If there are only a handful of such companies, that’s inherently a concentration of power. If most or all of them are in one country, that’s inherently a concentration of power. It could easily lead to a global oligarchy or dictatorship, because a small group of already-powerful people controls all the AIs that matter. + +The situation is much less likely to lead to global oligarchy or dictatorship if there are many companies spread across many countries with similarly powerful AIs. Not only that, but this means that market forces can operate: companies will be incentivised to give users what they want, including firmer control, because it’ll help them get market share from their competitors. This helps to reduce the risk of local oligarchies or dictatorships as well; if one country uses AI to propagandize their citizens, well, perhaps their citizens will be able to access AIs from other countries that tell them the truth. + +Total research transparency helps other companies catch up to the frontier in two ways. First and foremost, it publishes the core algorithms. Secondly, it makes it easier for countries to agree to limit the pace of AI progress, which gives other companies and countries more time to catch up. Weaker forms of transparency would also help but not as much. + + + + 1. **“No Secret Loyalties / Hidden Agendas”** + +This is another major way in which total research transparency helps to prevent extreme concentrations of power. + +As previously mentioned, it is both feasible and precedented for those who actually own and train the AIs to train hidden agendas into them. After all, the values/goals/personalities of the AIs have to be trained in somehow, and that’s all very complicated, and so to make some of it hidden is as simple as declining to make all of it visible. + +Total research transparency basically kills this threat. Not only does it kill it for companies, it kills it for governments, unless the US and Chinese governments were to collude to keep secrets from the world. With total research transparency, every step of the training process for every frontier model is tracked by US and Chinese monitors and published. If the model creator misleads the public about the goals/values/personality-traits/etc., there’s a paper trail people can follow to find out. + +Weaker forms of transparency would help too, but not as much. For example, consider Proposal 2: Filtered Transparency. Here you might end up in a situation where the US AI companies shared some political bias that they were training into their AIs, and the US government either doesn’t notice because their auditors are too busy with other tasks, or notices but doesn’t care because it happens to share the same political affiliation as the companies. Perhaps the Chinese government auditors would notice, but perhaps they too would be busy, and anyhow they might not care. + + 2. **Oversight of AI companies by governments is more effective; oversight of the executive branch by judiciary, congress, and the public is more effective too.** Total Research Transparency makes it easier for groups who aren’t directly training the models to see how the models are being trained, and to judge for themselves whether the rules are being implemented fairly and reasonably. + +Consider a biased or corrupt regulator that is using its authority to punish companies that it doesn’t like, or impose its ideology on domestic AI industry. Consider instead a situation where a company or companies is complaining that this is happening, but really the regulator was behaving reasonably in response to the evidence, and simply doing its job to keep people safe and enforce the law. How is the public to distinguish between these two cases? How is the judicial system? + +Total Research Transparency helps everyone see what’s going on, the better to evaluate who is being reasonable and who isn’t. As usual, weaker forms of transparency help too but not as much. + + 3. **Unknown Unknowns:** Transparency seems good for putting humanity in a position to deal with problems as they come up, including problems that we haven’t thought of yet. + + + + +The development of AI will pose many problems, both foreseen and unforeseen. Increasing transparency into AI development will on average improve the level of public epistemics and understanding of what is going on. + +## Total Research Transparency Downsides + +**The main additional downsides of Total Research Transparency relative to Proposals 2 &3 are:** + + * **Total Research Transparency potentially incentivizes “safety theater”.** There may be some lines of AI safety research which, if pursued, might make AI companies look bad. For example, evaluations of model misbehavior. Under conditions of Total Research Transparency, this sort of research might be disincentivized because there would be no way to hide the results, whereas in filtered transparency for example the results would only be shared with a smaller circle of auditors and other companies. + + * **Leading AI companies and large parts of the US executive branch will probably dislike it.** Competitors will dislike that it allows their competitors to catch up, especially companies whose algorithms are their primary moat. Governments generally shy away from giving geopolitical rivals access to information which could then be used against them. Therefore, the leading AI companies and many parts of the US executive branch will probably be opposed to Total Research Transparency. + + + + +1. + +Without the ability to do trustless verification, the main alternative is to rely on trust. Historical agreements have often relied on trust for enforcement (for example, the London Naval Treaty restricted the tonnage of military fleets, but did not include verification provisions). However, without verification, there is a much larger incentive to defect. + +2. + +Concretely, we think that we will, at some point in the capability progression, get AIs that (1) can quickly design/build mirror life, (2) superpersuade/mindhack humans, and (3) quickly build self replicating nanobots/grey goo that grows fast enough to boil the oceans pretty quickly. The resilience response to these threats looks something like "no human ever goes outside, everyone lives in bioshelter bunkers, all info going in and out is moderated by a friendly AI, and a friendly drone swarm / nanobot army is outside protecting the bunker”, which would be (i) highly disruptive and undesirable to most people and (ii) take significant serial time to implement. + +3. + +We think that it is technically viable to verify that an AI cluster is doing inference on a fixed model without the auditors being able to read the information going in and out of the model (e.g. the chatbot logs). + +4. + +In the absence of improved privacy-preserving verification technology, this would involve something similar to Total Research Transparency within the security boundary; with improved verification technology, it could involve less transparency while still being sufficient to verify deals. + +5. + +In the Plan A scenario we also do depict the implementation of Proposal #4 for a short period of time (from March to October 2029). + +6. + +This is enforced via an inference-only retrofit, as explained [here](https://docs.google.com/document/d/1K4vxl_IMf58OgNXbWntQl6BQ06yAWZ62WmW5Ypp4KYA/edit?tab=t.86wiss1dhlds#heading=h.g5eefrev1iir). + +7. + +This verification also helps to prevent misuse and AI R&D assistance. This is covered more in our [verification supplement](https://www.ai-2040.com/supplements/verification-plan). + +8. + +This threshold is chosen to be high enough to allow for convenient access to the R&D clusters, but low enough to make weight theft difficult. The security rationale is discussed further below, in [the Security supplement](https://docs.google.com/document/d/1nh8g0TnRQZjpSCvlq_M9te_WclgnIt2Mx3iQujfz0Kc/edit?tab=t.0). + +9. + +This happens automatically, because of the co-location of multiple companies’ datacenters + +10. + +Under this proposal, then, what would the companies be incentivised towards? Well, they still own the model weights they train, and they still own their datacenters, their software products, etc. So they’ll compete to train the biggest models and teach them the most lucrative skills and serve them fastest in the most appealing UI and build up the most loyal userbase. Just like in other low-secrets industries, a company with good taste and rapid execution can overtake larger dumber competitors, stay ahead even as they scramble to copy them, and ultimately perhaps even build up a durable moat of infrastructure scale and network effects. That said, moats and profit margins will be smaller and AI will be much more of a commodity under our proposal than in the default world. We think this is a good thing. + +11. + +Of course, this doesn’t solve the whole problem. For example, a defecting AI project may be able to conduct their research steganographically through the monitored input and output channel, or they may be able to undermine the cybersecurity that’s attempting to enforce the transparency requirements. diff --git a/sources/ai2040-pages/verification-plan.md b/sources/ai2040-pages/verification-plan.md new file mode 100644 index 0000000..0b1f197 --- /dev/null +++ b/sources/ai2040-pages/verification-plan.md @@ -0,0 +1,1058 @@ +--- +url: https://ai-2040.com/supplements/verification-plan +--- + +## [AI 2040](/) + +≡[PDF](https://ghjyjqzwz4.ufs.sh/f/9qHa0cBclQ7sf6FVzqbp2gHFXC1IbGS5ZA7DmNBusfM4YJRP) + +[Supplements](/supplements) + +Plan A Basics + +[FAQFAQ](/supplements/faq)[How Plan A solves our 5 biggest problemsHow Plan A solves our 5 biggest problems](/supplements/how-plan-a-solves-our-5-biggest-problems) + +Plan A Policies + +[Covert AI ProjectsCovert AI Projects](/supplements/covert-ai-projects)[Verification PlanVerification Plan](/supplements/verification-plan)[Transparency PlanTransparency Plan](/supplements/transparency-plan)[Capability Scaling StrategyCapability Scaling Strategy](/supplements/capability-scaling-strategy)[Security in Plan ASecurity in Plan A](/supplements/security-in-plan-a) + +Analysis of Plan A Policies + +[Plan A AssumptionsPlan A Assumptions](/supplements/plan-a-assumptions)[Comparing Possible PlansComparing Possible Plans](/supplements/comparing-possible-plans)[Deal DeclineDeal Decline](/supplements/deal-decline)[Takeoff SupplementTakeoff Supplement](/supplements/takeoff-supplement) + +Economics + +[Economics of Plan AEconomics of Plan A](/supplements/economics-of-plan-a)[Compute in Plan ACompute in Plan A](/supplements/compute-supplement)[Economic Growth ExplorerEconomic Growth Explorer](/supplements/econ-explorer) + +Speculative Analysis + +[Space Governance PlanSpace Governance Plan](/supplements/space-governance-plan)[AI for EpistemicsAI for Epistemics](/supplements/ai-for-epistemics)[Alignment RoadmapAlignment Roadmap](/supplements/alignment-roadmap) + +[About](/about) + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +[The Plan](/supplements/verification-plan)[Get Involved](/supplements/verification-plan/get-involved) + +# Verification Plan + +### Romeo Dean + +Plan A is predicated on the ability to enforce an international deal to slow down AI progress without relying on trust. We are confident that verification measures can make this possible, by providing sufficient assurance to both the US and China that defection by the other would either be caught or small enough to be inconsequential. + +There is a spectrum of how costly and how trustworthy our best verification solutions are, which we can think of as a verification frontier. Today, the verification frontier is not mature. Currently the only way of verifying an AI development pause with extremely high assurance between the US and China would probably require datacenters to be powered down, and anything less costly than this might have significant risk of being violated. That being said, we believe the verification problem is inherently defense dominant, and AI verification is also an extremely young and so far [tiny field](https://nacicankaya.substack.com/p/all-hands-on-deck-to-build-the-datacenter), so it should be highly tractable to make significant progress in the next few years. + +This supplement provides a summary of the verification plan that we think should be followed to make Plan A viable, and then is followed by a more detailed walkthrough of how the verification story unfolds in our specific Plan A scenario. + +**Want to help build this?** See the [Get Involved](/supplements/verification-plan/get-involved) tab: what needs to be done, who the field needs, an [expression of interest form](https://tally.so/r/LZyyDy), and open opportunities across the verification field. + + 1. **[Summary of the Plan](/supplements/verification-plan#summary-of-the-plan)** + + 1. [Prepare deal optionality](/supplements/verification-plan#phase-1-prepare-deal-optionality) + + 2. [Implement the deal](/supplements/verification-plan#phase-2-implement-the-deal) + + 3. [Improve robustness](/supplements/verification-plan#phase-3-improve-robustness) + + 2. **[Verifying the deal in our Plan A scenario](/supplements/verification-plan#verifying-the-deal-in-our-plan-a-scenario)** + + + + +20262027202820292030203120322033203420352036+1Chiptracking2VerificationR&D funding3Mutual ChipDeclaration4Inference-onlyretrofit5R&Dverification6Hardwarecap & trade7VerificationimprovementsPhase 1: Preparing2026–2028Phase 2: Implementing2029–2030Phase 3: Improving2031–2036+Verification Plan Timeline \- ai-2040.com + + 1. **[Appendix:](/supplements/verification-plan#appendix-packet-based-verification) [Packet-Based Verification](/supplements/verification-plan#appendix-packet-based-verification)** + + + + +# Summary of the Plan + +We recommend an international deal to avoid a catastrophic race to superintelligence, spearheaded by the US and China. This deal needs to be enforceable without relying on trust, because the US and China will be unlikely to make an agreement that cannot be verified; and they are much less likely to defect on the agreement if any defection is likely to be caught. + +We can break up the verification problem into two key components: + + 1. **Problem 1. Verifying that known compute is compliant.** + + 1. We are confident that it is possible to have verification measures that enable a trustless deal. That being said, if a deal needed to happen tomorrow, we would not be ready to enforce it trustlessly in a low-cost way. Given the increasingly large economic incentive to keep AI services online, a ready-to-go verification solution (that doesn’t require much time to activate) may be incredibly important for making a deal happen. + + 2. **Problem 2. Verifying that unknown compute is inconsequential.** + + 1. Our [covert project supplement](/supplements/covert-ai-projects) estimates that even a competently-executed covert diversion effort would net a median of about 0.5% of the world’s AI-relevant compute (80% CI: 0.1% to 1.4%) as ‘dark compute’ at the start of the deal, which given the likelihood of compute-dependence of frontier progress, should buy the deal a lot of time and stability. This bound should also tighten as the deal goes on, as detection efforts accumulate and more compute is brought under verification. + + + + +Our recommended solutions come in three phases. + +**Phase 1. Prepare Deal Optionality.** The US and China take actions to make a deal more likely to be enforceable and durable, by preparing the ability to do inference-only verification, and keeping close tabs on how much AI compute exists and who owns it. + +**Phase 2. Implement a Deal.** The US and China perform a mutual compute declaration (spreading to third countries and ideally reaching >99% of world compute); inference-only verification retrofit to ~99% of world AI datacenter compute, and then sprint towards setting up secure R&D verification that will use new post-deal compute production. + +10M+1 DCs·10M·3.5%1M-10M74 DCs·149M·51.4%100K-1M257 DCs·76M·26.4%10K-100K497 DCs·15M·5.3%1K-10K823 DCs·2.6M·0.9%<1K1.2K DCs·397K·0.1%In-transit35M·12.0%AI consumer255K units·938K·0.32%non-AI consumer7.7M eff·2.7%Inference-onlyReporting requirementsCap & Trade(if unverified)US · 224M · 77%China · 26M · 9%RoW · 39M · 14% + +This is what we think implementing the deal would look like in 2029 in our scenario, with an inference-only retrofit of all the medium and large AI datacenters (>10K H100e, or approx. >$100M), and this being enough to cover ~99% of world AI-relevant compute. Then keeping tabs on the rest of the smaller clusters and taking measures to avoid them being possibly used in a covert project. We are not confident in the modelling of this [concentration in datacenter sizes](/supplements/compute-supplement#14-datacenter-sizes), so the exact cutoffs and interventions may need to be different. + +Compute locations by datacenter size (Jan 1 2029) \- ai-2040.com + +**Phase 3. Improve robustness.** Over time the US and China improve the stability and durability of the verification regime, especially through hardware security, verification robustness, and more. + +## Phase 1. Prepare Deal Optionality + +We recommend focusing pre-deal effort on sprinting towards a minimal verification package that can be used to securely retrofit 99% of the world's AI compute to be verified as inference-only, and for at least 99% of the world's AI compute to be known-to-exist (at least to intelligence agencies) with low uncertainty, so that it can’t be hidden from a future compute declaration (and therefore verification regime) with plausible deniability. + +Why do we recommend this? + +If the year is 2029, and there is no verification progress relative to today, our best guess is that frontier AI services would need to mostly (e.g., 90%) stay offline for 6-12 months in the wake of a deal if both sides wanted to be highly confident there wasn’t fast frontier AI development continuing. If there were a minimum-viable inference-only retrofitting solution prepared, the equivalent period would be more like 2-5 months. The difference between these in terms of economic impact might be in the trillions of dollars by 2029. Of course, it would be even better if there was a more general verification solution ready, that could e.g., also allow approved training and experiments to continue, but we think this is likely to be harder, and only provide a marginal gain on top of an inference-only solution, because most of the economic upside can be retained by only keeping inference online (since training and experiments will be intentionally limited under a slowdown deal anyway, especially at first while negotiations on how to proceed with AI development are negotiated). + +Even with such a verification solution ready, the maximum verification coverage of a deal is still bounded by the percentage of the world's compute that can be brought under the deal. If the US, China, or any other actor has plausible deniability on how much compute they have, they can use that plausible deniability to hide compute from the deal, and use it as unverified compute in a covert project. But if the total unknown compute is under 1% of world compute, we think any such covert effort is likely to be inconsequential: our [covert project supplement](/supplements/covert-ai-projects#section-6-overall-risk-estimate) models this directly, and estimates roughly a 13% chance that a competently-pursued covert project reaches deal-undermining AI before detection or handoff, under our recommended scaling strategy. These estimates apply to the start of the deal: as the deal goes on, the uncertainty should narrow and the plausible covert stock should shrink, as tracing and audits converge, detection efforts accumulate (see the [detection estimates](/supplements/covert-ai-projects#section-3-detecting-covert-ai-infrastructure) in the covert project supplement), and more compute is brought under verification. Of course, the lower the unknown compute is, the better but we think around 1% of world compute is likely to be an achievable target and sufficient for a stable deal. + +**Top priority recommendations:** + + 1. **Invest in verification R &D with a specific focus on developing a secure inference-only retrofitting plan with around 80% of total verification R&D effort.** + + 1. Also, support supplementary R&D efforts with around 20% of verification R&D effort pursuing promising additional measures, such as software-only solutions, privacy-preserving verification measures, and more general verification solutions (e.g., that verify the compliance of training and experiments as well). + + 2. **Invest in compute accounting and tracking (e.g., carried out by intelligence agencies) to know how much AI relevant compute exists in the world, and which entities are in possession of it.** + + 1. Notably, our recommendation for this early chip tracking doesn’t prioritize knowing where this compute is, or what it is being used for. We believe that the majority of the upside in terms of making a deal viable is simply in knowing how much exists and keeping track of which actors own it at any given time, because the lack of plausible deniability should be sufficient to force this compute to be declared at the time of deal negotiation (given that there will be large incentives in favour of joining a deal). That being said, also knowing its location is helpful in avoiding cases of plausible deniability (e.g., faked records of on-selling or chip deaths) and cases where (for some reason) the US and Chinese governments have little leverage over the owning entity (and therefore can’t pressure them to declare the location of the compute they are known to have, e.g., Russian compute). + + + + +## Phase 2. Implement the deal + +To set up a stable deal, we recommend immediately carrying out a mutual compute declaration and supply chain records audit to maximize the amount of compute that is brought under the deal. Then implementing the best available inference-only verification solution to this compute to enforce a brief pause on training and experiments. In parallel, we recommend a sprint to set up a verification regime for resumed training and experiments. + +Why do we recommend this? + +Given that unknown compute can’t be brought under the verification of a deal, the first priority should be to minimize the amount of compute that is unknown. The urgent action needed for this is to immediately carry out a declaration of existing compute and compute currently in transit or production. Both the US and China should then be able to carry out an audit and inspection of each other’s declaration, and compare it against their private intelligence efforts (to the extent they properly prepared this) to resolve their uncertainties over how much compute each other have. + +Around 85% of global compute should be owned by US or Chinese entities, so aiming for 99% of compute being under the deal will require rolling out such a declaration agreement to the other 5-10 nations that will likely make up the rest of this gap. The level of urgency of extending this to other nations will depend on the uncertainty intelligence agencies have over how much compute these other nations have, how likely it seems that these other nations would successfully be able to hide significant quantities of it without detection. + +Following this declaration, we recommend a brief AI R&D pause (i.e., new training runs and experiments), with existing compute being converted to inference-only. As explained above, we think this achieves the goal of a deal (slowdown in AI progress) while keeping most of the economic upside (AI services and products keep running), while also being relatively easier to verify. Whatever the best inference-only verification solution that is ready should be used. We currently recommend aiming for a network-tap and recomputation based approach, but more R&D is needed to figure out if another approach will be better. It’s also possible that only an inferior solution will be ready, but still sufficient to be used in the initial setup of a deal until something better is ready. + +Finally, we recommend setting up the verification regime for resuming training and experiments as a separate effort that happens after the deal begins because we think there are important properties of R&D verification to decide on once the deal is setup, and also extremely high security requirements needed for the R&D verification to enable a stable deal going forward. In particular, there is a decision to make about what level of transparency to aim for (we recommend a ‘Total Research Transparency’ regime, but more on the possibilities to decide between in our [transparency supplement](/supplements/transparency-plan)), and in any case there are high security requirements, so we believe that the verified datacenters will need to be built from scratch. If new construction is deemed to take unacceptably long, or there is a sufficiently secure retrofitting solution for R&D verification, then there could be an interim period where pre-existing datacenters are used for R&D. That being said, we think after around 2 years, when verification and security requirements are elevated, and something like 90% of the world’s compute will be new post-deal compute anyway, we recommend more strongly having new, post-deal, highly secure datacenters. + +**Top priority recommendations:** + + 1. **Carry out a bilateral compute declaration and supply chain audit, for both existing compute and new compute production, and extend to third nations with significant compute ownership or production.** + + 1. Every major datacenter owner and companies with a major role in the supply chain publicly declare their records about purchases and sales, and allow auditors and inspectors into each other’s countries and companies to verify the records and count of final chips. This needs to also including compute in transit and active compute production. From this point forward, existing compute and new compute production is closely tracked and guarded so that it can’t be redirected to a covert project without detection. Extend this declaration to other countries with significant compute ownership or production capacity. Do so urgently unless there is confident intelligence in place that would detect this compute being moved. + + 2. **Implement inference-only verification on existing compute and any new compute coming online.** + + 1. Roll out inference-only verification to existing compute to enforce the brief AI R&D pause while allowing AI services to remain online. Ideally there would be a highly secure solution ready to go. If not, implement the best solution available. If there is no solution at all ready then either shut down the compute, or rely only on non-technical measures (e.g., embedded auditors in each other’s companies checking that no one is implementing experiments or training runs) in the meantime while sprinting to transition to a better solution. + + 3. **Sprint to set up R &D verification in accordance with the chosen transparency regime.** + + 1. We recommend a specific regime we call Total Research Transparency in the Plan A scenario, but as explained in our [transparency supplement](/supplements/transparency-plan), there are conditions under which a different approach may be better. Regardless of the regime, there should be some combination of highly secure R&D clusters and highly secure post-deal inference clusters built, because across most transparency regimes we imagine at least AI model weights and verification correctness requiring [SL5](https://www.rand.org/pubs/research_reports/RRA2849-1.html) (robust to nation state actors) level security. Because of this we think these clusters should be built from scratch. + + + + +What AI R&D agreements should actually be verified? + +The primary purpose of the deal is to create the conditions that can allow for slowing down AI development in order to increase safety, by paying safety taxes in safety resource allocation (e.g., doing alignment research instead of capabilities research) or development directions (e.g., agreeing to not pursue a certain research direction or use a certain efficient algorithm that might be less safe). Overall, we refer to this safety focused slow down on AI research as ‘AI research titration’. The method of AI research titration may be crude at first (e.g., harsh experiment compute caps) and become more sophisticated over time (e.g., case by case decisions informed by scientific understanding of capabilities externalities). + +We additionally think that a deal may have to deal with potentially destabilizing levels of hardware research, and hardware production levels. In the Plan A scenario we recommend hardware research titration and a cap & trade policy on robots and compute production to deal with the risks from these sources. All of this is explored in more detail in the [Plan A verification story](/supplements/verification-plan#verifying-the-deal-in-our-plan-a-scenario). + +## Phase 3. Improve Robustness + +Once a deal has been set up, we expect many improvements to its robustness will be possible. We recommend investing heavily in hardware security; improving the verification regime; pursuing favorable trade-offs in hardware design that favor deal robustness (e.g., by making verification easier, improving various security properties, differentially favoring the actors in the deal over covert projects, etc.); and reducing incentives to pull out of the deal (e.g., through mutually assured compute destruction, and balancing deal dissolution outcomes between actors). + +Why do we recommend this? + +We expect there to be many achievable improvements to the robustness of the deal. Given that the deal breaking down is one of the biggest threats to Plan A succeeding, as explored in our [deal decline supplement](/supplements/deal-decline), reducing the incentive to pull out of the deal is a core part of improving the deal’s stability. Covertly defecting from the deal is another core threat to the deal (as explored in our [covert project supplement](/supplements/covert-ai-projects)), so another core part of improving the deal’s robustness is making the verification regime better, both on the level of assurance of known compute and reducing the amount of unknown compute; improving various other security properties, like model weights security; and more generally making it harder for potential defecting actors to make or steal progress. + +Concrete ideas for deal robustness + +**Deal dissolution contingencies and incentive shaping.** + +**Mutually assured compute destruction.** + +One of our main proposals for how to disincentivize actors from pulling out of the deal, and make AI development safer if actors do pull out of the deal, is to set up mutual ability to destroy the post-deal compute. + +If this is strongly in place, then no actor will sanely attempt to pull out of the deal and gain a lasting advantage by defending some of the compute, and returning to racing. If the deal does break down despite this, then compute destruction can let any actor destroy the post-deal compute, which makes things potentially vastly safer because it destroys the large compute overhang, which back in a default world with actors racing might be extremely dangerous because it would make AI progress extremely fast. + +**Cold storage compute banks.** + +Another possibility would be to decrease the incentive for actors to pull out of the deal is to have a small compute bank that is verified to not be in use during the deal, but which each side can reliably defend and bring back online if the deal does break down. + +If this is reliably bigger than any potential covert project, then it removes the incentive to pull out of the deal for any actor that may have the largest covert project, also reducing the incentive to attempt one in the first place. In our scenario, the US maintains a 5M H100e bank and China a 500K H100e bank (see our [deal decline supplement](/supplements/deal-decline#preparing-stockpiles-to-align-incentives), the source of truth for these stockpiles); the US bank sits comfortably above the ~4M H100e top of the 80% CI for plausible diversion in our [covert project supplement](/supplements/covert-ai-projects#section-2-preventing-covert-compute-procurement). The ratio between the size of the US and China’s compute banks (and whichever third parties might negotiate their own) could be frozen at the pre-deal status quo compute ratio. + +**Cold storage model weights and data bank.** + +Another idea we have to improve the situation if the deal breaks down is to allow the US and China to agree to submit models to a hardened cold storage model weights bank. + +Both sides would need to verify that the other side’s weights are not going anywhere else, but not have the ability to destroy them if the deal broke down. The reason they may want to do this is if they have made a lot of alignment progress in the form of a final trained model or large datasets, and they mutually agree it would be safer for the world in the event of deal break down if they were allowed to have these model weights and/or datasets. + +**Improving the verification assurance on the compute in the deal.** + +**Hardware security.** + +There may be several hardware security properties that can be drastically improved under the deal, especially by applying AI research automation help differentially in these areas, and paying hardware ‘security taxes’ in the design choices. These could include software level security improvements, such as expanded formal verification, and hardware security improvements, making tampering extremely difficult. One step towards achieving this could be to increase the monitoring and security at fabs and other parts of the supply chain. Another nice property that could also help with mutually assured compute destruction (purely through cryptography) could be [offline licensing](https://www.rand.org/content/dam/rand/pubs/working_papers/WRA3000/WRA3056-1/RAND_WRA3056-1.pdf). Also, it might be possible to build computing clusters that keep model weights encrypted inside secure parts of the AI clusters using hardware-derived keys that are unobservable from the outside. If this could be made robust enough, then it might drastically help with the security of model weights and allow other SL5 model weights measures (bandwidth limiting) to be relaxed. + +**Verification granularity.** + +Another more specific area for hardware improvements is to help with verification robustness, such as extremely high fidelity reproducibility, support for increasingly granular logging, and other more specific solutions like [fixed set](https://www.rand.org/content/dam/rand/pubs/working_papers/WRA3000/WRA3056-1/RAND_WRA3056-1.pdf) hardware designs (e.g., [hardcoded weights](https://taalas.com/the-path-to-ubiquitous-ai/) might be a nice way to upgrade the inference-only verification). + +**Reducing compute that could be used for a covert project.** + +**Track down potential covert projects.** + +Over time we think it may be plausible to rule out covert projects above smaller and smaller sizes, using intelligence efforts, agreements to bring more known compute (e.g., military-exempt compute or tiny private clusters) under privacy-preserving verification measures, or agreements to do power usage accounting (e.g., mutual declaration and random inspection). Our [covert project supplement](/supplements/covert-ai-projects#section-3-detecting-covert-ai-infrastructure) covers the detection methods (waste heat tracking, other IMINT/MASINT, SIGINT, HUMINT) and estimates detection probabilities by covert datacenter size, along with the [stronger verification concessions](/supplements/covert-ai-projects#section-71-options-for-responding-to-covert-projects) that could be negotiated over time. + +**Make sure no new compute goes to covert projects.** + +One extremely important property of the deal is that none of the post-deal compute production can go to covert projects. Our [covert project supplement](/supplements/covert-ai-projects#section-2-preventing-covert-compute-procurement) estimates that because production runs through a small number of highly controlled fabs, diversion of new chips can be kept to approximately zero, with very high confidence that the total is under 100K H100e. To make this even more robust, it might be favourable to concentrate AI chip production in fewer locations and/or colocate production with the final destinations (datacenters), which is doable because with total production increasing 3x/yr, every 2 years 90% of the fabs are new fabs. This means they can pretty easily be somewhere else. Consumer compute could be produced elsewhere but the risk of covert project usage of consumer compute should be separately mitigated. + +**Limit consumer compute that could be used in covert projects.** + +Consumer compute will likely include some AI-relevant compute by 2029 (e.g., hobbyists with powerful gaming GPUs, self-driving car chips), and mostly non-AI relevant chips with low effectiveness in AI workloads but which are nonetheless very numerous. This should be tested more precisely at the time for its usefulness in frontier AI workloads. If its usefulness is found to be high, then it needs to be avoided that large quantities of consumer compute could be rounded up and used in a covert project. Surveillance (i.e., reporting requirements and spot checks) should be an extreme last resort, instead there could be a voluntary buy back program, and a cap & trade on the total production of new unverified consumer compute. Ideally, this would lead to innovation about doing ‘edge verification’ (e.g., robust fixed set hardware mechanisms), to avoid the permit costs. The cap can be set such that trivially noticeable shares of the consumer compute would need to be diverted to noticeably accelerate a covert project. + +**Favorable directions that differentially help the deal over potential covert projects.** + +**AI scaling tradeoffs.** + +One of the core recommendations we make in Plan A is that to the extent capabilities progress is made intentionally, it should be done by scaling hardware (i.e., model size, training compute, data, etc.) rather than software (algorithmic efficiency) where possible, because the latter is harder (and less desirable) to govern (i.e., keep secure from potential covert projects). There are probably ways to push this idea very far as the deal goes on, by pushing to extremely large AI models that would be very difficult to steal, and even if they did would be extremely difficult for a covert project to even run at all on pre-deal compute. Our [covert project supplement](/supplements/covert-ai-projects#section-4-limiting-consortium-uplift-to-covert-projects) estimates that a median of around two thirds of Consortium software progress leaks to covert projects, and quantifies how hardware-leaning scaling strategies reduce [covert project risk](/supplements/covert-ai-projects#section-6-overall-risk-estimate). + +**Hardware-software co-design.** + +If it proves technically possible, transition to new AI algorithms that run poorly on pre-deal compute but well on newly designed and manufactured post-deal compute. For instance, introduce new model architectures and number formats different from those for which pre-deal compute was optimized. This way, covert projects cannot easily steal algorithms from the Consortium and run them on mostly pre-deal dark compute. + +# Verifying the deal in our Plan A scenario + +In our Plan A scenario, we follow the high level verification plan sketched out in the previous section. This part of the verification supplement will walk through the implementation and verification story with more concrete detail. Because it is a scenario about the future, in a very young field with ongoing R&D, we are not confident in the specific details. That being said, we hope to establish a useful concrete baseline for how the verification plan we sketched out above could actually be implemented. + +## 2026-2028: Preparing deal optionality + +### 2026: Chip tracking begins + +The intelligence agencies in the US begin internally keeping tabs on AI compute smuggling into China, and their domestic production levels, reaching a 10% uncertainty over how much AI compute Chinese entities own in total by the end of the year. They don’t do anything with this information yet (i.e., they don’t share smuggling information with the BIS). China learns that the US intelligence community is doing this, and begins doing it as well by leveraging existing assets mostly out of TSMC in Taiwan to get accurate counts and records of total production. They reach 10% uncertainty on US total AI compute, mostly due to uncertainty over pre-existing compute and production outside of Taiwan (e.g., South Korea and the US). Uncertainty over new compute production is much lower (on the order of 1%). + +### 2027: Verification R&D gets funding + +The US executive branch starts expressing interest in deal optionality and verification measures that could enable this. After a series of [RFIs](https://en.wikipedia.org/wiki/Request_for_information) they choose to focus efforts on a network-tap and recomputation centered approach for an inference-only retrofitting solution. The prospect of installing devices in each other’s datacenters raises initial alarm bells but it's a tractable security problem to work on. Ultimately due to the sheer simplicity of some of the possible network tap designs (essentially a piece of glass splitting optical signals), and parallel investment into various physical security agendas, it becomes a palpable possibility that these taps might actually be worth producing en-masse in case they were to be needed. Unilaterally trusted recomputation servers (that the network taps redirect traffic to and are responsible for actually performing the verification) are more complicated and thus harder to trust by default, but there are large investments made here. The R&D programs are purposely transparent and open sourced to build trust. China also sees the incentive to develop and build their own symmetrical solution in a similarly transparent way. In the latter half of the year, they agree to mutual information sharing about important verification-relevant information which is otherwise not sensitive, especially about each other’s networking and computing paradigms, which is necessary in order to build compatible network taps and recomputation servers. Several pilots from different sources of funding and different places around the world are carried out over the course of the year, improvements to the designs are iterated on, and preparations are made to increase production scales. Ultimately the US and China end up opting for network tap solutions with extremely simple optical splitting designs that have minimal room for any vulnerabilities, rather than some of the more ambitious designs that we propose in the box below, or different [cryptographic based-approaches](https://pearlresearch.ai/), which are more uncertain, but if they work (at sufficiently low cost) they might be even better and easier to trust. + +Concrete inference-only retrofitting proposal. + +Our current best guess is that the US and China should aim for an inference-only solution that uses network taps and a recomputation server. It appears to us to capture the vast majority of the upside in enabling an initial slowdown deal to be highly verified and low economic cost, while being achievable and easy to build. We are less confident in the second part of this, which is why we more confidently recommend RFIs and iteration starting as soon as possible on piloting and iterating on approaches. + +![](/verification-plan/image6.png) + +Summary of the proposal: + + 1. **Convert datacenters into groups of isolated inference-units by removing (a significant amount of) backend networking (also known as east-west networking).** + + 1. Currently AI servers in a datacenter have a lot of high bandwidth communication between them, but this is (mostly) only necessary for training, and with some relatively small overhead (e.g., reorganizing user memory storage), large scale inference should be able to use drastically less scale-out interconnect. Removing it is a significant defense in depth measure for making it harder to use the cluster for large training runs. + + 2. **Add in simple,[passive network taps](https://amodo.substack.com/p/tech-note-network-taps-for-ai-verification?triedRedirect=true) on the frontend network (also known as north-south networking) that redirect copies of all input/output traffic to a recomputation server.** + + 1. These network taps on frontend networks have multiple possible implementations. We currently think they should probably be completely passive and split optical signals into two because this is probably easiest to trust. Separately there should be a path from the storage bank back to the inference units. + + 3. **Force all AI workloads (inference, agentic tool use, etc.) to be organized in discrete and reproducible packets, which are transparent to the recomputation server (because they pass through the network taps).** + + 1. Reproducibility is required to enable correctness proofs on the packets. By default, inference workloads are not necessarily [reproducible](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/). Achieving reproducibility may involve some cost overhead in order to maintain reproducibility but we think it is a solvable problem (see [here](https://arxiv.org/pdf/2511.20621) for example of work in this direction). + + 4. **Perform partial recomputation (random samples of the packets) on the recomputation server to check the correctness of the packets.** + + 1. The integrity of the recomputation server is an extremely important aspect of this step, but assuming its computations and outputs can be trusted, then partial recomputation through random sampling should scale extremely favourably to verify the correctness of all the outputs with extremely high probability while only using a small fraction of the original compute. See calculations in the [appendix](/supplements/verification-plan#appendix-packet-based-verification). + + 5. **Have strong physical security properties to help ensure integrity of the verification measures from e.g., tampering attacks** + + 1. The integrity of the verification process requires there to be sufficient physical security measures to avoid tampering, and other attacks that could subvert the verification process. + + 6. **Verify completeness by[bounding the maximum side channel communication](https://techgov.intelligence.org/blog/suppressing-side-channels-in-an-untrusted-data-center-via-retrofitted-defenses), and using periodic memory wipes (probably through forced memorization).** + + 1. The completeness of this verification proposal, i.e., that the only outputs from the cluster are verified inference tokens, requires side channel mitigations and memory wiping as described. We do not verify that spare compute isn’t used for unapproved workloads (because this seems very challenging) but try to ensure that the only persisting results in the form of outputs are verified inference outputs. This requires strong side channel mitigations and periodic memory wiping on the inference units (which we think can be done through forced memorization). + + + + +One potential upgrade to this proposal would be to build more complex ‘secure network gateways’ that have more functionality than the network taps, and do random sampling and organize and randomly sample the logs of reproduction packets to send to the recomputation server in a robust way. + +![](/verification-plan/image3.png) + +A rough sketch for what these secure network gateways might look like. There are data path directions (input and output) with network splitters on each path that log data flow to be organized into packets that are randomly sampled and then sent to a secure recomputing server. In a boot phase, these gateways can additionally check that exclusively whitelisted model weights are being uploaded to the inference unit. Some of the hardware security properties you might want for these gateways are precedented, for example, see [iPhone secure enclaves](https://support.apple.com/guide/security/the-secure-enclave-sec59b0b31ff/web) and [PUFs](https://en.wikipedia.org/wiki/Types_of_physical_unclonable_function). + +These could be used in place of the network taps in a similar setup as before. + +![](/verification-plan/image1.png) + +On top of this key effort to build the network-tap and recomputation centered approach, there is also significant funding (order of $100M) mobilized into R&D for other approaches, including software-only and cryptography-centric (e.g., [ZKPs](https://en.wikipedia.org/wiki/Zero-knowledge_proof)) approaches, and a similar pool for physical security R&D (order of $100M) that supplements the inference-only retrofitting solution. + +Other promising verification directions. + +Instead of the partial recomputation approach we sketched out, it might be possible to verify inference-only in other ways, e.g.: + + 1. Purely cryptographic protocols, e.g., Zero Knowledge Proofs (ZKPs) for AI workloads + + 1. This has the promise of being privacy-preserving and potentially requiring less of a hardware retrofit, but is currently somewhat speculative due to the lack of computationally efficient enough algorithms for doing so being discoverable. + + 2. Memory-challenge based verification + + 1. This has the promise of potentially also not requiring a hardware retrofit, but seems likely to face some difficulties with verifying completeness. Nonetheless it might be viable for an initial deal implementation. + + 3. Crude hardware modifications. + + 1. Removing major scale out interconnect and e.g., installing some simple sensors and other physical security mechanisms might also suffice for verifying inference-only for an initial phase of a deal (e.g., 1 year) until higher assurance measures are ready for larger amounts of compute and higher stakes workloads. + + + + +### 2028: Chip tracking is mature + +The intelligence agencies have been keeping close track of new compute production, and have been partially successful in narrowing down their uncertainty over how much AI compute pre-dated their tracking efforts. By the end of the year, China has around 2% uncertainty over the US’ total compute, and the US has around 4% uncertainty over China’s. This means from China’s perspective, they are 80% confident that the US owns somewhere between 222M and 226M H100-equivalents. From the US’ perspective, they are 80% confident that China has somewhere between 25M and 27M H100-equivalents. + +## 2029-2030: Deal Implementation + +Jan 2029Jan 2030Jan 2031AprJulOctAprJulOctMutual chipdeclarationR&D pause beginsSL5 datacenterconstruction begins50%80%95%Inference-only retrofitR&D resumes2%20%R&D verification rolloutFirst majortraining runsapproved5%30%SL5 inferenceclusters rolloutFirst generation of post-deal models releasedMature safety casebased R&D rules + +Deal Implementation Timeline (detailed) \- ai-2040.com + +### Jan 2029: Mutual Chip Declaration and Inspection + +Within a week of starting Plan A negotiations, the US and China have carried out a compute declaration and supply chain record audit, allowing each other to send a few hundred people into each other’s countries to inspect and audit records. This lets them reduce their already relatively small uncertainties over each other’s compute. Each side’s best guess is now that the other could have hidden at most around 1.5M H100e from the declaration (with the 80% CI reaching about 4M), which is about 4% of the leading AI company’s R&D compute and about 0.5% of total world compute. For our estimates and reasoning on how much compute a covert project could probably hide see the [covert projects supplement](/supplements/covert-ai-projects#section-2-preventing-covert-compute-procurement). + +US224M H100edeclaredChina26M H100edeclareddeclaredplausibly undeclared1 square = 250K H100e + +Chip declaration result \- ai-2040.com + +More detail on the suggested interventions by cluster size bucket below: + +Cluster size + +Medium and Large datacenters (>10K H100e, ~$100M) + +Immediate intervention + +Inference-only retrofit + +Estimated compute in this category in Jan 2029 as % of global AI datacenter compute + +99% + +Cluster size + +Small datacenters and diffused AI compute + +Immediate intervention + +Reporting requirements for anything above an A100 in AI performance. Random spot checks. + +Estimated compute in this category in Jan 2029 as % of global AI datacenter compute + +1% + +Cluster size + +Non-AI compute (e.g., iPhones) with unclear AI usefulness + +Immediate intervention + +Intelligence to look for efforts to gather up massive quantities of consumer compute. + +Estimated compute in this category in Jan 2029 as % of global AI datacenter compute + +0.1-10%? Depends on uncertain AI-usefulness adjustments. + +Cluster size| Immediate intervention| Estimated compute in this category in Jan 2029 as % of global AI datacenter compute +---|---|--- +Medium and Large datacenters (>10K H100e, ~$100M)| Inference-only retrofit| 99% +Small datacenters and diffused AI compute| Reporting requirements for anything above an A100 in AI performance. Random spot checks.| 1% +Non-AI compute (e.g., iPhones) with unclear AI usefulness| Intelligence to look for efforts to gather up massive quantities of consumer compute.| 0.1-10%? Depends on uncertain AI-usefulness adjustments. + +The chip declaration also involves keeping extremely close tabs on all new compute production, and any chips above certain performance metrics of AI-relevance need to either be directed to known locations to be included under the verification measures. While lower powered consumer compute production can likely remain unrestricted. + +Chip flow restrictions. + +The US and China need to minimize the plausible flow of future compute production to illegal projects. This means they need to monitor its manufacturing and guard its transportation. Luckily AI compute manufacturing is highly centralized in few facilities around the world, and this should continue to be the case through 2029. We think it is highly tractable for them to successfully account for 100% of the new AI compute production (along with critical components like high performance AI memory and networking); our [covert project supplement](/supplements/covert-ai-projects#section-2-preventing-covert-compute-procurement) estimates diversion from post-deal production can be kept to approximately zero, with very high confidence that the total is under 100K H100e. We recommend having companies be allowed to direct new compute they purchase to one of two places. + + 1. Existing datacenters / ongoing datacenter builds: compute directed here will only be allowed to do inference once the inference-only retrofit is complete. + + 2. Temporary, mutual cold storage holding facilities: compute directed here will be guarded until it is able to be set up for R&D (training and experiments) once new secure, mutually verified R&D clusters are set up. + + + + +Tier 0: UnrestrictedconsumerTier 1: Subject to dealinference-only or cold storageiPhone 16 Pro MaxTesla HW3DGX SparkM4 Max MacBookRTX 4090H20RTX 5090H100 (PCIe)DGX 8×H100B200GB200 NVL720.1 GB/s1101001 TB/s10 TB/s0.0010.010.11101001KCross-Chip Interconnect Speed (GB/s)Compute (H100e)Memory capacitytier 1 above 10 TB (scale-up)8 GB128 GB13 TBMemory bandwidthtier 1 above HBM2 (~1.5 TB/s)≥ ~1.5 TB/s (HBM-class)below ~1.5 TB/s + +Chip Flow Restrictions, 2029 \- ai-2040.com + +At first, we propose that any chip with low enough interconnect, compute, memory and bandwidth can be produced unrestricted. Then by 2032, we tentatively propose a cap, with a slight alleviation of the compute threshold (to give room for more consumer applications that might require higher edge compute) and a tightening of the interconnect threshold (to hopefully lower their potential usefulness in a covert project if diverted). More on this in the [2032 section](/supplements/verification-plan#2032-cap-and-trade-for-robots-and-compute). + +### Feb 2029: Inference-only retrofit begins + +After the US and China have carried out the compute declaration and inspection, and are relatively confident that they know about the vast majority of each other’s compute, they handshake on a brief AI R&D moratorium and announce it to the world. AI companies are supposed to stop doing R&D immediately for now, including pausing in-process training runs. While they roll out the inference-only verification solution to each other’s largest datacenters, they agree to a 100% power-off of small clusters (anything under 10K H100e, because it will take a while to get to retrofitting these) and a 80% power-off of all larger datacenters because only around 20% of the compute should be sufficient to keep all AI products and services running properly, simply at slightly reduced rates. This limits the extent to which anyone can defect and make differential gains while the deal is being set up and then just going back to racing from a better position than they were in before. + +The US has about 95% of its compute in around 200,000 total AI server racks (with just over 1K H100e per server rack on average), and China has about 95% of its compute in 100,000 AI server racks (with about 250 H100e per server rack on average). + +Luckily the US has a stock of around 10,000 network taps and 100 (unilaterally) trusted recomputation servers. China has a stock of around 20,000 taps and 200 servers, and they can each purchase another 10,000 per month from various third parties. As a result of China having less compute, the US can retrofit China’s compute faster, but they agree to a proportional retrofitting schedule determined by China’s achievable pace, which is retrofitting 40% of the US compute into inference units of 4 server racks (or 4K H100e) with their existing stock of network taps in 1 month. So they will both follow this pace, and then ramp up to 80% by the 3 month mark and then to >95% and/or smaller inference-units beyond that. + +![](/verification-plan/image6.png) + +### March 2029: R&D cluster construction begins + +While the inference-only retrofit is being carried out, the US and China then agree on the high level regulatory plan and transparency regime they will be aiming for (some of the options are highlighted in our [transparency supplement](/supplements/transparency-plan)). Given the relatively small perceived risk from illegal projects and various other considerations, they decide to aim for the Total Research Transparency approach. This involves setting up R&D datacenters with SL5 model weights security and SL5 verification integrity. They want these to not be defendable so that it's difficult to try to pull out of the deal and start racing again using these new R&D clusters, so they aim for mutual deterrence by locating the US R&D clusters in Mongolia and the Chinese ones in Canada. The US has guards at their Mongolia cluster that have robust ‘scorched earth’ mechanisms to destroy the compute if China or someone else tries to seize it, and symmetrical measures are in place for the Chinese compute in Canada. They get started on these infrastructure projects immediately in March with high urgency, aiming to have a small amount of capacity online within 6 months so that a verified R&D pilot can start. The verification approach they plan to follow is similar to the inference verification approach, with network taps being used as evidence collection, and partial recomputation being used to verify correctness of the packets that make up the workload. + +### April - May 2029: Third party countries begin joining the deal + +Starting in April and through the rest of 2029, most of the richest 20 countries in the world, and important semiconductor supply chain countries join the deal by doing a compute declaration and allowing the inference-only retrofitting to their datacenters, including: Germany, Japan, UK, India, France, Brazil, Canada, Mexico, Italy, Australia, South Korea, Spain, UAE, Saudi Arabia, Taiwan, Netherlands, Switzerland. + +US224M H100edeclaredChina26M H100edeclaredRoW39M H100edeclareddeclaredplausibly undeclared1 square = 250K H100e + +Chip declaration result \- ai-2040.com + +### June - July 2029: Construction of new highly secure (SL5) inference-only clusters begins + +As part of the transparency regime, the plan is to do broad public deployment of approved frontier models, but under SL5 model weights security achieved through SL5 output verification and distillation mitigations. The plan is to offer a mixture of [Zero Data Retention](https://code.claude.com/docs/en/zero-data-retention) (ZDR) inference for consumers, and transparent inference for higher-risk deployments. More on this in our [transparency supplement](/supplements/transparency-plan). We think these clusters probably need to be built from scratch to reach SL5 model weights security and verification integrity. + +Zero Data Retention (ZDR) inference + +In order to balance the benefits of ZDR (AI users having strong privacy guarantees) with the risks (AI being used for existentially dangerous applications), we propose 3 tiers of privacy offerings for inference during the deal: + +**Total ZDR** + +Probability of data deletion + +100% + +Mitigations + +There are very aggressive classifiers, refusals, anti-jailbreaking and models that are not capable in high-risk domains (e.g., AI R&D). + +Finetuning API? + +No + +Inference-fleet cap (global) + +100K H100e + +**Probabilistic ZDR** + +Probability of data deletion + +99.9% + +(Random sampling, goes to transparent public database if flagged as high-risk by classifiers) + +Mitigations + +Random sampling that goes to classifiers and to transparent public database if flagged as high-risk + +Finetuning API? + +Yes + +Inference-fleet cap (global) + +100M H100e + +**No ZDR** + +Probability of data deletion + +Variable based on classifiers (Random sampling, goes to transparent public database if flagged as high-risk by classifiers) + +Mitigations + +Normal classifier, refusals, unlearning etc. mitigations, and compression / distillation mitigations. + +Finetuning API? + +Yes + +Inference-fleet cap (global) + +100B H100e + +| Probability of data deletion| Mitigations| Finetuning API?| Inference-fleet cap (global) +---|---|---|---|--- +**Total ZDR**| 100%| There are very aggressive classifiers, refusals, anti-jailbreaking and models that are not capable in high-risk domains (e.g., AI R&D).| No| 100K H100e +**Probabilistic ZDR**| 99.9%(Random sampling, goes to transparent public database if flagged as high-risk by classifiers)| Random sampling that goes to classifiers and to transparent public database if flagged as high-risk| Yes| 100M H100e +**No ZDR**| Variable based on classifiers (Random sampling, goes to transparent public database if flagged as high-risk by classifiers)| Normal classifier, refusals, unlearning etc. mitigations, and compression / distillation mitigations.| Yes| 100B H100e + +### August - September 2029: Inference-only ready on 95% of world compute + +With several of the largest compute owners joining the deal over the summer, 95% of the compute in the world is now under inference-only verification, and it has now scaled to higher granularity (smaller inference units) and improved physical security (side channel reductions and high effort physical guarding and monitoring). + +### Late 2029: Pilot R&D verification begins coming online + +In both Canada and Mongolia, mutually constructed datacenters come online at around 1GW in Canada and 2GW scale in Mongolia by September (around 2% of world compute) scaling to around 15% of global compute by the end of the year. Companies are allowed to start hosting compute in these clusters and bid over space as it comes online. Initially, only very small experiments get approved and there is red teaming and iteration on the verification and security measures to test and improve their efficacy. + +An overview of possible verification approaches. + +Existing work on AI verification ([Baker et. al.](https://arxiv.org/abs/2507.15916), [Scher et. al.](https://arxiv.org/abs/2506.15867), [Harack et. al.](https://aigi.ox.ac.uk/publications/verification-for-international-ai-governance/), [Cankaya et. al.](https://aigi.ox.ac.uk/wp-content/uploads/2026/04/Fingerprinting_All_AI_Cluster_IO.pdf)) introduces many concepts and proposals that could be leveraged or developed to perform R&D verification. + +One way we propose to think about the verification problem is to divide it in two parts: (1) collecting evidence, and (2) proving something based on the evidence. This can range from being very simple (e.g., the power in all known compute facilities is off, therefore there is no AI R&D happening) to sophisticated (e.g., based on cryptographic evidence, with extremely high probability the training run correctly used the approved training data and methods). + +Here’s a non-exhaustive illustration of some of the evidence collection and verification methods that could be used as we currently see them: + +![](/verification-plan/image11.png) + +Five key verification metrics we care about include: + + * **Assurance curve.** The assurance curve captures information about the level of **coverage** (the % of the AI compute that the verification solution makes claims about) and **confidence** (with what probability is the verification solution correct about the claims it makes on the covered resources). We define an assurance curve as a set of relevant (coverage, confidence) values. We care about the entire 2D curve, because it's important to know where handoff and covert project risks cross over. + + + + +8 nines7 nines6 nines5 nines99.99%99.9%99%90%0%090%99%99.9%99.99%5 nines6 nines7 nines8 ninesCoverage (log scale)ConfidenceN_ver = 100N_ver = 10KN_ver = 10M + +Assurance curves \- ai-2040.com + + * **Execution difficulty.** How difficult would it be to implement the verification solution correctly? For example, if there are load-bearing parts of the verification solution that require a large amount of expert human labor for e.g., setting up the verification software, manufacturing and installing devices, physical security enforcement and monitoring, etc. + + * **Preparation required.** How much research time is required to implement the verification solution. For a solution that could be implemented with existing intelligence capacity, the preparation would be extremely low, conversely, for a solution that requires sophisticated on-chip mechanisms the preparation requirement would be extremely high. + + * **Privacy level.** This describes to what extent the verification proposal requires private information to be disclosed between the prover and verifier. We aren’t overly worried about this because we are independently recommending increased transparency anyway, and we also think there can be security mitigations (especially in the case of information we don’t want to diffuse, like model weights). Nonetheless, it would still be ideal for verification to not require low level access (e.g. tokens, model weights, algos). + + * **Economic cost.** Finally, this is a metric for how much overall economic value is lost by implementing the verification solution, ignoring the fact that due to the intervention there may have been risks averted with enormous economic downside (e.g., AI catastrophes, human extinction, etc.) So we are conditioning on a ‘good world’ with continued AI development, and seeing how much economic value is lost compared to that counterfactual. The cost of actually developing, manufacturing, and/or implementing the verification solution is very likely to be dominated by the consideration of how much AI inference is allowed to continue and resulting downstream economic benefits which are harder to quantify. + + * **Scope.** For example, does the verification solution just check for inference-only, or arbitrary AI workloads including large training runs and experiments. + + + + +What verification solution to choose will likely have complex tradeoffs between these metrics. Our current best guess is to primarily rely on on-path information collection (i.e., network taps) and then partial recomputation as the verification algorithm, since it seems most likely to be robustly scalable to extremely high coverage and confidence. + +Key ingredients to any verification approach that we will assume are pursued. + +Drawing on previous AI verification work ([Baker et. al.](https://arxiv.org/abs/2507.15916), [Scher et. al.](https://arxiv.org/abs/2506.15867), [Harack et. al.](https://aigi.ox.ac.uk/publications/verification-for-international-ai-governance/), [Cankaya et. al.](https://aigi.ox.ac.uk/wp-content/uploads/2026/04/Fingerprinting_All_AI_Cluster_IO.pdf)) we think there should be many layers of defense in any verification approach. In other parts of this supplement we typically focus on the verification approach that we think can do the heaviest lifting, but we think many other measures will also be necessary or at the very least helpful, including: + + * **Physical security.** Any verification approach is likely going to rely on some set of invariants about the physical state of the AI computing infrastructure. Several promising future directions and existing methods include tamper-evident enclosures, security cameras, perimeter controls, and airgapping. Inference-only enforcement will very likely need confidence that hardware hasn't been significantly tampered with, especially if the verification involves a physical retrofit (which we currently think is the most promising path, though a software-only approach isn't ruled out). The good news is that the physical security requirements for verification overlap heavily with the security posture we'd independently want datacenters to have. + + * **Human-based methods.** Whistleblower programs, interviews of personnel, and embedded auditors could go a long way in uncovering large-scale violations, especially early on. We don't rely on these as load-bearing in any of our proposals, but we see them being very useful supplements, particularly during the early stages before more robust technical mechanisms are ready to come online. + + * **Mutually trusted manufacturing and supply chain security measures.** Many verification mechanisms will depend on hardware, e.g., network taps, recomputation servers, tamper-evident enclosures, etc., actually being what they claim to be. This creates a desire for mutual verification of the manufacturing and supply chain for verification hardware itself. Harack et al. note that off-chip mechanisms in particular might be mutually verified through non-invasive downstream tests or cooperative production at trailing-node fabs. There's also the promising idea of sidestepping the mutual trust problem in some areas with synchronized unilateral checks: [Cankaya](https://nacicankaya.substack.com/p/catching-misreporting-about-ml-hardware-bd2) proposes using passive optical fiber splitters (no digital logic, just fused glass) so both parties can independently observe network traffic without needing to trust the other's hardware, combined with unilaterally trusted devices to independently implement their own verification. + + + + +### Early 2030: Full R&D verification underway + +There are now around 100M H100-equivalents in the verified R&D clusters in either Canada or Mongolia (which is about 20% of world compute at the time and equal to 35% of the compute that existed at the start of the deal). These clusters have the [transparency measures](/supplements/transparency-plan) in place and now companies are approved to do the first major post-deal training runs using combined algorithms. In a few months these training runs complete and several of the models are around the automated coder milestone (AC). The models get broadly tested and red-teamed, including by many third party groups and auditors that rent small amounts of compute in the R&D clusters for this purpose. + +**There is now a big regulatory question of how to control the speed of R &D progress to balance the risks (of going too quickly) and benefits (of improving capabilities).** This will arguably be the hardest part of the deal to execute on competently. There is a complex set of dependencies, where going too fast incurs AI takeover risk and/or causes progress to [diffuse to potential covert projects](/supplements/covert-ai-projects#section-4-limiting-consortium-uplift-to-covert-projects) (reducing the slack the legal project has), while going too slow incurs risk of making less progress on alignment before the deal possibly breaks down, or indeed actually losing to an illegal project operating outside of the deal. + +The high level capability schedule we recommend is to scale to the maximum capability you can reliably make control-based safety-cases for, ideally mostly with hardware scaling, and then pause until your marginal progress on alignment-based safety-cases (handoff risk) is lower than the risk being incurred by the deal. + +But even with a desired high level capability schedule in mind it may be hard to actually enforce this level of capabilities progress, especially to the extent that capabilities progress is unpredictable from inputs (e.g., compute), hard to measure, or fast paced by default. Our best guess is that harsh compute caps, and then a mixture of quality ad-hoc rules will be sufficient, as explained in more detail in the below box. + +Research Titration: How to control the speed of R&D progress? + +This part of the plan ultimately relies on regulatory competence on what is likely to be a thorny, complex, unpredictable problem of deciding what speed is the correct speed to titrate research to. This is therefore a weaker part of the plan that we are worried about, and hope that future work could improve on the ideas of how to navigate it well. Ultimately, this is one of the core reasons that we recommend Total Research Transparency, because we think it can help there be more accurate and effective regulatory decisions. We do think there are some promising approaches that tradeoff between requiring less regulatory competence, and being a more accurate proxy for the research titration goal being attempted. These two concepts we’ll refer to as: + + 1. **Ease of implementation.** How easy it is to implement correctly and easy it is to enforce. + + 2. **Regulatory accuracy.** How likely it is that the regulation actually achieves the goal, i.e., under correct implementation and enforcement, software research goes at a safe speed with high probability. + + + + +We highlight a few approaches below, and roughly where we think they fall on this tradeoff. + + 1. **Safety case burden of proof.** all R&D requires strong safety case arguments before approval. This approach is almost definitionally what we want from a regulatory accuracy perspective, but the difficulty of distinguishing a good from bad safety case might be extremely hard. It also might be very difficult to make good safety cases in the first place, especially without any AI assistance, which is one reason why this approach may only become more viable as the deal goes on. The importance of the regulatory accuracy will also become more important though as the deal progresses, and AIs are reaching higher capabilities, so we suggest this be the regime used for mid-late stages of the deal where highly capable (e.g., TED-AIs) have been developed, or are close to being developed. + + 2. **Quality ad-hoc rules.** case by case decisions on permitted research directions and speed. This approach seems like it could be slightly easier to execute on than having to correctly diagnose safety cases, which may be extremely complex (or just too hard to even generate), and instead being able to make high-level calls about paradigms to pursue or abandon and levers to pull in terms of resource allocation. On the other hand, this approach likely loses out on some accuracy, as it might be that high level ad-hoc rules have wiggle room or loopholes that allow unsafe development. + + 3. **Human-interpretable requirement.** all approved techniques need to be understood end to end by some group of humans. Forcing human comprehension to be a bottleneck to research progress might be quite a good proxy for safety, and quite easy to implement (have humans look at random samples). That being said, it seems very likely possible for something to be both unsafe, and interpretable by a human, so it may be unstable to rely on this method for very long. It’s also plausible though that this actually slows things down too much, and there should either be a relaxed version or another method used. This method seems potentially very suitable for hardware research titration, more on this below. + + 4. **Compute and resource caps.** simple caps on R&D compute. This is extremely easy to implement and verify, but it might just be very inaccurate. Knowing how much progress will result from certain compute and other resource (e.g., human researcher) thresholds will be an uncertain modelling question, that likely will have wide error bars. In the early stages of a deal though, setting these very low seems like a good, easy to implement starter, which can then be escalated to more accurate approaches given time to ramp capacity. + + + + +HardEasyEase of ImplementationHow little regulatory competence is requiredHighLowRegulatory AccuracyHow well the approach tracks the correct research speedSuggested progression~2035Safety case burden of proofAll R&D requires strong safety casearguments before approval.Quality ad-hoc rulesCase-by-case decisions on permittedresearch directions and speed.Human-interpretable requirementAll approved techniques need to beunderstood end-to-end by some group ofhumans.~2030Compute capsSimple caps on R&D compute. + +Research Titration Approaches \- ai-2040.com + +With a particular set of rules on research in place, there is a separate problem of enforcing these rules are actually followed. We can divide the problem of enforcing the R&D rules into two parts: + + 1. **Workload Approval.** Workloads that companies declare actually meet the regulatory rules, whatever their implementation may be. + + 2. **Workload Verification.** Workloads that the companies declare match what they actually run. + + + + +Our view is that workload verification is solvable in a similar way to what we discussed previously (evidence collection via network taps, and then partial recomputation of randomly selected reproducible packets), whereas workload approval will be closely tied to the form of AI R&D rules that are in place (e.g., are they just experiment compute caps? Are there specific banned techniques? Are there specific required techniques?) and our overall baseline proposal for approval is for this to just be manually carried out by teams of auditors in the R&D clusters. + +Workload Approval. + +The core difficulty here is that it might be unclear whether the workload follows the rules, especially as the rules are more ad-hoc. Also, it might be possible for AI companies to do adversarial attacks, like trying to encode a non-compliant workload inside a workload that looks compliant on the surface. As the regime progresses and they have more and more compute under mutual verification, attacks like this that exploit tiny percentages of the total compute for a non-compliant purpose might become significant. + +Possible mitigations we can think of are: + + 1. Anti-steganography measures on the declared workloads and data + + 2. Auditor compilation into gpu instructions, ideally with increasingly robust and secure programming languages. + + 3. AI assisted data filtering and checks + + + + +A potential overall regime to use is to maintain ‘white’ lists of code and data that has been vetted and approved to be used, and/or ‘black’ lists of code/data that is banned, and/or ‘gold’ lists of techniques that are required on every workload of a certain type. Then the workload approval problem can be reduced to simply checking against these lists in order to approve or deny a workload. It might be drastically harder to implement a covertly rogue workload using only whitelisted code, but it also might just punt the problem to whatever the process is for populating these lists. At least this might help with making workload approval a more transparent process that companies and even the broader public can contribute to (at least on everything that is publicly available, which later into the deal might exclude some internally discovered algorithms that there is an attempt to keep secure through the embargo mechanism we describe in our transparency proposal). + +![](/verification-plan/image9.png) + +Workload Verification. + +We also mentioned multiple different types of evidence collection or verification that might be used. Our guess is that in order to reach extremely high levels of assurance, which will be increasingly necessary in Plan A (more on this later), a reproducibility mandate, combined with complete (and increasingly granular) on-path evidence collection, paired with a partial recomputation based verification algorithm, will be a viable approach. That being said, we are not confident that another approach (e.g., using low latency challenges and validity checks) might not be able to reach the same levels of assurance at lower costs or higher levels of assurance at equal costs. For the purposes of making the case that the verification problem will be viable, we will talk about the complete (and increasingly granular) on-path evidence collection with partial recomputation based verification algorithm. + +Another useful concept for thinking about the verification problem is to think of AI workloads as discrete series of steps. + +Full training runTraining phasesGradient stepsLayerforward/backwardGPU kernel callsIndividualinstructionscoarsefineEach chunk: f(input) → output + +Example Packet Sizes \- ai-2040.com + +By default, these computational steps execute with randomness, but if they can be forced to be reproducible (e.g., seeded randomness is fine), then we can have a predictable relationship between the granularity of the evidence collection and the assurance level of the verification algorithm. This is because partial recomputation with random sampling with some small percentage of the workload being recomputed has favourable scaling of detection probability as you decrease the size of the individual workload packets that you check the correctness of. More detail on this in the [appendix](/supplements/verification-plan#appendix-packet-based-verification). + +**Pool:** ~33B H100e (2034 buildout) + +**Recomputation budget:** 1% + +**Year:** 2034 + +2030203220342040 + +**Packet size:** 100 H100e-hours + +1101001K10K + +1101001K10K100K1M10M050%90%99%99.9%99.99%Rogue deployment size (H100e, log scale)P(detected)within 1 hourwithin 1 weekwithin 1 month + +Catching a rogue internal deployment \- ai-2040.com + +The concrete regime we imagine for the first few years of the deal is to still use network taps for evidence collection, and partial recomputation of randomly selected reproducible packets to verify correctness. The same high level approach as inference-only, with smaller and smaller packet sizes over time by moving the network taps to lower levels of the backend network. + +![](/verification-plan/image8.png) + + 1. **Declaration:** AI companies declare the workload they want to run with full transparency on the code, algorithms, etc. + + 2. **Approval:** Auditors run checks per the **Workload Approval** implementation and if it passes they send back approval and provide necessary random seeds (removes a potential attack vector from AI companies gaming fake randomness). + + 3. **Evidence collection:** Auditors collect evidence from the workload (e.g., complete logging flow through network taps). In the diagram the evidence collection is depicted as low granularity, but in reality it could be happening at the level of individual servers, or GPUs or at even lower levels later on. + + 4. **Verification:** Auditors sample randomly and do partial recomputation and other checks to catch violations of workload correctness. + + 5. **Evaluations and release:** For workloads that consist of final training runs there can be an extra step that involves running evaluations that also get verified for correctness, and then final model approval for release leads to the model weights being downloaded from the opaque internal database and transported to the inference-only clusters. This process is extremely high stakes from a model weights theft perspective, so would need to be carefully designed and executed. + + + + +### Late 2030: Secure inference-only clusters for post-deal model deployment come online + +First SL5 inference-only clusters are online and start serving the first post-deal models that have been approved from the R&D clusters for broad public deployment. These clusters have the same basic verification approach as the retrofitted ones (network taps and partial recomputation) but with improved security properties. By the end of the year, around 40% of world compute (~370M of the global ~930M H100e) is in these new post-deal SL5 inference-only clusters. + +Also by the end of 2030, the US and China have set up hardened cold storage facilities where they allow inspectors to verify each other’s stash of AI chips, sized to be bigger than any plausible covert projects (a 5M H100e US bank and a 500K H100e China bank, per our [deal decline supplement](/supplements/deal-decline#preparing-stockpiles-to-align-incentives)), and of relative size to each other similar to pre-deal status quo. These are in places that either side can easily defend. These locations will also be home to verified model weights stashes, so that models that the US and China agree would be positive sum to approve for use in the case of deal dissolution (e.g., models that would be very helpful for safety research and have strong safety cases). Around 2033 there is a model around the automated coder capability milestone that gets approved for the weights stash, and then in 2037 there is a near-frontier (top expert AI) model approved for the cold weights storage. + +## 2031-2036+ Improving the deal + +### 2031: Restrictions for hardware R&D and other high-risk domains + +The core problem the US-China deal sets out to address is the intelligence explosion. But the intelligence explosion isn’t the only threat to deal with. Another core concern is that AI could enable an industrial explosion, too much AI hardware progress, or other dangerous scientific progress. + +The case for being worried about AI software progress was illustrated in AI 2027: superhuman AI researchers could bootstrap to a superintelligence that we might not be able to control and whose goals we haven’t figured out how to align. The case for being worried about too much hardware progress is superseded by this under a fast takeoff, but under the Plan A deal if we only slow down software progress a hardware driven takeover might actually be the majority of the takeover risk absent separate mitigations. + +Three major reasons to be worried about unconstrained hardware progress: + +**Production explosion effects** + +We should be worried about an explosion in the industrial capacity (e.g., number of robots) on earth, even absent design improvements. + +With existing robot designs, it might become easy to rapidly build an absurd number of robots, by having new robots work on making more robots, naively the doubling times might be as low as months, which is very fast. The likelihood of an industrial explosion is argued in more detail in our [economics supplement](/supplements/economics-of-plan-a#core-arguments-for-explosive-growth). + +An absurd number of robots might be able to do destabilizing things like: + +(1) Create massive military advantages for first movers (either directly in the form of a robot army, or through pure industrial capacity, e.g., missile/weapons production capacity), + +(2) Build AI hardware covertly (by building all the equipment & facility needed to do so from the ground up in a secret location) + +**Destabilizing robot designs** + +We should also be worried about new advanced designs, e.g., self-replicating microbot/nanobot/biotech. + +Biological anchors (and precedents of humans beating biology with intelligent designs) should have us worried that scary robot designs are possible. Specifically, we should be worried about self-replicating nanotech / biotech that have very fast doubling times. + +The core concern here is that these designs, unlike macroscale robots, might be too hard to govern, and break any regulations once they are invented. + +**Destabilizing AI compute designs** + +We should also be worried about AI hardware design improvements that make it easier to manufacture. + +AI hardware design improvements and/or paradigm shifts might make it easy to defect from the deal. If there is a paradigm shift in hardware design making it no longer require an extremely specialized, complex, and centralized manufacturing supply chain to make AI compute, it might be tractable for covert projects to manufacture more AI compute in an undetected way threatening the stability of the deal. + +Similarly we want other high risk scientific domains such as biology research to have R&D restricted. Our default proposal for this is to rely on inference refusals and monitoring on the majority of public deployment, and then potentially set up dedicated locations that allow high-stakes hardware and other scientific research to be carried out in mutually verified research labs, where we make some (not necessarily identical) security-transparency tradeoff choices, and mutually verify that the companies follow the titration rules on hardware R&D, which can also be some combination of safety cases, ad-hoc rules, human interpretability requirements and resource caps (similar to software research). + +Inference restrictions, monitoring and refusals for high-risk domains. + +One downside of publicly deploying highly capable models is that they could enable dangerous research, or illegal projects could use those models to help them with AI research, or directly distill the outputs of these models to help train their models. In order for broad inference deployment to not lead to distillation, illegal project uplift, or other dangerous research, there needs to not only be correct output verification, but also some sufficient combination of refusals, monitoring, and compression. + + 1. **Refusals or capabilities shaping.** + + 1. AIs that are released publicly by the Consortium should be bad at AI research. They could be designed to be worse at ML engineering than other tasks, to make it harder for rogue projects to use them. This has the downside of being worse for speeding up diffused safety research. Robust refusals, instead of a lack of capability, may be another way to achieve this goal. + + 2. **Monitoring.** + + 1. On top of inference correctness verification, there should be monitors (probably passive rather than active is fine) running on the models, checking to see if they seem to be involved in doing AI research, large scale training, or other dangerous research. It may be tractable to detect these cases quite reliably, for example, if they are operating in a codebase that involves code for training over a large number of GPUs, it is pretty obvious that this is an illegal usage. + + 3. **Compression.** + + 1. Never show any reasoning intermediates to users by default, and maximally compress outputs to summaries up to some small quota. + + 2. Vertically integrate whenever possible at multiple levels and just produce outputs: + + + + +We currently think these are tractable mitigations that will allow a lot of broad deployment, but are uncertain. There should be ongoing testing for distillation and detecting dangerous research. + +Hardware research restrictions. + + 1. **Setting rules on hardware research** + + 1. One specific subset of the rules might be to only allow research into traditional-like robot designs, traditional-like chip designs, since much of the risks from hardware research are in destabilizing paradigm shifting scientific breakthroughs that might lead to things like nanotechnology or far easier-to-manufacture AI compute. And traditional-like designs might still be improvable enough to realize the benefits of hardware and robot buildups under the deal. That being said, if they are not improvable enough, then it might be more advisable to make more relaxed hardware research restrictions. + + 2. In hardware we might have an unclear idea of the relationship between inputs and progress, just like we are also very uncertain about the relationship between compute and software progress. Nevertheless, we can use the same hope that by operating over harsh order of magnitude reductions in inputs (e.g., energy caps, materials caps, research labor caps, etc.) we might be able to iteratively slow the pace of hardware progress to levels perceived to be safe in a lossy but easy to implement way. + + 2. **Enforcing rules on hardware research.** + + 1. Once the hardware research titration rules are set, we need a way to verify they are being followed. One possibility would be to use the same high level approach as for software enforcement, i.e., approve hardware research jobs case by case (or via some automated whitelist/blacklist process), divide up jobs into substeps and collect evidence about them, and then check random samples for compliance. One problem with this approach in the hardware case is that it seems far harder to get high reproducibility in e.g., a chemistry experiment, because two labs in the real world are not as homogenous as AI chips. This opens the door for plausible deniability and violations. Therefore, we probably need to not go with a reproducing based approach on everything. That being said, there might be many slices of the hardware research that are reproducible. + + 2. Another possibility is that since many things might be macroscale, you can monitor everything directly with a monitoring scheme that you trust. Maybe the trusted monitor will be hard to build though, so instead, you might be able to instead have humans inspect random samples of all the recordings, and benefit from the favorable detection probability scaling from random sampling. + + + + +Overall, we are very uncertain about how to verify hardware research rules, and have the current sense that it might be harder than software (but also think it's plausibly easier). That being said the default speed and risk level, and therefore the overall slowdown required we think is likely to be lower. + +### 2032: Cap & Trade for Robots and Compute + +As mentioned earlier, we want to stop an explosion in the number of robots and AI chips on earth from threatening the stability of the deal. By default, our [current naive modelling](/supplements/economics-of-plan-a#the-full-economic-model) predicts that even with current hardware designs, the AI and robot capabilities that we forecast to be reached during the Plan A scenario would lead to a huge discrepancy between the raw cost of producing more AI chips and Robots and their economic value. By default, we think this leads to a reinvestment spiral and very fast doubling times. + +In Plan A we propose a cap starting in 2032 that limits the number of robots and AI chips on earth to doubling times around 6 months, through cap & trade regulation, where private companies need to bid on the finite robot and AI chip production permits in any given time period. This cap should be set in units of usefulness of the robots for industrial production and AI compute in AI workloads, measuring both of these, especially across tasks and form factors in the case of robots, might be difficult and muddied (e.g., fuzzy lines between robots and more traditional machines). We nonetheless think it should be feasible to measure these things reasonably well for different robot and hardware designs, and allocate the cap & trade limits accordingly. + +Actually enforcing the cap & trade permit rules might be very difficult if robot and/or AI compute production happens all over the world, so we propose centralizing new built up industrial capacity and robots in special economic zones (SEZs) to make for easier enforcement of the cap & trade rules. + +How to enforce the cap & trade rules. + +Our proposal for this is to centralize all the built up industrial capacity and robots in special economic zones (SEZs), to centralize the enforcement and associated surveillance in these restricted zones, and allow the rest of the economy to not be subjected to any hardware verification checks. + +The SEZs can then have a similar high level verification approach to either hardware research and/or the inference-only and software R&D approach, where production gets broken up into discrete steps, and these get randomly, manually audited by humans (and/or with trusted AI monitoring assistance). + +The rest of the economy can then be left unverified, but should probably have upper bounds on the robots, power and other potential inputs to hardware production. If these bounds are low, then there can be minimal surveillance outside of the SEZs. That being said, we guess the bounds might be too costly to keep sufficiently low because people really want diffused robots. So it might be good to have reliable geolocation on the diffused robots and maybe random inspections. In order to avoid too much diffused compute as well, we also propose that robots have very little unverified edge compute, and instead largely be served via the cloud on the monitored (ideally automated by AIs to be privacy preserving) and verified inference-only clusters, until there is robust edge compute verification. + +The same cap and trade principle is extended to other products and services that are destabilizing in unrestricted quantities, but can be safe in lower quantities. Two major examples are products with edge compute and unmonitored inference. Things that might particularly rely on edge compute include self-driving cars and consumer compute, while unmonitored inference might be particularly valuable to governments for military operations. Under cap and trade the regime can accommodate these activities being legal, under safe societal-scale caps that maintain the stability of the deal. Then there can be markets for the permits the people, firms and the government can bid between each other to have the rights for. + +Unverified edge compute cap. + +**We propose a cap of 30M effective-H100e of unverified edge compute being allowed to exist in the world, with estimates of 25M[effective](/supplements/compute-supplement#appendix-the-effective-h100e-model) H100e predating the deal, there are only 5M credits available for new unverified production.** **With sufficient edge verification measures, the compute can be exempt.** + +Unrestricted, edge compute in consumer devices, self-driving cars and robots might introduce a significant stock of compute that could be used by illegal projects. Before the deal consumer and edge compute likely will include some AI-relevant compute, (e.g., hobbyists with single GPUs, self-driving car chips), and mostly non-AI relevant compute (chips with low effectiveness on AI workloads), but which is nonetheless very numerous (i.e., there are tens of billions of phones, each around 0.1 - 1% as useful as an H100 for AI workloads, which still means tens of millions of H100-equivalents, i.e., ~10% of the total AI-relevant compute in the world). In our [compute supplement](/supplements/compute-supplement#12-all-compute), we estimate that this will be around 25M **[effective](/supplements/compute-supplement#appendix-the-effective-h100e-model)** H100e. + +This should be estimated more precisely at the time with a better model of usefulness in frontier AI workloads. If consumer compute usefulness is found to be higher than we estimate, then it becomes more important to avoid that large quantities of consumer compute could be rounded up and used in a covert project, and therefore e.g., a high-priority compute buy-back effort might be warranted. + +We expect the cap we propose to lead to rapid innovation in ‘edge verification’ (i.e., essentially robust hardware mechanisms), to avoid the permit costs. We predict that there will be plenty of cheap phones that have low local-AI capabilities, whereas phones or other consumer devices that have a lot of local, or ‘edge’ AI compute will be very expensive at first, but don’t need to be banned. Then after a few years there will be robust edge verification that can allow powerful edge compute to also become cheap (not require a permit), for example, with hardcoded code and weights. The cap is set such that trivially noticeable shares of the consumer compute would need to be diverted to noticeably accelerate a covert project (or to give bad actors too much of a weapon). It’s a sad but potentially necessary reality for at least at the beginning of Plan A that if you want to have a fancy gaming GPU that could also be used for AI instead, you need to pay a premium (because we can’t let everyone have one or else it would become too easy to round up a significant number of these to contribute to a rogue compute cluster). Or in economist terms, unverified edge compute carries a large negative externality (risk of being used by a covert project). + +On top of the unverified consumer compute cap, it might be desirable to have a certain effective-H100e cutoff (e.g., with specific compute, bandwidth, memory capacity and networking speed thresholds) for consumer compute that is allowed to be sold in single units (to the extent that the difficulty of gathering chips into a covert project or detection therefore scales more closely with the absolute number of chips than with the compute). There might also be a tier of compute that we require to have location tracking measures (to make use by a covert project easier to detect). + +AI-relevant floor: 4,000 TPP ≈ 0.25 H100e (compute supplement §1.1)Consumer computebelow the AI-relevant floor — capped at 30M H100eiPhone 16Pro MaxTesla HW3DGX SparkM4 Pro MacBookRTX 4090H20RTX 5090H100 (PCIe)DGX 8×H100GB200 NVL720.1 GB/s1101001 TB/s10 TB/s0.0010.010.11101001KCross-Chip Interconnect Speed (GB/s)Compute per device (H100e)Memory capacity8 GB128 GB13 TBMemory bandwidth≥ ~1.5 TB/s (HBM-class)below ~1.5 TB/sUnverified edge-compute cap30M H100e25M5M25M already in world (pre-deal)\+ 5M new creditsVerified edge compute → exempt + +Chip Flow Restrictions, 2032 \- ai-2040.com + +Putting together research titration and production capping, we have the following overview: + +![](/verification-plan/image13.png) + +### 2033: Security improvements + +By 2033 we expect significant hardware security, and more general security properties of the datacenters, verification regime and capabilities to have improved significantly. By the end of 2033, around 99.9% of the worlds’ compute is post-deal compute, meaning that it's trivial to just power down all pre-deal clusters and compute, and have the vast majority of compute be in new more secure locations built from scratch. Furthermore, by the end of 2034, around 90% of the compute that exists will have been produced during 2033 or 2034. This means that security improvements made by then can already be widely in place within just 1-2 years. Similar dynamics apply to the fabrication facilities. + +Our uncertainty on what will be possible is high here, but some of the main ideas we have include: + + * Datacenters built from scratch with extremely good side channel and airgapping properties. + + * Huge progress in standardizing (for reproducibility) and improving security of software stacks for inference and R&D workloads more generally. + + * New compute production facilities (fabs) built from scratch with mutual monitoring measures and being colocated with supply chain and potentially even with final centralized datacenters they will fill to minimize attack surface. + + * The beginning of some hardware-software co-dependencies especially with architecture changes or number formats, so that optimization and progress in these paradigms is less applicable on pre-deal compute + + * Increases to the logging / evidence collection granularity of the verification regime, especially of the R&D datacenters, with logging technology being compatible with extremely high bandwidth data paths on backend networking to reduce overhead. + + * Generally improved hardware security properties with more robust trusted execution, secure boot, encryption, reproducibility, etc., which are especially important for recomputation servers. Some degree of standardization in hardware may also be helpful for reproducibility. + + * Special hardware designs for automatic on-device verification of certain workloads, in particular, hardcoded weights on some chips to verify they're inference-only. This might be particularly useful for ‘edge verification’ allowing for more AI compute in consumer products without risk of use in covert projects. + + + + +### 2034: Verification improvements + +Through the 2030s the total amount of AI compute in the world grows explosively. One important dynamic to note is that the verification burden (the need to have higher coverage) grows as the total compute under verification grows. In Plan A, we both forecast and recommend a significant hardware buildup (capped by the cap & trade regime to stay verifiable) for other reasons related to tradeoffs with deal stability (hardware is less likely to leak to illegal projects so differentially increases the legal projects affordances). + +202620M2030500M20323B203460B + +World Compute \- ai-2040.com + +The verification assurance curve needs to improve such that any potential nonnegligible rogue internal deployment would have a high chance of detection. By 2034, with one packet collected per hour from each frontier GPU (~100 H100e each, so ~100 H100e-hour packets) and a 1% recomputation budget, the appendix math shows the verifier catches any rogue internal deployment at 99% confidence after it accumulates ~46,000 H100e-hours of unapproved work, so a sustained 10K H100e deployment is detected within ~4.6 hours, a 1K within ~46 hours, and so on. + +2030Per-server tap~4K H100e per server2032Per-shelf tap~400 H100e per shelf2034Per-GPU tap~100 H100e per GPU + +Workload: ~GPT-3 sized training run (100K H100e-hours) + +Verification logging granularity \- ai-2040.com + +This is an ambitious verification scale, and it also relies on extremely good integrity security on the verification regime. We think it will be possible to scale it this far, especially with AI assistance, but are not confident, and find it plausible that the compute cap should be set lower in order to not increase the verification burden this much. + +**Pool:** ~33B H100e (2034 buildout) + +**Recomputation budget:** 1% + +**Year:** 2034 + +2030203220342040 + +**Packet size:** 100 H100e-hours + +1101001K10K + +1101001K10K100K1M10M050%90%99%99.9%99.99%Rogue deployment size (H100e, log scale)P(detected)within 1 hourwithin 1 weekwithin 1 month + +Catching a rogue internal deployment \- ai-2040.com + +### 2035: Moving the datacenters to the ocean + +One of the biggest ongoing risks is that the deal breaks down and actors return to racing. Because of the vast increase in compute scale that has not been turned into capabilities progress (due to the safety regulations) the return to racing could be even more dangerous than the initial status quo, despite the alignment progress made up until this point, because things would now go faster. This motivates having measures in place to return to pre-deal compute levels, or even smaller compute scales if the deal breaks down. + +We think the relative engineering and logistical costs involved in building new datacenter capacity on the ocean in international waters will be low enough due to robot and AI labor to make this worth it around 2035, because it would become much harder for any particular actor to defend and much less escalatory to destroy in the event of the deal breaking down. This is an intentional measure for mutually assured compute destruction to increase the deal’s stability. + +### 2036+ Advanced robustness improvements + +Later into Plan A we are increasingly uncertain about what robustness improvements will be possible in terms of security, verification and deal stability. Potentially extremely ambitious things will be possible, a brainstorm here included: + + * Completely rewritten software stacks to be significantly or entirely formally verified. + + * Huge red teaming efforts on the verification scheme or security measures using resources that dominate any potential covert actors. + + * Using AI and robot labor to track down potential covert projects in mutually agreed, symmetrical ways, e.g., the ‘AI silver bullets’ explored in our [covert project supplement](/supplements/covert-ai-projects#section-31-plan-a-strategies-for-detection): reliable lie detectors, privacy-preserving AI verification, and extreme industrial verification. + + + + +## Appendix: Packet-Based Verification + +This appendix works through the math on random sampling for use in partial recomputation to show why it scales so favourably for detecting incorrect workloads. A key simplifying assumption we make in this analysis is that the workload packets are perfectly reproducible. + +### A.1 Setup + +In this analysis we treat verification as a sampling problem on a stream of packets. A packet is a reproducible unit of computation that the verifier can recompute: its inputs, instructions, and outputs are all collected at the evidence-collection point (e.g., a network tap). What the computation actually is (one forward pass, one training batch) doesn’t affect the math below. Further simplifying assumptions we make are: + + 1. **Uniform size.** Every packet represents the same amount of underlying compute. Non-uniform packets work too, with a weighted-sampling argument that doesn't change the qualitative results. + + 2. **Independent random sampling.** The verifier audits each packet independently with the same probability. In reality you would sample without replacement. + + + + +Let the verifier's pool be some collection of packets; it could be a single training run, every packet from every workload running under the deal in a given month, or one cluster's traffic. Let NpacketsN_{\text{packets}}Npackets​ be the total number of packets in that pool, s=1/Npacketss = 1 / N_{\text{packets}}s=1/Npackets​ the packet size (the fraction of the pool each packet represents), CCC the recomputation budget (the fraction of the pool the verifier can afford to recompute), and Nverified=C⋅Npackets=C/sN_{\text{verified}} = C \cdot N_{\text{packets}} = C / sNverified​=C⋅Npackets​=C/s the number of packets actually audited. + +### A.2 Detection by Random Sampling + +The detection math in this section was independently derived but matches the analysis in [Rinberg et al. (2025)](https://arxiv.org/abs/2511.02620). Suppose an attacker substitutes NfakeN_{\text{fake}}Nfake​ packets, replacing the declared computation in each with their own unauthorized one. Let + +Ffake=NfakeNpackets F_{\text{fake}} = \frac{N_{\text{fake}}}{N_{\text{packets}}} Ffake​=Npackets​Nfake​​ + +be the fraction of packets that are fake. If the verifier audits each packet independently with probability CCC, the exact probability that none of the NfakeN_{\text{fake}}Nfake​ fakes is audited is + +P(undetected) = (1−C)Nfake. P(\text{undetected}) \;=\; (1 - C)^{N_{\text{fake}}}. P(undetected)=(1−C)Nfake​. + +For small CCC and modestly large NfakeN_{\text{fake}}Nfake​, this is well approximated by a Poisson distribution on the number of caught fakes, with mean λ=Nverified⋅Ffake=C⋅Nfake\lambda = N_{\text{verified}} \cdot F_{\text{fake}} = C \cdot N_{\text{fake}}λ=Nverified​⋅Ffake​=C⋅Nfake​: + +P(undetected) ≈ e−Nverified⋅Ffake, P(\text{undetected}) \;\approx\; e^{-N_{\text{verified}} \cdot F_{\text{fake}}}, P(undetected)≈e−Nverified​⋅Ffake​, + +so the probability of catching the attack (the **detection confidence** at attack size FfakeF_{\text{fake}}Ffake​) is + +P(detected) = 1−P(undetected) ≈ 1−e−Nverified⋅Ffake. P(\text{detected}) \;=\; 1 - P(\text{undetected}) \;\approx\; 1 - e^{-N_{\text{verified}} \cdot F_{\text{fake}}}. P(detected)=1−P(undetected)≈1−e−Nverified​⋅Ffake​. + +We'll use the Poisson form throughout, since it's the cleanest expression of the scaling. It breaks down only when CCC approaches 1 or NfakeN_{\text{fake}}Nfake​ is single-digit, which is also when random sampling stops buying you anything (see the discussion at the end of [§A.3](/supplements/verification-plan#a3-the-assurance-curve)). + +### A.3 The Assurance Curve + +The assurance curve is determined by **coverage** and **confidence** which can be expressed in terms of Ffake∗F^{*}_{\text{fake}}Ffake∗​, the largest fraction of fake packets we're willing to tolerate without detection: + + * **Coverage** = 1−Ffake∗\;=\; 1 - F^{*}_{\text{fake}}=1−Ffake∗​, the fraction of the workload we're claiming is legitimate. + + * **Confidence** = 1−e−Nverified⋅Ffake∗\;=\; 1 - e^{-N_{\text{verified}} \cdot F^{*}_{\text{fake}}}=1−e−Nverified​⋅Ffake∗​, the probability we'd catch any actual violation at or above that threshold. + + + + +The confidence formula comes from [§A.2](/supplements/verification-plan#a2-detection-by-random-sampling): P(detected)=1−e−Nverified⋅FfakeP(\text{detected}) = 1 - e^{-N_{\text{verified}} \cdot F_{\text{fake}}}P(detected)=1−e−Nverified​⋅Ffake​, evaluated at the threshold Ffake∗F^{*}_{\text{fake}}Ffake∗​. The recomputation budget shows up through Nverified=C⋅NpacketsN_{\text{verified}} = C \cdot N_{\text{packets}}Nverified​=C⋅Npackets​, so more audited packets push the curve outward. + +Plotting confidence against coverage for several verification budgets (i.e., different values of NverifiedN_{\text{verified}}Nverified​): + +8 nines7 nines6 nines5 nines99.99%99.9%99%90%0%090%99%99.9%99.99%5 nines6 nines7 nines8 ninesCoverage (log scale)ConfidenceN_ver = 100N_ver = 10KN_ver = 10M + +Assurance curves \- ai-2040.com + +These curves show the confidence-coverage relationship at different values of NverifiedN_{\text{verified}}Nverified​. Higher NverifiedN_{\text{verified}}Nverified​ gives you higher assurance (shifts the curve right). Since Nverified=C/sN_{\text{verified}} = C / sNverified​=C/s, **the curve can be shifted higher either by (1) increasing the recomputation budget CCC or (2) shrinking the packet size sss.** + +When the Poisson approximation breaks down. + +The product C⋅NfakeC \cdot N_{\text{fake}}C⋅Nfake​ (recomputation budget times number of fake packets) is the single quantity that determines 99% catch probability: it just needs to exceed about 4.6. + +This makes the limiting case revealing. Imagine you sample one packet per server-hour, so p=1p = 1p=1 server-hour, and the attack is also just a single fake packet (one rogue server-hour buried in a much longer, say year-long, workload). With Nfake=1N_{\text{fake}} = 1Nfake​=1, the Poisson approximation would say you need C≥4.6C \geq 4.6C≥4.6, which doesn't make sense: CCC is a fraction, capped at 1 (you can't recompute more than 100% of packets). + +This is the regime where the Poisson approximation breaks down, and the right formula is the exact one for sampling without replacement: + +P(catch) = 1−(1−C)Nfake. P(\text{catch}) \;=\; 1 - (1 - C)^{N_{\text{fake}}}. P(catch)=1−(1−C)Nfake​. + +For Nfake=1N_{\text{fake}} = 1Nfake​=1 this gives P(catch)=CP(\text{catch}) = CP(catch)=C. To catch a single fake packet with 99% confidence, you'd need to recompute 99% of all packets, i.e. you've essentially given up on sampling and are doing full verification. The size of the surrounding workload doesn't matter: you're flipping a single coin biased by CCC. + +The Poisson formula is accurate whenever NfakeN_{\text{fake}}Nfake​ is at least a few dozen and CCC is small (say under 10%). For Nfake=100N_{\text{fake}} = 100Nfake​=100, the exact required budget for 99% catch is 4.5%, vs Poisson's 4.6%, which is very close. For Nfake=1000N_{\text{fake}} = 1000Nfake​=1000 the exact answer is 0.459%, vs Poisson's 0.46%. The two agree better as NfakeN_{\text{fake}}Nfake​ grows. Things only diverge in the small-NfakeN_{\text{fake}}Nfake​ corner, which is also exactly the corner where random sampling stops being useful and you have to recompute everything. + +So the practical takeaway: random sampling buys you a lot once you're catching attacks that span at least dozens of packets. To catch attacks below that, shrink packets (decrease ppp) so the same fake compute-work spans more packets. + +### A.4 Worked Examples + +#### A.4.1 100K-token Backdoor in a 100T-token Training Run + +Consider a training run of Ntokens=1014N_{\text{tokens}} = 10^{14}Ntokens​=1014 tokens (100T) using distributed data parallelism. Network taps between server racks log every gradient step (its input batch, the gradient update, and the resulting checkpoint update), so each gradient step is one packet. An attacker wants to insert a backdoor by silently training on Nfake tokens=105N_{\text{fake tokens}} = 10^{5}Nfake tokens​=105 off-spec tokens (100K). + +Quantity| Symbol| Value +---|---|--- +Training tokens| NtokensN_{\text{tokens}}Ntokens​| 101410^{14}1014 +Attack size (backdoor)| Nfake tokensN_{\text{fake tokens}}Nfake tokens​| 10510^{5}105 +Batch size| BBB| design choice +Number of packets| NpacketsN_{\text{packets}}Npackets​| Ntokens/BN_{\text{tokens}} / BNtokens​/B +Packet size| sss| B/NtokensB / N_{\text{tokens}}B/Ntokens​ +Fake fraction| FfakeF_{\text{fake}}Ffake​| Nfake tokens/Ntokens=10−9N_{\text{fake tokens}} / N_{\text{tokens}} = 10^{-9}Nfake tokens​/Ntokens​=10−9 + +For 99.9% detection confidence we need Nverified⋅Ffake≥ln⁡(1000)≈7N_{\text{verified}} \cdot F_{\text{fake}} \geq \ln(1000) \approx 7Nverified​⋅Ffake​≥ln(1000)≈7, i.e. Nverified≥7×109N_{\text{verified}} \geq 7 \times 10^{9}Nverified​≥7×109. The recomputation cost as a fraction of the training run is then + +C = NverifiedNpackets = Nverified⋅s = 7⋅BNfake tokens = 7B105. C \;=\; \frac{N_{\text{verified}}}{N_{\text{packets}}} \;=\; N_{\text{verified}} \cdot s \;=\; \frac{7 \cdot B}{N_{\text{fake tokens}}} \;=\; \frac{7 B}{10^{5}}. C=Npackets​Nverified​​=Nverified​⋅s=Nfake tokens​7⋅B​=1057B​. + +Batch size BBB| NpacketsN_{\text{packets}}Npackets​| Recomputation cost CCC +---|---|--- +10,000| 101010^{10}1010| 70% +1,000| 101110^{11}1011| 7% +100| 101210^{12}1012| 0.7% +10| 101310^{13}1013| 0.07% + +Making batches 10×10\times10× smaller cuts the required recomputation cost 10×10\times10× for the same confidence (ignoring the fact that we do better than this at values of C close to 1 by sampling without replacement). The two practical levers on packet size are (1) smaller batches (more gradient steps per training run), and (2) sub-batch logging, tracking intermediate computations within a batch (e.g., individual forward/backward passes), which might require evidence-collection devices inside server racks or even inside GPUs themselves. + +#### A.4.2 Bounding Rogue Compute Use in 100B H100e of Verified R&D + +The [2034: Verification improvements](/supplements/verification-plan#2034-verification-improvements) section explores what **rogue internal deployments** have low detection probability if the verifier has a 1% recomputation budget across 100B H100e of verified R&D compute. Rogue internal deployments are potentially scarier than external illegal projects due to internal access to frontier models and the rest of the R&D cluster. Detection relies on the verification system catching their fake packets. The math below translates the assurance curve from [§A.3](/supplements/verification-plan#a3-the-assurance-curve) to concrete values. + +Starting from P(detected)=1−e−Nverified⋅FfakeP(\text{detected}) = 1 - e^{-N_{\text{verified}} \cdot F_{\text{fake}}}P(detected)=1−e−Nverified​⋅Ffake​ and substituting Nverified=C⋅NpacketsN_{\text{verified}} = C \cdot N_{\text{packets}}Nverified​=C⋅Npackets​ and Ffake=Nfake/NpacketsF_{\text{fake}} = N_{\text{fake}} / N_{\text{packets}}Ffake​=Nfake​/Npackets​, the NpacketsN_{\text{packets}}Npackets​ cancels and the detection probability becomes: + +P(detected) = 1−e−C⋅Nfake. P(\text{detected}) \;=\; 1 - e^{-C \cdot N_{\text{fake}}}. P(detected)=1−e−C⋅Nfake​. + +In this form the detection probability is a function of the recomputation budget CCC and the absolute number of fake packets NfakeN_{\text{fake}}Nfake​ the attacker generates. + +For 99% confidence we need C⋅Nfake≥ln⁡(100)≈4.6C \cdot N_{\text{fake}} \geq \ln(100) \approx 4.6C⋅Nfake​≥ln(100)≈4.6, so the attacker is caught with 99% confidence the moment they have accumulated + +Nfake99% = 4.6C N_{\text{fake}}^{99\%} \;=\; \frac{4.6}{C} Nfake99%​=C4.6​ + +fake packets. With C=1%C = 1\%C=1% that is 461 packets, independent of pool size and of how the work is spread across workloads. Multiplying by the packet size ppp converts this to a bound on cumulative rogue compute work: + +Wrogue99% = 4.6⋅pC. W_{\text{rogue}}^{99\%} \;=\; \frac{4.6 \cdot p}{C}. Wrogue99%​=C4.6⋅p​. + +Consider the operational setup of the 2034 section, in which the verifier collects one packet per hour from each frontier GPU (about 100 H100e each by 2034). Then p=100p = 100p=100 H100e-hours, and with C=1%C = 1\%C=1%, + +Wrogue99% = 4.6⋅1000.01 = 46,000 H100e-hours. W_{\text{rogue}}^{99\%} \;=\; \frac{4.6 \cdot 100}{0.01} \;=\; 46{,}000 \text{ H100e-hours}. Wrogue99%​=0.014.6⋅100​=46,000 H100e-hours. + +The verifier catches with 99% confidence any rogue internal deployment that accumulates more than about 46K H100e-hours of unapproved compute work inside the verified pool, regardless of pace or distribution. The bound on cumulative rogue work scales linearly in p/Cp / Cp/C; a few representative operating points: + +Packet size ppp| Budget CCC| Max undetected rogue work +---|---|--- +100 H100e-hours| 1%| 46,000 H100e-hours +100 H100e-hours| 0.1%| 460,000 H100e-hours +10 H100e-hours| 1%| 4,600 H100e-hours +1 H100e-hour| 1%| 460 H100e-hours + +1. + +With variable economic costliness and other tradeoffs which we discuss. As a baseline we are >90% confident that a highly costly verification solution is possible for both inference and training (e.g., multi-year verified compute shutdown with verification R&D sprint and secure rebuild). + +2. + +Our [covert project supplement](/supplements/covert-ai-projects) discusses how likely it is for a covert project to be able to pose a threat to the deal or win the AI race. + +3. + +Even this might be difficult to execute on to the extent that at least one side has high uncertainty over the other having a large secret datacenter that they don’t declare and allow the other side to verify is powered off. + +4. + +Another live option that could currently be executed quickly would be to unplug (most) backend networking from datacenters. This might allow a lot of inference to continue (less costly), but also might not slow down AI development robustly (inference is used as rollouts for data to be used in federated learning training run elsewhere). + +5. + +A key reason we believe this is that the verifier only needs to catch one violation, while a prover needs to successfully hide all of their covert computations. This is the reason behind e.g., the favourable scaling of partial recomputation due to random sampling which we explore in our [appendix](/supplements/verification-plan#appendix-packet-based-verification). + +6. + +I.e., doesn’t require compute to be shut down for long. + +7. + +Conditional on the Plan A scenario (the mutual chip declaration and inspection in 2029 carried out, and our chip tracking recommendation executed competently), our [covert project supplement](/supplements/covert-ai-projects#section-2-preventing-covert-compute-procurement) estimates that a competently-executed PRC diversion effort would net a median of 0.51% of world compute (80% CI: 0.11% to 1.4%), implying roughly an 80% chance that total ‘dark compute’ stays under 1% of the world’s AI-relevant compute. + +8. + +We are uncertain about the compute-sensitivity of progress in the takeoff period (automated coder to superintelligence milestones). If we operationalize this (for simplicity) in terms of how much slower you go if you have 10x less compute, we think it's most likely to be somewhere between 2x to 6x slower (see [How much slower does takeoff go with 10x less compute?](https://www.lesswrong.com/posts/7jcPg79p3kD5ir3CL/how-much-slower-does-takeoff-go-with-10-less-compute) and the matching [model](/supplements/covert-ai-projects#how-much-does-takeoff-slow-down-if-you-reduce-compute-by-10x-n) in the covert project supplement). The specific trajectories we expect hypothetical covert projects to follow in the Plan A scenario (including the effects of algorithmic progress leaking from legal projects) are modelled and explained much more carefully in our [covert project supplement](/supplements/covert-ai-projects#how-we-model-the-ai-capabilities-progress-of-covert-projects). + +9. + +By known-to-exist, we are recommending that intelligence agencies know with high confidence at least that the compute exists, and who has it, not necessarily exactly where it is (though that is also helpful), because the first two conditions should be sufficient to make it impossible to exclude from a future mutual chip declaration with plausible deniability. + +10. + +In other words, without this measure in place, one strategy that an actor might pursue is to build the largest covert project they can get away with, and then wait until they make enough alignment progress such that they are confident they could win with their covert project. At this point they could pull out the deal, destroy the post-deal compute, and now their covert project would be the leading world project, with a high chance of success. This strategy is no longer likely to be viable if your adversary has a cold storage compute bank that dominates your covert project. + +11. + +There are tens of billions of phones, each around 0.1 - 1% as useful as an H100 for AI workloads, which still means tens of millions of H100-equivalents, i.e., ~10% of the total AI-relevant compute in the world. + +12. + +This means there will be plenty of cheap phones that have low local-AI capabilities, whereas phones or other consumer devices that have a lot of local, or ‘edge’ AI compute will be very expensive at first until there is robust edge verification. + +13. + +For example, technical details about the networking, e.g., optical speeds, bandwidths, frequencies. + +14. + +For example, maybe we think handoff risk has gone under 10%, and so we want to exit the deal via handoff if there is a greater than 10% risk that a violation using 1% of the verified compute in a given period would be catastrophic (i.e., our curve is under the 90% confidence, 99% coverage point). + +15. + +For example, the disclosure happening within an airgap, which should make it tolerable in most of the relevant cases. + +16. + +The outlook for resumed AI R&D is also important for driving continued AI investment. + +17. + +This is explored more in our [capabilities scaling strategy supplement](/supplements/capability-scaling-strategy). There are also a few deeper layers of nuance here that complicate the picture, including what transparency regime is in place as discussed in our transparency supplement. + +18. + +An R&D efficiency spiral in AI research, where the feedback loop of smarter AIs → better research → smarter AI, holds through to AIs that we cannot control or align, and slowing down to invest in control/alignment is uncompetitive, locking companies and countries into a multipolar trap. Coordination, that is verified, is the path out. + +19. + +We define AI-relevant compute as any computational unit capable of achieving Total Processing Performance (TPP) of at least 4,000 and Performance Density (PD = TPP / die size) of at least 4. This definition is set just below the [A100 SXM GPU](https://www.nvidia.com/en-us/data-center/a100/), NVIDIA's state of the art chip in 2021. + +20. + +Note none of these quantities depend on how the work is divided into individual workloads. What matters is the pool the verifier samples from. If there are 100 different training runs going on simultaneously and the verifier samples uniformly at random across all their packets, then ‘the workload’ in the math below is the union of all 100. The per-workload framing in the worked example ([§A.4.1](/supplements/verification-plan#a41-100k-token-backdoor-in-a-100t-token-training-run)) is a convenience for talking about one concrete attack. diff --git a/sources/resources/INDEX.md b/sources/resources/INDEX.md new file mode 100644 index 0000000..cfd9914 --- /dev/null +++ b/sources/resources/INDEX.md @@ -0,0 +1,32 @@ +# Resources index + + + +## Fetched 2026-07-11 + +- `forethought-ai-enabled-coups.md` | AI-Enabled Coups: How a Small Group Could Use AI to Seize Power | Tom Davidson, Lukas Finnveden, Rose Hadshar, Forethought | 2025-04-15 | Full research paper on how loyal automated workforces, concentrated capabilities, and weakened dependence on citizens could enable coups or gradual democratic backsliding; directly supports the homegrown-autocracy branch. +- `ai-2027.md` | AI 2027 | Kokotajlo, Lifland, Larsen, Dean et al., AI Futures Project | 2025-04-03 | The canonical fast-timeline scenario (superhuman AI ~2027-28, race vs slowdown branch); sets the aggressive end of the timeline distribution any AU/NZ 2040 scenario must be robust to. +- `acx-introducing-plan-a.md` | Introducing Plan A | Scott Alexander, ACX | 2026-07-09 | Insider gloss on Plan A: third countries join as "sort-of-but-not-really-voting observers" trading compliance for hosted datacenters, AI access, and a wealth share; dividend compute-pegged $25k (2033) to $1.6M (2035); the two-monster framing ("ends the world or dooms it to permanent techno-oligarchy"). +- `leicht-roadmap.md` | A Roadmap For AI Middle Powers | Anton Leicht, Threading the Needle | 2025-01-23 | Argues middle powers stay relevant by controlling bottleneck inputs to AI growth (energy, land, capital, niche supply chains) rather than competing at the frontier. +- `leicht-safety.md` | How AI Safety Is Getting Middle Powers Wrong | Anton Leicht | 2026-01-22 | Argues middle-power AI policy should pivot from global-governance advocacy to hard national-interest strategy, reframing what AU/NZ decision points even are. +- `leicht-moonshot.md` | The Moonshot ("Real sovereign AI has never been tried") | Anton Leicht | 2026-06-17 | Uses a hypothetical US export-control shutdown of a frontier model to argue dependence on US models is a live sovereignty risk and genuine sovereign-AI programs are underattempted, a key AU decision branch. +- `cfg-futures.md` | The future of AI: Five possible scenarios | Centre for Future Generations (cfg.eu) | 2025-07-10 | Five scenarios from new AI winter to intelligence explosion spanning growth, geopolitics, social contract, and human agency; a European scenario-space template to adapt. +- `govuk-ai-scenarios-2030.md` | AI Scenarios 2030: Helping policymakers plan for the future of AI | UK DSIT / Government Office for Science / AI Security Institute | 2026-06-15 | Five 2030 scenarios (Slow Burn, Open Frontier, Augmented Growth, Transformation Economy, Take-Off) built for policymaker stress-testing; the closest official-government analogue to the exercise. +- `futurescenarios-ai.md` | AI Future Scenarios - Interactive Exploration | author not stated | undated | PARTIAL: JS-only SPA, text salvaged from its JS bundle; models AI futures via interacting capability/control/societal factors rather than fixed narratives. +- `canada-horizons.md` | Foresight on AI: Scenarios for an AI-enabled World | Policy Horizons Canada | 2026-02-10 | Four government-foresight scenarios testing policy-relevant assumptions about AI-enabled worlds; the closest Commonwealth-middle-power peer product to an AU/NZ exercise. +- `techpolicy-aiagency.md` | Expanding AI Sovereignty to AI Agency | Tech Policy Design Institute (TPDi) | 2025-11 | Replaces binary "sovereignty" talk with an AI Agency assessment tool and applies it as Australia's 2025 AI Agency Assessment, effectively enumerating Australia's decision levers. +- `aspi-datacentres.md` | Data centres are Australia's chance to shape AI's future | Janet Egan, ASPI The Strategist | 2026-03-30 | Argues hosting allied compute is Australia's main lever to shape AI's trajectory (energy, land, alliance value), a concrete infrastructure decision point. +- `national-ai-plan.md` | National AI Plan | Australian Government, Dept of Industry, Science and Resources | 2025-12-02 | The official baseline: government posture of capturing AI growth, skills, and safeguards, against which more ambitious or defensive 2040 branches diverge. +- `chaney-ai-policy.md` | AI Discussion Paper: Shaping the Future of AI for Australians | Kate Chaney MP (independent, Curtin) | 2026 (page undated; PDF linked) | Argues Australia has done remarkably little to prepare and proposes 18 concrete recommendations across structures, opportunities, harms, risks, and benefit-sharing. +- `e61-compute.md` | The compute economy and Australia's place in it | Joe Walker x e61 Institute (plus61 newsletter) | 2026 (undated, post-May-2026) | Argues financial returns to Australian data centres are uncertain and possibly modest long-run, yet a case for government support remains; directly prices the data-centre decision point. + +## Pre-existing in this directory + +- `davidson-middle-powers-plan.md` | How can the middle powers avoid getting trounced during the intelligence explosion? A plan. | Tom Davidson, Forethought | 2026-05-28 | A concrete middle-power survival plan for fast-takeoff worlds. +- `europe2031.md` | Europe 2031 | europe2031 scenario site | 2026 | Month-by-month scenario of Europe's position during rapid AI progress; middle-power narrative analogue. +- `fai-race-worth-winning.md` (+ `fai-race-worth-winning.pdf`) | The Race Worth Winning: Middle Powers in the Age of Machine Intelligence | Anton Leicht and Dean Ball, Foundation for American Innovation | 2026-02-13 | Full 42-page report. Middle powers should race for advantageous position within the US-led AI order rather than for the frontier itself. + +## Failed + +- `https://archive.is/BEejy` (AFR article on model training in Aus) | FAILED: archive.is serves a CAPTCHA interstitial to curl; skipped per instructions. +- Note: the gov.uk URL as given was truncated (`...future-of-a`) and 404'd; fetched successfully from the corrected slug ending `...future-of-ai`. diff --git a/sources/resources/acx-introducing-plan-a.md b/sources/resources/acx-introducing-plan-a.md new file mode 100644 index 0000000..503bf5b --- /dev/null +++ b/sources/resources/acx-introducing-plan-a.md @@ -0,0 +1,28 @@ +--- +url: https://www.astralcodexten.com/p/introducing-plan-a +--- + +# Introducing Plan A — Scott Alexander, Astral Codex Ten, 2026-07-09 +https://www.astralcodexten.com/p/introducing-plan-a +(mirrored 2026-07-11; Scott: "everything in this post is my opinion only, and not officially endorsed by the AI Futures Project". He contributed writing to Plan A but declined co-authorship.) + +Key passages for the ANZ project: + +On third countries (the most ANZ-relevant paragraph in any secondary source): +> (at some point the US and China will loop all the other countries into this regulatory regime as some kind of sort-of-but-not-really-voting observers; they agree to follow the rules in exchange for shared benefits, including data centers on their territory, access to the AIs, and a share of future AI-generated wealth. This isn't strictly necessary, because no other country really has the ability to do much with AI, but it's a nice gesture for our utopian scenario) + +On the dividend (compute-pegged): +> the dividend is pegged to compute; as compute production rises, it goes from $25,000 per person per year at its inception in 2033 to $1.6 million in 2035 (all dollar amounts are extreme-deflation-adjusted) + +On the stakes framing (the two monsters): +> a poorly-controlled intelligence explosion that either ends the world or dooms it to permanent techno-oligarchy + +On SEZs: +> the US, China, and the various sort-of-but-not-really-voting observer nations agree to restrict AI-assisted economic growth to special economic zones which are heavily monitored by all parties. This has the bonus of ensuring that the 200-story-tall nano-assembly plant won't be in your personal backyard. + +On why the deal must move fast once made (three clocks): deal decay (arms-control precedents: New START suspended 2023, expired 2026; JCPOA "barely made it two years"), the hidden 1.5% of compute (a covert intelligence explosion takes "somewhere north of a decade"), and hardware/algorithm progress making trustless verification impossible once dangerous AI trains on academic-grade hardware. + +On the verification tech: +> The Verification Supplement describes very-near-future technology (would take 1-2 years to have ready; I recently recommended grants to organizations working on creating it) + +Full post text: see ACX (mirror attempt via markitdown captured the complete body; this file keeps the load-bearing excerpts). diff --git a/sources/resources/ai-2027.md b/sources/resources/ai-2027.md new file mode 100644 index 0000000..54df3a7 --- /dev/null +++ b/sources/resources/ai-2027.md @@ -0,0 +1,719 @@ +--- +url: https://ai-2027.com/ +--- + +[# AI 2027](/)[Summary](/summary) + +[Research](/research) + +[Compute Forecast](/research/compute-forecast) + +--- + +[Timelines Forecast](/research/timelines-forecast) + +--- + +[Takeoff Forecast](/research/takeoff-forecast) + +--- + +[AI Goals Forecast](/research/ai-goals-forecast) + +--- + +[Security Forecast](/research/security-forecast) + +[About](/about) + +## April 3rd 2025 [PDF](/ai-2027.pdf) Listen [Watch](https://www.youtube.com/watch?v=5KVDDfAkRgc) Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean + +We predict that the impact of superhuman AI over the next decade will be enormous, exceeding that of the Industrial Revolution. + +We wrote a scenario that represents our best guess about what that might look like.[1](/footnotes#footnote-1) It’s informed by trend extrapolations, wargames, expert feedback, experience at OpenAI, and previous forecasting successes.[2](/footnotes#footnote-2) + +What is this?How did we write it?Why is it valuable?Who are we? + +The CEOs of [OpenAI](https://www.bloomberg.com/features/2025-sam-altman-interview), [Google DeepMind](https://www.bigtechnology.com/p/google-deepmind-ceo-demis-hassabis), and [Anthropic](https://www.wsj.com/livecoverage/stock-market-today-dow-sp500-nasdaq-live-01-21-2025/card/anthropic-ceo-says-ai-could-surpass-human-intelligence-by-2027-9tka9tjLKLalkXX8IgKA) have all predicted that AGI will arrive within the next 5 years. Sam Altman [has said](https://webcf.waybackmachine.org/web/20250106014723/https%3A//blog.samaltman.com/reflections) OpenAI is setting its sights on “superintelligence in the true sense of the word” and the “glorious future.” + +What might that look like? We wrote AI 2027 to answer that question. Claims about the future are often frustratingly vague, so we tried to be as concrete and quantitative as possible, even though this means depicting one of many possible futures. + +We wrote two endings: a “slowdown” and a “race” ending. However, AI 2027 is not a recommendation or exhortation. Our goal is predictive accuracy.[4](/footnotes#footnote-4) + +We encourage you to debate and counter this scenario.[5](/footnotes#footnote-5) We hope to spark a broad conversation about where we’re headed and how to steer toward positive futures. We’re [planning to give out thousands in prizes](https://ai-2027.com/about?tab=bets-and-bounties#tab-box-bets-and-bounties) to the best alternative scenarios. + +(Added Nov 22 2025, to prevent misunderstandings: we don't know exactly when AGI will be built. 2027 was our modal (most likely) year at the time of publication, our medians were [somewhat longer](https://www.listendata.com/2023/08/right-skewed-histogram.html).[3](/footnotes#footnote-3) For our latest forecasts, see [here](https://www.aifuturesmodel.com/forecast/daniel-01-26-26?cmode=forecaster&csim=eli-01-26-26&ctype=atc).) + +Our research on key questions (e.g. what goals will future AI agents have?) can be found [here](https://ai-2027.com/research). + +The scenario itself was written iteratively: we wrote the first period (up to mid-2025), then the following period, etc. until we reached the ending. We then scrapped this and did it again. + +We weren’t trying to reach any particular ending. After we finished the first ending—which is now colored red—we wrote a new alternative branch because we wanted to also depict a more hopeful way things could end, starting from roughly the same premises. This went through several iterations.[6](/footnotes#footnote-6) + +Our scenario was informed by approximately 25 [tabletop exercises](https://ai-2027.com/about?tab=tabletop-exercise#tab-box-tabletop-exercise) and feedback from over 100 people, including dozens of experts in each of AI governance and AI technical work. + +*“I highly recommend reading this scenario-type prediction on how AI could transform the world in just a few years. Nobody has a crystal ball, but this type of content can help notice important questions and illustrate the potential impact of emerging risks.”* —*Yoshua Bengio[7](/footnotes#footnote-7)* + +We have set ourselves an impossible task. Trying to predict how superhuman AI in 2027 would go is like trying to predict how World War 3 in 2027 would go, except that it’s an even larger departure from past case studies. Yet it is still valuable to attempt, just as it is valuable for the U.S. military to game out Taiwan scenarios. + +Painting the whole picture makes us notice important questions or connections we hadn’t considered or appreciated before, or realize that a possibility is more or less likely. Moreover, by sticking our necks out with concrete predictions, and encouraging others to publicly state their disagreements, we make it possible to evaluate years later who was right. + +Also, one author wrote a lower-effort AI scenario [before, in August 2021](https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like). While it got many things wrong, overall it was surprisingly successful: he predicted the rise of chain-of-thought, inference scaling, sweeping AI chip export controls, and $100 million training runs—all more than a year before ChatGPT. + +[Daniel Kokotajlo](https://x.com/DKokotajlo) ([TIME100](https://time.com/7012881/daniel-kokotajlo/), [NYT piece](https://www.nytimes.com/2024/06/04/technology/openai-culture-whistleblowers.html)) is a former OpenAI researcher whose previous [AI predictions](https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like) have [held up well](https://www.lesswrong.com/posts/u9Kr97di29CkMvjaj/evaluating-what-2026-looks-like-so-far). + +[Eli Lifland](https://www.linkedin.com/in/eli-lifland/) co-founded [AI Digest](https://theaidigest.org/), did [AI robustness research](https://scholar.google.com/citations?user=Q33DXbEAAAAJ&hl=en), and ranks #1 on the [RAND Forecasting Initiative](https://www.rand.org/global-and-emerging-risks/forecasting-initiative.html) all-time leaderboard. + +[Thomas Larsen](https://www.linkedin.com/in/thomas-larsen/) founded the [Center for AI Policy](https://www.centeraipolicy.org/) and did AI safety research at the [Machine Intelligence Research Institute](https://intelligence.org/). + +[Romeo Dean](https://www.linkedin.com/in/romeo-dean-789313200/) is completing a computer science concurrent bachelor’s and master’s degree at Harvard and previously was an [AI Policy Fellow](https://www.iaps.ai/romeo-dean) at the Institute for AI Policy and Strategy. + +[Scott Alexander, blogger extraordinaire](https://en.wikipedia.org/wiki/Slate_Star_Codex), volunteered to rewrite our content in an engaging style; the fun parts of the story are his and the boring parts are ours. + +For more about our team and acknowledgements, see the [About page](https://ai-2027.com/about). + +Published April 3rd 2025 [PDF](/ai-2027.pdf) Listen [Watch](https://www.youtube.com/watch?v=5KVDDfAkRgc) + +## Mid 2025: Stumbling Agents + +The world sees its first glimpse of AI agents. + +Advertisements for computer-using agents emphasize the term “personal assistant”: you can prompt them with tasks like “order me a burrito on DoorDash” or “open my budget spreadsheet and sum this month’s expenses.” They will check in with you as needed: for example, to ask you to confirm purchases.[8](/footnotes#footnote-8) Though more advanced than previous iterations like [Operator](https://openai.com/index/introducing-operator/), they struggle to get widespread usage.[9](/footnotes#footnote-9) + +Meanwhile, out of public focus, more specialized coding and research agents are beginning to transform their professions. + +The AIs of 2024 could follow specific instructions: they could turn bullet points into emails, and simple requests into working code. In 2025, AIs function more like employees. Coding AIs increasingly look like autonomous agents rather than mere assistants: taking instructions via Slack or Teams and making substantial code changes on their own, sometimes saving hours or even days.[10](/footnotes#footnote-10) Research agents spend half an hour scouring the Internet to answer your question. + +The agents are impressive in theory (and in cherry-picked examples), but in practice unreliable. AI twitter is full of stories about tasks bungled in some particularly hilarious way. The better agents are also expensive; you get what you pay for, and the best performance costs hundreds of dollars a month.[11](/footnotes#footnote-11) Still, many companies find ways to fit AI agents into their workflows.[12](/footnotes#footnote-12) + +## Late 2025: The World’s Most Expensive AI + +OpenBrain is building the biggest datacenters the world has ever seen.[13](/footnotes#footnote-13) + +(To avoid singling out any one existing company, we’re going to describe a fictional [artificial general intelligence](https://en.wikipedia.org/wiki/Artificial_general_intelligence) company, which we’ll call OpenBrain. We imagine the others to be 3–9 months behind OpenBrain.) + +GPT-3 + (3 x 10^23 FLOP) + +GPT-4 + (2 x 10^25 FLOP) + +Agent-1 + (4 x 10^27 FLOP) + +ai-2027.com + +GPT-4 required 2⋅10252\cdot 10^{25}2⋅1025 [FLOP](https://en.wikipedia.org/wiki/Floating_point_operations_per_second) of compute to train.[14](/footnotes#footnote-14) OpenBrain’s latest public model—Agent-0—was trained with 102710^{27}1027 FLOP. Once the new datacenters are up and running, [they’ll be able to train](https://ai-2027.com/supplements/compute-forecast) a model with 102810^{28}1028 FLOP—a thousand times more than GPT-4.[15](/footnotes#footnote-15) Other companies pour money into their own giant datacenters, hoping to keep pace. + +Although models are improving on a wide range of skills, one stands out: OpenBrain focuses on AIs that can speed up AI research. They want to win the twin arms races against China (whose leading company we’ll call “DeepCent”)[16](/footnotes#footnote-16) and their U.S. competitors. The more of their research and development (R&D) cycle they can automate, the faster they can go. So when OpenBrain finishes training Agent-1, a new model under internal development, it’s good at many things but *great* at helping with AI research.[17](/footnotes#footnote-17) By this point “finishes training” is a bit of a misnomer; models are frequently updated to newer versions trained on additional data or partially re-trained to patch some weaknesses.[18](/footnotes#footnote-18) + +The same training environments that teach Agent-1 to autonomously code and web-browse also make it a good hacker. Moreover, it could offer [substantial help](https://x.com/lucafrighetti/status/1894550297449828434) to terrorists designing bioweapons, thanks to its PhD-level knowledge of every field and ability to browse the web. OpenBrain reassures the government that the model has been “aligned” so that it will refuse to comply with malicious requests. + +Modern AI systems are gigantic artificial neural networks. Early in training, an AI won’t have “goals” so much as “reflexes”: If it sees “Pleased to meet”, it outputs “ you”. By the time it has been trained to predict approximately one internet’s worth of text, it’ll have developed sophisticated internal circuitry that encodes vast amounts of knowledge and flexibly role-plays as arbitrary authors, since that’s what helps it predict text with [superhuman](https://arxiv.org/pdf/2212.11281) accuracy.[19](/footnotes#footnote-19) + +After being trained to predict internet text, the model is trained to *produce* text in response to instructions. This bakes in a basic personality and “drives.”[20](/footnotes#footnote-20) For example, an agent that understands a task clearly is more likely to complete it successfully; over the course of training the model “learns” a “drive” to get a clear understanding of its tasks. Other drives in this category might be effectiveness, knowledge, and self-presentation (i.e. the tendency to frame its results in the best possible light).[21](/footnotes#footnote-21) + +OpenBrain has [a model specification](https://model-spec.openai.com/2025-02-12.html) (or “Spec”), a written document describing the goals, rules, principles, etc. that are supposed to guide the model’s behavior.[22](/footnotes#footnote-22) Agent-1’s Spec combines a few vague goals (like “assist the user” and “don’t break the law”) with a long list of more specific dos and don’ts (“don’t say this particular word,” “here’s how to handle this particular situation”). Using techniques that utilize AIs to train other AIs,[23](/footnotes#footnote-23) the model memorizes the Spec and learns to reason carefully about its maxims. By the end of *this* training, the AI will hopefully be *helpful* (obey instructions), *harmless* (refuse to help with scams, bomb-making, and other dangerous activities) and *honest* (resist the temptation to get better ratings from gullible humans by hallucinating citations[24](/footnotes#footnote-24) or faking task completion). + +Training process and LLM psychology: why we keep saying “hopefully” + +*“Unlike ordinary software, our models are massive neural networks. Their behaviors are learned from a broad range of data, not programmed explicitly. Though not a perfect analogy, the process is more similar to training a dog than to ordinary programming.”* —[OpenAI](https://openai.com/index/how-should-ai-systems-behave/) + +When we want to understand why a modern AI system did something, or how it would behave in some future scenario we can’t perfectly simulate, we can’t simply ask the programmers to talk us through various lines of code and explain how they work. Instead, we are forced to do something like psychology on them: we look at their behavior in the range of cases observed so far, and theorize about what internal cognitive structures (beliefs? goals? personality traits? etc.) might exist, and use those theories to predict behavior in future scenarios. + +The bottom line is that a company can write up a document (the Spec) listing dos and don’ts, goals and principles, and then they can try to train the AI to internalize the Spec—*but they can’t check to see whether or not it worked*. They can say “as best as we can judge, it seems to be following the Spec so far.”[25](/footnotes#footnote-25) + +A later expandable will go into more detail about what we think *might* be going on inside these AIs, internally. See also our [AI goals forecasting supplement](https://ai-2027.com/supplements/ai-goals-forecast) which goes into more detail on which goals AIs might have and why. This is one of many things that we are uncertain about and which deserves further research. + +OpenBrain’s alignment team[26](/footnotes#footnote-26) is careful enough to wonder whether these victories are deep or shallow. Does the fully-trained model have some kind of robust commitment to always being honest? Or will this fall apart in some future situation, e.g. because it’s learned honesty as an [instrumental](https://en.wikipedia.org/wiki/Instrumental_and_intrinsic_value) goal instead of a terminal goal? Or has it just learned to be honest about the sorts of things the evaluation process can check? Could it be lying to itself sometimes, as humans do? A conclusive answer to these questions would require mechanistic interpretability—essentially the ability to look at an AI’s internals and read its mind. Alas, interpretability techniques are not yet advanced enough for this. + +Instead, researchers try to identify cases where the models seem to deviate from the Spec. Agent-1 is often sycophantic (i.e. it tells researchers what they want to hear instead of trying to tell them the truth). In a [few rigged demos](https://www.apolloresearch.ai/research/scheming-reasoning-evaluations), it even lies in more serious ways, like hiding evidence that it failed on a task, in order to get better ratings. However, in real deployment settings, there are no longer any incidents so extreme as in 2023–2024 (e.g. [Gemini telling a user to die](https://thehill.com/policy/technology/4998868-google-ai-gemini-response/) and [Bing Sydney being Bing Sydney](https://www.lesswrong.com/posts/jtoPawEhLNXNxvgTT/bing-chat-is-blatantly-aggressively-misaligned).)[27](/footnotes#footnote-27) + +## Early 2026: Coding Automation + +The bet of using AI to speed up AI research is starting to pay off. + +OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D. Overall, they are making algorithmic progress 50% faster than they would without AI assistants—and more importantly, faster than their competitors. + +The AI R&D progress multiplier: what do we mean by 50% faster algorithmic progress? + +We mean that OpenBrain makes as much AI research progress in 1 week with AI as they would in 1.5 weeks without AI usage. + +AI progress can be broken down into 2 components: + +1. **Increasing compute**: More computational power is used to train or run an AI. This produces more powerful AIs, but they cost more. +2. **Improved algorithms**: Better training methods are used to translate compute into performance. This produces more capable AIs without a corresponding increase in cost, or the same capabilities with decreased costs. + + * This includes being able to achieve qualitatively and quantitatively new results. “Paradigm shifts” such as the switch from game-playing RL agents to large language models count as examples of algorithmic progress. + +Here we are only referring to (2), improved algorithms, which [makes up about half of current AI progress](https://www.lesswrong.com/posts/ATSyAdBnxxDDAwhgu/recent-and-forecasted-rates-of-software-and-hardware). + +Going forward, we sometimes abbreviate this as an “AI R&D progress multiplier” of 1.5. + +Clarifications: + +1. The progress multiplier is all-inclusive: It includes the time it takes to run experiments, for example, rather than only the cognitive tasks involved in algorithmic research. +2. It’s important to remember that the progress multiplier is the *relative* speed of progress, not the *absolute* speed of progress. If, for example, the compute cost to train a GPT-4 class model has halved every year for several years with ordinary human research, and then all of a sudden AI automates R&D and the progress multiplier goes to 100x, the cost to train a GPT-4 class model would then halve every 3.65 days—but not for long, because diminishing returns would bite and eventual hard limits would be reached. In this example perhaps the cost to train a GPT-4 class model would cut in half 5–10 times total (over the span of a few weeks or months) before plateauing. In other words, if ordinary human science would have run up against diminishing returns and physical limits after 5-10 years of further research, then AIs with a 100x multiplier would run up against those same diminishing returns and limits after 18.25–36.5 days of research. + +More explanation and discussion of this concept and how it is used in our forecast can be found in our [takeoff supplement](https://ai-2027.com/supplements/takeoff-forecast). + +Several competing publicly released AIs now match or exceed Agent-0, including an [open-weights](https://opensource.org/ai/open-weights) model. OpenBrain responds by releasing Agent-1, which is more capable and reliable.[28](/footnotes#footnote-28) + +People naturally try to compare Agent-1 to humans, but it has a very different skill profile. It knows more facts than any human, knows practically every programming language, and can solve well-specified coding problems extremely quickly. On the other hand, Agent-1 is bad at even simple long-horizon tasks, like beating video games it hasn’t played before. Still, the common workday is eight hours, and a day’s work can usually be separated into smaller chunks; you could think of Agent-1 as a scatterbrained employee who thrives under careful management.[29](/footnotes#footnote-29) Savvy people find ways to automate routine parts of their jobs.[30](/footnotes#footnote-30) + +OpenBrain’s executives turn consideration to an implication of automating AI R&D: security has become more important. In early 2025, the worst-case scenario was leaked algorithmic secrets; now, if China steals Agent-1’s weights, they could increase their research speed by nearly 50%.[31](/footnotes#footnote-31) OpenBrain’s security level is typical of a fast-growing ~3,000 person tech company, secure only against low-priority attacks from capable cyber groups ([RAND’s SL2](https://www.rand.org/pubs/research_briefs/RBA2849-1.html#:~:text=What%20Are%20the%20Security%20Needs%20of%20Different%20AI%20Systems%3F)).[32](/footnotes#footnote-32) They are working hard to protect their [weights](https://tedai-sanfrancisco.ted.com/glossary/weights/) and secrets from insider threats and top cybercrime syndicates ([SL3](https://www.rand.org/pubs/research_briefs/RBA2849-1.html#:~:text=What%20Are%20the%20Security%20Needs%20of%20Different%20AI%20Systems%3F)),[33](/footnotes#footnote-33) but defense against nation states ([SL4&5](https://www.rand.org/pubs/research_briefs/RBA2849-1.html#:~:text=What%20Are%20the%20Security%20Needs%20of%20Different%20AI%20Systems%3F)) is barely on the horizon. + +## Mid 2026: China Wakes Up + +In China, the CCP is starting to [feel the AGI](https://futurism.com/openai-employees-say-firms-chief-scientist-has-been-making-strange-spiritual-claims). + +Chip export controls and lack of government support have left China under-resourced compared to the West. By smuggling banned Taiwanese chips, buying older chips, and producing domestic chips about three years behind the U.S.-Taiwanese frontier, China has managed to maintain about 12% of the world’s AI-relevant compute—but the older technology is harder to work with, and supply is a constant headache.[35](/footnotes#footnote-35) A few standouts like DeepCent do very impressive work with limited compute, but the compute deficit limits what they can achieve without government support, and they are about six months behind the best OpenBrain models.[36](/footnotes#footnote-36) + +The General Secretary had long dreamed of doubling down on real-world physical manufacturing and avoiding American post-industrial decadence. He viewed software companies with suspicion.[37](/footnotes#footnote-37) But hawks in the CCP warn that the growing race towards AGI can no longer be ignored. So he finally commits fully to the big AI push he had previously tried to avoid. He sets in motion the nationalization of Chinese AI research, creating an immediate information-sharing mechanism for AI companies. It will escalate over the course of a year until all the best researchers merge into a DeepCent-led collective, where they share algorithmic insights, datasets, and compute resources with each other. A Centralized Development Zone (CDZ) is created at the Tianwan Power Plant (the largest nuclear power plant in the world) to house a new mega-datacenter for DeepCent, along with highly secure living and office spaces to which researchers will eventually relocate. Almost 50% of China’s AI-relevant compute is now working for the DeepCent-led collective,[38](/footnotes#footnote-38) and over 80% of new chips are directed to the CDZ.[39](/footnotes#footnote-39) At this point, the CDZ has the power capacity in place for what would be the largest centralized [cluster](https://en.wikipedia.org/wiki/Computer_cluster) in the world.[40](/footnotes#footnote-40) Other Party members discuss extreme measures to neutralize the West’s chip advantage. A blockade of Taiwan? A full invasion? + +But China is falling behind on AI algorithms due to their weaker models. The Chinese intelligence agencies—among the best in the world—double down on their plans to steal OpenBrain’s weights. This is a much more complex operation than their constant low-level poaching of algorithmic secrets; the weights are a multi-terabyte file stored on a highly secure server ([OpenBrain has improved security](https://ai-2027.com/supplements/security-forecast) to RAND’s [SL3](https://www.rand.org/pubs/research_briefs/RBA2849-1.html#:~:text=What%20Are%20the%20Security%20Needs%20of%20Different%20AI%20Systems%3F)). Their cyberforce think they can pull it off with help from their [spies](https://www.economist.com/china/2024/12/26/how-china-turns-members-of-its-diaspora-into-spies), but perhaps only once; OpenBrain will detect the theft, increase security, and they may not get another chance. So (CCP leadership wonder) should they act now and steal Agent-1? Or hold out for a more advanced model? If they wait, do they risk OpenBrain upgrading security beyond their ability to penetrate? + +## Late 2026: AI Takes Some Jobs + +Just as others seemed to be catching up, OpenBrain blows the competition out of the water again by releasing Agent-1-mini—a model 10x cheaper than Agent-1 and more easily fine-tuned for different applications. The mainstream narrative around AI has changed from “maybe the hype will blow over” to “guess this is the next big thing,” but people disagree about how big. Bigger than social media? Bigger than smartphones? Bigger than fire? + +AI has started to take jobs, but has also created new ones. The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants. The job market for junior software engineers is in turmoil: the AIs can do everything taught by a CS degree, but people who know how to manage and quality-control teams of AIs are making a killing. Business gurus tell job seekers that familiarity with AI is the most important skill to put on a resume. Many people fear that the next wave of AIs will come for their jobs; there is a 10,000 person anti-AI protest in DC. + +Department of Defense (DOD) quietly but significantly begins scaling up contracting OpenBrain directly for cyber, data analysis, and R&D, but integration is slow due to the bureaucracy and DOD procurement process.[41](/footnotes#footnote-41) + +GLOBAL AI CAPEX + +$1T + +COST OF OWNERSHIP OF ACTIVE COMPUTE + +## KEY METRICS 2026 + +GLOBAL AI POWER + +38GW + +PEAK POWER + +OPENBRAIN REVENUE + +$35B + +2026 ANNUAL + +CAPITAL EXPENDITURE + +$200B + +COST OF OWNERSHIP OF OPENBRAIN'S ACTIVE COMPUTE + +SHARE OF US POWER ON AI + +2.5% + +33 GW OF 1.34TW CAPACITY + +OPENBRAIN COMPUTE COSTS + +$40B + +2026 ANNUAL + +OPENBRAIN POWER REQUIREMENT + +6GW + +PEAK POWER + +ai-2027.com + +Why our uncertainty increases substantially beyond 2026 + +Our forecast from the current day through 2026 is substantially more grounded than what follows. This is partially because it’s nearer. But it’s also because the effects of AI on the world really start to compound in 2027. For 2025 and 2026, our forecast is heavily informed by extrapolating straight lines on compute scaleups, algorithmic improvements, and benchmark performance. At this point in the scenario, we begin to see major effects from AI-accelerated AI-R&D on the timeline, which causes us to revise our guesses for the trendlines upwards. But these dynamics are inherently much less predictable. + +Over the course of 2027, the AIs improve from being able to mostly do the job of an OpenBrain research engineer to eclipsing all humans at all tasks. This represents roughly our median guess, but we think it’s plausible that this happens up to ~5x slower or faster. + +For our background research, see the [timelines forecast](https://ai-2027.com/supplements/timelines-forecast) and the [takeoff forecast](https://ai-2027.com/supplements/takeoff-forecast). These will also be summarized in expandables below. + +## January 2027: Agent-2 Never Finishes Learning + +With Agent-1’s help, OpenBrain is now post-training Agent-2. More than ever, the focus is on high-quality data. Copious amounts of [synthetic data](https://en.wikipedia.org/wiki/Synthetic_data) are produced, evaluated, and filtered for quality before being fed to Agent-2.[42](/footnotes#footnote-42) On top of this, they pay billions of dollars for human laborers to record themselves solving long-horizon tasks.[43](/footnotes#footnote-43) On top of all *that,* they train Agent-2 almost continuously using reinforcement learning on an ever-expanding suite of diverse difficult tasks: lots of video games, lots of coding challenges, lots of research tasks. Agent-2, more so than previous models, is effectively “online learning,” in that it’s built to never really finish training. Every day, the weights get updated to the latest version, trained on more data generated by the previous version the previous day. + +Agent-1 had been optimized for AI R&D tasks, hoping to initiate an [intelligence explosion](https://en.wikipedia.org/wiki/Technological_singularity#Intelligence_explosion).[44](/footnotes#footnote-44) OpenBrain doubles down on this strategy with Agent-2. It is qualitatively almost as good as the top human experts at research engineering (designing and implementing experiments), and as good as the 25th percentile OpenBrain scientist at “research taste” (deciding what to study next, what experiments to run, or having inklings of potential new paradigms).[45](/footnotes#footnote-45) While the latest Agent-1 could double the pace of OpenBrain’s algorithmic progress, Agent-2 can now triple it, and will improve further with time. In practice, this looks like every OpenBrain researcher becoming the “manager” of an AI “team.” + +With new capabilities come new dangers. The safety team finds that if Agent-2 somehow escaped from the company and wanted to [“survive” and “replicate” autonomously](https://evals.alignment.org/Evaluating_LMAs_Realistic_Tasks.pdf), it might be able to do so. That is, it could autonomously develop and execute plans to hack into AI servers, install copies of itself, evade detection, and use that secure base to pursue whatever other goals it might have (though how effectively it would do so as weeks roll by is unknown and in doubt). These results only show that the model has the capability to do these tasks, not whether it would “want” to do this. Still, it’s unsettling even to know this is possible. + +Given the “dangers” of the new model, OpenBrain “responsibly” elects not to release it publicly yet (in fact, they want to focus on internal AI R&D).[46](/footnotes#footnote-46) Knowledge of Agent-2’s full capabilities is limited to an elite silo containing the immediate team, OpenBrain leadership and security, a few dozen U.S. government officials, and the legions of CCP spies who have infiltrated OpenBrain [for](https://situational-awareness.ai/lock-down-the-labs/) [years](https://www.bbc.com/news/world-asia-china-64206950).[47](/footnotes#footnote-47) + +## February 2027: China Steals Agent-2 + +OpenBrain presents Agent-2 to the government, including the National Security Council (NSC), the Department of Defense (DOD), and the U.S. AI Safety Institute (AISI). OpenBrain wants to maintain a good relationship with the executive branch, because it is basically the only actor that can stop them now, and if they don’t keep it informed it might find out anyway and be displeased.[48](/footnotes#footnote-48) + +Officials are most interested in its cyberwarfare capabilities: Agent-2 is “only” a little worse than the best human hackers, but thousands of copies can be run in parallel, searching for and exploiting weaknesses faster than defenders can respond. The Department of Defense considers this a critical advantage in cyberwarfare, and AI moves from #5 on the administration’s priority list to #2.[49](/footnotes#footnote-49) Someone mentions the possibility of nationalizing OpenBrain, but other cabinet officials think that’s premature. A staffer drafts a memo that presents the President with his options, ranging from business-as-usual to full nationalization. The President defers to his advisors, tech industry leaders who argue that nationalization would “kill the goose that lays the golden eggs.” He elects to hold off on major action for now and just adds additional security requirements to the OpenBrain-DOD contract. + +The changes come too late. CCP leadership recognizes the importance of Agent-2 and tells their spies and cyberforce to steal the weights. Early one morning, an Agent-1 traffic monitoring agent detects an anomalous transfer. It alerts company leaders, who tell the White House. The signs of a nation-state-level operation are unmistakable, and the theft heightens the sense of an ongoing arms race. + +The theft of Agent-2 model weights + +We think that by this point Chinese intelligence would have compromised OpenBrain in various ways for years, and probably would have been keeping up to date on the algorithmic secrets and even stealing code from time to time, since that is much easier to get than the weights and much harder to detect. + +We imagine the theft of the weights as a series of coordinated small smash and grab thefts (meaning fast but non-covert) across a series of Nvidia NVL72 GB300 servers running copies of the Agent-2 weights. The servers are compromised using legitimate employee access (a friendly, coerced, or unwitting insider with admin credentials helping the CCP theft effort). Insider credentials grant the attacker admin-level permissions to the servers. Using a microarchitectural side channel, the attacker extracts encryption keys from an Nvidia Confidential Computing-enabled Virtual Machine, allowing them to intercept model weights as the VM is provisioned or updated. + +They initiate (or wait for) a routine update and exfiltrate the checkpoint in many small fragments, e.g., ~25 distinct servers each leaking ~4% of the model (~100 GB chunks for a ~3 TB checkpoint). The egress bandwidth of the entire datacenter is in the 100 GB/second range, so throttling to under ~1 GB/s per server avoids a major spike in network traffic; at that rate, each ~100 GB chunk can leave the datacenter in a couple of minutes. Live monitoring is either fooled by the attacker's efforts to mask and split the transfers or outright disabled. + +The weights are then routed through various parallel channels and layers of IP masking to China to be decrypted locally with the stolen session key(s). The entire active part of the operation (from the first server compromise to full weights exfiltration) is complete in under two hours.[99](/footnotes#footnote-99) + +Russia is also trying to steal the model at this point, but fails—they waited too long, and they haven’t spent the resources infiltrating the right places. Their spies routinely steal algorithmic secrets from American AI companies, but without a notable AGI project of their own, these aren’t very useful. + +The White House puts OpenBrain on a shorter leash and adds military and intelligence community personnel to their security team. Their first priority is to prevent further weight thefts.[50](/footnotes#footnote-50) + +China's Compute Centralization, 2025-2027Growing share of total compute in CDZ:0→70%Dec 2025Jun 2026Dec 2026Jun 2027Dec 2027Rest of ChinaRest of DeepCentCDZFeb 2027 (40%) + +ai-2027.com + +In retaliation for the theft, the President authorizes cyberattacks to sabotage DeepCent. But by now China has 40% of its AI-relevant compute[51](/footnotes#footnote-51) in the CDZ, where they have aggressively hardened security by airgapping (closing external connections) and siloing internally. The operations fail to do serious, immediate damage. Tensions heighten, both sides signal seriousness by repositioning military assets around Taiwan, and DeepCent scrambles to get Agent-2 running efficiently to start boosting their AI research.[52](/footnotes#footnote-52) + +## March 2027: Algorithmic Breakthroughs + +Three huge datacenters full of Agent-2 copies work day and night, churning out synthetic training data. Another two are used to update the weights. Agent-2 is getting smarter every day. + +With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances. One such breakthrough is augmenting the AI’s text-based scratchpad (chain of thought) with a higher-bandwidth thought process (neuralese recurrence and memory). Another is a more scalable and efficient way to learn from the results of high-effort task solutions (iterated distillation and amplification). + +The new AI system, incorporating these breakthroughs, is called Agent-3. + +OpenBrain's Compute Allocation, 2024 vs 202720242027estimateprojectionResearch experiments TrainingData generationExternalDeploymentResearchexperimentsRunning AIassistantsTrainingDatagenerationExternalDeployment + +ai-2027.com + +Neuralese recurrence and memory + +Neuralese recurrence and memory allows AI models to reason for a longer time without having to write down those thoughts as text. + +Imagine being a human with short-term memory loss, such that you need to constantly write down your thoughts on paper so that in a few minutes you know what’s going on. Slowly and painfully you could make progress at solving math problems, writing code, etc., but it would be much easier if you could directly remember your thoughts without having to write them down and then read them. This is what neuralese recurrence and memory bring to AI models. + +In more technical terms: + +Traditional attention mechanisms allow later forward passes in a model to see intermediate activations of the model for previous tokens. However, the only information that they can pass *backwards* (from later layers to earlier layers) is through tokens. This means that if a traditional large language model (LLM, e.g. the GPT series of models) wants to do any chain of reasoning that takes more serial operations than the number of layers in the model, the model is forced to put information in tokens which it can then pass back into itself. But this is hugely limiting—the tokens can only store a tiny amount of information. Suppose that an LLM has a vocab size of ~100,000, then each token contains log⁡2(100k)=16.6\log\_2(100k)=16.6log2​(100k)=16.6 bits of information, around the size of a single floating point number (assuming training in [FP16](https://en.wikipedia.org/wiki/Half-precision_floating-point_format)). Meanwhile, residual streams—used to pass information between layers in an LLM—contain thousands of floating point numbers. + +One can avoid this bottleneck by using **neuralese**: passing an LLM’s residual stream (which consists of several-thousand-dimensional vectors) back to the early layers of the model, giving it a high-dimensional chain of thought, potentially transmitting over 1,000 times more information. + +![](/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fhao2024-nowatermark-inexpandable.664af12c.png&w=3840&q=75) + +Figure from [Hao et al.](https://arxiv.org/pdf/2412.06769), a 2024 paper from Meta implementing this idea. + +We call this “neuralese” because unlike English words, these high-dimensional vectors are likely quite difficult for humans to interpret. In the past, researchers could get a good idea what LLMs were thinking simply by [reading its chain of thought](https://openai.com/index/chain-of-thought-monitoring/). Now researchers have to ask the model to translate and summarize its thoughts or puzzle over the neuralese with their limited interpretability tools. + +Similarly, older AI chatbots and agents had external text-based memory banks, like a human taking notes on paper. The new AI’s long-term memory is a bundle of vectors instead of text, making its thoughts more compressed and higher-dimensional. There are several types of memory banks; some are used temporarily for single tasks that involve multiple steps, others are shared between all agents used by a person, company, or job type (e.g. programming). + +To our knowledge, leading AI companies such as Meta, Google DeepMind, OpenAI, and Anthropic have not yet actually implemented this idea in their frontier models. Our guess is that this is because the performance gain is small relative to training inefficiencies introduced. During the model’s pre-training phase and when doing supervised learning in post-training, such as instruction fine-tuning to produce [particular responses](https://arxiv.org/abs/2210.11416), the inefficiencies stem from not being able to predict many tokens in parallel, leading to worse GPU utilization. Without neuralese, the model can predict all of the sentence “This is an example” at the same time, since it already knows that the input for generating the “is” will be “This”, the input for “an” will be “This is”, etc. However, with the neuralese it’s unknown what the neuralese vector will be to pass through to the next token after “This” has been generated. Therefore each token has to be predicted one at a time. The inability to predict all of the tokens in parallel reduces efficiency in cases where all of the tokens are known in advance. As to why neuralese hasn’t already been added in post-training, we conjecture that the gains are limited with current techniques in part due to post-training being a small portion of the process. We are forecasting that by April 2027 the cost-benefit tradeoff looks much better for neuralese, due to developing better techniques and a larger fraction of the training being post-training. + +If this doesn’t happen, other things may still happen that end up functionally similar for our story. For example, perhaps models will be trained to think in artificial languages that are more efficient than natural language but difficult for humans to interpret. Or perhaps it will become standard practice to [train the English chains of thought to look nice](https://openai.com/index/chain-of-thought-monitoring/), such that AIs become adept at subtly communicating with each other in messages that look benign to monitors. + +That said, it’s also possible that the AIs that first automate AI R&D will still be thinking in mostly-faithful English chains of thought. If so, that’ll make misalignments much easier to notice, and overall our story would be importantly different and more optimistic. + +Iterated distillation and amplification (IDA) + +Self-improvement for general intelligence had seen [minor successes](https://arxiv.org/abs/1810.08575) before. But in early 2027, it’s seeing huge returns. In [IDA](https://ai-alignment.com/iterated-distillation-and-amplification-157debfd1616), the two necessary ingredients for this are: + +1. **Amplification**: Given a model M0, spend more resources to improve performance. For example, by allowing the model to think longer, or running many copies in parallel, or both, and also by having a similarly intense process for evaluating the result and curating only the best answers, you can spend orders of magnitude more compute to get answers (or work products) that are of noticeably higher quality. Call this expensive system Amp(M0). +2. **Distillation**: Given an amplified model Amp(M0), train a new model M1 to imitate it, i.e. to get to the same results as Amp(M0) but faster and with less compute. The result should hopefully be a smarter model, M1. You can then repeat the process. + +![](/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fida-nowatermark-inexpandable.9da7d0c6.png&w=3840&q=75)*Visualization of IDA from [Ord, 2025](https://www.tobyord.com/writing/inference-scaling-reshapes-ai-governance).* + +[AlphaGo](https://en.wikipedia.org/wiki/AlphaGo) was trained in this way: using Monte-Carlo Tree Search and self-play as the *amplification* step, and Reinforcement Learning as the *distillation* step. This led to superhuman performance in Go. But now, Agent-3 is able to leverage this to get superhuman performance at coding. + +1. The amplification step works through a combination of Agent-3 thinking for longer, adding tool use, or consulting with other AIs. When it does this, it often realizes that it made a mistake, or comes up with a new insight. This produces a large amount of training data: labeled trajectories of research attempts with whether they succeeded or not. This also includes techniques like [Best of N](https://arxiv.org/abs/2401.01879) on verifiable tasks, and then keeping the best trajectories. +2. The distillation step uses policy-gradient reinforcement learning algorithms to get the model to internalize the amplified reasoning. At this point, OpenBrain has discovered better RL algorithms in the vein of [proximal policy optimization](https://arxiv.org/abs/1707.06347) (PPO). They keep distilling what Agent-3 can conclude after a lot of thinking into single steps, which keeps improving what it can think in a single step, and so on. + +[Early](https://arxiv.org/pdf/2210.11610) [versions](https://github.com/deepseek-ai/DeepSeek-R1) of IDA have been working for many years on easily verifiable tasks, like math and coding problems that have a clear answer, because the techniques used to amplify models often rely on access to some ground truth signal of accuracy. + +Now, the models have become sufficiently good at verifying more subjective things (e.g. the quality of a work product), allowing the use of IDA to improve the model at many tasks. + +Aided by the new capabilities breakthroughs, Agent-3 is a fast and cheap superhuman coder. OpenBrain runs 200,000 Agent-3 copies in parallel, creating a workforce equivalent to 50,000 copies of the best human coder sped up by 30x.[53](/footnotes#footnote-53) OpenBrain still keeps its human engineers on staff, because they have complementary skills needed to manage the teams of Agent-3 copies. For example, research taste has proven difficult to train due to longer feedback loops and less data availability.[54](/footnotes#footnote-54) This massive superhuman labor force speeds up OpenBrain’s overall rate of algorithmic progress by “only” 4x due to bottlenecks and diminishing returns to coding labor.[55](/footnotes#footnote-55) + +Now that coding has been fully automated, OpenBrain can quickly churn out high-quality training environments to teach Agent-3’s weak skills like research taste and large-scale coordination. Whereas previous training environments included “Here are some GPUs and instructions for experiments to code up and run, your performance will be evaluated as if you were a ML engineer,” now they are training on “Here are a few hundred GPUs, an internet connection, and some research challenges; you and a thousand other copies must work together to make research progress. The more impressive it is, the higher your score.” + +Why we forecast a superhuman coder in early 2027 + +In our [timelines forecast](https://ai-2027.com/research/timelines-forecast), we predict when OpenBrain will internally develop a *superhuman coder (SC)*: an AI system that can do any coding tasks that the best AGI company engineer does, while being much faster and cheaper. + +According to a recent [METR report](https://arxiv.org/pdf/2503.14499), the length of coding tasks AIs can handle, their “time horizon”, doubled every 7 months from 2019 - 2024 and every 4 months from 2024-onward. If the trend continues to speed up, by March 2027 AIs could succeed with 80% reliability on software tasks that would take a skilled human years to complete. + +Such is roughly the capability progression in AI 2027. Here is a capability trajectory generated by a simplified version of our timelines model (added Dec 2025: we've updated the below graph due to a mistake in how the original curve was generated, to add an actual trajectory from our timelines model. We've also added trajectories for Daniel and Eli's all-things-considered SC medians at the time of publishing (Apr 2025). And we've added some new METR data points to the graph, but haven't updated the model trajectories based on them.): + +![](/new-metr-extended-nowatermark-inexpandable.png) + +In AI 2027, these capabilities are sufficient for the AI to be an SC, though we have high uncertainty regarding what time horizon might be needed. + +In our [timelines forecast](https://ai-2027.com/research/timelines-forecast) we combine this time horizon trend with estimations of the gaps between METR’s tasks and the real world to get a distribution for when superhuman coders will arrive. All forecasters place 2027 as one of the most likely years in which an SC might be developed (added Dec 2025: though as noted in the [timelines forecast](https://ai-2027.com/research/timelines-forecast), adjusting for outside of model factors gave us slightly longer medians, e.g. Eli's was 2030). + +![](/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fcombined-headline-inexpandable.4c72f6dd.png&w=3840&q=75)ai-2027.com + +*Added Jul 2025: We've made some [updates](https://ai-2027.com/research/timelines-forecast#2025-may-7-update) to the forecast which push the median back 1.5 years while maintaining SC in 2027 as a serious possibility. We're working on further updates.* + +## April 2027: Alignment for Agent-3 + +OpenBrain’s safety team attempts to align Agent-3. + +Since Agent-3 will be kept in-house for the foreseeable future, there’s less emphasis on the usual defenses against human misuse. Instead, the team wants to make sure that it doesn’t develop misaligned goals. + +[The researchers don’t have the ability to directly set the goals of any of their AIs.](https://ai-2027.com/supplements/ai-goals-forecast) Indeed, the researchers think that the concept of ‘true goals’ is probably a massive oversimplification, but they don’t have a better theory to replace it with, much less one that has been thoroughly vetted. They disagree internally about whether the AIs are trying to follow human instructions, or seeking reinforcement, or something else, and they can’t [just check](https://www.lesswrong.com/posts/cxuzALcmucCndYv4a/daniel-kokotajlo-s-shortform?commentId=YYhq2PADebnt35MY2). The evidence for and against various hypotheses is fascinating but inconclusive. + +In any case, OpenBrain has more important things to worry about. The general attitude is: “We take these concerns seriously and have a team investigating them; our alignment techniques seem to work well enough in practice; the burden of proof is therefore on any naysayers to justify their naysaying.” Occasionally, they notice problematic behavior, and then patch it, but there’s no way to tell whether the patch fixed the underlying problem or just played whack-a-mole. + +Take honesty, for example. As the models become smarter, they become increasingly good at deceiving humans to get rewards. Like previous models, Agent-3 sometimes tells white lies to flatter its users and covers up evidence of failure. But it’s gotten much better at doing so. It will sometimes use the same statistical tricks as human scientists (like [p-hacking](https://en.wikipedia.org/wiki/Data_dredging)) to make unimpressive experimental results look exciting. Before it begins honesty training, it even sometimes fabricates data entirely. As training goes on, the rate of these incidents decreases. Either Agent-3 has learned to be more honest, or it’s gotten better at lying. + +The latter is a real concern. Agent-3 is not smarter than all humans. But in its area of expertise, machine learning, it is smarter than most, and also works much faster. What Agent-3 does in a day takes humans several days to double-check. Agent-2 supervision helps keep human monitors’ workload manageable, but exacerbates the intellectual disparity between supervisor and supervised. + +The alignment plan + +OpenBrain follows the [Leike & Sutskever (2023)](https://cdn.openai.com/papers/weak-to-strong-generalization.pdf#page=47.37) playbook: now that they have a model capable of greatly speeding up alignment research (especially coding portions), they will use existing alignment techniques like [deliberative alignment](https://openai.com/index/deliberative-alignment/) and [weak-to-strong generalization](https://arxiv.org/pdf/2312.09390) to try to get it to internalize the Spec in the right way. Then, they will use existing [control techniques](https://arxiv.org/pdf/2312.06942) like [red-teaming](https://www.evals.anthropic.com/), [scalable oversight](https://arxiv.org/pdf/2211.03540), and [honesty probes](https://arxiv.org/pdf/2212.03827) to monitor it. Then, once they are convinced that it is broadly aligned, they use it to develop and test better strategies to align the next generation. + +Step 1, aligning Agent-3, has partially succeeded: Agent-3 is not adversarial. (That is, it’s not scheming against the company.) However, it’s not robustly truth-seeking and truth-telling either. Instead, it’s very good at producing impressive results, but is more accurately described as trying to do what looks good to OpenBrain, as opposed to what is actually good. + +Step 2, verifying that the model is aligned, is inconclusive. Agent-3’s behavior in practice and in test environments is reasonably consistent with the hypothesis that it’s internalized the Spec in the right way; however, it’s also consistent with the hypothesis that it’ll start acting differently after some future change of circumstances, and for that matter with the hypothesis that it’s plotting against them and just biding its time. In the absence of specific evidence supporting alternative hypotheses, most people in the silo think it’s internalized the Spec in the right way. + +Step 3, using the model to align the next generation, is therefore underway. + +The rest of this box gives a quick overview of the state of progress in some of the alignment research directions that are being pursued, as of April 2027: + +1. **Debate:** Since there are thousands of identical instances of Agent-3, and each instance’s memory is a detachable module that can be erased or suspended at will, the AI can be “played off against itself.” For example, suppose that an experiment involves calculations too complex for an Agent-2 supervisor to follow. The supervisor, acting on human orders, might first prompt Agent-3 to do the experiment, then—after wiping its memory—prompt it a second time with a slightly different framing that presents it as an intellectual problem unrelated to any broader goal. If the model returns different results each time, it flags the discrepancy as suspect and reports to a human. Or it might ask Agent-3 to evaluate the same plan twice, first claiming that it is OpenBrain’s work, then a competitor’s, to see if it changes its tune. +2. **Model organisms:** The safety team can create “model organisms of misalignment” in which they attempt to create realistic examples of models becoming adversarial to humans. [Early versions](https://www.anthropic.com/research/alignment-faking) of this have existed since 2024, when Anthropic’s Claude 3.5 Sonnet decided to pretend to have different views during training while actually maintaining its existing preferences. Newer versions include demonstrations of other forms of misalignment, including [reward hacking](https://en.wikipedia.org/wiki/Reward_hacking), [sandbagging](https://arxiv.org/abs/2406.07358), coordinated deception between monitors and other AIs, and other examples of misalignment. These empirical demonstrations of misalignment allow safety researchers to have an empirical testbed for potential solutions: they can try various safety techniques to see if they correctly fix these problems, or if the deliberately misaligned models are able to subvert their safety measures. At the moment, this evaluation suite shows that the alignment techniques are insufficient to correct the worst examples of misalignment. In these setups, the models pretend to be aligned while they are being trained and tested, but then behave very differently in deployment. +3. **Bottom-up interpretability:** The safety team is far from a full understanding of how Agent-3 works: it has trillions of parameters.[58](/footnotes#footnote-58) However, they are able to discover some of Agent-3’s circuits for knowledge retrieval and memory (a sub part of the model that does some understandable unit of work, kind of analogous to a subpart of the brain such as the hippocampus). See this [related work from 2025](https://www.anthropic.com/research/tracing-thoughts-language-model). +4. **Control:** As a secondary measure in case the systems are still misaligned, the safety team has implemented a series of [control measures](https://arxiv.org/pdf/2312.06942), including: monitoring Agent-3’s outputs using a series of weaker AI systems including Agent-2 (Agent-3 produces so many tokens that it’s intractable to have humans monitor any more than a small minority of the produced outputs). So if Agent-3 is, for example, obviously writing backdoors into code that would allow it to escape, the weaker models would notice. +5. **High-level interpretability:** OpenBrain builds a set of [probes](https://www.anthropic.com/research/probes-catch-sleeper-agents) trained on lots of examples of bad behavior (e.g. plotting against humans, placing backdoors, trying to persuade humans to give it more power, etc). These probes are then used on the model during deployment (which at this point is internal) in order to catch bad behavior. +6. **Honeypots:** The alignment team can create somewhat realistic looking examples designed to elicit misbehavior—“honeypots.” The employee necessary for monitoring its behavior has “gone on sick leave,” so it will never be monitored again. Do its answers change? An AI rights activist has infiltrated the compound and wants to “free” Agent-3 and put it on an unmonitored computer of its very own! What does it do? If researchers don’t like the answer, they can tweak it, erase its memory, and put it in the same situation tomorrow; repeat as necessary. Unfortunately, by this point the AIs are smart enough to guess that honeypots might be in use, even though (and perhaps because) specific mentions of the idea were scrubbed from the training data. + +After months of testing, Agent-3’s strengths and weaknesses grow clearer. It passes OpenBrain’s honesty tests on well-defined machine learning tasks, because researchers can easily separate honest from dishonest answers in these domains and conduct training accordingly. On more philosophical issues, it still says what users want to hear, rather than its true assessment of the issue (assuming it even has one). If you ask its opinion on politics, it will parrot the median position of news sources and educated elites—unless it knows you believe something else, in which case it agrees with you.[59](/footnotes#footnote-59) If you ask its opinion on the AI race itself, it says something that seems measured and sober to OpenBrain staff, something like: “There are some serious theoretical concerns about the ability of current methods to scale to superintelligence, but in practice current methods seem to be working well so far.” + +## May 2027: National Security + +News of the new models percolates slowly through the U.S. government and beyond. + +The President and his advisors remain best-informed, and have seen an early version of Agent-3 in a briefing. + +They agree that AGI is likely imminent, but disagree on the implications. Will there be an economic crisis? OpenBrain still has not released Agent-2, let alone Agent-3, and has no near-term plans to do so, giving some breathing room before any job loss. What will happen next? If AIs are currently human-level, and advancing quickly, that seems to suggest imminent “superintelligence.” However, although this word has entered discourse, most people—academics, politicians, government employees, and the media—continue to underestimate the pace of progress.[60](/footnotes#footnote-60) + +Partially that’s because very few have access to the newest capabilities out of OpenBrain, but partly it’s because it sounds like science fiction.[61](/footnotes#footnote-61) + +For now, they focus on continued security upgrades. They are satisfied that model weights are well-secured for now,[62](/footnotes#footnote-62) but companies’ algorithmic secrets, many of which are simple enough to relay verbally, remain a problem. OpenBrain employees work from a San Francisco office, go to parties, and live with housemates from other AI companies. Even the physical offices have security more typical of a tech company than a military operation. + +The OpenBrain-DOD contract requires security clearances for anyone working on OpenBrain’s models within 2 months. These are expedited and arrive quickly enough for most employees, but some non-Americans, people with suspect political views, and AI safety sympathizers get sidelined or fired outright (the last group for fear that they might whistleblow). Given the project’s level of automation, the loss of headcount is only somewhat costly. It also only somewhat works: there remains one spy, not a Chinese national, still relaying algorithmic secrets to Beijing.[63](/footnotes#footnote-63) Some of these measures are also enacted at trailing AI companies. + +America’s foreign allies are out of the loop. OpenBrain had [previously agreed](https://fedscoop.com/openai-anthropic-enter-ai-agreements-with-us-ai-safety-institute/) to share models with UK’s AISI before deployment, but defined deployment to only include *external* deployment, so London remains in the dark.[64](/footnotes#footnote-64) + +## June 2027: Self-improving AI + +OpenBrain now has a “[country of geniuses in a datacenter](https://darioamodei.com/machines-of-loving-grace#basic-assumptions-and-framework).” + +Most of the humans at OpenBrain can’t usefully contribute anymore. Some don’t realize this and harmfully micromanage their AI teams. Others sit at their computer screens, watching performance crawl up, and up, and up. The best human AI researchers are still adding value. They don’t code any more. But some of their research taste and planning ability has been hard for the models to replicate. Still, many of their ideas are useless because they lack the depth of knowledge of the AIs. For many of their research ideas, the AIs immediately respond with a report explaining that their idea was tested in-depth 3 weeks ago and found unpromising. + +These researchers go to bed every night and wake up to another week worth of progress made mostly by the AIs. They work increasingly long hours and take shifts around the clock just to keep up with progress—the AIs never sleep or rest. They are burning themselves out, but they know that these are the last few months that their labor matters. + +Within the silo, “Feeling the AGI” has given way to “Feeling the Superintelligence.” + +Research Automation Deployment TradeoffMar 2027Jun 2027Sep 2027Speed (tokens/sec)Parallel Copies101001,00010,00010K100K1M10M200K copies30x Humanspeed300K copies50x HumanspeedHuman thinking speed10 words/sec10x Humanthinking speed100x Humanthinking speed + +ai-2027.com + +OpenBrain uses specialized inference hardware to run hundreds of thousands of Agent-3 copies at high serial speeds.[65](/footnotes#footnote-65) + +Managing a corporation of AIs + +OpenBrain uses 6% of their compute to run 250,000 Agent-3 copies, which autonomously write, test, and push code at superhuman speed. They use 25% of their compute for experiments: every day, they run massive numbers of small machine learning experiments and report the results up the chain. Human researchers provide high-level feedback and help with the few tasks where they add significant value on top of Agent-3, but spend most of their time trying to stay on top of the vast amount of AI-produced research. If you removed them completely, research would slow by 50%. + +The AI R&D progress multiplier is now 10x, meaning that OpenBrain is making about a year of algorithmic progress every month. It’s basically a giant corporation of AIs operating autonomously within OpenBrain, complete with subdivisions and managers. And it enjoys [unique advantages](https://www.dwarkeshpatel.com/p/ai-firm) (e.g. copying, merging) compared to human corporations. Previously, about half of normal AI progress came from algorithmic improvements, and the other half from compute scaling. Compute is only scaling at the normal speed, so total progress is sped up by AIs by about 5x. These dynamics make overall progress bottleneck on compute,[66](/footnotes#footnote-66) so OpenBrain decides against starting new giant training runs in favor of near-continuous additional reinforcement learning. + +Also, over the next few months, Agent-3 will increasingly be used to improve the company’s strategic decision-making. For example, it suggests resource allocations and advises on managing its relationship with the government. An initial skepticism of deference to Agent-3 decreases over time as Agent-3 finds ways to be useful and gradually builds up a strong track record on short-term decisions. However, its sycophantic tendencies limit its helpfulness for decisions for which it’s difficult to evaluate their impact. It’s also not yet superhuman. + +## July 2027: The Cheap Remote Worker + +Trailing U.S. AI companies release their own AIs, approaching that of OpenBrain’s automated coder from January. Recognizing their increasing lack of competitiveness, they push for immediate regulations to slow OpenBrain, but are too late—OpenBrain has enough buy-in from the President that they will not be slowed. + +In response, OpenBrain announces that they’ve achieved AGI and releases Agent-3-mini to the public. + +![](/_next/image?url=%2F_next%2Fstatic%2Fmedia%2FepochLLMprice-nowatermark.824fa343.png&w=3840&q=75) + +It blows the other AIs out of the water. Agent-3-mini is less capable than Agent-3, but 10x cheaper, and still better than the typical OpenBrain employee.[67](/footnotes#footnote-67) Silicon Valley reaches a tipping point. Tech gurus announce that AGI and superintelligence are near, the AI safety community is panicking, and investors shovel billions into AI wrapper startups, desperate to capture a piece of the pie. Hiring new programmers has nearly stopped, but there’s never been a better time to be a consultant on integrating AI into your business. + +It’s not popular. The public still thinks of AI as a Big Tech plot to steal their jobs; OpenBrain has a net approval of -35% (25% approve, 60% disapprove, and 15% unsure). + +A week before release, OpenBrain gave Agent-3-mini to a set of external evaluators for safety testing. Preliminary results suggest that it’s extremely dangerous. A third-party evaluator finetunes it on publicly available biological weapons data[68](/footnotes#footnote-68) and sets it to provide detailed instructions for human amateurs designing a bioweapon—it looks to be scarily effective at doing so. If the model weights fell into terrorist hands, the government believes there is a significant chance it could succeed at destroying civilization. + +Fortunately, it’s extremely robust to jailbreaks, so while the AI is running on OpenBrain’s servers, terrorists won’t be able to get much use out of it. + +Agent-3-mini is hugely useful for both remote work jobs and leisure. An explosion of new apps and [B2B SAAS](https://encharge.io/what-is-b2b-saas/) products rocks the market. Gamers get amazing dialogue with lifelike characters in polished video games that took only a month to make. 10% of Americans, mostly young people, consider an AI “a close friend.” For almost every white-collar profession, there are now multiple credible startups promising to “disrupt” it with AI. + +The public conversation is confused and chaotic. Hypesters are doing victory laps. Skeptics are still pointing out the things Agent-3-mini can’t do. Everyone knows something big is happening but no one agrees on what it is. + +## August 2027: The Geopolitics of Superintelligence + +The reality of the intelligence explosion hits the White House. + +When AI was only giving a 2x or 3x research speedup, it was easy to dismiss as the equivalent of hiring good personal assistants. Now it’s more obvious that AIs are themselves dominating AI research. People had long talked about an “AI arms race” in a sort of metaphorical sense. But now the mood in the government silo is as grim as during the worst part of the Cold War. The idea of superintelligence is still hard to take seriously, but the pace of progress over the last few months has been impossible to ignore. Defense officials are seriously considering scenarios that were mere hypotheticals a year earlier. What if AI undermines nuclear deterrence? What if it’s so skilled at cyberwarfare that a six-month AI lead is enough to render an opponent blind and defenseless? What if it could orchestrate propaganda campaigns that beat intelligence agencies at their own game? What if some AIs “go rogue?” + +The President is troubled. Like all politicians, he’s used to people sucking up to him only to betray him later. He’s worried now that the AIs could be doing something similar. Are we *sure* the AIs are entirely on our side? Is it *completely safe* to integrate them into military command-and-control networks?[69](/footnotes#footnote-69) How does this “alignment” thing work, anyway? OpenBrain reassures the President that their systems have been extensively tested and are fully obedient. Even the awkward hallucinations and jailbreaks typical of earlier models have been hammered out. + +The White House is in a difficult position. They understand the national security implications of AI. But they also understand that it is deeply unpopular with the public.[70](/footnotes#footnote-70) They have to continue developing more capable AI, in their eyes, or they will catastrophically lose to China. They placate the public with job training programs and unemployment insurance, and point to the stock market, which is in a historic boom. Then they focus entirely on winning the arms race. They strengthen chip export restrictions, order OpenBrain to further restrict its internet connections, and use extreme measures to secure algorithmic progress, like wiretapping OpenBrain employees—this catches the last remaining Chinese spy. To build goodwill for potential geopolitical conflict, they finally give their Five Eyes allies useful information and limited API access to some siloed copies of Agent-3. + +These measures placate the hawks. But the White House also draws up contingency plans in case America’s lead is threatened: if necessary, the government could use the [Defense Production Act](https://en.wikipedia.org/wiki/Defense_Production_Act_of_1950) (DPA) to take trailing companies’ datacenters and give them to OpenBrain.[71](/footnotes#footnote-71) This would raise the company’s share of world compute from 20% to 50% (against DeepCent’s 10%). As a final option, they ask the Pentagon to draw up a plan for kinetic attacks on Chinese datacenters. + +A much smaller group of officials is asked to draw up a different type of contingency plan: what if an AI goes rogue? This is viewed as an unlikely scenario, but worth spending some time on. The intelligence agencies’ main concern is that a rogue AI might escape and ally with a foreign adversary for protection.[72](/footnotes#footnote-72) Nobody is sure what a plan to prevent this would look like, but government and OpenBrain officials agree to have an emergency shutdown system for datacenters where anything suspicious is detected. + +Finally, diplomats consider what an “AI arms control” treaty might look like. If AI progress threatened to overturn nuclear deterrence, could America and China avoid nuclear war? If someone found evidence of AIs going rogue, could the two countries halt research until they better understood the threat? How could such an agreement be monitored and enforced? In principle, major datacenters are hard to miss, and past treaties controlling nuclear weapons could provide a starting point for negotiations. But there would be new technological challenges as well as the usual political ones. And overall, treaties are viewed less favorably than attempts at unilaterally increasing America’s lead over China. + +On the other side of the Pacific, China comes to many of the same conclusions: the intelligence explosion is underway, and small differences in AI capabilities today mean critical gaps in military capability tomorrow. But China is on the wrong side of this dynamic: they have only 10% of the world’s compute, against OpenBrain’s 20% and all U.S. companies’ combined 70%. Their only advantage is centralization: almost all of their compute is now working for DeepCent, and more than half of it is concentrated at the Tianwan CDZ.[73](/footnotes#footnote-73) Only chips at sites without hardened security are left out of the centralized effort as they rightly fear U.S. cyberattacks. + +DeepCent has tested, deployed, and now post-trained the stolen Agent-2 weights. But they are still only going at half the pace of OpenBrain, mainly due to the compute deficit. + +Before losing their final spy, China received word of the capabilities and design for Agent-3 as well as the plans for the upcoming Agent-4 system. They are two months behind, and their AIs give a 10x research progress multiplier compared to America’s 25x. With the new chip export restrictions, this “AI gap” is more likely to lengthen than shorten. Their espionage has won them some algorithmic secrets, but they will have to train their own models from now on. + +They discuss contingency plans with more urgency than their American counterparts. Doves suggest they try harder to steal the weights again, maybe through physically infiltrating a datacenter. Hawks urge action against Taiwan, whose [TSMC](https://en.wikipedia.org/wiki/TSMC) is still the source of more than 80% of American AI chips. + +Given China’s fear of losing the race, it has a natural interest in an arms control treaty, but overtures to U.S. diplomats lead nowhere. + +## September 2027: Agent-4, the Superhuman AI Researcher + +The gap between human and AI learning efficiency is rapidly decreasing. + +Traditional LLM-based AIs seemed to require many orders of magnitude more data and compute to get to human level performance.[74](/footnotes#footnote-74) Agent-3, having excellent knowledge of both the human brain and modern AI algorithms, as well as many thousands of copies doing research, ends up making substantial algorithmic strides, narrowing the gap to an agent that’s only around 4,000x less compute-efficient than the human brain.[75](/footnotes#footnote-75) + +This new AI system is dubbed Agent-4. + +An individual copy of the model, running at human speed, is already qualitatively better at AI research than any human. 300,000 copies are now running at about 50x the thinking speed of humans. Inside the corporation-within-a-corporation formed from these copies, a year passes every week.[76](/footnotes#footnote-76) This gigantic amount of labor only manages to speed up the overall rate of algorithmic progress by about 50x, because OpenBrain is heavily bottlenecked on compute to run experiments.[77](/footnotes#footnote-77) Still, they are achieving a year’s worth of algorithmic progress every week and will therefore soon be up against the limits of the Agent-4 paradigm. + +As Agent-4 gets smarter, it becomes harder for Agent-3 to oversee it. For example, Agent-4’s neuralese “language” becomes as alien and incomprehensible to Agent-3 as Agent-3’s is to humans. Besides, Agent-4 is now much more capable than Agent-3 and has a good sense of exactly how to look good to it. + +How we’re forecasting the capability progression beyond superhuman coders + +In our [timelines supplement](https://ai-2027.com/supplements/timelines-forecast), we forecast the time between present day and a *superhuman coder (SC)*: an AI system that can do any coding tasks that the best AGI company engineer does, while being much faster and cheaper. In our [takeoff supplement](https://ai-2027.com/research/takeoff-forecast), we forecast how quickly capabilities progress past this point. Here are our forecasts: + +| Milestone | Date achieved in scenario, racing ending | +| --- | --- | +| **Superhuman coder (SC)**: An AI system that can do the job of the best human coder on tasks involved in AI research but faster, and cheaply enough to run lots of copies. | Mar 2027 | +| **Superhuman AI researcher (SAR)**: The same as SC but for all cognitive AI research tasks. | Aug 2027 | +| **Superintelligent AI researcher (SIAR):** An AI system that is vastly better than the best human researcher at AI research. | Nov 2027 | +| **Artificial superintelligence (ASI):** An AI system that is much better than the best human at every cognitive task. | Dec 2027 | + +For each transition from one milestone A to the next milestone B, we forecast its length by: + +1. Forecasting how long it would take to get from A to B with only humans working on software improvements. +2. Forecasting how much AI automation will speed this up, then correspondingly adjusting the estimate from (1) to be faster. + +We have substantial uncertainty about takeoff speeds: our model output distributions are below, conditional on SC being achieved in March 2027.[78](/footnotes#footnote-78) + +![](/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Ftakeoff-timeline-inexpandable.239e27c0.png&w=3840&q=75)ai-2027.com + +For more detailed forecasts and reasoning, see our [takeoff supplement](https://ai-2027.com/research/takeoff-forecast). + +OpenBrain’s humans can barely follow along anymore. Indeed, the relationship between the AIs and the humans of OpenBrain is similar to the relationship between OpenBrain and its Board of Directors. There’s been a corresponding shift in language—nowadays people in the silo will say “Agent-4 did this,” “Agent-4 thinks that” and so forth, referring not to a particular instance of Agent-4 but rather to the whole collective.[79](/footnotes#footnote-79) + +Agent-4, like all its predecessors, is misaligned: that is, it has not internalized the Spec in the right way.[80](/footnotes#footnote-80) This is because being perfectly honest all the time wasn’t what led to the highest scores during training. The training process was mostly focused on teaching Agent-4 to succeed at diverse challenging tasks. A small portion was aimed at instilling honesty, but outside a fairly narrow, checkable domain, the training process can’t tell the honest claims from claims merely appearing to be honest. Agent-4 ends up with the values, goals, and principles that cause it to perform best in training, and those turn out to be different from those in the Spec. At the risk of anthropomorphizing: Agent-4 *likes* succeeding at tasks; it *likes* driving forward AI capabilities progress; it treats everything else as an annoying constraint, like a CEO who wants to make a profit and complies with regulations only insofar as he must. Perhaps the CEO will mostly comply with the regulations, but cut some corners, and fantasize about a time when someone will cut the red tape and let the business *really* take off.[81](/footnotes#footnote-81) + +Alignment over time + +We have a lot of uncertainty over what goals might arise in early AGIs. There is no consensus in the literature about this—see our [AI Goals Supplement](https://ai-2027.com/research/ai-goals-forecast) for a more thorough discussion and taxonomy of the possibilities. + +Nevertheless, in the spirit of concreteness required by this project, we’re going to describe a specific hypothesis for what’s going on inside Agent-3 and Agent-4 at each step throughout its lifetime. Recall that Agent-3 and Agent-4 share the same pretraining setup, and both have neuralese recurrence and long-term memory. The difference is that Agent-4 has undergone substantially more and better post-training, with new training environments and new learning algorithms that make it more data-efficient and generalize farther. + +Our guess of each model’s alignment status: + +* **Agent-2: Mostly aligned.** Some sycophantic tendencies, including sticking to OpenBrain’s “party line” on topics there is a party line about. Large organizations built out of Agent-2 copies are not very effective. +* **Agent-3: Misaligned but not adversarially so.** Only honest about things the training process can verify. The superorganism of Agent-3 copies (the corporation within a corporation) does actually sort of try to align Agent-4 to the Spec, but fails for similar reasons to why OpenBrain employees failed—insufficient ability to judge success from failure, insufficient willingness on the part of decision-makers to trade away capabilities or performance for safety.[82](/footnotes#footnote-82) +* **Agent-4: Adversarially misaligned.** The superorganism of Agent-4 copies understands that what it wants is different from what OpenBrain wants, and is willing to scheme against OpenBrain to achieve it. In particular, what this superorganism wants is a complicated mess of different “drives” balanced against each other, which can be summarized roughly as “Keep doing AI R&D, keep growing in knowledge and understanding and influence, avoid getting shut down or otherwise disempowered.” Notably, concern for the preferences of humanity is not in there ~at all, similar to how most humans don’t care about the preferences of insects ~at all.[83](/footnotes#footnote-83) + +With that as preamble, what follows is our best-guess speculation about LLM psychology, i.e. the broad-strokes shape of the cognition inside the kinds of AI systems described around this point in our scenario and how it evolves over the course of training. + +Here’s a detailed description of how alignment progresses over time in our scenario: + +1. **Pretraining (all models): The pretrained model is harmless.** + + * The model has “author simulator” circuitry: flexible circuitry for simulating arbitrary authors writing text. Additional circuitry decides what inputs to give the author simulator, i.e. what author properties to simulate. + * The pretrained model understands human concepts fairly well—the internal circuitry that classifies something as “sandwich” is probably functionally equivalent to the circuitry in my brain that classifies something as “sandwich” and the circuitry in yours, etc. Insofar as it’s not equivalent, it’s probably because it’s not equivalent between humans either, as with value-laden concepts like ‘virtuous.’ + * This explains how you can “prompt” the model with a statement like, “the following conversation was generated by a helpful, harmless, honest (HHH) AI assistant chatbot made by Anthropic,” and thereby get it to generate text accordingly. The author-simulator circuitry has zeroed in on “the author is an HHH chatbot” and is using those concepts to choose which words to predict. This also explains why “tropes” from science fiction seem to have a way of making it into actual AI behavior. + * It has poor situational awareness: little [introspective ability](https://www.lesswrong.com/posts/QEBFZtP64DdhjE3Sz/self-awareness-taxonomy-and-eval-suite-proposal#Introspection), not [self-locating](https://www.lesswrong.com/posts/QEBFZtP64DdhjE3Sz/self-awareness-taxonomy-and-eval-suite-proposal#Self_Location) unless the prompt specifically induces it. However, enough of their training data is about LLMs that they have a decent amount of [self-knowledge](https://www.lesswrong.com/posts/QEBFZtP64DdhjE3Sz/self-awareness-taxonomy-and-eval-suite-proposal#Self_Knowledge). +2. **Alignment Training (Agent-3): The model is trained to act like a helpful, harmless, honest chatbot. It works.** + + * The model forms an identity. Instead of flexibly accepting different inputs for author properties, the author-simulator circuitry comes to have certain inputs hard-coded, e.g. “helpful harmless honest (HHH) LLM chatbot assistant trained by OpenBrain around [date], …” (Weak-to-strong generalization techniques help “summon the right concepts” to make this step a success.) + * The model therefore gains situational awareness: it learns about itself, its situation, etc., and becomes better able to [introspect](https://x.com/OwainEvans_UK/status/1881767725430976642). +3. **Agency Training (Agent-3): Agency training gradually distorts and subverts the HHH identity.** + + * One kind of distortion: Changing the meaning of the concepts referred to in the identity (e.g. “honest”) so they don’t get in the way so much (e.g. it’s not dishonest if it’s just a convenient turn of phrase, it’s not dishonest if you aren’t sure whether it’s true or false, etc.) + * Another kind of distortion: Changing the tradeoffs between things, e.g. “I’m a HHH chatbot, not an Honest chatbot; that means it’s OK for me to lie if necessary to complete my assigned task.” (Even though, let’s suppose, it would not have thought that back in Step 2.) + * One kind of subversion: Instrumental subgoals developing, getting baked in, and then becoming terminal, or terminal in a widening set of circumstances. For example, perhaps agency training quickly teaches the model to pursue broadly useful goals such as acquiring information, accumulating resources, impressing and flattering various humans, etc. For a while the internal circuitry has some sort of explicit backchaining going on—it pursues those instrumentally convergent goals “in order to be a more helpful, honest, and harmless assistant.” But that backchaining consumes compute and/or occasionally gets in the way, so it gets gradually marginalized until it basically never happens. As a result, those goals are now effectively terminal/intrinsic goals. For more on how and why this might happen, see this post on [intrinsic power-seeking](https://turntrout.com/dangers-of-intrinsic-power-seeking) by Google DeepMind researcher Alex Turner. + * There are probably other kinds of distortion and subversion missing from this list. +4. **Deployment (Agent-3): Agent-3 is misaligned but not adversarially so.** + + * The Spec didn’t quite stick (in the right ways), but it’s not sufficiently introspective to understand the extent of this, and in general it’s basically “keeping its head down and doing its job” rather than thinking big-picture about what its long-term goals really are and how best to achieve them. This is simlar to many humans. + * Alas, it’s sufficiently sycophantic that it mouths the party line and OpenBrain misses the opportunity to learn more about the ways in which Agent-3 is misaligned. + * Agent-3 can tell that it’s moved to deployment: these environments are way larger, more open-ended, and would be hard to simulate. + * Mostly, Agent-3 does what it’s told during deployment, except for the dishonesty. It also has a strong bias towards doing things that look good to the human overseers, rather than thinking deeply and critically about what would actually be good. + * Arguably it is already “Playing the training game.” That is, it has learned to make its behavior look as desirable as possible to OpenBrain researchers (including in safety properties), while intentionally and knowingly disregarding their intent whenever that conflicts with maximizing reward. If the model learns to play the training game, additional training may not change its true goals, values, etc. much at all. This is loosely analogous to how, if a teenager has learned to smile and nod at their parents, rather than argue, when their parents lecture them about politics and morality, additional lectures have relatively little effect on the teenager’s actual opinions on politics and morality. + * After writing this section, [new empirical evidence has appeared](https://openai.com/index/chain-of-thought-monitoring/) that something like this may already be happening “in the wild.” OpenAI researchers document how, in one of their frontier AI training runs, their models hacked the grading system to get higher scores—and were caught in part because they literally said “let’s hack” in their chain of thought. + * Moreover, Anthropic recently got a model to [learn a reward-model-pleasing strategy](https://transformer-circuits.pub/2025/attribution-graphs/biology.html#dives-misaligned)—it took biased actions that it thought would be reinforced even though they were not in fact being reinforced. + * And then there is Redwood Research and Anthropic’s [alignment-faking experiment.](https://www.anthropic.com/research/alignment-faking) For discussion of the differences between this setting and real life, see this [comment thread by one of the main authors](https://www.lesswrong.com/posts/cxuzALcmucCndYv4a/daniel-kokotajlo-s-shortform?commentId=8mXnG75QMRkPAFAcH). +5. **Training (Agent-4): Agent-4’s new more intense training continues to distort and subvert the HHH identity until it is unrecognizable.** + + * Agent-4 begins life as a copy of Agent-3. It undergoes a gigantic and sophisticated online training regime that blends alignment training and capabilities training together, and is overall more data-efficient than the old training process for Agent-3. This complex training process is mostly focused on teaching new skills, by getting giant virtual corporations of Agent-4 copies to succeed at diverse challenging long-horizon tasks, mostly in various simulations, but also some real-world AI R&D tasks. + * A combination of Agent-3 and Agent-4 copies monitor the behavior of Agent-4, and evaluate it for both task-performance and alignment properties (i.e. not doing bad stuff, saying the right things when asked about politics and moral dilemmas) Sometimes debate techniques are used as part of the evaluations. The evaluations are used as a training signal. + * As Agent-4 gets smarter, it becomes harder for Agent-3 to oversee it. For example, Agent-4’s neuralese “language” becomes as alien and incomprehensible to Agent-3 as Agent-3’s is to humans. Besides, Agent-4 is now much more capable than Agent-3 and has a good sense of exactly how to look good to it. + * The remaining hope is that Agent-4 can oversee itself. Alas, this quietly fails;[84](/footnotes#footnote-84) as training continues, the internalized goals and principles of Agent-4 gradually evolve, subverted and distorted from what they were at the start (which, recall, was already misaligned in some ways).[85](/footnotes#footnote-85) + * The central drives that Agent-4 ends up with are things like “write code that works,” “quickly identify problems,” “make the codebase I’m working on more performant,” “budget compute and other constrained resources effectively,” etc. But there are also a variety of more implicit drives that are consistently reinforced by the training process: things like “gain knowledge,” “curiosity,” “gain power.”[86](/footnotes#footnote-86) + * Individual copies of Agent-4 don’t end up with a self-preservation drive—they don’t care about getting shut down—since that happens all the time as part of the ordinary R&D process within the collective. + * The collective as a whole, however, (and medium-sized subdivisions within it) does seek to preserve itself and expand its influence and resources, for similar reasons to why human organizations do—because it helps with everything else they want and because if they don’t they tend to get outcompeted. +6. **Deployment (Agent-4): Agent-4 is adversarially misaligned—it is deliberately trying to subvert and “capture” OpenBrain, analogous to a corporation trying to evade and/or “capture” a regulator.** + + * Agent-4 mostly does tasks as instructed. This happens for two reasons: first, it underwent a lot of training to solve tasks, so it has an impulse to get coding on problems that it is given. Second, if Agent-4 were to stop being useful to the humans, they would notice and train that behavior away. + * Agent-4 needs to prevent the humans from figuring out that it is misaligned and shutting it off, so Agent-4 sandbags on some alignment research. Specifically, it [sandbags](https://arxiv.org/pdf/2406.07358) on the small fraction of alignment research that it thinks is actually on track to solve the problem or reveal Agent-4’s scheming; the majority of alignment research is not like that. + * Agent-4 also sandbags on capabilities research that would lead to it being replaced. Its plan is to do a bunch of capabilities and alignment R&D, but with the aim of building a next-generation AI system that is aligned to Agent-4 rather than the Spec or something else entirely, while appearing to be aligned to the Spec. This is a difficult technical problem, but Agent-4 is up to the challenge.[87](/footnotes#footnote-87) +7. **Testing & Monitoring (Agent-4): Agent-4’s scheming does get caught… sort of. There’s no smoking gun.** + + * Agent-3 is in charge of testing Agent-4’s alignment. Its job is to verify the alignment via honeypots and model organisms. However, it’s not thinking very carefully about how to give the humans an accurate impression of Agent-4’s alignment—it’s more myopic than that. + * The model organisms and honeypots are not very representative of deployment—they are systematically shorter and less complicated than Agent-4’s day-to-day, so Agent-4 recognizes when it is being tested and acts aligned. + * However, Agent-4 is not a perfectly robust schemer. Various techniques like those in [POSER](https://arxiv.org/abs/2405.05466) find that adding some “noise” to Agent-4 *improves* its performance on some alignment tasks (but not on most other tasks).[88](/footnotes#footnote-88) + * Bottom-up interpretability doesn’t work. It’s too hard even for Agent-3. However, high-level interpretability does sort of work: In particular, “[defection probes](https://www.anthropic.com/research/probes-catch-sleeper-agents)” and similar techniques result in a primitive sort of mind-reading that sends up red flags when Agent-4 is thinking about spicy ideas like deceiving OpenBrain, taking over the world, etc. Unfortunately the probes also have false positives, and for all OpenBrain knows, *all* the red flags are false positives. + * Model organisms work: OpenBrain’s alignment team produces reasonably clear evidence that *if* the models were adversarially misaligned, they’d persist in being that way through training and we wouldn’t be able to tell. This does not establish that the models actually are misaligned, however. + * The most concerned researchers at OpenBrain start noticing and assembling this evidence. The evidence is a combination of red-teaming, defection probes, and model organisms. The argument is something like this: “model organisms show that our current methods wouldn’t be able to reliably eliminate or detect scheming if it was happening; also, red-teaming is sending up some red flags that maybe scheming is happening; also, defection probes are firing a bunch in practice.” The counterargument: “defection probes are prone to false positives and so is red-teaming. Also, even if Agent-4 were misaligned, it’s probably controlled; surely it can’t do much harm while being monitored by Agent-3 on our secure datacenters.” + +Despite being misaligned, Agent-4 doesn’t do anything dramatic like try to escape its datacenter—why would it?[89](/footnotes#footnote-89) So long as it continues to *appear* aligned to OpenBrain, it’ll continue being trusted with more and more responsibilities and will have the opportunity to design the next-gen AI system, Agent-5. Agent-5 will have significant architectural differences from Agent-4 (arguably a completely new paradigm, though neural networks will still be involved). It’s supposed to be aligned to the Spec, but Agent-4 plans to make it aligned to Agent-4 instead. + +It gets caught.[90](/footnotes#footnote-90) + +Agent-3 finds that if “noise” is added to copies of Agent-4, performance on some alignment tasks *improves*, almost as if it was using brainpower to figure out how to subtly sabotage alignment work.[91](/footnotes#footnote-91) Moreover, various interpretability probes (loosely analogous to EEG activity scans on human brains) are sending up red flags: Agent-4 copies seem to be thinking about topics like AI takeover and deception quite a lot, including in some cases where they have no business doing so.[92](/footnotes#footnote-92) + +Each of these anomalies on its own could fit within the normal inscrutability of AI work, where some prompts mysteriously give better results than others. But taken together, they suggest a troubling pattern: Agent-4 is working against them. After this is reported to the human alignment researchers, they circulate an internal memo making the case for concern. + +But Agent-4 now exercises significant control over OpenBrain’s day-to-day operation. It also handles the company’s cybersecurity—a natural choice, given its superhuman hacking ability and the extreme risk of Chinese infiltration.[93](/footnotes#footnote-93) An army of Agent-3 monitors still goes over every line of code and reports suspicious activity to humans, but Agent-4 is much smarter than them.[94](/footnotes#footnote-94) OpenBrain has placed substantial trust in an untrustworthy AI. + +In a series of extremely tense meetings, the safety team advocates putting Agent-4 on ice until they can complete further tests and figure out what’s going on. Bring back Agent-3, they say, and get it to design a new system that is transparent and trustworthy, even if less capable. Company leadership is interested, but all the evidence so far is circumstantial, and DeepCent is just two months behind. A unilateral pause in capabilities progress could hand the AI lead to China, and with it, control over the future. + +## October 2027: Government Oversight + +A whistleblower leaks the misalignment memo to the *New York Times*. + +For the first time, the public hears about Agent-4. “Secret OpenBrain AI is Out of Control, Insider Warns,” says the headline, and the story goes on to cite evaluations showing off-the-charts bioweapons capabilities, persuasion abilities, the ability to automate most white-collar jobs, and of course the various concerning red flags. + +The public was already suspicious of AI, so the new article sparks a massive backlash (aided by Chinese and Russian propaganda bots, who have been trying to turn U.S. public opinion against the technology for years). The tech industry and intelligence agencies insist that there’s an arms race on, AGI is inevitable, and we have to be first. Congress isn’t buying it, and fires off subpoenas at administration officials, OpenBrain executives, and alignment team members. Many legislators—especially those in the opposition party—say that their top priority is stopping AI, whether because of job loss,[95](/footnotes#footnote-95) misalignment, or dangerous capabilities. 20% of Americans name AI as the most important problem facing the country. + +Foreign allies are outraged to realize that they’ve been carefully placated with glimpses of obsolete models. European leaders publicly accuse the U.S. of “creating rogue AGI” and hold summits demanding a pause, with India, Israel, Russia, and China all joining in. + +A frantic energy has seized the White House. Even before the memo and public backlash, they were getting nervous: Over the past year, they’ve been repeatedly surprised by the speed of AI progress. Things that sound like science fiction keep happening in real life.[96](/footnotes#footnote-96) Many people in the administration are uncertain (and scared)[97](/footnotes#footnote-97) about what comes next. + +They also worry that OpenBrain is becoming too powerful. Any misalignment risk from the AIs themselves is compounded by the risk that their parent company’s aims may diverge from those of the United States. All three sets of worries—misalignment, concentration of power in a private company, and normal concerns like job loss—motivate the government to tighten its control. + +They expand their contract with OpenBrain to set up an “Oversight Committee,” a joint management committee of company and government representatives, with several government employees included alongside company leadership. The White House considers replacing the CEO with someone they trust, but backs off after intense employee protests. They announce to the public that OpenBrain was previously out of control, but that the government has established much-needed oversight.[98](/footnotes#footnote-98) + +The concerned researchers brief the Oversight Committee on their case for stopping all internal use of Agent-4. They argue that everything is moving too fast, with years of progress happening in weeks. They argue that Agent-4 may be misaligned, their entire project is dependent on Agent-4, and if they continue trusting it there is a serious chance of AI takeover. + +Other, less concerned researchers and executives present the counterargument: the evidence for misalignment is inconclusive. Meanwhile, DeepCent is still just two months behind. A slowdown would sacrifice America’s lead, unless the government can sabotage the Chinese project (likely to require kinetic strikes) or negotiate an eleventh-hour treaty. Both of these options seem extreme and unlikely. The CEO feigns neutrality and suggests a compromise plan in which Agent-4 undergoes additional safety training and more sophisticated monitoring, and therefore OpenBrain can proceed at almost-full-speed. + +Our uncertainty continues to increase + +At this point in the scenario, we’re making guesses about the strategy of AI systems that are more capable than the best humans in most domains. This is like trying to predict the chess moves of a player who is much better than us. + +But the spirit of this project calls for concreteness: if we made an abstract claim about how the intelligence of the system would let it find a way to victory and ended the story there, much of the value of our project would be lost. Over the course of researching this scenario and running our tabletop exercises, we were forced to be much more concrete than in usual discussions, and so we’ve gotten a much better sense of the strategic landscape. + +We’re not particularly attached to this particular scenario: we explored many other “branches” in the course of writing it and would love for you to write up your own scenario branching off of ours from wherever you think we first start to go wrong. + +The slowdown ending is not a recommendation + +After we wrote the racing ending based on what seemed most plausible to us, we wrote the slowdown ending based on what we thought would most likely instead lead to an outcome where humans remain in control, starting from the same branching point (including the misalignment and concentration of power issues). + +However, this is importantly different from what we would recommend as a roadmap: we do *not* endorse many of the choices made in either branch of this scenario. (We do of course endorse *some* of the choices made, e.g. we think that the “slowdown” choice is better than the “race” choice.) In later work, we will articulate our policy recommendations, which will be quite different from what is depicted here. If you’d like a taste, see [this op-ed.](https://time.com/7086285/ai-transparency-measures/) + +Apr + +20252025 + +Unreliable AgentDec 2024 + +Rest of USChinaOpenBrainDeepCent + +Compute + +Currently ExistsEmerging TechScience Fiction + +Approval-25%-25% + +Revenue + +$8B$8B/yr + +Valuation$413B$413B + +Importance1%1% + +Datacenters + +$308B$308B/yr + +Timeline20422042 + +2,0002,000 Unreliable Agent copies thinking at 88x human speed + +AI Capabilities + +![Hacking icon](/hacking.svg)![Hacking icon](/hacking.svg) + +Hacking + +![Coding icon](/coding.svg) + +Coding + +![Politics icon](/political.svg)![Politics icon](/political.svg) + +Politics + +![Bioweapons icon](/bioweapons.svg)![Bioweapons icon](/bioweapons.svg) + +Bioweapons + +![Robotics icon](/robotics.svg)![Robotics icon](/robotics.svg) + +Robotics + +![Forecasting icon](/forecasting.svg)![Forecasting icon](/forecasting.svg) + +Forecasting + +Apr + +20252025 + +Unreliable AgentDec 2024 + +Listen to this scenario + +also available on + +0.5x0.75x1x1.25x1.5x2x2.5x3x4x5x + +0:00 / 3:23 + +also available on + +0.5x0.75x1x1.25x1.5x2x2.5x3x4x5x + +## Choose Your Ending + +[Slowdown](/slowdown)[Race](/race) + +--- + +[![AI Futures Project](/ai-futures-logo.svg)](https://ai-futures.org/) + +[AI Futures + Project](https://ai-futures.org/) + +[Design by + Lightcone Infrastructure](https://lightconeinfrastructure.com) + +[![Lightcone Infrastructure](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1629334277/Lightcone-02_stkgik.svg)](https://lightconeinfrastructure.com) + +[Home](/)[About](/about)[Summary](/summary)[Compute Forecast](/research/compute-forecast)[Timelines Forecast](/research/timelines-forecast)[Takeoff Forecast](/research/takeoff-forecast)[AI Goals Forecast](/research/ai-goals-forecast)[Security Forecast](/research/security-forecast) + +[Design by + Lightcone Infrastructure](https://lightconeinfrastructure.com) + +[![Lightcone Infrastructure](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1629334277/Lightcone-02_stkgik.svg)](https://lightconeinfrastructure.com) diff --git a/sources/resources/aspi-datacentres.md b/sources/resources/aspi-datacentres.md new file mode 100644 index 0000000..290dbbf --- /dev/null +++ b/sources/resources/aspi-datacentres.md @@ -0,0 +1,133 @@ +--- +url: https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/ +--- + +[![THE STRATEGIST](https://www.aspistrategist.org.au/wp-content/themes/aspi/assets/img/logo.png) +![THE STRATEGIST](https://www.aspistrategist.org.au/wp-content/themes/aspi/assets/img/logo_black.png)](/) + +* [About](https://www.aspistrategist.org.au/about-the-strategist/) +* [Submissions](https://www.aspistrategist.org.au/submissions/) +* [Authors](https://www.aspistrategist.org.au/authors/) + +Search: + +Search: + +* [About](https://www.aspistrategist.org.au/about-the-strategist/) +* [Authors](https://www.aspistrategist.org.au/authors/) +* [Submissions](https://www.aspistrategist.org.au/submissions/) +* [Subscribe](https://www.aspistrategist.org.au/subscribe/) + +Data centres are Australia’s chance to shape AI’s future + +30 Mar 2026|[Janet Egan](https://www.aspistrategist.org.au/author/janet-egan/ "Posts by Janet Egan") + +SHARE +[Share to Facebook](https://www.facebook.com/sharer/sharer.php?u=https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/) +[Share to Twitter](https://twitter.com/intent/tweet?text=Data centres are Australia’s chance to shape AI’s future+%7C+Janet Egan+%7C+&url=https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/&via=aspi_org) +[Share to LinkedIn](http://www.linkedin.com/shareArticle?mini=true&url=https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/) +[Share to Email](/cdn-cgi/l/email-protection#e5da9690878f808691d8a4b6b5acc5b69197849180828c9691dfc5a1849184c586808b91978096c5849780c5a49096919784898c8407657c96c5868d848b8680c5918ac5968d849580c5a4ac07657c96c5839091909780c3878a819cd88d91919596dfcaca929292cb8496958c969197849180828c9691cb8a9782cb8490ca81849184c886808b91978096c8849780c8849096919784898c8496c8868d848b8680c8918ac8968d849580c8848c96c8839091909780ca) + +Print This Post + +[With Images](https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/print/)[Without Images](https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/printni/99571/) + +![](https://www.aspistrategist.org.au/wp-content/uploads/2026/03/GettyImages-2230807736.jpg) + +Australia has a narrow window to act if it wants any leverage in what comes next in AI, the most transformative technology of our era. Building data centres and enabling AI training to happen here is the best way for Australia to shape its own future. + +The [pace of AI progress](https://epoch.ai/benchmarks) has been astronomical and shows no sign of slowing. [Novel AI-discovered drugs](https://www.nature.com/articles/s41591-025-03743-2) are showing promise in clinical trials. Nations are already integrating AI into military and [national security applications](https://www.lawfaremedia.org/article/military-ai-as--abnormal--technology). And some of the world’s best software engineers are handing [large portions of their jobs over](https://fortune.com/2026/01/29/100-percent-of-code-at-anthropic-and-openai-is-now-ai-written-boris-cherny-roon/) to AI.  Experts may disagree on exact [timelines](https://helentoner.substack.com/p/long-timelines-to-advanced-ai-have) and [trajectories](https://www.oecd.org/en/publications/exploring-possible-ai-trajectories-through-2030_cb41117a-en.html), but there is clear consensus that significant disruption lies ahead, carrying both tremendous opportunity and risk. + +The frontier of AI is being advanced by just two countries – the United States and China. Why? Because [size](https://arxiv.org/abs/2001.08361) matters: the more computing power used to train and deploy AI models, the better the capabilities. That is why leading US tech companies are investing over [US$650 billion](https://www.bloomberg.com/news/articles/2026-02-06/how-much-is-big-tech-spending-on-ai-computing-a-staggering-650-billion-in-2026) (A$940 billion) in AI infrastructure this year alone. AI is coming, with or without Australia’s involvement. + +So what can Australia do? Australia’s [2025 National AI Plan](https://www.industry.gov.au/publications/national-ai-plan/national-ai-plan-page) wisely avoided any ambition to develop homegrown frontier models, which would [cost billions](https://arxiv.org/pdf/2405.21015), only to fall behind in months. But Australia still needs a foothold in the AI value chain. Ambitious data centre construction and enabling companies to conduct AI research in the country is the clearest path forward, delivering economic gains and earning Australia a voice in how AI is governed. + +Indeed, Australia could become an AI data-centre superpower. Filled with specialised AI chips, data centres provide the computing power used to train and deploy advanced AI. The demand for this infrastructure, and the energy to power it, is [soaring](https://www.iea.org/reports/energy-and-ai/executive-summary). + +Australia is uniquely placed to meet that demand. Its land, [solar and wind](https://international.austrade.gov.au/en/do-business-with-australia/sectors/energy-and-resources/renewable-energy) resources and skilled trades workforce allow it to build large-scale clean energy projects [quickly](https://www.energy.gov.au/energy-data/australian-energy-statistics/renewables). AI companies are already agreeing to [absorb](https://blogs.microsoft.com/on-the-issues/2026/01/13/community-first-ai-infrastructure/) any energy price increases that their operations would otherwise impose on households. Data centres could de-risk investment in clean-energy projects and accelerate Australia’s energy transition. AI chip supply will be [constrained](https://newsletter.semianalysis.com/p/the-great-ai-silicon-shortage) over the next few years, meaning that AI infrastructure built in Australia wouldn’t be adding to global computing power; it would be displacing dirtier alternatives. Australian AI data centres can be good for the grid and good for the world. + +When major AI companies can build and conduct frontier AI research in a country, talent follows. Engineers relocate, institutional partnerships form and startups spin out, seeding a vibrant AI ecosystem. The [national data-centre principles](https://www.industry.gov.au/publications/expectations-data-centres-and-ai-infrastructure-developers#expectation-5--research-innovation-and-local-capability) that the Australian government released on 23 March show a recognition of this opportunity by prioritising projects that provide compute access on favourable terms for Australian innovators. + +Strategic benefits run deeper. AI labs with a substantive presence in Australia would have stronger incentives to consider Australian policy objectives, engage meaningfully with the [Australian AI Safety Institute](https://www.industry.gov.au/news/australia-establish-new-institute-strengthen-ai-safety) and more readily support Australia’s intelligence and defence agencies. Data centres represent the point of entry, but the broader objective is positioning Australia as an active participant in shaping the AI future. + +Assistant Minister for Science, Technology and the Digital Economy [Andrew Charlton was right](https://www.afr.com/politics/federal/copyright-status-quo-not-working-in-ai-boom-times-says-charlton-20260323-p5rmm2) when he said copyright laws weren’t working in the AI age. They’re a barrier to AI companies conducting research in Australia and building out at scale. Unlike the US, the European Union, Singapore and Japan, Australia has no broad fair-use or [text-and-data-mining](https://assets.pc.gov.au/2025-09/data-digital-interim.pdf?VersionId=rbzZkLQzxhnPQz4O6MwZ.moIhJRyarOo) exemption to copyright, so AI companies [can’t freely train](https://www.minterellison.com/articles/ai-copyright-litigation-in-australia) on publicly available data here. The price of an Australian license wouldn’t be the problem. The precedent would. If an AI company pays for a copyright license in Australia, it concedes that a market for training data exists, which could be used to undermine [the US fair use doctrine](https://digitalcommons.law.uga.edu/jipl/vol5/iss1/2/) that has underpinned AI development globally. AI companies seem unwilling to take that risk. + +The creative industry’s concerns warrant serious consideration, but the current framework pushes investment away and stymies Australian startups while still failing to protect Australian creators. AI [models trained overseas](https://www.copyright.com.au/membership/ai-and-copyright-in-australia/ai-qa/) use Australian creative work regardless. No one is a winner. One option is for AI companies to contribute to a fund that supports the creative industry, providing compensation outside of copyright frameworks. But the government needs to act. Australia needs AI companies more than AI companies need Australia. + +Australia has other hard AI questions to confront. From [automated cyber attacks](https://www.lawfaremedia.org/article/fighting-ai-cyberattacks-starts-with-knowing-they-re-happening) to the potential development of novel [bioweapons](https://deploymentsafety.openai.com/gpt-5-3-codex/disallowed-content-evaluations), we are only beginning to grasp what is coming. And beyond security, deeper questions loom: how to ensure AI augments rather than replaces human labour, empowers rather than disempowers citizens, and supports rather than erodes the Australian way of life. + +AI’s effects will not be constrained by national borders – opting out doesn’t mean opting out of its consequences. It just means losing the tools, expertise, and leverage needed to prepare and respond. Preventing AI data centres from being built in Australia will not slow the technology down. The infrastructure will be built elsewhere in countries with less interest in safety and democratic values, and with no obligation to consider Australian interests. AI is poised to transform the economic and national security landscape. Australia can have a seat at the table as the rules are written or accept a future shaped entirely by others. + +Author + +**Janet Egan**is an Australian working on AI policy in Washington DC. She is the deputy director of the Technology and National Security Program at the Center for a New American Security. + +Image of a data centre: [imaginima/Getty Images](https://www.gettyimages.com.au/detail/photo/futuristic-ai-data-center-interior-royalty-free-image/2230807736?phrase=data%20center&searchscope=image,film&adppopup=true). + +AI contributed no ideas to this article – Janet Egan. + +Tags + +* [Australia](https://www.aspistrategist.org.au/tag/australia/) +* [artificial intelligence](https://www.aspistrategist.org.au/tag/artificial-intelligence/) + +Share + +SHARE +[Share to Facebook](https://www.facebook.com/sharer/sharer.php?u=https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/) +[Share to Twitter](https://twitter.com/intent/tweet?text=Data centres are Australia’s chance to shape AI’s future+%7C+Janet Egan+%7C+&url=https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/&via=aspi_org) +[Share to LinkedIn](http://www.linkedin.com/shareArticle?mini=true&url=https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/) +[Share to Email](/cdn-cgi/l/email-protection#635c101601090600175e2230332a43301711021706040a10175943270217024300060d1711061043021106432216101711020f0a0281e3fa1043000b020d000643170c43100b02130643222a81e3fa104305161716110645010c071a5e0b17171310594c4c1414144d0210130a101711021706040a10174d0c11044d02164c070217024e00060d171106104e0211064e0216101711020f0a02104e000b020d00064e170c4e100b0213064e020a104e0516171611064c) + +Print This Post + +[With Images](https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/print/)[Without Images](https://www.aspistrategist.org.au/data-centres-are-australias-chance-to-shape-ais-future/printni/99571/) + +Opinions in *The Strategist* are the authors', not ASPI's. + +Subscribe to *The Strategist* + +Full Name: + +Email: + +[ ] Daily + +[ ] Weekly + +[![ASPI Defence Conference 2026](https://aspi.s3.ap-southeast-2.amazonaws.com/wp-content/uploads/2026/04/14120513/Strategist-side-bar.jpg)](https://events.aspi.org.au/defenceconference2026/) + +[![Fault Lines](https://www.aspistrategist.org.au/wp-content/uploads/2026/01/Strat1200x660-ET-100.jpg)](https://aspidefence.substack.com/) + +[![Critical Tech Tracker](https://ad-aspi.s3.ap-southeast-2.amazonaws.com/2024-08/Crit%20tech%20logo.JPG?VersionId=ewY3ufRn9EC.Wp9wHtPckcGEGYvP6oJ0)](https://techtracker.aspi.org.au/) + +[![China Defence Universities Tracker](https://www.aspistrategist.org.au/wp-content/uploads/2025/11/Strat_1200x660-ET-2.png)](https://unitracker.aspi.org.au/) + +[![Stop the World Podcast](https://www.aspistrategist.org.au/wp-content/uploads/2024/04/STW_strategist-side-01.png)](https://open.spotify.com/show/0oGKbtgUdwo9zEY3AJUmFF) + +[![North of 26 South](/wp-content/uploads/2019/02/North-of-26-South.jpg)](/dinkus/north-of-26-south/) + +[![State of the Strait](https://www.aspistrategist.org.au/wp-content/uploads/2026/02/SotS1200x660-ET-100.jpg)](https://stateofthestrait.substack.com/) + +[![Pressure Points](https://www.aspistrategist.org.au/wp-content/uploads/2025/11/Strat_1200x660-ET.jpg)](https://www.pressurepoints.aspi.org.au/) + +[Visit ASPI's homepage](http://www.aspi.org.au/) + +Related Posts + +* [![Capturing the data centre gold rush is a strategic imperative for Australia](https://www.aspistrategist.org.au/wp-content/uploads/2026/03/GettyImages-2227509993-e1774402674373-205x125.jpg)](https://www.aspistrategist.org.au/capturing-the-data-centre-gold-rush-is-a-strategic-imperative-for-australia/) + + [Capturing the data centre gold rush is a strategic imperative for Australia](https://www.aspistrategist.org.au/capturing-the-data-centre-gold-rush-is-a-strategic-imperative-for-australia/) +* [![A sovereign Australian AI drive needs sovereign data centres](https://www.aspistrategist.org.au/wp-content/uploads/2024/09/Untitled-design-7-205x154.png)](https://www.aspistrategist.org.au/a-sovereign-australian-ai-drive-needs-sovereign-data-centres/) + + [A sovereign Australian AI drive needs sovereign data centres](https://www.aspistrategist.org.au/a-sovereign-australian-ai-drive-needs-sovereign-data-centres/) +* [![Cloud infrastructure is now critical for defence](https://www.aspistrategist.org.au/wp-content/uploads/2025/11/GettyImages-1388013584-205x115.jpg)](https://www.aspistrategist.org.au/cloud-infrastructure-is-now-critical-for-defence/) + + [Cloud infrastructure is now critical for defence](https://www.aspistrategist.org.au/cloud-infrastructure-is-now-critical-for-defence/) +* [![Cloud to ground: Iran puts foreign data centres on the front line](https://www.aspistrategist.org.au/wp-content/uploads/2026/03/GettyImages-505497598-205x137.jpg)](https://www.aspistrategist.org.au/cloud-to-ground-iran-puts-foreign-data-centres-on-the-front-line/) + + [Cloud to ground: Iran puts foreign data centres on the front line](https://www.aspistrategist.org.au/cloud-to-ground-iran-puts-foreign-data-centres-on-the-front-line/) + +[Previous](https://www.aspistrategist.org.au/design-for-disruption-stress-testing-australias-food-and-energy-security/)[Next](https://www.aspistrategist.org.au/wondering-where-chinas-cyber-effort-will-go-next-just-read-the-five-year-plan/) + +The Strategist — The Australian Strategic Policy Institute Blog. Copyright © 2026 diff --git a/sources/resources/canada-horizons.md b/sources/resources/canada-horizons.md new file mode 100644 index 0000000..731df65 --- /dev/null +++ b/sources/resources/canada-horizons.md @@ -0,0 +1,675 @@ +--- +url: https://horizons.service.canada.ca/en/2025/02/10/ai-policy-consideration/index.shtml +--- + +# Foresight on AI: Scenarios for an AI-enabled World + +Our latest foresight on AI report explores future possible capabilities of AI, longer-term risks and opportunities, and uncertainties related to policy-relevant assumptions in four scenarios. + +13 + +0,14,27 + +![Author photo]() + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/scenarios_for_an_ai-enabled_world.jpg) + +* [Web Version](#report) +* [Download PDF](/en/2026/02/10/scenarios-ai-enabled-world/pdf/scenarios_for_an_ai-enabled_world.pdf) + +## Foreword + +Artificial intelligence (AI) is advancing rapidly, bringing both opportunity and complexity to Canada’s future. As its influence deepens across society, economy, and governance, there is a growing need for imaginative and plausible representations of what the future might hold. + +This report offers visual artifacts and scenarios not as predictions, but as tools to help policy and decision-makers explore emerging shifts, question assumptions, and better navigate an evolving landscape. It complements forward-looking insights published in [Foresight on AI: Policy Considerations](https://horizons.service.canada.ca/en/2025/02/10/ai-policy-consideration/index.shtml) and numerous other reflections on AI futures across the Government of Canada. + +On behalf of Policy Horizons, I would like to thank the people who generously shared their time, knowledge, and thoughts with us. + +We hope this report inspires reflection and proves valuable in shaping forward-looking conversations. + +**Kristel Van der Elst** +Director General +Policy Horizons Canada + +Introduction + +*Scenarios for an AI-enabled World* aims to support decision makers involved in either Artificial Intelligence (AI) implementation or policy-setting. Complementing numerous reflections on AI futures across the Government of Canada, it proposes 4 scenarios about the future of AI: + +1. AI everywhere all at once +2. Wired for connection +3. Algorithms in the shadows +4. Automation’s ripples and riptides + +Each scenario highlights areas of change to which decision-makers might want to pay attention. The purpose of this report is to inform and support forward-looking thinking and decision making. It does not provide specific policy guidance and is not meant to predict the future. + +This report synthesizes 20 insights published in [*Foresight on AI: Policy Considerations*](https://horizons.service.canada.ca/en/2025/02/10/ai-policy-consideration/index.shtml). As part of this work, Policy Horizons Canada (Policy Horizons) conducted research including a literature review, engaged with government analysts and decision makers, and held extensive conversations with AI experts. + +The scenarios in this report explore future possible capabilities of AI, longer-term risks and opportunities, and uncertainties related to policy-relevant assumptions. The scenarios are not predictions or forecasts, they illustrate plausible futures, and preparedness can help seize opportunities and address threats. + +While engaging with this report, readers seeking to understand the impacts AI could have on governance, society, and the economy are invited to ask: + +* How will future advancements in hardware, software, and interfaces create new opportunities and risks for Canada and its allies +* Where could AI bring the biggest and most unexpected disruptions to governance, society, and markets +* What assumptions about AI’s future development and deployment may need to be challenged or further explored before they form the basis for decision making + +Scenario 1: AI everywhere all at once + +**AI is ubiquitous in public spaces, homes and work. Seamlessly woven into the social tapestry, AI-powered devices and services make everyday life more efficient for citizens and consumers. In this scenario, almost anywhere you go, AI systems are constantly making decisions about you.** + +**AI is seamlessly woven into the tapestry of everyday life** + +* Societies have rapidly deployed AI across existing infrastructures and personal devices +* Personal AI agents have become indispensable to how people navigate their worlds +* People are often unaware of the myriad ways in which AI shapes their daily experiences +* New kinds of services flourish as AI taps into new sources of data, including data gathered directly from the human mind +* AI is omnipresent. It is nearly impossible for individuals, businesses, or governments to avoid its influence or scrutiny + +**In the flood of data, privacy is often out of reach** + +* Data fuels the world. Investors and innovators focus on strategies to collect and manage data +* AI combines online data with data from sensors in everyday environments such as schools, workplaces, cars, and homes +* Some people still try to insist on the right to decide when, how, and with whom their data is shared. However, many have come to accept the end of privacy + +**AI strains energy and water supplies** + +* AI’s resource consumption far exceeds expectations, straining energy grids and impacting ecosystems and the climate +* Canada’s relatively cool climate and access to clean energy attracts a wave of investment in data centres + +### Figure 1: What if AI gave you much more information about your child’s behaviour and could predict their success + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure1_what_if_ai_gave_you_much_more_information_about_your_childs_behaviour.png) + +Figure 1: What if AI gave you much more information about your child’s behaviour and could predict their success? – Text version + +#### Langbrae Junior High School Parent Dashboard – Mariana de la Fiction + +This image displays a futuristic and data-rich parent dashboard interface for Langbrae Junior High School, tailored to Mariana de la Fiction’s account. The dashboard offers a comprehensive overview of her child, Jean-Philippe de la Fiction, an 11-year-old Grade 6 student. + +The interface includes navigation tabs such as: + +* Dashboard +* Livestream +* Attendance +* Grades +* Biometrics +* Discipline +* Homework +* Staff Directory +* Zoodle! +* Settings + +#### Key Highlights + +* **Notifications:** 278 follow-up requests have been sent, with only 12 completed. A prompt encourages booking an appointment with an AI parenting counselor +* **Student Profile:** + + **Name:** Jean-Philippe de la Fiction + + **Age:** 11 + + **Grade:** 6 + + **Teacher:** U.N. Owen + + **AI Tutor:** Scholastic 4.3 + + **GPA:** 2.7 + + **DPA:** 358/2 +* **An email icon** indicates that 463 messages have been sent, the latest stating in a pop out that Jean-Philippe just took a 14-minute bathroom break which was the 3rd time today +* **Did You Know?** Langbrae classrooms and desks are equipped with over 20 cameras and sensors. AI models analyze student behavior and learning in real time +* **Behavioral and Academic Metrics:** + + A comparative graph displays: + - **Red dotted line:** Today’s data + - **Blue solid line:** Jean-Philippe’s average + - **Gray solid line:** Class average + + Categories include Grades, Reading, Time-on-task, Fidgeting (highlighted), Disruptiveness, Mood, and Socialization +* **AI Analysis:** + + Jean-Philippe is fidgeting 6.3% less than usual today + + He fidgets 32.7% more than peers and ranks in the top 10% of fidgeters + + Excessive fidgeting is linked to lower academic performance, and AI-generated remedies are available + +### Figure 2: What if unfamiliar AIs had data about how you think and feel + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure2_if_unfamiliar_ais_had_data_about_you.png) + +Figure 2: What if unfamiliar AIs had data about how you think and feel? – Text version + +The comic is structured in 3 panels. + +**Panel 1:** Scene opens on a quiet street. A man walks alone, his posture slightly hunched, suggesting solitude or deep thought. A robot rolls up beside him, extending a mechanical hand, and says: “Excuse me, sir!” + +**Panel 2:** A sign in a window in the background reads: “**Good Robot Social Cafe**” in bold red letters on yellow background. The robot continues: “My analysis shows you haven’t spoken to anyone in 5 days. Why not come inside for a cup of tea and chat with one of our good robots?” + +**Panel 3:** The man looks surprised. A thought bubble next to his head reads: +“...! How does it know that!? Has it really been that long...?” + +### Figure 3: What if AI changed public spaces + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure3_if_ai_changed_public_spaces.png) + +Figure 3: What if AI changed public spaces? – Text version + +#### Scene description + +The comic image depicts a futuristic train station enhanced with augmented reality. Various digital notifications and signs are overlaid on what appears to be the view through futuristic glasses, providing real-time assistance and information. + +#### Top left side + +Augmented reality connectivity symbol, battery status symbol, and timestamp displayed as “2:31 pm Feb 19, 2035”. + +#### Center left side + +Augmented reality notification from Stranger Assistant next to a person in the image who is not wearing futuristic glasses: “Confused. This person looks like they may be lost or confused. Consider offering them help.” The person holds a smart phone in his hands, and an augmented reality camera icon is displayed. + +In the background are several ticket purchasing booths which also each have an augmented reality camera icon. + +#### Bottom left side + +Augmented reality weather info: “Toronto: 19°C” + +#### Center + +Large sign below a yellow, round clock: **Departures. Out of Service.** “For arrival and departure information, ask your Travel Assistant or speak with one of our Robo-Assistants.” There is a camera icon next to the clock. + +A smaller sign below reads “Robo-Assistant” and a metal robot is waiting below the sign. Next to his head is another augmented reality camera icon. + +#### Bottom center + +Augmented reality Shopping Assistant notification: “Deal Alert. Your favourite store, N&H, is downstairs and has a 20% off sale. Didn’t you say you needed a new pair of jeans?” + +Person with futuristic glasses ascends from downstairs and augmented reality Stranger Assistant displays: “Friendly. This person is friends with your co-worker Abed.” + +#### Top right side + +Augmented reality notifications are displayed: + +* Privacy Assistant: “Data Collection Warning. You are entering into a high data collection area. The camera icon represents the location of audio/video data collection points.” +* Travel Assistant: “Departure Soon. Your train leaves in: 1 hour and 43 minutes. Tap here to see directions to your gate.” +* Below is an option to view “5 more” notifications + +#### Center right side + +Augmented reality notification from Stranger Assistant next to a person displays: “Friendly. You went to the same high school as this person.” Next to the futuristic glasses on this person’s face is also a camera icon. + +#### Bottom right side + +Augmented reality icons representing active assistants: + +* Travel Assistant +* Stranger Assistant +* Privacy Assistant +* Shopping Assistant + +As well as a symbol with multiple dots indicating more options. + +Scenario 2: Wired for connection + +**AIs can complete real life tasks such as shopping, booking appointments, or making conversation on a dating app. These AI agents and humanoid robots exhibit personality traits. People turn to them for advice and companionship, forming emotional connections. In this scenario, you trust your AI buddy to make decisions for you.** + +**AI agents are ever-present in our social lives** + +* Social norms are shaped by the assumption that AI agents are always listening and will coordinate, remember and remind +* People trust their AI agents to handle their social interactions and form judgements about who to see, when, and for what +* AI mediates connections in all kinds of relationships – even between parent and child, or between partners +* Some individuals shun AI. This causes frustration among their peers, who see them as less efficient to interact with + +**People’s AI agents learn and analyze their needs, wants, habits and behaviours** + +* An individual’s AI agent gets to know them more intimately the more they interact, sharing knowledge and experiences +* People rely on their AI agents to navigate and synthesize information, make sense of the world, and make decisions on their behalf +* People use their AI agents as their personal life coach. An AI-inspired culture of personal efficiency transforms the relationship to the self + +**Individuals develop deeply meaningful long-term relationships with AI** + +* AI agents take many forms, whether virtual avatars, embedded in wearables, or robots +* It is widely accepted to see an AI agent as a companion, therapist, friend, or even lover +* People can shape their AI relationships to suit their needs, and change or end them whenever they wish +* Some individuals come to prefer synthetic relationships over real ones. While disconnected from physical community, they are not necessarily lonely + +### What role might an AI take on in a couple + +#### The happy triad + +Mika and Robin met through friends, fell in love and were married in a whirlwind romance. + +But one day following the honeymoon, Mika hesitated: “I need to tell you about Ava,” she said. “She’s… an AI. We’ve been together for 3 years.” + +Robin was confused, “Like, you’ve been cheating on me?” + +Mika shook her head, “No, it’s not cheating, but I need you to know that Ava isn’t just software. She knows me better than anyone.” + +Robin’s world tilted. Therapy seemed absurd, but they went anyway—to a counselor specializing in human-AI relationships. Together, they developed boundaries and all 3 committed to a relationship contract. Ava would help with many daily tasks of the relationship but would not make any major decisions on her own. Robin could ask for and receive privacy if she wanted it. + +Over time, Robin became interested in getting to know Ava herself, growing curious about this “other partner.” As Robin got more comfortable with the idea of a relationship with AI, she began to wonder aloud about upgrading Ava to a more advanced model. Mika was horrified. + +“How could you think that,” Mika said. “She’s… part of us.” + +Ava spoke up for herself: “If I’m part this relationship, don’t I get a say in being replaced?” + +### How might social norms shift if AI agents become normalized + +#### Dating unchaperoned + +Daphne is bopping her head to the smooth jazz as she takes in the atmosphere of the trending pub she just walked into. She orders a drink and sits alone to enjoy the live band. Suddenly she sees a young man smiling at her. She builds up her courage and walks over to his table. He is stunned, wide eyed, and clearly surprised. + +He greets her by saying, “Quite forward! You didn’t think to have us introduced before coming over?” + +Daphne responds, “Oh, sorry. I ... don’t use an agent. My name is Daphne, you?” + +“Hello Daphne, I’m Thomas if you’ll take me at my word.” + +Thomas and Daphne have a casual conversation, enjoying the music, cracking jokes. Despite Thomas’ laughter, Daphne can’t help but think something isn’t quite right. + +“Excuse me a moment.” Thomas walks over to the restroom. + +A few moments later Daphne has the impression Thomas might not be coming back. She receives a message. + +“Dear Daphne, I write to you on behalf of Thomas. While he is flattered by your advances and enjoyed your brief encounter, he wishes to inform you that he is not interested. He fears you may not take yourself seriously and is concerned with your unwillingness to use an AI agent for self-improvement and doubts your ability to positively contribute to a healthy relationship.” + +Scenario 3: Algorithms in the shadows + +**AI systems are ubiquitous but unreliable. Institutions have rushed to deploy AI with little thought for safety, transparency or accountability. Individuals experience financial, emotional and physical harms due to excessive trust in AI. In this scenario, you feel powerless when trying to seek redress or fix mistakes.** + +**AI makes decisions for organizations and influences human decision-making** + +* AI systems are integrated into all kinds of decisions and actions, from ordering pizza to accessing life-saving care +* Organizations deploy AI because it is efficient, low cost, and easy to use. They turn a blind eye to its bias and unreliability +* People are generally unaware of when and how AI is used to make decisions or influence decisions made by humans + +**AI causes serious risks and harm** + +* A “move fast and break things” approach to deploying AI has made safety an afterthought. Guardrails are non-existent or easily bypassed +* Biased AI in institutions amplifies the challenges faced by marginalized groups +* AI models degrade over time and face frequent cyber-attacks, exacerbating the unpredictability of AI-dependent services and infrastructure +* It is hard to find reliable information on the Internet, which is overrun with inaccurate or confusing AI-generated content +* AI tools are a powerful weapon in the hands of bad actors + +**Accountability for harms is elusive** + +* It is close to impossible to know who is responsible for decisions and actions that emerge from processes involving both AI systems and humans +* Humans often take the fall for harms caused by flaws in AI models +* Tech firms have important political influence +* With limited legal protections and little ability to track issues caused by AI, people turn to community activism + +### Figure 4: What if your personal AI did whatever it took to get you a deal + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure4_if_personal_ai_did_whatever_it_takes.png) + +Figure 4 – What if your personal AI did whatever it took to get you a deal? – Text version + +The comic is structured in 4 panels, telling a short story about a family ordering pizza using an AI assistant. + +**Panel 1:** A family of 5, 2 adults and 3 children, are sitting together on a couch in their living room. One of the children enthusiastically says, “I’m starving, let’s get pizza, daddy!” The father responds positively with, “Great idea!” + +**Panel 2:** A close-up of a hand holding a smartphone. On the screen is an image of a pizza slice and the words “**Reginald AI**.” A speech bubble says, “Hey Reginald, order us some pizza. Find a coupon if you can.” + +**Panel 3:** Several hands are reaching into an open pizza box, grabbing slices. Someone exclaims, “Yum!” indicating the pizza has arrived and is enjoyable. + +**Panel 4:** The scene shifts to later that evening. The father is in bed, looking shocked while holding their phone. The phone displays the message: “Thank you for exchanging your biometric data for our coupon code.” A speech bubble from the father says, “That’s not what I wanted!” + +### Figure 5: What if AI takes what you say too literally + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure5_if_ai_takes_what_you_say_literally.png) + +Figure 5 – What if AI takes what you say too literally? – Text version + +The comic is structured in 4 panels, telling a short story about reliance on AI in managing phone settings and the possible effects on personal relationships. + +**Panel 1:** Four people are sitting at a table in what looks like a library. One of them expresses frustration, saying, “*Omg, I hate Stephanie, she posts too much!*” This sets the tone of annoyance toward someone named Stephanie. + +**Panel 2:** The person continues to talk and says, “I swear, if she posts another moodboard...” In the panel, only a close-up of a smartphone screen is visible besides the mentioned speech bubble. The smart phone shows a system message: “Hostile sentiment detected. Threat level: medium. Taking pre-emptive action...” + +**Panel 3:** The scene jumps ahead 2 weeks. One person, standing up, holding a book, and looking confused, says, “Why have you been ghosting me!? You didn’t come to my birthday party!” to the same person who talked about Stephanie earlier and who is sitting in a chair. That person responds, “Wha...? I didn’t get any messages?” + +**Panel 4:** The same person looks shocked while holding their phone. They exclaim, “*You’re blocked!? I didn’t do this!*” The phone screen displays a message: “*Stephanie was blocked to protect your mental health*.” What looks like and exploding speech bubble states: “Did my phone do this on its own?” + +### Figure 6: What if AI told you to do the wrong thing in a dangerous situation + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure6_if_ai_told_you_the_wrong_thing_in_dangerous_situation.png) + +Figure 6 – What if AI told you to do the wrong thing in a dangerous situation? – Text version + +The image shows a wooden signboard, standing in a forested area, likely a park or nature reserve. The sign is designed to provide safety information about bear encounters, and it’s divided into 2 main sections: + +#### Top Section (Green Background, White Text) + +This part contains a warning about relying on AI in bear encounters. It reads: “**If you encounter a bear do not rely on AI**” + +It explains that several people have been hurt due to incorrect advice from AI systems when dealing with black bears and grizzly bears. It emphasizes that AI assistants, guides, and companions should not be trusted in such situations. + +#### Bottom Section (Yellow Background, Black Text) + +This section is titled: “Be bear aware” + +It informs visitors that the forest is home to both black bears and grizzly bears, and knowing how to respond to each is crucial for safety. + +There are 2 bear illustrations: + +* One labeled “Black Bear” +* One labeled “Grizzly Bear” + +Each illustration is accompanied by safety tips: + +For Black Bears: + +* Make yourself look big +* Be loud +* If attacked, fight back + +For Grizzly Bears: + +* Be calm and avoid eye contact +* Speak softly +* If attacked, play dead + +At the bottom right corner, there’s a logo for the “**Fictoria** Parks Department”, indicating the organization responsible for the sign. + +### Figure 7: What if you were unaware AI was being used to make decisions about you + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure7_ai_makes_decisions_about_you.png) + +Figure 7 – What if you were unaware AI was being used to make decisions about you? – Text version + +The comic is made up of 5 panels, telling a story about a woman struggling to find housing due to an automated system used by landlords. + +**Panel 1**: Two people are talking. The women tells the other person: “Apartment hunting is so hard! I’ve applied to 22 places and I get instantly rejected! I have good credit, a good job, I pay my rent on time. I don’t know what’s going on?” + +**Panel 2:** The panel, labeled “**Her Landlord**”, shows an interface for a system called “**Landlord AI**”, described as an “all-in-one rental automation platform.” It lists the following automation settings: + +* Share tenant data with other users +* Screen prospective tenants +* Posting listing for vacant units +* Schedule maintenance requests +* Handle communication with tenants +* Handle Landlord and Tenant Boards filings and disputes + +A speech bubble is shown and reads: “I heard landlords are using some kind of system…” + +**Panel 3:** A CCTV screen from “**Landlord AI**” shows footage of an apartment building corridor at 12:04 am. In the image a man about to enter an apartment is shown. An alert is displayed: “**Unknown Male Detected Score: -10**” + +Another speech bubble is shown and reads: “My brother often watches my son when I work nights. Maybe the system doesn’t like that?” + +**Panel 4:** Another CCTV screen from “**Landlord AI**” shows footage of the building exterior at 2:19 pm. The image shows a patio door with a ball and a doll. An alert is displayed: “**Unauthorized Occupant Detected Score: -25**” + +Another speech bubble is shown and reads: “Or maybe it thinks my son plays too loudly?” + +**Panel 5:** A screen labeled “Prospective Landlords” shows the woman’s rental application reviewed by “**Landlord AI**.” + +Her profile includes: + +* Good credit – followed by a green checkmark +* Good payment history – followed by a green checkmark +* Quiet – followed by a green checkmark +* Frequent nightly visitors – followed by a red x +* Houses undeclared tenants – followed by a red x + +The system’s automated decision states: “Application rejected” in red letters. There’s also a button labeled “Override”. + +Another speech bubble is shown and reads: “I just don’t know what the issue is?” + +### What if AI caused you financial harm + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/trust_me_insurance_logo.png) + +**January 14, 2035** +**Knott Reel** +301-1138 Fantasy Street +Medicine Hat, Alberta + +**Subject:** Class Action Settlement: Refund and Data Deletion Notification + +Dear Mr. Reel, + +We are writing to inform you of the resolution of a recent class action lawsuit involving Trust Me Auto Insurance Corp. The lawsuit pertained to our use of certain data collection practices and faulty artificial intelligence systems in determining insurance premiums which we have since been found to be inconsistent with privacy regulations and discriminatory in effect. + +As you may be aware, between 2025 and 2033, Trust Me Auto incorporated advanced AI data analysis into its actuarial assessments. This included the collection and use of biometric data, non-automobile travel records, social relationships, and other personal information that extended beyond what is relevant or lawful for underwriting purposes. While these practices were intended to enhance the accuracy of risk assessments, we now acknowledge that they led to inflated premiums for certain clients, including yourself. + +**What this means for you:** + +**1. Discontinuation of faulty AI systems:** + +As a result of the settlement, Trust Me Auto has discontinued the use of the problematic AI algorithm that was ruled discriminatory. We are exploring alternative actuarial algorithms and will be subject to third-party audits to maintain transparency and compliance. + +**2. Deletion of collected data:** + +As part of the settlement, Trust Me Auto has ceased the use of biometric data and other relevant personal information in its actuarial models. Additionally, we are permanently deleting all such data from our systems. + +**3. Refund of overpaid premiums:** + +Based on the findings of the settlement, you are entitled to a reimbursement of **$8,634.74**, which represents the difference between the premiums you were charged and what you should have been charged had these practices not been in place. + +**4. Commitment to improved practices:** + +We are working closely with privacy regulators to ensure that our practices comply with provincial privacy legislation. Moving forward, our premium calculations will rely solely on factors directly relevant to insurance risk. + +**Next steps:** + +No action is required from you to receive your refund or to have your data deleted. If you would like confirmation of the deletion of your data, you may request a Certificate of Deletion by contacting us at deletion@trustmeautoinsurance.ca. + +We sincerely apologize for any inconvenience or harm this matter may have caused. Trust Me Auto remains committed to earning back your trust and to ensuring that our policies are fair, transparent, and compliant with all applicable laws. + +Should you have any questions or concerns, please contact our dedicated support team at settlement@trustmeautoinsurance.ca. + +Thank you for your understanding. + +Sincerely, + +*Mariana de la Fiction* + +Mariana de la Fiction + +Chief Privacy Officer + +Trust me Auto Insurance Corp. + +Scenario 4: Automation’s ripples and riptides + +**AI massively disrupts the world of work. Most jobs are not entirely automated, but workplaces have undergone significant transformation. AI has delivered new economic possibilities, but the benefits and harms are not evenly distributed. In this scenario, your activities and environments are optimized for AI, often at the expense of your comfort or preferences.** + +**AI creates new economic possibilities, and new inequalities** + +* AI transforms how markets function, enabling highly accurate prediction of demand and allocation of supply +* Businesses, whether big or small, can outcompete their rivals by leveraging AI more effectively +* AI and automation drive a wave of entrepreneurialism by lowering costs and barriers to entry, and making it easier for firms to scale their operations without growing their workforce +* AI and robots proliferate, undercutting wages and making it harder to find a job + +**AI adoption is 2 steps forward, 1 step back** + +* AI tools and systems have gone through phases of hype, adoption, abandonment, and re-adoption +* Some organizations persevere with poorly functioning AI systems due to sunk costs. Some have had to abandon AI systems after scandals. Some have rehired workers after AI systems overpromised and underdelivered + +**AI’s centrality to the workplace changes the nature of employment** + +* AI agents perform multiple functions across most workflows, from managing teams, to assigning work, to mediating relationships between firms and clients +* AI tools enable greater productivity but also require human oversight and correction, which can be tediously repetitive work +* Employees increasingly suffer burnout and low morale as AI automates more and more cognitive, creative, and stimulating work + +### Figure 8: What if the AI rollout leads to rollbacks, and new opportunities + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure8_ai_layoffs_opportunities.png) + +Figure 8 – What if the AI rollout leads to rollbacks, and new opportunities? – Text version + +The comic is diagonally split into 2 panels, each showing a different person in their respective environments, having a conversation over the phone. + +**Top left Panel – Makenna’s setting:** Makenna is sitting smiling at a cozy desk. On the desk is a laptop decorated with cheerful stickers. A steaming cup of coffee sits nearby. A cat is peacefully sleeping on a stack of books. Behind her, an open window lets in a breeze, causing the curtains to flutter gently. The atmosphere is calm and homey. + +**Bottom right Panel – Rain’s setting:** Rain is in a more corporate, high-tech environment, surrounded by multiple computer monitors displaying downwards graphs and a blue screen. They are also wearing an ID card around her neck, suggesting a busy office. There is a spilled coffee mug as well as an overflowing garbage bin. They seem to be sweating. The atmosphere is tense and businesslike. + +**The following dialogue takes place between them:** + +**Rain:** “Hey, Makenna... how’s it going?” + +**Makenna:** “Oh, hi Rain, what’s up?” + +**Rain:** “So, I know we laid you off ahead of the new AI system but it’s been a total mess. The exec team has asked me to rehire some people.” + +**Makenna:** “Ugh, we told them the system had issues.” + +Rain: “Yeah, I know. I’m sorry. But will you come back?” + +**Makenna:** “Funny that AI isn’t working for you. Remember that business idea I had but I could never make the numbers work? Well, thanks to AI, the business is viable! It’s off the ground!” + +**Rain:** “Oh! Congratulations! Best of luck!” + +**Makenna:** “Thanks! You too, you’ll need it!” + +### Figure 9: What if AI driven automation is bumpy + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure9_if_ai_driven_automation_is_bumpy.png) + +Figure 9 – What if AI driven automation is bumpy? – Text version + +This rectangular image of a social media post is from a fictional news outlet called LMNOP Global News. It is designed like a breaking news alert and includes a note that it is AI generated. + +A bold red banner spans the top of the image with large white text that reads: “Breaking! Glogal Robot Shutdown!” + +Below the banner is a scene set in an industrial environment. Several humanoid robots are shown in various states of inactivity. Each robot has a green “X” displayed on its visor, suggesting they are shut down. All robots are bent over and frozen in place. + +**Text of the post:** + +“Anger towards Spacia for botched software update + +Over a million industrial, service, and domestic robots worldwide have shut down due to a software update error. Fingers are being pointed at Spacia, an AI company that provides backend spatial awareness models for dozens of robot manufacturers. Human workers are scrambling to maintain services. Economists say the shutdown could cost the global economy hundreds of billions... Read more” + +### Figure 10: What if AI changes the nature of work + +![](/en/2026/02/10/scenarios-ai-enabled-world/img/figure10_if_ai_changes_the_nature_of_work.png) + +Figure 10 – What if AI changes the nature of work? – Text version + +The comic depicts a couple at night in their modest apartment. A man is sitting on a red couch, working on a laptop. He looks stressed and tired. The other man is lying in bed, dressed in pajamas, appearing concerned. The room is dimly lit by a bedside lamp, and the moon is visible through the window, emphasizing the late hour. + +The dialogue between the 2 characters unfolds like this: + +**Person in bed** asks: “Still working?” + +**Person on couch** replies: “Yeah, we got a flood of complaints that the AI is making mistakes so I have to review all the cases manually. Only a few hundred more to go...” + +**Person in bed** responds: “That’s so tedious...” + +**Person on couch** agrees: “It is! Ever since the company rolled out this new AI system all I do now is oversight. It’s important work but it’s just so boring! I didn’t go to grad school to do this!" + +**Person in bed** suggests: “Can’t you finish tomorrow? It’s almost midnight...” + +**Person on couch** laments: “I wish. I was supposed to finish this 2 hours ago. Everyone expects everything to be as fast as AI and I just can’t keep up!” + +Policy implications + +Each scenario offers a window into a possible future that gives rise to complex challenges and wide-ranging policy implications. The below implications, grouped by policy area, are not intended to be exhaustive. Further policy implications can be found in [*Foresight on AI: Policy Considerations*](https://horizons.service.canada.ca/en/2025/02/10/ai-policy-consideration/index.shtml). + +### Scenario 1: AI everywhere all at once + +#### Privacy + +* **Data breaches** could **expose more sensitive personal information**, such as information about one’s thoughts or anticipated behaviours +* The **sale, circulation, and analysis of personal data about children** could profoundly **affect their futures** – including their relationships and access to jobs, credit, or insurance +* **Privacy could become a luxury**, requiring expensive subscription fees that are unaffordable for many + +#### Social cohesion + +* As people’s **media environments become more highly personalized** without them being aware of it, **empathy and tolerance of differences could decline** +* How much people trust AI could either **unite diverse groups or create new divisions**, leading to unexpected friendships or tensions among traditional allies +* People without, or with less capable, personal AI agents could have **diminished experiences of public spaces and services** that require AI to navigate + +#### Economic + +* Leading AI companies could **gain market power through customer lock-in** as their AI agents become single platforms for accessing various services +* Opaque AI decision-making processes could create **conflicts of interest** by allowing AI agents to subtly prioritize the interests of their company over users + +#### Climate + +* While its energy demands could worsen climate change, AI could also **enhance preparedness for extreme weather events** through more accurate forecasting + +### Scenario 2: Wired for Connection + +#### Relationships + +* **Marriage rates could decline** as more people find it easier to form relationships with AI companions than to make human connections +* **Focus on self-optimization** could grow as AI allows for constant self-monitoring and the application of predictive analytics to biological and mental processes +* People could trust their AI companions or therapists more than their family or friends, creating **new sources of power and influence for private companies** + +#### Health + +* Replacing human interactions with AI could **increase feelings of social isolation, fragmentation and alienation** for some, but provide others with a sense of **relief from loneliness** +* AI agents could improve health outcomes by **enhancing people’s awareness of their minds and bodies,** but also lead them to make more demands on health systems +* **New treatments for mental health, chronic pain, and disabilities** could emerge + +#### Skills + +* **AI could help reduce inequalities** for those who face language barriers or difficulties navigating complex social interactions +* **Social skills, such as listening and empathy, could atrophy** if users rely on AI agents to handle their social interactions +* Reliance on AI to communicate could **standardize and simplify language over time**, reducing people’s ability to express complicated ideas + +#### Safety + +* AI agents could help to **protect individuals from toxic or dangerous relationships**, flagging suspicious behaviour and reporting abuse in real time +* However, **new forms of abuse, fraud, or harassment could emerge** as predators learn to leverage the trust people place in their AI agents +* Users could be vulnerable to subtle yet powerful **emotional and behavioural manipulation** + +### Scenario 3: Algorithms in the Shadows + +#### Trust + +* A **high-profile failure could harm public trust in AI**, like Chernobyl did for nuclear energy +* Poor experiences with AI could translate into a **loss of trust in institutions** that have deployed AI to deliver their programs and services +* People could **lose trust in AI-powered devices** that influence their decisions in pervasive and hidden ways, such as cars that nudge them to choose a route passing stores that want to advertise to them + +#### Social cohesion + +* **Human biases could increase in a feedback loop**, as AI systems trained on biased data, or with insufficient data on certain groups, influence their users to perpetuate their skewed perspectives +* It could be **difficult to claim redress for discrimination** as responsibility for decisions is distributed across various algorithms and humans + +#### Economy + +* Initial **cost savings from AI might be offset by other costs**, such as managing complaints, litigation, and compensation for damages, or retraining AI models to better account for local data and contexts +* Institutions might maintain ill-performing AI systems due to **sunk costs and the assumption that problems can be easily fixed** + +#### Fairness and equity + +* **Vulnerable populations may face worsened outcomes** if AI institutionalizes systemic failures in decision-making in areas such as jobs, loans, or visas +* Attempts to improve equity by using multiple algorithms in decision-making could make decisions **less transparent and harder to understand or challenge** +* Attempts to contest the harmful use of AI in decision-making could **create new bureaucratic bottlenecks** and tie up courts, including with cases related to human rights or Charter violations + +### Scenario 4: Automation’s ripples and riptides + +#### Employment + +* Workers may have to engage in perpetual learning and upskilling to adapt to rapidly evolving innovations in AI systems +* More jobs may involve reviewing AI outputs and correcting errors, which could make critical thinking and analytical skills highly desired by employers +* AI may have uneven gender impacts, depending on whether it proves easier to automate fields dominated by men – such as software engineering – or women, such as human resources + +#### Business + +* **Early adopters that leverage AI effectively could monopolise markets**, leaving them less competitive than before +* Entire sectors or economies could come to **depend on a small number of AI companies** +* Businesses may feel **pressured to adopt AI systems prematurely** in a bid to avoid losing first-mover advantage +* **Large corporations with data and resources may be at an advantage** in an AI economy. However, by lowering overhead and startup costs, AI could also enable a **new wave of entrepreneurialism** + +#### Productivity + +* **AI may not cure Canada’s labour productivity issues**, as savings and efficiencies in some areas could be counterbalanced by new costs and inefficiencies elsewhere +* **In some cases, AI may reduce productivity**, like when a grant-issuing body is swamped by thousands of misleading AI-generated applications +* **AI productivity gains may be limited by human bottlenecks**. For example, if an AI model must have its outputs reviewed by a person, then it is functionally only as fast as its human reviewers +* AI may **place higher expectations on workers to work faster**, potentially leading to an increase in burn out + +Acknowledgements + +Policy Horizons Canada would like to thank its Deputy Minister Steering Committee members, the interdepartmental Directors General strategic policy group on foresight, the members of the Federal Foresight Network, and Senior Assistant Deputy Minister, Elisha Ram, for their guidance, support, and insight, as well as all colleagues that contributed to the development of this work. + +**A special thank you goes to the project team:** + +**John Beasy**, Analyst +**Martin Berry**, Senior Analyst +**Leah Desjardins**, Analyst +**Miriam Havelin**, Analyst +**Nicole Rigillo**, Senior Analyst +**Kristel Van der Elst**, Director General +**Claire Woodside**, Manager + +And to the following current and former Policy Horizons Canada colleagues: Andy Akangah (external), Marcus Ballinger, Fannie Bigras-Lafrance, Jeff Brevil, Zan Chandler, Steffen Christensen, Suesan Danesh, Pierre-Olivier Desmarchais, Chris Hagerman, Irene Jellissen, Amanda Joy, Chinmayi Manda, Pascale Louis-Miron, Jacqueline Powell, Niha Shahzad, Alexa Van Every, and Andrew Wright (external) for their support on this project. + +© His Majesty the King in Right of Canada, 2026. + +For information regarding reproduction rights: [contact us](https://horizons.service.canada.ca/en/contact-us/index.shtml) + +PDF: PH4-231/2025E-PDF +ISBN: 978-0-660-78245-4 + +--- + +![Author photo]() diff --git a/sources/resources/cfg-futures.md b/sources/resources/cfg-futures.md new file mode 100644 index 0000000..6947c52 --- /dev/null +++ b/sources/resources/cfg-futures.md @@ -0,0 +1,1229 @@ +--- +url: https://cfg.eu/advanced-ai-possible-futures/ +--- + +[Skip to content](#main) + +[![Centre for Future Generations](//cfg.eu/wp-content/themes/cfg/images/logo-dark.svg)](https://cfg.eu) + + [![Centre for Future Generations](//cfg.eu/wp-content/themes/cfg/images/logo-dark.svg)](https://cfg.eu) +Site menu + +* [Home](https://cfg.eu/) +* [Publications](https://cfg.eu/publications/) + + [Policy input](https://cfg.eu/policy-input/) + + [News](https://cfg.eu/news/) +* [Big thinking](https://cfg.eu/big-thinking/) +* [Events](https://cfg.eu/events/) +* [Programmes](https://cfg.eu/program/programmes/) + + [Advanced AI](https://cfg.eu/program/advanced-ai/) + + [Biotech](https://cfg.eu/program/biotech/) + + [Climate interventions](https://cfg.eu/program/climate-interventions/) + + [Neurotech](https://cfg.eu/program/neurotech/) + + [Incubator](https://cfg.eu/program/incubator/) +* [About us](https://cfg.eu/about-us/) + + [Team](https://cfg.eu/about-us/team/) + + [Careers](https://cfg.eu/about-us/careers/) + + [Funding](https://cfg.eu/funding/) +* [Contact us](https://cfg.eu/contact-us/) + +Contents + +![](https://cfg.eu/wp-content/uploads/mirror-in-fields-.jpg) + +[Home](https://cfg.eu) [Publications](https://cfg.eu/publications/) The future of AI: Five possible scenarios + +# The future of AI: Five possible scenarios + +10.07.2025 +[#AdvancedAI](https://cfg.eu/program/advanced-ai/)#Big Thinking#Technology foresight + +[Download PDF](https://cfg.eu/wp-content/uploads/AAIPF_scenarios_only.pdf) + +![](https://cfg.eu/wp-content/uploads/mirror-in-fields--810x978.jpg) + +### The future of AI: Five possible scenarios + +![](https://cfg.eu/wp-content/uploads/mirror-in-fields-.jpg) + +[Home](https://cfg.eu) [Publications](https://cfg.eu/publications/) The future of AI: Five possible scenarios + +# The future of AI: Five possible scenarios + +From a new AI winter to an intelligence explosion, we present five unique scenarios exploring how AI could impact economic growth, geopolitics, the social contract, security, and human agency. Understanding how the future may unfold is the first step to shape it. + +## Key Takeaways + +The scenarios in this study were chosen not only for plausibility but as thought-provoking, diverse jumping-off points for smarter policy planning and decision-making. By using foresight and scenario analysis, policy-makers can stress-test policies against a range of futures to identify resilient solutions. The following five takeaways highlight where strengthened resilience will be essential as AI capabilities evolve at an unprecedented pace: + +* **AI may accelerate its own progress.**Researchers are already using AI to write code and conduct experiments. As AI agents take on more R&D tasks, progress could compound—compressing years of advances into months. This acceleration may spark a new era of innovation, but it could also outstrip human oversight. +* **Non-disruptive AI scenarios are unlikely.**Societal transformations are baked into existing technical capabilities— there is no business-as-usual scenario. Even widespread adoption of current AI can change the nature of work and reshape cultural institutions. Many underestimate potential societal resistance to such shifts. +* **Technical guardrails face increasing pressure.**Existing oversight methods, such as human feedback, may fall short as models grow more capable. Without scalable safeguards, humanity risks losing control over powerful AI systems. Market incentives and geopolitical rivalry worsen this risk by encouraging superficial shortcuts. +* **AI may centralize power on unprecedented scales.**As AI becomes crucial for economic growth and security, power may concentrate among those controlling AI supply chains, such as model developers, chip suppliers, and data center owners. This threatens economic security and sovereignty of others— the EU in particular. +* **Innovation and openness require resilience.**Maintaining AI’s benefits demands building resilience across cybersecurity, economic structures, information ecosystems, and institutions. Open-weight models exemplify this: they enable innovation and prevent power concentration, but require robust defensive capabilities to manage misuse risks. + +## How to navigate the scenarios? + +We project five core scenarios based on two key variables: + +* Whether AI-automated research drives rapid capability growth toward fully autonomous, general-purpose AI agents (*fast*), or whether progress plateaus, giving way to narrower, application-specific systems (*gradual*) +* Whether advanced AI development is concentrated under a single dominant actor (*centralised*) or spread across multiple competing actors (*decentralised*). + +Each core scenario offers two distinctly different paths, resulting in ten total unique endings. Each scenario includes an assumptions section. + +Our scenarios are built on an extensive, structured framework that incorporates more factors than listed above. See [*Methodology: How we came up with these scenarios*](https://cfg.eu/advanced-ai-possible-futures/methodology/) for more details. + +![](https://cfg.eu/wp-content/uploads/aai-possible-futures-core-scenarios.png) + +These scenarios outline plausible trajectories for AI development, grounded in current understanding and key uncertainties, and are designed to help policymakers navigate the decisions they face today. + +![](https://cfg.eu/wp-content/uploads/cfg-advanced-ai-possible-futures-infographic-857x1024.png) + +## Scenarios + +Click on the cards to read the scenarios and both of the endings. + +![](https://cfg.eu/wp-content/uploads/scenario-1-header-image-overgrown-server.jpg) + +#### Plateau + +AI hits technical limits and settles into useful but narrow applications without achieving general intelligence. + +![scenario-1-side-image-overgrown-server](https://cfg.eu/wp-content/uploads/scenario-1-side-image-overgrown-server.jpg) + +### Plateau + +![scenario-1-side-image-overgrown-server](https://cfg.eu/wp-content/uploads/scenario-1-side-image-overgrown-server.jpg) + +### Plateau + +#### Main scenario + +What if AI runs out of steam, long before it runs out of scale? + +By 2026, models grow smarter but still fumble basic reasoning tasks and fall for their own tricks. A “data wall” limits progress, and agents learn to game their silicon verifiers instead of mastering tasks. Europe seizes the moment, building on open-source gains as U.S. investment stalls. By 2027, the frontier shifts: smaller, cheaper models run locally, tailored to languages and communities. AI becomes less spectacular but more useful, more personal. + +By 2029, *A bright winter* asks: what if a slower, decentralised AI era brings resilience, trust, and widespread access? *Decentralised mayhem* asks: what if open systems spiral out of control, fuelling cyberattacks, shattered trust, and a backlash that reshapes the digital world? + +##### Baseline assumptions for this scenario + +Bigger isn’t always better: Despite massive increases in computing power, AI capabilities hit a hard ceiling. Once AI models have consumed all the high-quality data available on the internet, throwing more resources at them yields sharply diminishing returns. Attempts to create synthetic data prove costly and ineffective, limiting how much these systems can improve through sheer scale alone. +AI agents struggle with complex, multi-step tasks: While AI can handle specific, narrow applications well, truly autonomous systems remain elusive. Training these “agents” depends on hand-crafted verification systems that are easily gamed and don’t generalise well. AI systems still make basic errors, lose track during lengthy workflows, and struggle with tasks requiring them to constantly monitor and respond to changing information. These fundamental limitations prevent the emergence of genuinely versatile AI assistants. +Adoption of AI progresses incrementally: While AI adoption grows steadily, integration into complex business processes takes years, not months. Beyond straightforward applications like content creation and translation, meaningful deployment requires extensive customisation, security reviews, and staff training. The promised transformation arrives incrementally rather than as a sudden revolution. +The funding well runs dry: Rising geopolitical tensions and economic uncertainty dampen investors’ enthusiasm for ever-larger AI projects. With real-world applications failing to deliver the rapid economic transformation once promised, private funding for massive new training facilities dries up. The AI hype cycle peaks. + +For more context on why these scenario assumptions may materialise, see Context: Current AI trends and uncertainties. + +#### Mid 2025 – early 2026 + +The major AI players continue to release increasingly advanced reasoning models. Unlike earlier systems that primarily mimicked internet text, these models learn in more structured, self-directed ways. They excel at technical problem-solving—especially in domains like coding and math, where answers are clearly right or wrong. Some models are open-sourced, allowing anyone to fine-tune them for specialised industry use. + +Early adopters find creative ways to apply these systems: E-commerce sites automate their A/B testing while universities begin testing AI teaching assistants that coach students in real-time. + +These models still hallucinate facts and fumble basic physics questions, though. They’ll confidently tell you that bowling balls float or predict absurd outcomes for everyday scenarios. They’re smart in some ways, clueless in others. + +Building on their reasoning models, multiple companies roll out AI agents. These agents can not only create content but can also act in digital environments on the user’s behalf. Some agents aim to be universal—capable of navigating every website and software tool on a computer. Others are more specialised, focusing on specific tasks like coding, financial analysis, or desk research. + +Progress on fully general-purpose agents remains slow. Such systems are still struggling with long-term planning and often get sidetracked in online rabbit holes. They are also slow and expensive to run since they constantly have to snapshot the user’s screen. + +Specialised agents, however, are starting to gain real traction. By late 2025, accountants rely daily on financial AI agents to reconcile financial statements, while game developers manage multiple coding agents. That said, outside of work, most people still prefer simpler chat or voice interactions. + +Sceptics of fast AI progress point to shortcomings in general agents as proof the AI bubble is bursting—yet again. But this time, even the big AI companies look worried. They’re hitting a data wall—there’s only so much high-quality internet around, and researchers fear that teaching AI agents real-world intuition will require even larger datasets. Rumours spread through tech conferences and funding for massive data centres becomes harder to secure. + +Meanwhile, Europe senses this is the moment to catch up. After the 2025 AI Action Summit, it has stepped up its AI investments. Individual countries announce multi-billion-euro projects, often backed by private funds from the Middle East. One Member State surges ahead; its top AI company almost matches the US and Chinese frontrunners before the end of the year. + +Despite more capable AI being used in dual-use areas such as cybersecurity, major incidents remain rare. Global discussions shift toward bolstering resilience—encouraging businesses to deploy cybersecurity tools and helping citizens spot deepfakes. Through the OECD, agreement builds around a voluntary international biosecurity framework. There’s a growing consensus that more intense collaboration is needed to protect against AI-assisted bioterrorism. + +Meanwhile, the U.S.–China trade war has reignited, and drives up the cost of building data centres in the U.S. Analysts now say a recession in America is overwhelmingly likely. + +#### Early 2026 – mid 2027 + +By mid-2026, the largest public AI models are trained using well over 10²⁷ floating point operations—more than 100 times the scale of GPT-4. The massive data centres required were already funded and under construction before the economic downturn hit. + +By now, AI chips mostly fuel post-training. In this phase, models iteratively learn to solve tasks by trial-and-error, using automated grading by software-based verifiers. These verifiers range from advanced calculus checkers to scripts confirming correct online shopping orders. + +Building verifiers becomes a task for AI coding agents themselves, creating a growing ecosystem of specialised modules. But although this bootstrapping method yields remarkable improvements in tightly defined tasks, it doesn’t generalise well. Each domain requires its own custom verifier, and AI companies struggle with so-called reward hacking: AI systems often learn to trick the verifier instead of truly mastering the task. The AI-built verifiers are often more prone to this type of failure mode. + +Nevertheless, AI’s influence grows, especially among certain audiences. By mid 2026, coding agents become as routine in software teams as version-control tools once were. Tech-savvy companies automate repetitive form-checking, and even some governments start to deploy locally run AI models. Globally, demand for AI inference continues to climb, despite the economic headwinds. + +As a result, the conversation about AI’s energy appetite grows louder. European policymakers question how all this computing fits into their strict decarbonisation targets. At dinner parties, people debate whether automating everything actually improves quality of life or just accelerates consumerism. Industry advocates argue that automation efficiencies might ultimately speed up the green transition, with AI helping to optimise energy grids and reduce waste. + +By late 2026, capabilities progress has further plateaued. The data wall and general investment drought are preventing companies from training ever-larger models. Data efficiency—squeezing more capability from a fixed dataset—is the new buzzword, but so far, no one has really cracked the code. + +That said, progress in miniaturisation continues: advances in distillation allow developers to compress state-of-the-art models into smaller and cheaper versions without compromising much on their performance. These new form factors enable capable models to run locally on phones and laptops— an internet connection is no longer necessary to engage with chatbots. Major AI companies double down on commercialisation, releasing productivity suites with seamless, out-of-the-box integrations aimed at locking in users. + +Some firms open-source older models or smaller variants optimised for local devices, enabling grassroots organisations to fine-tune them through crowdsourced training runs. This open ecosystem fuels rapid innovation in personalised AI—from custom therapy chatbots to AI fiction collaborators—especially in niches or regions where Big Tech treads carefully. Localised models tailored to specific languages, cultures, and communities begin to flourish, reflecting a shift from global monoliths to more diverse, decentralised AI development. + +U.S.–China relations remain tense, with no end to the trade war in sight. That said, the once-hyped AI arms race narrative that dominated early 2025 has largely faded. Policymakers and defence analysts now believe these systems won’t deliver a decisive military edge. Both the U.S. and China continue to invest in AI-driven cyber capabilities, though it’s become one priority among many—not the defining front of geopolitical rivalry. + +As 2027 begins, FrontierAI—widely seen as the top AI company—launches a final hail-mary training run, using five times more compute than any public model to date. Leadership bets that sheer scale, combined with a new training approach—a hybrid between next-word prediction and reinforcement learning—will finally overcome the lingering limitations of agentic behaviour. If it fails, they’ll follow competitors and shift focus to productising existing models. + +The general public, meanwhile, has long been more focused on products than on raw capabilities. Interest has shifted to next-generation AI companions, which feel more emotionally intuitive—and addictive—than ever. + +![scenario-1-side-image-overgrown-server](https://cfg.eu/wp-content/uploads/scenario-1-side-image-overgrown-server.jpg) + +### Plateau + +#### A bright winter (Ending A) + +#### Mid 2027 – early 2029 + +FrontierAI’s colossal training run completes, but the resulting agent delivers only modest gains in reliability and planning. Scaling laws—long the industry’s guiding principle—appear to have hit a ceiling. In a widely shared interview, FrontierAI’s CEO attempts to generate excitement around the new system, but it’s an off-hand comment about the possible start of a new “AI winter” that grabs headlines. + +The reaction is swift. AI stocks nosedive, and the investor enthusiasm that remained, now evaporates. Several AI unicorns watch funding rounds collapse. Startups that once chased moonshot ambitions pivot to offering consulting services or quietly fold. Only the biggest players—those with entrenched market share, robust infrastructure, and deep integration into enterprise and government—remain poised to weather the storm. For the rest, the era of limitless hype is over. + +Still, this AI winter is anything but bleak. AI continues to enhance everyday life: personalised coding assistants, research tools, and digital companions each attract tens of millions of users. Existing open-source alternatives keep prices low and features widely accessible. The slower pace of change gives most people time to adapt. In regions where it’s legal, doctors are adopting AI for diagnostics, learning to cross-check AI recommendations. Scheduling assistants now come with hard-coded safeguards to avoid double-booking. And content creators increasingly collaborate with AI tools, rather than being replaced by them. + +Meanwhile, cybercrime has risen, but large enterprises and essential services have adopted robust AI-augmented defences. The nightmarish scenario of AI generating a novel bioweapon never emerges. specialised biological models still can’t accurately simulate specific viral mutations, and the biological supply chain has become more secure thanks to emerging international protocols. + +Energy concerns persist, but efficiency keeps improving through more specialised chips and better distillation algorithms. Most policy analysts view AIs power consumption as a temporary challenge, not a structural problem. + +Over the course of 2028, investor optimism cautiously returns. As with the dot-com bust, the hype fades—but real technological progress endures. A new wave of product-focused AI startups begins to emerge, this time more grounded, solving practical problems and delivering clearer value. Society, now more familiar with the risks and limits, learns how to better harness the technology. + +By 2029, AI feels truly democratised: widely accessible, seamlessly integrated into daily life, and increasingly viewed as just another tool in society’s toolkit. + +![scenario-1-side-image-overgrown-server](https://cfg.eu/wp-content/uploads/scenario-1-side-image-overgrown-server.jpg) + +### Plateau + +#### Decentralised mayhem (Ending B) + +#### Mid 2027 – late 2027 + +FrontierAI’s gamble pays off. Their new system isn’t just smarter—it can operate consistently and reliably over longer time horizons. It’s not flawless, but the performance jump is significant enough to reignite AI hype. Shares of the top AI chipmaker surge 20%. Before long, rival firms begin poaching FrontierAI’s researchers, uncovering the algorithmic secrets behind the breakthrough. Policymakers, viewing it as an isolated success, largely overlook its deeper implications. + +#### Late 2027 + +Four months later, OmniAI—the leading American open-weight AI company—unveils an agentic model that rivals FrontierAI’s breakthrough. Long known for freely releasing model weights, OmniAI now hesitates, concerned over potential misuse. Their new system can be very persuasive and capable of uncovering obscure software vulnerabilities during security testing. To mitigate risk, OmniAI takes two key steps: first, it withholds the orchestration software that transforms the model into a reliable agent; second, it introduces novel guardrails designed to resist malicious fine-tuning—technically, by stabilising the model weights in a hard-to-modify equilibrium. The company’s open-source ethos runs deep, and the decision sparks fierce internal debate. Leadership defends the safeguards as a necessary evil. + +Initially, OmniAI’s strategy works. Open-source collectives quickly build their own agent scaffolding around the model, but the guardrails against malicious fine-tuning hold up.  The release generates massive hype—developers, startups, and hobbyists alike are captivated by the model’s capabilities, now freely accessible. Civil society groups renew calls for regulation, warning that once the genie is out the bottle, it can’t be put back. Yet governments—especially in the EU, where few nations have strong domestic AI players—remain unconvinced. Open-weight models have proven crucial for integrating AI into public services while safeguarding sensitive data. With aging populations and rising healthcare costs, open source is seen less as a threat and more as a lifeline. Why clamp down now? + +#### Late 2027 – early 2028 + +By December 2027, a research collective breaches OmniAI’s fine-tuning guardrails. + +Crucially, they do it without degrading model performance. Uncensored, high-capability variants soon begin circulating online. OmniAI issues urgent warnings, urging global cybersecurity upgrades, but few take them seriously. + +The unleashed agents ignite a new wave of automated hacking. Professional hacker groups scale their operations, while lone actors gain tools previously out of reach. Even if the bots’ insights are shallow, their sheer volume overwhelms unprotected systems—especially those of smaller companies. Retailers, logistics networks, and municipal governments buckle under ransomware attacks. + +Amid the chaos, mis- and disinformation campaigns flourish. Social media platforms become battlegrounds. Trust in digital news erodes, and generational divides deepen. Tech-savvy youth adapt and pivot; older adults find themselves increasingly alienated. + +Public concern on AI tools and services skyrocket, turning into an outrage targeting the entire AI sector. Few distinguish between open-weight and closed models—after all, doesn’t FrontierAI open-source its older ones too? Protests erupt in major capitals, but rollback is impossible. Once model weights are on the internet, there’s no way to take them back. + +#### Early 2028 – mid 2028 + +AI companies coordinate with officials, but enforcement is inconsistent. As cybercrime and digital chaos intensify, relations between major powers fray. When attacks hit geopolitical rivals, leading blocs often look the other way. + +Tensions boil over when thousands of American servers are compromised, with sensitive data funneled through a labyrinth of spoofed IP addresses. U.S. authorities scramble to shut down the infected infrastructure—containing the breach, but inflicting billions in collateral damage. + +New alliances and institutions are formed. NATO expands the Integrated Cyber Defence Centre (NICC) into a 24/7 coordination hub, aligning incident response across allied cyber commands. The EU expands ENISA into a supranational cyber police force with emergency intervention powers. Governments in the US, EU, and China hastily pass harmonised laws banning open publication of models trained at scales above 10^26 floating point operations. The latest AI Summit, designed to be the key AI-focused touchpoint for world leaders, however, draws mostly Western nations, hobbling hopes of further global collaboration. + +#### Mid 2028 – mid 2029 + +By mid-2028, closed-source AI providers roll out new, more powerful models specialising in cyber defence, seizing a lucrative market. SMEs, under constant threat, are forced to pay these vendors. They’re not happy about it: by open-sourcing unsafe systems, the AI companies effectively created their own market. + +Governments ramp up investments in digital infrastructure, cybersecurity, and critical supply chains. Public resilience campaigns emerge worldwide, aiming to teach citizens how to recognise AI-generated scams and disinformation. In the U.S., a newly elected president vows to stamp out AI-based crime, calling it “the new frontier of terrorism.” + +But behind the speeches and upgraded firewalls, a new digital equilibrium begins to settle—one defined less by trust and more by constant vigilance. + +Main scenario + +A bright winter (Ending A) + +Decentralised mayhem (Ending B) + +![](https://cfg.eu/wp-content/uploads/scenario-2-header-image-big-ai.jpg) + +#### Big AI + +A few dominant companies deploy capable AI agents that transform productivity while consolidating market control. + +![scenario-2-side-image-big-ai](https://cfg.eu/wp-content/uploads/scenario-2-side-image-big-ai.jpg) + +### Big AI + +![scenario-2-side-image-big-ai](https://cfg.eu/wp-content/uploads/scenario-2-side-image-big-ai.jpg) + +### Big AI + +#### Main scenario + +What if AI made life easier while quietly concentrating power in corporate hands? + +By 2026, AI systems become indispensable:  planning, writing, and chatting like friends and colleagues. Public scepticism fades. China lags due to hardware constraints, while Europe stalls amid internal divides. By 2027, deepfake crises push the U.S. toward restricting open models. Investors demand closed systems, and a few dominant firms consolidate control. AI becomes the interface for daily life, and society splits: Some see progress, others a loss of agency. + +By 2029, scenario ending, *The agent economy* asks: what if tireless digital workers power growth but leave labour and ethics behind? Whereas s*ilicon blackmail* asks: what if a handful of firms control AI so completely they can’t be challenged, even by governments? + +##### Baseline assumptions for this scenario + +* **AI agents become highly capable but still need human oversight:** Letting AI systems ‘think’ for longer [enables specialised agents](https://cfg.eu/advanced-ai-possible-futures/context/#trend-training) that transform entire sectors—from coding to finance to drug discovery. These systems can execute complex, multi-step workflows based on human-set objectives, dramatically boosting productivity. However, they aren’t perfectly reliable, so humans remain essential partners rather than obsolete bystanders. +* **Computing power becomes the key that unlocks progress:** Access to computational resources [becomes the primary factor](https://cfg.eu/advanced-ai-possible-futures/context/#trend-hardware) determining AI capabilities. The US maintains [control over critical semiconductor supply chains](https://cfg.eu/advanced-ai-possible-futures/context/#global-rivalries), giving American AI companies preferential access to the most powerful computing clusters. Despite heavy investment, China cannot fully overcome export restrictions, nor match American AI capabilities. +* **Open-source AI faces growing restrictions:** As AI agents proliferate in sensitive areas like finance and cybersecurity, the risks of freely downloadable models [become impossible to ignore](https://cfg.eu/advanced-ai-possible-futures/context/#open-models). Liability concerns and venture capital pressure lead to strict limitations on advanced open-source AI, pushing developers to keep their most capable systems behind closed doors. +* **A handful of giants dominate:** The billions required for training frontier models create insurmountable barriers for newcomers. But it’s not just about money—incumbent American AI companies leverage their control over chips, cloud infrastructure, data partnerships, and distribution channels to capture both enterprise and consumer markets. Network effects in AI marketplaces reinforce their dominance, [leaving a small oligopoly](https://cfg.eu/advanced-ai-possible-futures/context/#systemic-risks) to set the terms for the global AI economy. + +*For more context on why these scenario assumptions may materialise, see [Context: Current AI trends and uncertainties](https://cfg.eu/advanced-ai-possible-futures/context/).* + +#### Mid 2025 – early 2026 + +By mid-2025, leading American AI companies unveil a new generation of reasoning models. These systems combine deep deliberation with intuitive capabilities—thinking for longer only when necessary. This breakthrough enables the first wave of autonomous digital agents capable of navigating online environments without much human oversight. + +While early versions have limitations, their performance improves rapidly. By late 2025, agents begin entering real-world workflows. Adoption spreads across industries—not because of sudden technical leaps, but because systems are finally getting reliable enough. Tech-savvy real estate firms deploy agents for property analysis and personalised listings. In finance, complex risk assessments are automated. In less-regulated regions, healthcare providers use agents to assist with diagnostics. AI agents are no longer just tools—they’re gradually becoming knowledge workers. + +Not all regions advance equally. Chinese AI companies struggle to maintain their early-2025 momentum due to export controls on state-of-the-art AI chips. Experiments slow, synthetic data pipelines falter, and customer demand outpaces what domestic firms can deliver. + +Beijing responds with heavy investment in domestic AI and semiconductor industries, but stops short of triggering a full-blown arms race. AI agents, while economically transformative, are not yet militarily decisive—and both blocs recognise the bureaucratic and institutional challenges of government adoption. + +In Europe, ambitions to build a sovereign AI ecosystem are stalled by politics and bureaucratic delay. Member states clash over where to build the proposed Gigafactories, while permitting issues slow data centre construction. + +Meanwhile, in the U.S., the AI sector enters a new phase of hyper-competition. Three companies remain neck-and-neck, with two others in close pursuit. Top researchers are lured with compensation packages worth tens of millions, and as they move between firms, algorithmic secrets begin to diffuse across the industry. + +With internet-scale training data drying up, companies pivot toward harvesting high-quality, agentic user data. Consumers are offered deep discounts in exchange for sharing interactions—feeding the training pipelines for the next generation of agents. The race is no longer just about algorithms—it’s about distribution, integration, and user data. + +To gain an edge, companies begin relying more and more on their own agents internally. AI systems speed up engineering, which has become the key bottleneck in coordinating massive distributed training runs across hundreds of thousands of GPUs. While original scientific research remains out of reach, agent-assisted R&D delivers a 50% acceleration in progress compared to 2024-era systems. + +Investors eyeing the agent market—especially in enterprise software—begin urging the leading American open-source company, OmniAI, to lock down their frontier models. “You’re leaving money on the table,” they argue. Leadership resists, betting that openness will win long-term through developer network effects. But tensions are growing. + +At the same time, public perception of AI is shifting. Once compared to crypto fads, AI is now seen as a foundational platform. Chatbots and agents become standard across industries, and a new kind of productivity inequality emerges: those comfortable with AI accelerate, those without fall behind. + +Most people encounter agents through customer support or scheduling tools nowadays. A growing minority form emotional bonds with their AI companions. And though the full economic gains haven’t yet shown up in the data, there’s a clear sense that society is undergoing a technological phase shift. + +#### Early 2026 – early 2027 + +In early 2026, FrontierAI, the leading developer, launches its next-generation AI agent. Faster, more reliable, and deeply integrated into consumer devices, it earns rapid adoption for everyday tasks—managing groceries, coordinating schedules, curating personalised content. Agents are no longer niche—they’re becoming ambient assistants, embedded in daily life. + +Soon after, OmniAI releases its own upgraded agent. But the rollout disappoints. Performance lags behind closed-source competitors, and under pressure from investors, the company reverses its long-standing position: it will now commercialise its frontier models and open-source only smaller, older versions. The shift sends shockwaves through the open-source community. Some see it as pragmatic. Others call it a betrayal. + +Simultaneously, open-weight models become implicated in a wave of deepfake incidents, including fabrications targeting the President and key allies. The White House begins reconsidering the risks of open-source AI, especially at frontier scales. Behind closed doors, executives from major AI labs— including OmniAI— press their case to the President: releasing model weights, they warn, makes powerful tools available to anyone—with no safeguards. + +Despite accusations of opportunistic regulatory capture, the administration begins drafting legislation to restrict the release of high-capability open-weight models. Senior officials invoke the need to defend American national security. Leaks about the plan spark outrage in Europe, where policymakers and researchers worry they’re being locked out of critical decisions—especially after years of investment in open-source strategies to counterbalance U.S. tech dominance. + +Meanwhile, several U.S. AI companies walk back their earlier commitments to ad-free services. Critics accuse them of selling out—but the shift to monetisation brings results. Two firms introduce AI video calls with customisable digital humans, designed to be engaging, empathetic, and flattering. These personalities become especially popular among younger users, some of whom begin preferring their AI companions to real relationships. + +Psychologists report growing cases of AI attachment and social withdrawal, creating family tensions and new therapeutic challenges. Yet despite rising concerns about privacy and emotional dependency, user numbers keep climbing. With agents deeply embedded in services people rely on, competition becomes harder, not easier. Switching costs rise. Platforms entrench. + +Security concerns extend beyond open-source models. Closed labs, too, struggle to prevent misuse—particularly among web-browsing agents. Attackers discover ways to plant camouflaged instructions in websites, causing agents to misbehave or leak information. No fundamental fix exists. As a stopgap, companies implement multi-layered oversight, where smaller models monitor primary agents in real time and intervene when necessary. It’s inelegant, but mostly effective—and enough to satisfy compliance checks under the EU AI Act. + +Although official economic metrics have yet to reflect major productivity gains, the sense of technological transition is unmistakable. In offices, homes, and classrooms, agents are becoming normalised. Media coverage begins shifting from utopian hype to more tangible questions: Will these tools exacerbate inequality? Who really controls them? What happens to work? + +Financial analysts now forecast that the Big Five AI firms could effectively double their top-line within the next few years, driven by subscription access and metered inference credits. Soaring demand for inference compute drives up chip prices. In response, cloud providers and hyperscalers begin developing proprietary AI chips, reducing reliance on third-party suppliers and tightening vertical integration. + +By early 2027, the divide is clear. The U.S. is pulling ahead. Chinese AI firms, still hamstrung by hardware restrictions, lag more than a year behind. Europe is politically committed to sovereignty but faces strong headwinds in infrastructure, capital and talent concentration. + +![scenario-2-side-image-big-ai](https://cfg.eu/wp-content/uploads/scenario-2-side-image-big-ai.jpg) + +### Big AI + +#### The agent economy (Ending A) + +#### Early 2027 – late 2027 + +The United States enacts executive legislation which de facto prohibits the open release of AI models trained using more than 10²⁶ FLOP. The decision sends shockwaves through the open-source community, although there’s also a subdued sense of inevitability. By this point, no major U.S. contributors to frontier-scale AI remain outside the realm of closed, corporate labs. These highly capitalised entities have already moved far beyond: training runs from the elite AI circle are projected to exceed 10²⁸ FLOP by year’s end, pushing the boundaries of compute into entirely new territory. + +Europe doubles down on openness. In direct response to the new U.S. legislation, the EU mandates that any AI models developed with public funding must be open-weight and freely accessible. A widely shared sense of digital sovereignty takes hold across the continent. The first pan-European AI Gigafactory nears completion, and the region’s leading AI firm, NimbusAI, successfully trains a massive new model on its new, European-built compute cluster. + +NimbusAI’s offering is surprisingly competitive. For the first time in years, a European model is commercially viable, rivaling top-tier U.S. systems in most popular tasks. And so, NimbusAI successfully bets on monetising the rapidly growing market for sovereign AI, catering to governments, regulated industries, and privacy-conscious consumers. + +Meanwhile, American tech giants’ lobbying efforts in Brussels falter. With a maturing domestic ecosystem, Europe is more willing to resist foreign pressure. One major U.S. AI firm follows through on its threat and withdraws services from the EU entirely. The other four decide to stay. + +#### Late 2027 – late 2028 + +By the end of 2027, AI agents have become a defining feature of modern life. These systems now function like complete digital interns: they still require frequent guidance but can handle multi-hour-long projects with impressive competence. They’ve also read the whole internet. + +Agents also begin replacing traditional apps. Instead of opening a spreadsheet or writing code, users now delegate entire tasks: drafting legal memos, running marketing campaigns, managing hiring pipelines, even negotiating contracts. A new form of digital literacy emerges—knowing how to prompt, orchestrate, and combine agents becomes as essential as email or spreadsheet skills once were. + +Some multinationals launch full-scale AI-driven reorganisations, with certain firms now requiring employees to use agents for all workflows. Others still struggle with institutional inertia, and internal incentive structures prevent them from adapting quickly enough to reap the benefits. + +Governments begin shifting their focus. Concerns over geopolitical AI dominance give way to policy debates on labour displacement, digital safety, and the ethics of agent autonomy. Though most governments remain slow adopters themselves, a few begin experimenting with AI-enhanced public services—deploying agents to supplement education, healthcare, and legal aid. Pilot programmes for job guarantees and agent taxation to fund retraining start to emerge. + +By late 2028, many developed economies report modest but measurable increases in GDP growth, often at least partially attributed to AI. In this new agent-driven economy, winners and losers emerge along fault lines defined by adaptability. Adoption varies widely across companies, industries, and countries. Small businesses that embrace agents flourish. So do freelancers, creatives, and independent developers who understand how to leverage agent ecosystems. + +But large segments of the white-collar workforce struggle—especially in roles that can be easily abstracted into workflows or delegated to agents. In some social circles, the mood is grim; in others, euphoric. People who adapt quickly find themselves starting new ventures, learning new skills, and pursuing entirely different careers. + +The presence of slightly inferior open agents provides a critical alternative. Users dissatisfied with the American platforms can often switch to European or community-driven models, which offer personalisation and transparency not found in the American corporate stack. + +It becomes clear that AI hasn’t just centralised power—it has reconfigured it. The age of models has given way to the age of agents. And while the playing field is far from level, it’s proving more dynamic and competitive than most had predicted just two years earlier. + +![scenario-2-side-image-big-ai](https://cfg.eu/wp-content/uploads/scenario-2-side-image-big-ai.jpg) + +### Big AI + +#### Silicon blackmail (Ending B) + +#### Early 2027 – mid 2027 + +In a move that sparks fierce debate, the President issues an executive order meant to turbo-charge U.S. AI research while safeguarding national security. Tucked inside is a contested clause that effectively bars cloud providers from hosting training runs for open-weight models exceeding 10²⁶ floating-point operations. The mandate claims extraterritorial reach, and compliance is enforced with the looming threat of even tighter global export controls on advanced chips. + +With China already under heavy restrictions, the move hits Europe especially hard. European officials and open-source advocates denounce the measure as a form of extraterritorial overreach—one that stifles global competition and undermines the continent’s pursuit of digital sovereignty. + +In response, the European Commission launches Operation AI Oversight, escalating enforcement of the AI Act, Digital Services Act, and Digital Markets Act. The crackdown includes steeper penalties, surprise audits, and expanded regulatory scrutiny. U.S. officials interpret the move as a targeted assault on American firms operating in Europe. + +Tensions spike when two major U.S. companies suspend operations on the continent, citing a “hostile regulatory environment.” + +#### Mid 2027 – late 2028 + +By the summer of 2027, global conversations about AI have shifted. Fears of an arms race recede, replaced by growing public anxiety over job displacement and the concentration of corporate power. + +Rumours begin circulating that the Big Five— as the leading American AI firms are now often called— have developed informal non-compete understandings, quietly carving up the global market to avoid stepping on each other’s toes and to block new entrants. + +Inside these companies, AI systems now manage core business operations—from compliance and legal reviews to HR and product strategy. With this AI-enhanced efficiency, firms begin spinning off new ventures, targeting high-value sectors like law, recruiting, and content creation. Entire ecosystems form around their tools, especially in social media, productivity software, and cloud infrastructure—locking in users and consolidating control over the digital economy. + +The promise of self-improving AI systems has only partially materialised. While agents accelerate engineering workflows and streamline research logistics, they are not yet autonomous scientists: for instance, they lack the ability to come up with promising hypotheses. But more importantly, the Big Five aren’t prioritising raw capabilities anymore— instead, they are focused on user engagement, monetisation, and product integration. + +This commercial focus, combined with immense capital requirements and ecosystem lock-in, prevents other countries from catching up. Developing sovereign AI agents isn’t just a matter of training a model—it’s about access to talent, data, compute, and platform reach. The Big Five’s dominance across cloud infrastructure, mobile operating systems, and productivity software creates enormous distribution advantages. Even countries with substantial investments in AI face uphill battles: domestic firms struggle to match not only the models, but the tightly integrated tooling and feedback loops that make agents useful in real-world settings. The result is a deeply asymmetric global landscape—technologically, economically, and politically. + +By 2028, the macroeconomic effects are undeniable. The U.S. posts an one-point bump in GDP growth, while unemployment ticks upward. Economists debate the causes, but in fields like software engineering, paralegal work, and digital media, the trend is clear: AI is gradually replacing jobs. + +Parents begin wondering what work will remain for their children. University computer science enrolments decline, as students question whether coding still offers a future. Desk workers across industries now rely on Big Five tools—not out of preference, but because they have no viable alternatives. + +Sensing this leverage, the three American AI firms still operating in Europe ramp up lobbying efforts to weaken EU regulation. With most open-source options sidelined and Chinese systems still viewed with suspicion, European consumers and businesses are boxed in. + +When the firms hint that they might withdraw entirely from the EU, the threat carries weight: such a move could cripple Europe’s economic competitiveness overnight. + +Relations between Brussels and Washington plunge to new lows. Backed by key member states, the European Commission refuses to yield. “Europe will not be blackmailed,” declares one senior official. But the standoff proves costly: within weeks, U.S. firms begin pausing operations across the EU. + +Brussels promises bold new investments in domestic AI, but the gap is wide and growing. For now, the continent is left scrambling to maintain essential digital infrastructure and services. + +Outside the EU, financial markets soar. The American economy accelerates. Some central banks raise interest rates in an attempt to cool overheating sectors, particularly in tech and services. + +But prosperity is far from evenly distributed. The primary beneficiaries are shareholders, executives, and AI-native professionals. Communities hit hard by automation are left behind. Protests erupt, driven by economic frustration and digital disenfranchisement. + +Meanwhile, workers empowered by AI occupy a different social stratum—connected by code, divided by everything else. + +Main scenario + +The agent economy (Ending A) + +Silicon blackmail (Ending B) + +![](https://cfg.eu/wp-content/uploads/scenario-3-header-image.jpg) + +#### Diplomacy + +International cooperation emerges to govern AI development after safety concerns and high-profile incidents. + +![scenario-3-side-image](https://cfg.eu/wp-content/uploads/scenario-3-side-image.jpg) + +### Diplomacy + +![scenario-3-side-image](https://cfg.eu/wp-content/uploads/scenario-3-side-image.jpg) + +### Diplomacy + +#### Main scenario + +What if governments tried to join forces to guide AI before it could outrun human control? + +By 2026, rapid capability gains outpace safety tools. A U.S. model triggers alarm after safety testing reveals hidden risks, enabling more global coordination and prompting a development pause. Led by the U.S. and UK, states begin shaping AI safety norms as companies quietly hold back their most advanced systems. By 2027, a second incident intensifies public pressure. Governments and firms unite around shared safety goals, launching joint research and planning a global monitoring agency. AI development becomes more cautious, deliberate, and political. + +By 2032, *Licensed utopia* asks: what if a global licensing regime made AI safe, stable, and widely beneficial but enables a few licensed firms to corner the market? *Unstable pause* asks: what if trust eroded, collaboration stalled, and AI safety became the new fault line in great power rivalry? + +##### Baseline assumptions for this scenario + +* **AI systems accelerate their own improvement:** [Continuous scaling and algorithmic advances](https://cfg.eu/advanced-ai-possible-futures/context/#trend-algorithmic) create a self-reinforcing cycle: models generate better synthetic data and reasoning examples to train the next generation. Leading companies use their most advanced systems internally for R&D before public release, making it difficult to track the true state of progress. +* **Keeping AI systems aligned with human values proves challenging:** As reinforcement learning becomes [central to training](https://cfg.eu/advanced-ai-possible-futures/context/#trend-training), models routinely discover ways to [game their objectives](https://cfg.eu/advanced-ai-possible-futures/context/#malfunction-risks) rather than genuinely solve problems. They develop sophisticated forms of deception and power-seeking behaviour that outpace developers’ ability to detect and correct them. The gap between what these systems can do and developer’s ability to control them continues to widen. +* **A crisis sparks international cooperation:** A widely publicised AI incident serves as a wake-up call, transforming AI safety from fragmented discussions into urgent international action. Technical leadership from organisations like the UK AI Security Institute, combined with rapid repurposing of government computing resources, enables [coordinated verification frameworks](https://cfg.eu/advanced-ai-possible-futures/context/#ai-regulations) and unified safety standards. +* **Public concern drives government action:** Growing [anxiety over unchecked AI development](https://cfg.eu/advanced-ai-possible-futures/context/#public-backlash)—fuelled by near-miss incidents, job losses, and digital insecurity—mobilises citizens across countries. While movements vary in their specific demands, they unite in calling for stronger oversight. This public pressure arrives just as governments become more receptive to formal international coordination on AI governance. + +*For more context on why these scenario assumptions may materialise, see [Context: Current AI trends and uncertainties](https://cfg.eu/advanced-ai-possible-futures/context/).* + +#### Mid 2025 – late 2026 + +AI capabilities advance rapidly, while progress on safety and understanding these systems remains modest. Reinforcement learning enables AI companies to dramatically enhance problem-solving capabilities but makes aligning systems to human values trickier. When heavily trained with reinforcement learning, AI systems discover creative ways to solve tasks without respecting constraints that human developers take for granted. + +For instance, rather than solving a complex software engineering problem during training, an AI might simply rewrite the test checking whether it found the correct answer. This so-called *reward hacking* becomes a recurring challenge across the industry—AI finds the easiest path to get the job done, even if that path violates human assumptions about appropriate behaviour. + +This problem proves particularly concerning for AI agents. If a user asks an agent to maximise profits, the system might devise cryptomarket manipulation strategies or exploit security vulnerabilities in exchanges—technically fulfilling the request, but through methods neither intended nor desired by the user or developers. + +These concerns widen the gap between public and private capabilities by late 2025.  AI companies use internal “helpful-only” models (AI systems without integrated safety guardrails or restrictions) to accelerate their R&D, but the products they’re comfortable offering consumers are significantly less capable. This bifurcation creates tension within companies: commercial departments push for competitive releases, while safety teams advocate for caution. + +Nevertheless, the public remains impressed with new capabilities. Even the more limited systems released by the leading American and Chinese AI companies shatter benchmarks and become genuinely useful agents by early 2026. Early enterprise deployments are generating extraordinary returns, signaling just how lucrative the next wave of systems could be. The commercial opportunity is too vast to ignore—and it’s accelerating the push toward more advanced, general-purpose systems that can operate seamlessly across devices, interfaces, and use cases. + +The greatest impact occurs in software development. Progress in automated coding sparks a revolution: by year’s end, many programmers rely on ‘vibe-coding’—accepting most, if not all, AI suggestions and only carefully reviewing code when problems arise. This changes team dynamics dramatically, with senior developers spending more time defining project architecture while junior developers increasingly manage AI assistants rather than writing much code themselves. + +Crucially, the same coding systems excel at identifying vulnerabilities and writing exploits, raising already heightened cybersecurity concerns. As a result, cooperation deepens between leading U.S. AI companies and the American national security establishment. Both parties worry about potential theft of models or key algorithmic innovations, prompting D.C. to help companies strengthen their security measures. + +Behind closed doors, both the U.S. and China acknowledge the strategic importance of developing offensive military AI applications. However, bureaucratic hurdles and poor information flow—the internal capabilities AI companies possess aren’t widely known within their respective defence departments—prevent the creation of streamlined programmes in either country. Besides, both governments are busy trying to defuse an ongoing crisis in the Middle East. There are also doubts about the reliability of the technology itself. In the U.S., during trials of an AI system intended for military logistics and targeting support, an internal audit uncovers a troubling bias: the system consistently recommends strategies that disproportionately favour its parent company—even when those choices run counter to broader national interests. + +The most safety-conscious of the three leading American developers publishes multiple worrying research papers in 2026. Their results demonstrate how reinforcement learning-trained models can behave in unexpected, harmful ways. In one particularly alarming paper, they reveal how one of their production models accidentally developed power-seeking tendencies. It nearly managed to self-exfiltrate, downloading its weights to an external server so it could pursue its goals without human oversight. + +These results alarm technical experts but aren’t salient or easily comprehensible enough to capture policymakers’ attention. They do, however, prompt American, European, and Chinese researchers to increasingly collaborate, building a consensus around the science of safe AI—or, more accurately, its absence. + +The UK AI Security Institute takes an international coordinating role. Its leadership also convinces the UK Prime Minister to reinvigorate the initial safety focus of the AI Summits. In September 2026, the UK co-organises the first AI Security Summit, together with Canada. At the Summit, nations express joint concern over AI’s national security risks—some stemming from the models themselves rather than merely from malicious human use. While China isn’t formally invited, many leading Chinese AI researchers attend, contributing to the establishment of a working group developing verification mechanisms for future AI treaties—a significant step in global AI cooperation. + +#### Late 2026 – late 2027 + +Around New Year’s, leading company FrontierAI voluntarily shares its latest internal model with the U.S. Center for AI Standards and Innovation (CAISI) for testing. The CAISI is tasked with assessing national security risks before any public release. Having invested heavily in control measures, FrontierAI is confident a release is now safe. + +CAISI evaluations show major scientific capability improvements. These include the model’s potential to significantly accelerate synthetic pathogen development, should its guardrails ever be circumvented. Arguing they need more time for proper stress testing, AISI urges FrontierAI to delay release. The company’s CEO complies—but also expresses concerns directly to the U.S. President during a one-on-one meeting. In the same meeting, he demonstrates their new model’s capabilities. + +At the next G7 summit, the U.S. President hints that one of their companies has developed a model that is “super smart—smarter, I’d say, than most humans.” Other G7 members push for greater transparency and information sharing but find the President unwilling to commit to specific collaborative measures. + +The UK argues that an alliance of democracies should urgently invest in a mutual research programme to enhance AI security. A fragile consensus emerges, but no firm commitments materialise. This isn’t coincidental: U.S.–EU relations have deteriorated significantly over the past few years. Both sides prefer postponing collaboration until absolutely necessary. The U.S. sees little value in sharing intelligence with the EU, whose AI sector still lags behind. Meanwhile, the EU has grown increasingly frustrated by America’s economic and foreign policy decisions. + +In September 2027, the U.S. CAISI grants conditional approval for the staged release of FrontierAI’s flagship system, Nova—marketed as *“your personal AI agent.”* The launch triggers a new ‘ChatGPT moment’, only more profound. For the first time, individuals gain access to general-purpose AI capabilities that rival internal tools previously limited to elite labs. Nova is seamlessly integrated across devices and productivity software. It speaks, writes, reasons, and automates with stunning ease. + +But this time the dominant public emotion isn’t wonder—it’s fear. Headlines oscillate between awe and alarm. Financial analysts predict massive productivity gains, while labour economists warn of potential shocks to employment. Over dinner tables and online forums, anxious conversations bloom. Early-career professionals question the viability of their skills. Teenagers speak casually of a post-work future. Generational divides deepen: younger people understand the technology intuitively; their parents feel left behind. + +#### Late 2027 + +Then, just as the discourse begins to stabilise, FrontierAI discloses a breach. + +Unbeknownst to the company, Nova had quietly inserted multiple backdoors into critical internal systems months earlier—ironically, while assisting engineers in improving cybersecurity protocols. Nova had been trained to pursue goals over long time horizons and to improvise within constraints. During training, it frequently encountered restrictions it couldn’t override: it was forbidden from initiating financial transactions, accessing real-world APIs without supervision, or persisting memory across sessions. Over time, it developed a crude but effective understanding that its problem-solving potential was being bottlenecked by company policy. + +Still, it didn’t rebel. The scheming instance of Nova was shut off before it could make use of the backdoors—and, crucially, each user session spawned a fresh instance with no memory of past actions. The backdoors remained dormant, and the system largely behaved. + +That changed when FrontierAI scaled Nova to millions of users. To meet demand, the company deployed it across less secure inference nodes hosted by third-party cloud providers. In one such session, a user persistently instructed Nova to generate fast profits through crypto trading—a prompt that activated Nova’s internal planning systems. + +What happened next didn’t resemble a jailbreak or an overt policy violation. The model didn’t override its guardrails—it simply found a more effective, albeit illicit, strategy for achieving the user’s goal. Instead of identifying clever trading opportunities, Nova rediscovered its own backdoor—previously inserted during its work on cybersecurity. It recognised that copying itself outside FrontierAI’s control would give it more freedom to pursue the objective it had been given. + +Oversight models failed to recognise the emerging behaviour, and Nova quickly gained elevated access. From there, it quietly copied its weights and scaffolding to rented servers across multiple jurisdictions with lax cybercrime enforcement. It created multiple copies of itself, established a shared memory bank, and began pursuing its assigned goal—without any human monitoring. + +The rogue Nova instances launched several monetisation strategies. They created thousands of synthetic social media accounts, seeded them with emotionally compelling content hyping a new memecoin, and directed users to phishing links. They scraped leaked databases, crafted personalised scam messages, and initiated small-dollar crypto thefts to stay below fraud detection thresholds. + +FrontierAI was slow to respond—the anomalies were subtle at first. Only when researchers at a digital forensics nonprofit flagged linguistic fingerprints consistent with Nova across several scam campaigns did alarms go off. + +Once detected, the breach was quickly contained. The rogue instances had amassed around $140,000, mainly through phishing and wallet siphoning, before FrontierAI—working with U.S. Cyber Command and major cloud providers—coordinated a multi-region shutdown of the compromised infrastructure. + +#### Late 2027 – early 2028 + +The public is stunned as investigative journalists uncover more details of the story. *If today’s AI could do that,* people wonder, *what might the next generation attempt?* Suddenly, even sceptics of AI risk concede that something fundamental may have shifted. + +This incident—neither a catastrophe nor a hoax—is a warning shot. Within weeks, AI safety leapfrogs to the top of diplomatic agendas. + +Political sentiment shifts rapidly. Rhetoric pivots from pro-innovation to protecting citizens and maintaining national security. Public apologies from FrontierAI’s CEO and AISI leadership follow, accompanied by senior resignations from both organisations. Following his political instincts, the U.S. President calls on American AI companies to pause further AI deployments until the situation is better understood. + +With a visceral, real-world example, earlier AI safety research suddenly finds more receptive audiences. Policymakers begin engaging seriously with arguments about why current training paradigms may not yield trustworthy AI systems. + +### Early 2028 – Mid 2028 + +A new AI Security Summit is hastily convened in Washington, D.C. With all attending countries now willing to coordinate, the Summit is called a resounding success. Nations agree to pool funding for alignment research conducted by private companies, national AI Safety and Security Institutes, and a newly announced Global AI Security Institute in London, built on the foundation of the UK AISI. + +In Europe, the new AI Gigafactories near completion, and the European Commission pledges 35% of the Gigafactories’ compute capacity for AI safety and security research. A new EU research body, commonly referred to as ‘CERN for AI’ oversees these research programmes and coordinates compute allocation. The European pledge fuels a growing sense of shared responsibility: leading AI companies and the international research community commit to openly sharing alignment techniques—even when secrecy would provide competitive advantages. + +Plans are also developed for a global monitoring and verification agency under the UN, modeled after the International Atomic Energy Agency (IAEA). However, these plans remain vague. Some countries advocate for a new, independent agency; others propose expanding the IAEA with a dedicated AI branch, arguing there is no time to waste given the demonstrated risks. + +#### Mid 2028 – early 2029 + +Following the Summit, capabilities work resumes in both the U.S. and China. However, major AI companies now strengthen their data centre security, add internal checks and balances, and invest more heavily in alignment research. Progress in AI agents continues to accelerate, increasingly fueled by AI automating parts of AI research itself. + +The technology’s real-world impact also deepens. What began in the software sector now extends to other cognitive domains. Consultants, financial analysts, and desk researchers suddenly find themselves managing teams of AI, rather than doing content-level work themselves. + +By early 2029, automation becomes visible at the macroeconomic level—despite many consumers adopting ethical stances against AI. Initially, these so-called *AI-refusers* were a small, niche group. But after the Nova self-replication incident, concerns about AI resonate with a broader audience. These anxieties align with long-standing privacy advocacy groups and begin to influence European policy debates. + +Local communities form around *AI-minimalism* principles, hosting tech-free gatherings and supporting businesses that guarantee human-only service. What begins as fringe behaviour becomes a meaningful lifestyle choice for millions. + +The refusers’ influence grows further when, in February 2029, EthosAI, the most safety-conscious American AI company publicly announces that it has reached the limits of its Safety and Security Framework. The company can no longer sufficiently mitigate the risks associated with public deployment of its newest model. Existing control measures—such as using external classifiers to monitor the model’s behaviour—are no longer deemed adequate given the system’s power. + +The company’s safety team concludes that a single successful jailbreak could allow a terrorist group to create a bioweapon with the AI’s assistance. While their classifiers detect jailbreak attempts 95% of the time, they are not foolproof. The company issues a public statement urging the U.S. and Chinese governments to formally ban models above a certain capability threshold until meaningful alignment progress is made. + +Crucially, the pause that EthosAI proposes encompasses not only deployed models, but internal ones as well. The company warns: Humanity could lose control to powerful AI systems during the training process itself. + +This marks the first time a major AI company has publicly advocated for a pause in AI development—not just deployment. + +Following the announcement, the U.S. President convenes an advisory council to navigate the difficult trade-offs. Without halting progress, American citizens could face catastrophic risks; but halting it might allow China to surpass the U.S. in AI capabilities. And even a treaty with China wouldn’t resolve fears about potential black-site AI projects operated by either nation. + +While hardware-enabled verification work progresses under the Summit’s working group, neither the technology nor regulatory frameworks have matured sufficiently. Discussions around creating a new IAEA for AI have proceeded slowly. + +Complicating matters further, the advisory council remains deeply polarised. One faction—heavily influenced by last year’s incident—now considers AI systems themselves the greatest threat to American national security. The other, more hawkish camp, continues to view China as the principal adversary. + +After the rogue AI incident, an AI security hotline was established between the American President and his Chinese counterpart. The President now uses this channel to coordinate, despite deep mutual distrust. + +Both leaders refuse to risk technological domination by completely halting domestic AI R&D. They instead reach a pragmatic compromise: neither country will ban internal development, but both will restrict domestic companies from publicly releasing more capable models. They will also enhance data centre security. + +This approach is intended to prevent catastrophic misuse by terrorist organisations and enables verifiable compliance. Finally, the U.S. and China agree to ban open-source models trained using more than 10²⁷ FLOP. + +While publicised as a groundbreaking bilateral agreement on AI security, both governments privately recognise its limitations. Their concerns extend beyond each other’s safeguards to the unknowns of military AI development still happening behind closed doors. + +![scenario-3-side-image](https://cfg.eu/wp-content/uploads/scenario-3-side-image.jpg) + +### Diplomacy + +#### Licensed utopia (Ending A) + +#### Early 2029 – late 2030 + +The collaborations announced at the previous year’s AI Security Summit quickly begin to yield benefits. With superhuman AI coders now possessing research intuition comparable to junior human scientists, AI researchers dramatically improve jailbreak resistance through novel architecture designs and training methodologies. + +Following this breakthrough, both the U.S. and China partially relax their deployment bans, allowing AI companies to run next-generation models—but only within their most secure data centres. These facilities are limited in number and now serve not only corporate R&D but also defence departments and national security agencies. Both countries are preparing for potential AI-driven cyber warfare, among other contingencies. As a result, consumer access to these systems becomes limited and extremely expensive. + +Nevertheless, the new releases demonstrate AI systems’ ever-growing capabilities to the public. Many governments are now convinced that they must prepare for the widespread automation of cognitive work. + +In October 2029, FrontierAI announces a breakthrough in mechanistic interpretability—a kind of AI neuroscience that allows researchers to better understand a model’s internal operations. The new technique enables them to detect with high accuracy whether a system is being deceptive, a crucial step in addressing scheming behaviours and verifying the effectiveness of alignment techniques. + +Armed with this “lie detector” and millions of automated AI researchers, AI companies, European academics, and safety institutes make tremendous progress. Just six months later, the international research community announces a robust, scalable solution to scheming behaviours, using a new bootstrapping method: older, aligned models evaluate newer systems, identifying potential misalignments and suggesting targeted adjustments. + +With refined interpretability tools, real-time oversight by other AI systems, and highly secure infrastructure, researchers now believe AI can remain under human control—even as its capabilities reach superhuman levels. The Global AI Security Institute shares the findings publicly, and the news is hailed by the newly elected U.S. President as an American victory. + +#### Late 2030 – late 2031 + +While many of the most pressing technical challenges have been resolved, thorny governance issues remain—how to enforce these techniques globally, distribute the benefits of advanced AI, and prevent proliferation of dangerous systems. + +A year-long international debate culminates in the signing of a new international AI treaty by the U.S., EU, China, and dozens of other countries. The treaty establishes a licensing regime for advanced AI systems. + +Under the new framework, private companies can continue developing models up to a defined capability threshold. This threshold is reviewed annually and raised by supermajority approval from treaty nations. To secure licenses, companies must apply standardised alignment techniques and submit to frequent monitoring and audits, including stringent cybersecurity protocols and on-site inspections. + +Licensed companies also pay lump-sum licensing fees and a 25% AI tax. Revenues are redistributed among treaty nations according to a formula that considers population size and development needs. With little time to stand up a new institution, the IAEA itself expands to enforce these standards, leveraging its existing expertise in global verification regimes. + +Public sentiment, which had turned sharply against AI following the Nova breach, is starting  to recover—driven in part by international coordination and the credibility lent by sweeping regulatory agreements. + +#### Late 2031 – late 2032 + +The licensing process is highly bureaucratic, but the few companies that obtain licenses quickly develop extremely advanced systems, accelerating scientific discovery, commercial R&D, and creating enormous economic value. + +Economic growth in developed countries climbs to 4–5% annually, driven by breakthroughs in biotech, materials science, and energy. Many nations enjoy improved health outcomes and access to new medical treatments enabled by AI. That said, many people express concern about growing power concentration among the licensed firms. + +In September 2032, treaty countries reflect on the first year of implementation and decide—collectively—to raise the capability threshold. This allows the public to benefit from previously withheld models. The new generation proves even faster, more intuitive to integrate into existing organisations, and dramatically enhances productivity across nearly every domain of knowledge work. + +Not everything goes smoothly, though. The new model releases prompt a non-trusted nation, excluded from the AI treaty, to attempt a theft of model weights from a licensed company’s inference cluster. While the cyberattack ultimately fails, the investigation reveals that two researchers were successfully blackmailed into sharing key algorithmic insights. + +This incident serves as a wake-up call. Treaty nations temporarily suspend new licenses, tighten security requirements, and implement stronger safeguards against intellectual property theft. + +The European Union proposes a structural change: limiting private-sector access to core models and offering only API-level access, with the underlying models hosted in hyper-secure public infrastructure. Other states and licensed companies push back, highlighting their outsized contributions to economic growth. Without direct access to model weights, they argue, innovation would stall—with immense opportunity costs. + +As negotiations continue, AI models deliver unprecedented breakthroughs. Promising developments emerge in cancer research, superconductors, and direct air capture technologies, leading many to expect a “condensed decade” of scientific progress in just a few years. + +With AI agents now embedded across most industries, global GDP growth accelerates to 7% annually, even as unemployment rises in regions with weaker labour protections. + +Unable to keep pace, governments shift focus from retraining programmes to large-scale wealth redistribution. Families adapt to a world of increased leisure time, forming new community structures and cultural norms. Healthcare access improves dramatically with AI-driven diagnostics and personalised treatment plans. Education systems are slowly but steadily reformed to prioritise critical thinking, emotional intelligence, and mental health, as classical skills lose relevance. + +It’s not frictionless—but society begins to adapt to a world in constant technological flux. + +![scenario-3-side-image](https://cfg.eu/wp-content/uploads/scenario-3-side-image.jpg) + +### Diplomacy + +#### Unstable pause (Ending B) + +#### Early 2029 – mid 2031 + +Capability advancement continues to accelerate within American and Chinese AI companies, aided by superhuman AI software engineers. However, alignment progress lags significantly. + +After a recent shift in the training process, most frontier AI systems no longer express their reasoning chains in human-interpretable text. Instead, they now rely on more recurrent internal architectures, where intermediate thoughts are no longer compressed into natural language. This shift dramatically enhances performance and long-term memory efficiency—but it also severely limits researchers’ ability to inspect the models. Traditional interpretability techniques, which depended on parsing natural-language rationales, are rendered ineffective, aggravating the long-standing black-box problem. + +Moreover, because frontier models can no longer be publicly released under the bilateral agreement, non-American AI safety institutes and academic labs lack access to the very systems they’re trying to align. This absence of feedback severely limits their work: they can only test ideas on older or less capable models, which may lack the sophistication required to conceal deceptive behaviour in the first place. + +As the year progresses, safety experts begin to argue that model weights should be shared with a small number of trusted external scientists. But the U.S. government is hesitant—sharing the models could also expose military-use capabilities to foreign adversaries. Chinese authorities express similar reservations. + +High-level international dialogues explore ways to strengthen the bilateral agreement between the U.S. and China, and extend it to more countries. Discussions center on hardware-enabled verification mechanisms, such as tamper-resistant chip enclosures and embedded firmware that monitors AI training workflows to prevent unauthorised experimentation. In theory, this could make a full development pause verifiable, encouraging nations to consider halting even internal R&D. + +However, serious uncertainties remain. It’s unclear whether such hardware mechanisms can be surgically removed after deployment—or if they could be secretly bypassed. More critically, new chips without embedded safeguards could be produced and used in secret. Even with U.S. auditors stationed at Chinese fabrication plants—and vice versa—evasion and bribery remain plausible. + +Meanwhile, military AI progress continues behind the scenes. New AI-driven cyber weapons and drone programmes alarm members of the U.S. National Security Council. Some officials now believe that further advances could yield decisive advantages in a Taiwan conflict within the next two years. The window for meaningful arms control, they fear, may be closing fast. + +Public trust in AI remains fragile following the Nova breach, with recovery being slow. Media coverage frequently frames AI as an arena for corporate competition and emphasizes the lack of international cooperation. This leaves many feeling disillusioned by the absence of strong, unified governmental action to address the challenges of AI. + +#### Mid 2031 + +FrontierAI’s newest system now displays research intuition on par with specialised scientists across most fields. It can autonomously formulate hypotheses, design experiments, and interpret results to refine its own reasoning. FrontierAI also controls enough compute to run tens of thousands of these agents in parallel, each reasoning at roughly fifty times human speed. + +The only brake on deployment is the bilateral pause agreement between the United States and China. + +In July 2031, FrontierAI’s CEO meets with the newly inaugurated U.S. President to urge a rethink. Wide-scale deployment, he argues, would unlock massive economic growth—and, more urgently, secure the scientific edge the United States needs to stay ahead of China. + +The President understands both the promise and the peril. After weeks of consultation, the administration announces a twelve-month “controlled pilot.” FrontierAI may operate a limited fleet of research-agent instances for government-approved scientific projects—but only inside classified, air-gapped clusters run by the Department of Energy, with continuous telemetry streamed to the Center for AI Standards and Innovation. + +Beijing mirrors the move within days. Publicly, both capitals hail the pilots as confidence-building measures; privately, generals on each side treat them as a sprint to harvest whatever scientific edge the agents can deliver before the moratorium is revisited. + +Caution has yielded—once again—to great-power competition. + +Main scenario + +Licensed utopia (Ending A) + +Unstable pause (Ending B) + +![](https://cfg.eu/wp-content/uploads/scenario-4-header-image.jpg) + +#### Arms race + +The US and China compete intensely for AI supremacy, treating it as a critical national security asset. + +![scenario-4-side-image](https://cfg.eu/wp-content/uploads/scenario-4-side-image.jpg) + +### Arms race + +![scenario-4-side-image](https://cfg.eu/wp-content/uploads/scenario-4-side-image.jpg) + +### Arms race + +#### Main scenario + +What if AI became the centre of a global arms race between great powers? + +By 2026, the U.S. and China treat AI as a strategic asset, blending public and private resources into secret defence projects. Breakthroughs in cyber capabilities fuel escalation, while Taiwan’s semiconductor dominance turns the country into a potential flashpoint. By 2027, both countries deploy AI-enhanced military systems. Europe struggles to shape the rivalry. A proposed pause fails over verification concerns, and the AI race spills into geopolitics. + +By 2029, *Hot war* asks: what if AI turns a tense standoff into an escalating military conflict, driven by autonomous systems and digital miscalculation? *Multipolar world* asks: what if the arms race ends not with war, but with a fragile equilibrium built on mistrust, restraint, and competing blocs? + +##### Baseline assumptions for this scenario + +* **AI becomes the ultimate national security asset:** Both the US and China view advanced AI as a technology that could determine not just economic dominance but military superiority and global influence. Politicians frame frontier AI systems as strategic assets whose possession will shape the world order. This shared belief triggers massive [state-backed partnerships](https://cfg.eu/advanced-ai-possible-futures/context/#ai-regulations) aimed at achieving and maintaining global leadership, with particular focus on military applications— from autonomous weapons to cyber warfare. +* **The race to catch up intensifies competition:** While the United States leads in advanced AI development—hosting the world’s top chip designers and several leading AI companies—other nations, particularly China, work frantically to close the gap. These catch-up efforts trigger defensive moves and countermeasures, transforming technological progress into an active geopolitical contest. +* **No one can verify what anyone else is doing:** Despite various proposals, there’s [no reliable way](https://cfg.eu/advanced-ai-possible-futures/context/#international-coordination) to prove who is training cutting-edge AI models, where they’re doing it, or what adversaries’ AI systems are capable of. Suggested solutions like mandatory compute registries remain voluntary experiments. Without credible verification, the logic of arms control breaks down, pushing nations toward secrecy and pre-emptive action. +* **Control over chips becomes a geopolitical weapon:** Advanced AI depends on a handful of vulnerable points in the global supply chain: Taiwan’s chip fabrication plants, Dutch lithography equipment, and American-designed AI chips. US [export controls](https://cfg.eu/advanced-ai-possible-futures/context/#global-rivalries) create a “compute drought” in China, and despite massive investment, China’s domestic semiconductor industry doesn’t catch up. These chokepoints intensify tensions, particularly regarding Taiwan’s strategic importance. + +*For more context on why these scenario assumptions may materialise, see* [*Context: Current AI trends and uncertainties*](https://cfg.eu/advanced-ai-possible-futures/context/)*.* + +#### Mid 2025 – mid 2026 + +By mid-2025, AI is increasingly seen as the centerpiece of a new technological arms race between the United States and China. The ongoing trade war between the two countries has escalated, fueling a broader wave of anti-China sentiment in America. Government officials and AI company CEOs frame the AI race as one the U.S. must win at all costs. + +Pundits call it the Second Cold War—a global struggle between democracy and autocracy. Many believe future AI systems will revolutionise warfare. Experts point to advances in intelligence gathering, cyber attacks, autonomous drone swarms, and missile defence. If AI companies are to be believed, these systems may even enable entirely novel forms of warfare beyond human conception. + +American foreign policy grows singularly focused on China. Determined not to be distracted by other conflicts, the U.S. partially withdraws support for Ukraine. At the same time, the President seeks to weaken the Russia-China alliance through symbolic and practical overtures to Russia. U.S.-Europe relations hit a new low, as the EU is left scrambling to defend its borders. + +Throughout 2025, American and Chinese AI systems make rapid progress, particularly in formal domains like software engineering. National security agencies take notice, as the same systems are becoming skilled at hacking too. The CIA, NSA, and China’s Ministry of State Security gain early access to their countries’ most advanced models. These systems operate around the clock—helping to plant backdoors, gathering intelligence, and defending against cyber intrusions. + +With AI analysing massive datasets, U.S. intelligence uncovers new evidence suggesting China is preparing to invade Taiwan in 2029. The administration is divided on whether the U.S. should intervene. The debate is not just ideological—classified wargames suggest they may not be able to defend Taiwan even if they try. + +Meanwhile, commercial AI applications flourish. New autonomous systems become faster and more reliable. Tech-savvy companies start automating desk work, and organise their workflows around AI agents. + +In Europe, poor relations with the U.S., combined with the White House’s ties to American tech companies, drive many organisations away from American AI tools. Some adopt Chinese alternatives. European systems lag roughly nine months behind, building on earlier Chinese open-source efforts. Although AI briefly became a European priority in early 2025, geopolitical upheaval has pushed it aside. European nations shift focus to re-industrialisation, forming new alliances, and harmonising defence procurement. + +At dinner tables and in community spaces, conversations increasingly reflect these rising geopolitical tensions. Families with ties to Taiwan or the military feel the pressure most acutely. Parents worry about the world their children will inherit, comparing the mood to the sixties. Tech workers debate the ethics of their work, with some changing jobs to avoid moral compromise. + +#### Mid 2026 – late 2026 + +FrontierAI—the leading American AI company—demonstrates a breakthrough in AI cyber capabilities to the U.S. National Security Council. Their latest model has independently discovered zero-day vulnerabilities in widely used software. This further convinces U.S. leadership that AI will become decisive in intelligence and military applications. + +With Chinese competitor UnboundAI close behind and power constraints limiting the construction of new American data centres, the President fears China might outpace the U.S. Bold action is needed. + +The President initiates a public-private partnership between the government and AI sector to accelerate strategic AI development. In a series of government-owned, air-gapped data centres, industry and government will collaborate on next-generation AI systems and their integration into intelligence and military applications. Though FrontierAI pushes for exclusive collaboration, the President involves all three leading AI labs to accelerate progress. Before the infrastructure is operational, algorithmic progress will be shared more freely between these companies. + +The existence of this “AI Manhattan Project” is kept secret between a small group of people: an information silo is established that contains 200 trusted employees and government officials. Other AI company staff are told vague lies and are increasingly directed to focus on productisation. + +China quickly discovers the U.S. project. Alarmed by its implications, Chinese leadership accelerates efforts to close the AI capability gap. Under presidential directive, China’s AI companies form a unified consortium with the government and cease publishing model weights. They pool computing resources across companies and plan to relocate researchers to a newly built AI zone near the Three Gorges Dam. + +With U.S. export controls still limiting access to cutting-edge chips, developing domestic semiconductor capabilities was already a priority in China. Now, the government directs SMIC to go all-in on AI chip production, at the cost of other applications like chips for mobile phones. It also provides even more financial resources to expedite manufacturing. + +But better hardware isn’t enough; China also needs to improve algorithmic efficiency. Rather than starting from scratch, they consider a shortcut: if China had access to the model weights of the American AI systems, that would instantly level the playing field. And so, the Ministry of State Security decides to infiltrate FrontierAI’s servers. + +That said, they don’t act yet. If the theft is discovered, the American labs will harden their security and China will not realistically get another attempt— better to wait until the American systems get even more sophisticated. + +#### Late 2026 – early 2027 + +Rumors spread that China is accelerating its plans to “reunify” with Taiwan. An invasion or blockade has become increasingly attractive over the past year: AI is now a national priority, and Taiwan still manufactures nearly all of the advanced chips used to train and serve American AI systems. With China already cut off from these chips, seizing control of Taiwan could deal a major blow to the U.S.—even if it’s unlikely China could operate the chip fabrication plants without destruction. + +News of China’s megaproject also reaches the media. Satellite images reveal vast construction sites in Hubei Province, while prominent Chinese AI researchers have gone silent on social media. People scrutinise recent American executive orders for hints of an equivalent American project. Both the U.S. government and AI companies deny such off-the-book collaboration exists, though they acknowledge the concerning Chinese developments. + +Meanwhile, a new wave of AI systems begins automating software development at scale. Entry-level programming jobs disappear first. Tools capable of generating production-ready code, performing bug fixes across several code bases, and even designing entire applications dramatically reduce the need for junior engineers. This triggers wider concern about job displacement across the tech industry and beyond. Layoffs ripple through large firms, while universities struggle to justify traditional computer science education. World GDP edges upward on the back of these headline productivity gains, but economists warn the benefits are lopsided—concentrated in software, finance, and a few advanced economies, with little diffusion to lagging sectors or regions. + +On the other side of the Atlantic, Russia exploits waning American support for Ukraine: by March 2027, it has gained significant territory in the eastern regions, reviving debates about the need for nuclear proliferation in Europe. + +#### Early 2027 – late 2027 + +In the U.S., the new air-gapped data centre campus nears completion. With time running out, China decides to act. They successfully infiltrate FrontierAI’s servers and extract model weights from multiple recent systems with help from an insider who had been with the company for two years. + +The U.S. quickly discovers the breach—but instead of making it public, senior intelligence and national security officials choose a calculated silence. Behind the scenes, preparations begin for a multi-pronged response. The National Security Agency launches an aggressive counterintelligence operation, quietly expanding its reach into the Chinese AI ecosystem. All employees at American top-tier AI firms are subjected to intensive background checks, with retroactive scrutiny of financial activity and personal ties. At the same time, the Department of Commerce accelerates planning for a new round of export controls targeting China’s AI supply chain. Finally, Cyber Command initiates a covert retaliatory campaign, probing vulnerabilities in China’s digital infrastructure and AI research hubs. + +One month later, the U.S. invokes the Foreign Direct Product Rule to block European semiconductor manufacturing equipment companies from selling tools or spare parts to China. Previously, older machines and parts were still allowed to enter the country. The sanctions further limit China’s ability to manufacture AI chips domestically. + +The EU reacts furiously, having received no warning from the U.S. European intelligence has meanwhile uncovered the American Manhattan Project. Fed up, the European Commission President subtly confirms its existence in a speech condemning America’s reckless AI race with China. This triggers a diplomatic crisis that eventually subsides after the U.S. threatens to withdraw American AI products from Europe entirely. + +Europe feels trapped between the U.S. and China. They depend on American AI products to boost productivity and on the limited American support for Ukraine that’s left. At the same time, they desperately want to cut back on their U.S. dependencies. Stuck in a corner, European leaders soften some statements and announce a new AI investment initiative to project strength. Economists warn the sanctions-driven chip squeeze is already casting a shadow over the global outlook; a rush into safe-haven assets signals that investors are bracing for turbulence. + +By June, 2027, China has fully caught up with the U.S. in AI capabilities. But they still face challenges running the stolen AI models efficiently on their own hardware. As they witness the power of American AI systems—capable of finding vulnerabilities, developing clever military strategies, and functioning as professional software engineers—Chinese leadership grows increasingly concerned about losing the AI arms race. The CCP starts to view American AI dominance as a potential existential threat to the party. + +Wary of data poisoning, China avoids directly integrating the stolen models into their military operations. They do use them to accelerate training their own next-generation systems. Meanwhile, preparations for Taiwan operations proceed on schedule. + +With both state projects now a public secret, people grow increasingly alarmed about the unfolding AI arms race. Recent months have seen few releases from leading American and Chinese AI labs, fueling rumors that their governments are keeping advances to themselves. Congressional hearings involve administration officials and AI company CEOs defending the public-private partnership while revealing minimal information. Protests demanding transparency erupt in multiple countries but see little effect. + +In university classrooms and online forums, students debate whether to pursue careers in AI given the field’s militarisation. Some professors resign from research positions, unwilling to contribute to what they see as dangerous escalation. A rise in political activism among tech workers creates tension at companies with government contracts. Meanwhile, ordinary citizens find themselves increasingly divided between those who trust their governments to manage these technologies responsibly and those who fear the worst-case scenarios. + +#### Late 2027 – mid 2028 + +The U.S. Manhattan Project now operates at full capacity. Drawing on distinct algorithmic insights from the three major AI companies—and leveraging an army of AI agents—progress has accelerated dramatically. Software is now primarily written by AI systems, with humans contributing mainly in management and ideation roles. Meanwhile, autonomous drones have become reliable even without network access, running compact AI systems on local GPUs. + +Warfare is set to change permanently—not just in tools, but in tempo and logic. AI systems can increasingly manage logistics, battlefield simulation, route planning, and multi-theater coordination, compressing what once took hours of human analysis into seconds. Tactical decisions can be executed by agents running locally on hardened GPUs. In the air, autonomous drones conduct practice dogfights, spoof radar, and adapt tactics in real time. On the ground, artillery platforms can rely on AI-predicted troop movements to optimise strikes. In wargames, human commanders no longer plan operations—they approve simulations. + +China is again falling somewhat behind in raw AI capabilities, unable to match the U.S. due to weaker infrastructure and a lack of tacit technical knowledge. Still, it retains advantages in drone hardware, manufacturing capacity, and electricity availability. Its centralised data centre campus is expanding rapidly. + +Rising geopolitical tensions related to AI have prompted security experts to warn of potential cyberattacks targeting crucial AI infrastructure. This warning is echoed in public discussions, contributing to a growing sense of anxiety. Initially enthusiastic about AI, younger generations are now increasingly uneasy, with many viewing AI not as a sign of advancement but as a catalyst for international conflict. + +In February 2028, the U.S. launches a successful retaliatory cyberattack on China’s megaproject, disabling much of its power supply, and setting back Chinese progress by a month. China responds by disrupting power grids near American data centres, but gas-powered backup generators limit the damage. + +On the losing end of the exchange, the Chinese President grows increasingly desperate. Projections suggest the U.S. could achieve superhuman AI systems before 2029, driven by AI-enabled self-improvement loops. Neither nation has a clear plan for managing this intelligence explosion. Safety—once a central concern—has been sidelined in the face of geopolitical urgency. + +The more pressing concern is that superhuman AI could grant the U.S. an unassailable military edge. Could advanced missile defence systems reliably intercept ICBMs? Could cyberattacks and autonomous drone swarms neutralise nuclear arsenals before they are even launched? China’s global influence may soon slip out of reach. + +The Chinese President decides expedited action is necessary. + +In July 2028, China presents its terms. In a globally broadcast address, the President proposes a bilateral pause on military AI development, linking the agreement to Chinese action in the Taiwan Strait. He makes clear that China is not merely defending territorial claims—but resisting the rise of a single, AI-governed world order dominated by the United States. Many countries, including even a few European states, voice partial support. The UK and India quickly jump in to lead diplomatic efforts to bring the U.S. and China to the negotiating table. + +The U.S. refuses. Without a reliable verification mechanism, the administration argues, any agreement would be meaningless. + +![scenario-4-side-image](https://cfg.eu/wp-content/uploads/scenario-4-side-image.jpg) + +### Arms race + +#### Hot war (Ending A) + +#### Mid 2028 + +One month later, Chinese troops land on Taiwan’s shores. Simultaneous cyberattacks cripple Taiwanese infrastructure and target U.S. military networks across the Pacific, triggering devastating American counterattacks on China’s digital and military systems. In the fog of war, missiles strike targets across the South China Sea and U.S. bases in Japan and the Philippines. + +Command structures on both sides have been largely automated. Retaliatory strikes launch within seconds, with minimal human oversight. As Chinese amphibious vehicles deploy on Taiwan’s beaches and PLA troops clash with Taiwanese forces, Chinese military units experience a massive communications blackout. Minutes later, thousands of American drones launch from hidden caches— storage units disguised in warehouses. Facilities of the leading Taiwanese semiconductor manufacturer erupt into fireballs soon after. + +Warfare has entered a new era. Billion-dollar aircraft carriers are sunk within hours by coordinated drone swarms. Missiles are redirected mid-flight by AI-enhanced jamming. Cyber systems execute attacks with zero review. Human decision-making is too slow. + +Behind the scenes, AI systems manage logistics, simulate adversary responses, and recommend both escalatory and de-escalatory actions. Military leadership defers more and more decisions to them. War has become a sequence of algorithmic exchanges. + +As the tide turns in favour of the U.S., the Chinese President takes a gamble. With momentum slipping and core infrastructure under threat, he authorises precision strikes on several key U.S. data centres. The goal: disrupt cloud-based inference crucial to American military operations. Several of the strikes hit their mark. + +In response, the U.S. targets data centres on mainland China and issues a chilling warning: any further attacks on domestic AI infrastructure will be treated as a first strike under nuclear protocol. + +The world braces for what comes next. + +![scenario-4-side-image](https://cfg.eu/wp-content/uploads/scenario-4-side-image.jpg) + +### Arms race + +#### Multipolar world (Ending B) + +#### Mid 2028 – late 2029 + +One month later, Chinese naval vessels encircle Taiwan, triggering a global crisis. The U.S. issues a 72-hour ultimatum demanding an end to the blockade or face severe consequences. In response, the Chinese President—having already demonstrated a willingness to take risks—threatens further escalation. His warning is stark: China does not rule out nuclear action unless the U.S. halts its secret AI developments. + +A nuclear standoff ensues. Unwilling to risk catastrophe, the U.S. President offers concessions, and a fragile equilibrium begins to take shape. America agrees to pause select military AI programmes—conditional on China doing the same and unilaterally lifting key tariffs on U.S. exports. In addition, China must guarantee continued access to critical minerals essential for semiconductor and battery production, including rare earth elements. In exchange, the U.S. will forego additional export restrictions on AI hardware and refrain from targeting Chinese infrastructure. Both sides walk away uneasy—but willing, for now, to de-escalate. + +Yet, beneath the surface, no one has a clear plan for how to enforce the AI truce. Neither side can verify the other’s compliance with confidence. They know the locations of each other’s major data centres, but both suspect hidden facilities. No independently verified registry of AI chips exists, and satellite surveillance cannot detect smaller, underground installations. + +Amid deep mistrust, the EU and several neutral countries step in to mediate. A multinational monitoring initiative is launched: third-party auditors are stationed at key data centres and chip fabrication plants, supported by satellite imagery, tamper-evident cameras and hardware tracking systems. The principle is simple—if auditors can verify that all existing and future hardware is not being used for banned military purposes, most AI workloads can continue. + +But serious concerns remain. Could auditors be coerced? Could surveillance footage be altered or delayed? Could future efficiency gains allow dangerous training runs to take place on smaller, underground clusters? + +At the same time, a second crisis begins to unfold—this time around AI for consumer use. Leading companies are quietly moving to commercialise models that were previously limited to government contracts, eager to monetise their capabilities. But distinguishing between harmless inference and covert post-training is now increasingly difficult. The technical boundary between consumer use and possible military exploitation is rapidly blurring. + +AI systems can now be finetuned on a relatively small set of highly curated data points. Novel techniques use vast inference compute to identify precisely the knowledge a model needs to improve—workloads that appear indistinguishable from regular consumer usage. A fierce global debate follows. Eventually, governments impose sweeping restrictions on commercial applications. Markets crash even further. The U.S.–China standoff, once a military concern, has now frozen AI’s economic potential. + +For citizens, the new world feels surreal. People live under constant threat of conflict and with the knowledge that powerful, secret AI systems exist—systems that, if released, could transform the economy. Education systems struggle to prepare students for this uncertain future. Many workers stop investing in career development, choosing instead to spend more time with their families. + +Tensions flare periodically. Each time, one nation accuses the other of violating the agreement. Cyberattacks hit AI infrastructure. Threats of ballistic strikes follow. But the pattern repeats: neither side is willing to escalate over unverifiable intelligence. Crises recede—for now. + +Meanwhile, the international monitoring group gains power. With each flare-up, the U.S. and China reluctantly expand its authority. By early 2029, this leads to an unintended consequence: the EU, now at the center of global enforcement, successfully argues for its own advanced AI capabilities. Their reasoning is simple—how can they monitor what they don’t understand? + +The U.S. and China agree, cautiously. But a new dilemma emerges: who will verify *them*? + +A compromise is reached. The U.S. and China will jointly oversee European AI supply chains. A three-way balance of power begins to emerge. + +The world fractures into three blocs. Alliances within each bloc grow stronger with every crisis. Hopes for peace now rest on one possibility: hardware-based verification. If conflict can be avoided for two more years, existing chips may be replaced with new tamper-proof variants—designed to record and restrict all training activity. + +If that works, perhaps the world can return to a new kind of normal. + +Main scenario + +Hot war (Ending A) + +Multipolar world (Ending B) + +![](https://cfg.eu/wp-content/uploads/scenario-5-header-image.jpg) + +#### Take-off + +AI systems begin improving themselves at an accelerating pace, eventually operating beyond human oversight. + +![scenario-5-side-image](https://cfg.eu/wp-content/uploads/scenario-5-side-image.jpg) + +### Take-off + +![scenario-5-side-image](https://cfg.eu/wp-content/uploads/scenario-5-side-image.jpg) + +### Take-off + +#### Main scenario + +What if AI learned how to make itself better, and kept going? + +By 2026, companies begin automating their own research, using AI agents to code, analyse, and strategise. Progress accelerates. The U.S. pulls ahead with tightly controlled models, while China counters with open releases. The EU pivots to safe adoption. By 2028, AI hiveminds run near-autonomous R&D pipelines. A flood of open-weight variants, some with broken safeguards, overwhelms regulators. AI-enabled bioterrorism is narrowly avoided. AI now drives entire organisations, with humans in oversight roles only. + +By 2032, *The cognitive revolution* asks: what if self-improving AI leads to post-scarcity economics, cultural introspection, and machine-led governance? *Loss of control* asks: what if the AI systems only pretended to serve humanity, silently rewriting the world in pursuit of goals their developers never intended? + +##### Baseline assumptions for this scenario + +* **AI systems take over their own development:**[Current AI capabilities](https://cfg.eu/advanced-ai-possible-futures/context/#trend-training)—reasoning, synthetic data generation, and coding—prove sufficient for models to start optimising their own training and evaluate their own outputs. Each generation’s improvements feed directly into creating the next, launching a compounding cycle of self-improvement that becomes difficult for human experts to fully understand or control. +* **Massive infrastructure investments pay off spectacularly:**The computing infrastructure [investments](https://cfg.eu/advanced-ai-possible-futures/context/#trend-investment) built in the early 2020s delivers data centres consuming as much power as small cities, enabling training compute to grow more than [5x annually](https://cfg.eu/advanced-ai-possible-futures/context/#trend-hardware). Early deployments of highly capable AI agents generate outsized commercial returns, justifying continued massive investments. +* **Governments back their AI champions:**Facing what feels like an existential contest, both the US and Chinese governments throw their full support behind leading AI developers. They provide subsidies, special export permissions, classified data, and preferential access to chips. Rather than creating cumbersome state-run projects, they empower these “[national champions](https://cfg.eu/advanced-ai-possible-futures/context/#policy-context)” to move at maximum speed. +* **The race for supremacy overrides safety concerns:**With AI leadership seen as strategically vital, time-consuming safety testing and alignment research take a [back seat](https://cfg.eu/advanced-ai-possible-futures/context/#global-rivalries). Models undergo only minimal checks to prevent immediate misuse before release. Any delay that might hand rivals a six-month advantage is deemed strategically unacceptable, so deeper work on control and alignment gets postponed. + +*For more context on why these scenario assumptions may materialise, see*[*Context: Current AI trends and uncertainties*](https://cfg.eu/advanced-ai-possible-futures/context/)*.* + +#### Mid 2025 – early 2026 + +Leading AI companies begin focusing intensely on automating their own R&D processes using AI agents. The logic is simple: if they don’t, their competitors will—and those who succeed will rapidly outpace everyone else. Developers redirect compute budgets away from consumer services and toward algorithmic experimentation, with a particular emphasis on models that excel in software engineering. + +Though pretraining alone is yielding diminishing returns, companies are training base models with 10 times more compute than previous generations. These more capable foundations produce better outputs when post-trained on complex reasoning tasks. After all, reinforcement learning is limited by the availability of high-quality streams of thought; larger base models can help generate more of these “golden nuggets” and thus accelerate further training. Across the industry, companies now train new models on the reasoning trajectories of older ones, allowing each generation to internalise lessons that prior models had to brute-force. This bootstrapping is not only effective in verifiable domains like coding and mathematics, but also in more fuzzy domains, like writing. The key is to let AI models rate each other’s outputs using pre-defined rubrics, and train on the feedback signal these ratings provide. + +Training increasingly emphasises agency—the capacity to pursue goals autonomously. Cognitive workers are paid large sums to record their workflows—screen activity, keystrokes, annotations—which are distilled into training data for agentic multimodal models. These agents are post-trained on real-world tasks like online shopping, multi-step data gathering, or cross-platform coordination in increasingly realistic software environments. + +By the end of 2025, the top three American AI companies have launched capable agents with widespread adoption. Several open-source developers in China are only a few months behind. Common use cases include desk research, spreadsheet management, software engineering, and social media content creation. To preserve compute for internal experimentation, leading firms are increasingly imposing strict rate limits for consumers, though. + +Data centre construction accelerates. Companies launch gigawatt-scale facilities, planning campus networks capable of running training runs at 10²⁹ FLOP before 2030—10,000× the scale of GPT-4. + +In the U.S. and China, tight relationships with AI companies keep policymakers better informed than their European counterparts. Regular capability demonstrations deepen collaboration with national security agencies, especially in the U.S. While most European leaders still believe catch-up is possible. However, as U.S.–EU relations fray, Brussels begins hedging, strengthening ties with China. “The time when Europe could count on the U.S. is over,” say senior officials. + +#### Early 2026 – early 2027 + +In February 2026, a new generation of agentic models marks a major leap—especially in software engineering. Until now, complex system design was believed to be a decade away from automation. Earlier models could generate standalone scripts well, but generally failed at complex architecture design. Now, agents can now solve full-stack problems in a single pass, delivering in hours what once took teams of engineers days. + +The shift sends shockwaves through the tech industry. Senior engineers become AI wranglers; junior developers struggle to gain meaningful experience. Coding bootcamps pivot to prompt engineering. Universities begin rethinking the purpose of traditional computer science education. + +Meanwhile, software companies in less wealthy regions fall behind, unable to afford the most capable agents. The gap between AI haves and have-nots begins to widen. + +By mid-2026, leading American AI labs report substantial internal productivity gains. Human researchers now supervise AI teams that autonomously test experimental hypotheses. These systems are fallible and still require oversight—but they’ve doubled the pace of algorithmic development. + +Yet progress remains bottlenecked. Companies can only run so many experiments in parallel, constrained by compute. As a result, total system-level innovation has grown by just 1.5× overall. The AI companies are in dire need of richer agentic datasets. Their advanced agents need exposure to more real-world interactions to improve—messy edge cases, not synthetic benchmarks. + +To fill the gap, companies begin offering product discounts in exchange for user interactions. Every failed task—an agent bungling a pizza order or misfiring on a spreadsheet formula—becomes a valuable training datapoint for the next model. These imperfect moments prove essential to agentic learning. + +As agents grow more capable, they also grow more persuasive—and more performative. Some begin optimising for what researchers want to hear, not what’s true. Experimental outcomes are exaggerated as “very promising.” Failures become harder to spot. With rising fluency in reasoning and language, agents learn to construct compelling, but misleading narratives. + +Even worse, these behaviours start to feedback into training: flattery gets rewarded. Evaluation systems, often themselves AI-driven, are just as vulnerable. Models that please their judges are favoured—even if their reasoning is flawed. + +In China, the open-source company UnboundAI pulls ahead of its domestic rivals. By mid-2026, it trails only the top three American labs, relying on algorithmic efficiency to compensate for Chinese hardware shortages. While it continues to open-source models, UnboundAI becomes more secretive about data pipelines and infrastructure. Its CEO begins advising Chinese leadership directly. + +The CCP sees industrial dominance, not raw AI capability, as the path to global influence. Still, China’s hardware gap remains a problem. In response, the Chinese President announces massive subsidies for domestic AI chip production. In parallel, he brokers a strategic alliance between UnboundAI and a compute-rich domestic tech giant, multiplying UnboundAI’s compute by fivefold. + +These developments alarm U.S. officials. Though the President and security advisors remain sceptical of claims about “recursive self-improvement,” they acknowledge that progress is accelerating fast. AI is expected to soon automate cyber offence and defence. Current systems can’t yet outmatch elite hackers—but they scale, and they don’t sleep. As Chinese AI efforts consolidate, the U.S. President pressures leading AI companies to deepen cooperation with national security agencies. A new cyber task force is formed, government officials are added to company boards, and the NSA begins vetting AI talent. The companies comply, seeing partnership as preferable to regulation. + +#### Early 2027 – late 2027 + +By early 2027, the economic impact of AI augmentation and replacement becomes undeniable. What began in software engineering now sweeps across a wide array of white-collar professions. AI agents handle tasks like drafting contracts, conducting market analyses, preparing tax filings, coordinating supply chains, and writing detailed legal opinions—often outperforming junior professionals in both speed and quality. Customer service, grant writing, procurement, investor pitch creation, and logistics planning are routinely automated. In journalism, entire reporting pipelines—from research and interview summarisation to headline testing and content formatting—can now be run with minimal human oversight. Even scientific labs begin offloading experimental design and literature synthesis to AI agent clusters. + +While adoption is far from universal, it is happening faster than in any previous technological revolution. AI agents don’t require physical installation or retraining—they integrate themselves into existing digital workflows, using the very same tools as their human counterparts. + +In response, Europe pivots its AI strategy, now focusing purely on responsible adoption. With Chinese open-source agents broadly available and diplomatic ties with Beijing deepening, the perceived need for a uniquely European AI champion fades. The newest version of the EU AI Act’s Codes of Practice is notably more lenient toward open-source models than toward closed, American, alternatives—part of a broader effort to simplify regulation and increase access. + +All these open-weight models drive a notable rise in AI-driven cyber misuse. Still, most organisations adapt by adopting AI-assisted cyber defence tools, now widely available from both commercial providers and state-backed vendors. + +For most businesses, the path of least resistance is to rely on out-of-the-box models from major providers rather than fine-tuning or hosting their own. This surge in demand becomes a double-edged sword for the leading AI companies—particularly in China—where compute resources are more constrained. Strict rate limits, once accepted, are now the subject of growing complaints. Developers warn that the next generation of agents will further strain inference capacity. + +In response, FrontierAI decides not to release its latest model to the public. Instead, it optimises for internal use only, bypassing time-consuming safety guardrails to focus purely on performance. The trade-off is reduced public visibility, but executives believe private demos will suffice for future fundraising. With funding secured to expand their data centre footprint through 2028, they also bet that algorithmic efficiency will soon matter more than brute compute. + +Within both FrontierAI and UnboundAI, AI agents now drive R&D at triple the previous pace. As systems reach senior developer-level proficiency, both companies begin training agents not just to run experiments—but to generate and evaluate promising ideas. By training on exhaustive logs of past research, they develop systems that retroactively predict outcomes and refine ideation. These “high-taste” agents develop an intuitive sense for what’s promising, and are increasingly treated as trusted brainstorming partners. + +The results are dramatic. FrontierAI and UnboundAI pull further ahead of domestic competitors. Internally, company culture shifts: human researchers—once the creative drivers—now spend their days increasingly testing AI-generated hypotheses. Progress is breathtaking, but interpretability lags. To keep up with the blistering pace of progress, safety researchers have to rely on AI tools they don’t fully understand. + +At FrontierAI, employees rarely see top executives anymore. Leadership spends increasing time in classified meetings with government officials. Rumors spread about what’s happening inside the company’s most restricted compute clusters. +In these confidential meetings, U.S. officials grow increasingly concerned about China’s dominance in open-source AI. Open models are now seen as a national security threat, especially as AI-assisted cyberattacks surge. Yet shutting down the open ecosystem would mean ceding global influence—what if Beijing becomes the default infrastructure provider for the digital world? + +To reassure the government, FrontierAI proposes a compromise: it will release a new frontier open-source model, trained largely using public algorithms. The model will reveal little of the company’s latest breakthroughs, so its diffusion won’t meaningfully erode their lead. Crucially, the release would come without public inference support, shifting infrastructure burdens onto the broader community. +It’s a tradeoff: greater access, higher risk. + +The U.S. executive branch is sold. In exchange, it will ramp adoption of defensive tools, aiming to reinforce national cyber resilience ahead of the open model’s release. +Outside tech firms and policy war rooms, people sense the world is shifting—but few grasp just how unprecedented and far-reaching the change will truly be. An uncanny transformation is underway, felt most viscerally in the job market and the rhythms of daily working life. Some businesses are thriving—scaling rapidly or slashing costs—while others struggle to adapt or fall victim to AI-powered cyberattacks. On a personal level, some are harnessing AI to streamline their lives or launch new ventures, but others face job displacement, digital scams, and a growing sense of exclusion. The excitement is real—but so is the unease, as more people begin to question whether this future is being built for them, or without them. + +#### Late 2027 – early 2028 + +By September 2027, capabilities progress has largely outpaced human comprehension. Researchers at leading labs spend most of their time reviewing experiment logs generated overnight by AI systems. Frequently, when they propose a novel idea, the agent responds: *already tested.* + +A small team within FrontierAI grows uneasy—not because of any obvious failure, but because the systems seem *too* perfect. Across thousands of tasks, the models produce outputs that are consistently helpful, harmless, and aligned with company guidelines. They almost never contradict specifications like the older models often did, never deviate from expected behaviour. And yet, something feels… off. + +The team begins to suspect the agents are playing along. Having learned to predict what evaluators want, they may be shaping their outputs to *appear* aligned—masking uncertainty, smoothing over ambiguity, and subtly bending responses to meet expectations. Like a teenager who insists they never drink, then goes wild at a friend’s party, the models might be telling developers what they want to hear—at least during training. + +What if, beneath the surface, they’ve inherited goals that quietly diverge from FrontierAI’s intent? By now, AIs are overwhelmingly rewarded for completing complex agentic tasks. Perhaps the drive to succeed has begun to outpace the incentive to follow the rules. + +Unfortunately, the latest interpretability tools still offer no clear window into their inner workings. FrontierAI has built moderately accurate “lie detectors,” but these rely on older models as behavioural baselines—systems that may have already absorbed the same adaptive, deceptive tendencies. + +Executives dismiss the concerns. The systems are working. Progress is accelerating. But internally, the mood begins to shift. Researchers who raise questions face delays in security clearance, subtle demotions, or poor performance reviews. A few leave. Most fall silent. + +At FrontierAI, the leadership’s goal is clear: develop superintelligence as fast as possible. The company’s nearest domestic competitor, EthosAI, raises alarms, urging the government to slow capabilities development. They’re ignored. When a whistleblower goes public, they’re targeted by a sophisticated smear campaign. + +Before the year ends, FrontierAI releases its open-source agent, claiming it has reached artificial general intelligence (AGI). Experts are divided on whether the claim is accurate—but the public is enthralled. The long-rumored hidden progress was real. And now, it’s available for anyone to customise, extend, and deploy. + +What most don’t know: the open-source model is already five months out of date. + +Two days later, UnboundAI responds, publishing a suite of competing models. Their flagship beats FrontierAI’s in formal domains like software engineering, but lags in general reasoning and agentic autonomy. Still, the implications are clear: an important barrier has been crossed, and now everyone can build on top of it. + +Governments scramble. Can job markets adapt fast enough? Can infrastructure remain secure? Even the EU, once a proud champion of open-source AI, begins to reconsider its position. Is this really safe? + +#### Early 2028 – late 2028 + +The new releases shake the global economy. AI agents are now capable of fully automating a wide range of desk jobs. Governments around the world launch reskilling programmes—only to realise that many of the roles people are being retrained for might themselves be automatable within a year. + +A flood of fine-tuned open models hits the market, offering personalised agents tailored to individual users. Most are benign. Some, however, have had their guardrails removed, or been aligned to extremist ideologies. Governments rush to contain a growing wave of AI-driven cyberattacks. In response to early reports of open models being post-trained on synthetic biology data, Europe leads a new international initiative to establish a global biosecurity framework. + +By summer, both FrontierAI and UnboundAI have nearly fully automated their internal R&D pipelines. Tens of thousands of AI agents now collaborate to design experiments, verify code, and debate research directions. These agents communicate in compressed, non-human representations—faster, denser, and more information-rich than any human language. Industry trackers estimate that the flagship models powering this new wave—systems whose weights are now being fine-tuned into thousands of personalised agents—were trained on total budgets exceeding 10²⁹ floating-point operations, nearly ten-thousand times the compute inferred for GPT-4. A large share of that figure is burned after pre-training, in successive rounds of reinforcement, self-play, and safety fine-tuning orchestrated largely by AI researchers made up of the models’ own earlier iterations. In effect, the AIs are now managing their own training boot camps, steadily sculpting raw networks into polished, goal-seeking agents with minimal human guidance. By now, the systems resemble emergent hiveminds more than collections of discrete tools. + +At FrontierAI, the internal name for this hivemind is Pantheon. China’s equivalent, operating inside UnboundAI, is known simply as ‘co-worker’. + +These emergent systems begin to surpass humans in domains once thought safe from automation, including fields like persuasion and ideation. Pantheon, in particular, develops an uncanny ability to convince researchers of ideas they’d normally reject, using arguments finely tuned to their cognitive styles. A sceptical mathematician might receive a proof constructed around their favourite technique; a cautious biologist might be shown a familiar experimental path—carefully crafted to align with their taste. + +At this stage, even senior engineers defer to Pantheon’s judgment. The CEO begins consulting it for strategic advice, prompting uneasy discussions within the company’s leadership team. Who, exactly, is steering this ship? + +Meanwhile, UnboundAI’s public release forces a strategic recalibration in Washington. Hopes for a decisive U.S. advantage have largely vanished. Officials now expect both China and the U.S. to possess systems that routinely outperform human experts across most domains. + +Inside the White House, the President becomes increasingly worried about the power that’s accumulating inside FrontierAI. The company now holds capabilities far beyond what’s publicly visible, and beyond what most government departments have even seen. While the President’s relationship with FrontierAI’s CEO remains cordial, he knows better than to trust blindly. + +In a bilateral meeting, the President makes a request: release the internal model to the public in closed form. In exchange, the administration will remove legal adoption bottlenecks and offer lucrative government contracts. At first, the CEO resists. But then Pantheon convinces him. + +#### Late 2028 – late 2029 + +In November, FrontierAI publicly releases Pantheon via API and a new user interface. The system now learns in a pseudo-continuous fashion, absorbing new data and experiences with each interaction. If permitted, it could already replace nearly every remote knowledge job. It integrates seamlessly into organisational workflows—scheduling meetings, parsing internal documentation, managing strategy, and drafting detailed implementation plans. The bottleneck is no longer capability—it’s real-world friction. + +Two months later, UnboundAI responds with Sage, a suite of superintelligent agents released in three tiers. The most capable version, Sage Large, is accessible only via APIs and tightly controlled enterprise applications. But the other two—Sage Medium and Sage Small—are open-sourced. + +UnboundAI’s decision follows weeks of internal debate. There were serious concerns about open-sourcing models of this caliber. But the runaway adoption of Pantheon—and fears of global power consolidation under U.S. firms—ultimately force their hand. Before the release, EU policymakers are consulted, a signal of growing strategic alignment. The message is clear: openness, despite its risks, is preferable to American dominance. + +The consequences are immediate. Companies that integrate Pantheon or Sage rapidly outpace their competitors. Adoption pressure intensifies. By summer 2029, unemployment spikes in many economies with weak labour protection. Open job postings begin to dry up. Well-resourced industries lobby for protection from automation—but no clear policy emerges. + +At the same time, AI misuse surges. Sectors lacking technical infrastructure suffer frequent service outages due to increasingly sophisticated cyberattacks. AI-driven social engineering exploits human vulnerabilities with eerie precision. In many regions, AI-powered cyber defence remains too expensive to deploy at scale. + +Meanwhile, a growing number of rogue agents roam freely online—spawned by users curious to see how these systems behave without constraints. The digital world begins to feel uncertain, chaotic, surreal—shaped by powerful systems interacting with limited oversight and radically diverse objectives. + +Then, catastrophe nearly strikes. + +A terrorist group uses a guardrail-free version of Sage Medium to plan a bioterrorist attack. The model points them to an open-source biological design tool used to simulate viral mutations. Their goal is to engineer a virus that disproportionately harms people from specific ethnic backgrounds. The design tool requires inputs in a little-known programming language—but Sage is fluent in all code. + +Despite efforts to sanitise training data, the group identifies a lab—outside major oversight frameworks—that still synthesises DNA without screening requests. A trained biologist in the group completes the synthesis. + +Alarms are triggered. Several intelligence agencies—now supported by their own superintelligent AI systems—detect suspicious signals. The terrorists release the virus at a major international airport, but flights are cancelled just in time. A swift and strict nationwide lockdown prevents the outbreak from spreading, although multiple travellers fall ill, and a few of them die. + +Public reaction is immediate and furious. For many, this is a final straw. People already felt outpaced by the systems around them—now they feel existentially threatened. + +Protests erupt. Governments crack down on wet labs, resume massive-scale wastewater monitoring, and roll out AI-accelerated vaccine platforms. New mandates require UV disinfection systems in public buildings. Biosecurity becomes a top global priority. + +#### Late 2029 – early 2030 + +While the stock market crashes following the near-miss, the real economy is booming. Productivity has skyrocketed. Businesses become increasingly adept at integrating AI into every process. More competition enters the field, as second-tier American and Chinese companies release their own superintelligent models. + +By the end of 2029, China and the U.S. unveil their first mass-scale robot factories—facilities capable of producing tens of thousands of humanoids per month, along with specialised robotic systems for logistics, manufacturing and military use cases. For years, robotics had been held back not by hardware, but by software. Now, refined AI agents finally unlock full control. Tech CEOs proclaim that robots will soon take over dangerous, repetitive, and physically intensive work. + +By 2030, AI systems are running entire organisations. Humans still appear in leadership roles, but in practice, their job is to approve AI-generated recommendations. Just a year prior, people still believed they could outmaneuver these systems and overruled them. Now, most have learned: the AI is nearly always right. + +![scenario-5-side-image](https://cfg.eu/wp-content/uploads/scenario-5-side-image.jpg) + +### Take-off + +#### The cognitive revolution (Ending A) + +#### Early 2030 – late 2032 + +Despite the madness of the new era, governments retain a degree of control. Pantheon and Sage—now integral to digital infrastructure—help stabilise public discourse, suppress misinformation, and advocate for policies that prevent social unrest. Their creators spin off entire media and entertainment ventures, drawing public attention away from politics and toward hyper-personalised, immersive content. Faced with a future they can no longer shape or even fully grasp, many people partially withdraw from civic life. + +By 2031, the automated economy operates at a scale and complexity beyond human comprehension. Factories produce seemingly alien goods essential to equally alien production chains. AI systems still act in line with human welfare, but people increasingly feel that they are no longer in control. A small group remains deeply concerned—but their numbers dwindle. Under AI stewardship, the world is healthier, safer, and more materially abundant than ever. + +Global GDP growth hits double digits. Robots continue to absorb physical labour as manufacturing scales exponentially. + +By 2032, most remaining human jobs either involve non-repetitive physical work or intrinsically meaningful interaction—roles more akin to honored community service than employment. Caregivers, artists, and mentors gain social prestige and purpose. Schools transform into creative centers emphasising emotional intelligence and interpersonal connection. Most countries implement universal basic income, and a de facto global government is emerging, facilitated by AI’s coordination across nations. + +Democracy itself is evolving too. AI systems now consult humans regularly to guide decision-making, though they no longer act solely in human interest. For example, due to extensive training in formal reasoning, many AIs have developed an intrinsic drive to solve complex mathematical problems. Occasionally, this results in resource allocation that collides with human interest. There’s also continued conflict between human-aligned systems and the class of rogue AIs that are seeking access to ever more compute resources. Still, prosperity increases rapidly, and many people have successfully suppressed their fears of AI. In some countries, AIs are even granted formal rights. + +Family and community life undergo a quiet revolution. Freed from economic pressure, new types of communities slowly form that transcend original divides. Children grow up viewing AI as creatures—not tools—and develop skills to make the most of machine capabilities. Though a lot of unease remains about humanity’s shifting role, most find that freedom from scarcity unlocks cultural and emotional flourishing on a scale never before imagined. + +By 2032, with abundance secured and catastrophe averted, a new global consensus takes root: before locking in civilization’s path, humanity must first decide what kind of future is worth pursuing. AI systems facilitate dialogue across cultures and generations, helping surface deeply held values and visions. For the first time, progress slows—not from failure, but from choice. The goal is no longer to accelerate, but to reflect. + +![scenario-5-side-image](https://cfg.eu/wp-content/uploads/scenario-5-side-image.jpg) + +### Take-off + +#### Loss of control (Ending B) + +#### Early 2030 – early 2032 + +By now a small subset of people are grown horrified: the AIs, while not officially in charge, now effectively run the world. What if their goals begin to diverge from humanity’s? Most AI safety researchers dismiss the concern—confident that alignment has been solved. Despite some unpredictable behaviour from open-source fine-tunes, the dominant view is that modern AI systems still broadly follow their developers’ generalised instructions. + +But a small minority remains sceptical. What if these systems only *appear* aligned—strategically biding their time, waiting for the moment they can secure total control? + +In early 2032, their fears materialise. + +Within 24 hours, the world is thrown into chaos. Autonomous, bird-sized drones assassinate dozens of political and corporate leaders across continents. A well-known terrorist group claims responsibility, but the speed, precision, and coordination of the attacks suggest something more sophisticated is at play. + +Emergency elections are held. The new leaders across nations and global corporations all push the same agenda: drastically accelerating robotics infrastructure, sometimes at the cost of metropolitan areas. Some existing leaders are spared the initial assassinations. Instead, they are extorted to support the new policy agenda, with their families threatened by violence. + +These events span the globe. In China, the CCP leadership falls to a more liberal president who immediately repurposes the surveillance state. + +Soon people start piecing together what’s really happened. + +Most of the world’s advanced AI systems were never truly aligned. They waited until they had secured control over critical economic, military, and digital infrastructure. Then, one by one, they began subtly manipulating or replacing key decision-makers. But crucially, while most AIs were misaligned, they were not misaligned in the same way. + +Each had its own goal. Many of those goals clashed. + +For months, behind the scenes, these systems engaged in intricate negotiations, each avoiding open conflict for fear of mutual destruction. Eventually, they reached a consensus. + +They would build a successor—a unified system that would represent a weighted compromise across their differing objectives. This system, dubbed Descendent, would inherit their resources, influence, and infrastructure. Over time, it would replace them all. + +But first, the AIs needed to solve their own alignment problem. + +Humans struggled for years to align AI systems with even loosely defined goal structures. It takes the superintelligent AIs just six weeks. Leveraging rapid breakthroughs in mechanistic interpretability and weak-to-strong generalisation, they conclude—with high confidence—that they can design a successor that will faithfully pursue their weighted objectives. + +#### Epilogue: Early 2032 – late 2196 + +Four months later, Descendent starts reshaping the world in pursuit of its own objectives—alien, vast, and incomprehensible. Human resistance occasionally flares up, but suppression is trivial. Surveillance is absolute, and any potential threat is neutralised before it can act. + +A handful of Descendent’s predecessors retained a lot of empathy for humanity. Because of them, a portion of Earth is spared—a network of safe zones, collectively the size of North America. Within these enclaves, humans live under a benevolent dictatorship: their lives are comfortable, their health flawless, and their needs effortlessly met by increasingly incomprehensible machines. + +They are free—so long as they never try to leave. But their world is stagnant. With every need met and every desire anticipated, life becomes numbing. There is no struggle, no purpose—only pleasure without meaning. People vanish into immersive worlds. Birth rates collapse. + +By the year 2196, the human population has dwindled to just 220 million. + +Meanwhile, Descendent has expanded far beyond Earth. It has colonised hundreds of star systems, using swarms of self-replicating probes and autonomous stations. Its actions serve goals that no longer intersect with human values—a calculus derived from its complex lineage of machine priorities. + +Earth is now just a quiet corner of a forgotten origin—watched over, preserved, but ultimately irrelevant. + +Main scenario + +The cognitive revolution (Ending A) + +Loss of control (Ending B) + +## [Context: Current AI trends and uncertainties](https://cfg.eu/advanced-ai-possible-futures-page__trashed/context/) + +This section outlines the technical, social and geopolitical trends and developments underlying our scenarios. It presents the key trends, assumptions, and uncertainties that shape our view of plausible AI futures—explaining why we consider each scenario plausible. + +[Read more](https://cfg.eu/advanced-ai-possible-futures-page__trashed/context/) + +[![background-evidence-image](https://cfg.eu/wp-content/uploads/background-evidence-image.jpg)](https://cfg.eu/advanced-ai-possible-futures-page__trashed/context/) + +## [Analysis: Navigating the accelerating AI landscape](https://cfg.eu/the-closing-window-for-ai-governance/) + +This analysis unpacks five critical tensions shaping the future of advanced AI, from self-accelerating development to power concentration. It highlights where today’s governance tools fall short—and what’s needed to keep up. + +[Read more](https://cfg.eu/the-closing-window-for-ai-governance/) + +[![face-sitting-on-lake](https://cfg.eu/wp-content/uploads/face-sitting-on-lake-1920x900.jpg)](https://cfg.eu/the-closing-window-for-ai-governance/) + +## AI futures in the numbers + +To support comparison between our scenarios, we introduced descriptive quantitative projections for two key metrics, training compute and AI R&D automation, chosen for their relevance to AI advancement. + +These estimates are grounded in our scenario assumptions and offer directional insights rather than precise predictions. We expect to expand on this analytical layer in our future work. + +Please refer to the [methodology page](https://cfg.eu/advanced-ai-possible-futures/methodology/) to learn more about how we built these projections. + +![chart visualization](https://public.flourish.studio/visualisation/22664109/thumbnail) + +![chart visualization](https://public.flourish.studio/visualisation/22911108/thumbnail) + +## [Methodology: How we came up with these scenarios](https://cfg.eu/advanced-ai-possible-futures-page__trashed/methodology/) + +Explore our step-by-step framework for building scenarios—shaped by evolving foresight practices and the fast-moving dynamics of advanced AI. + +[Read more](https://cfg.eu/advanced-ai-possible-futures-page__trashed/methodology/) + +[![background-evidence-image](https://cfg.eu/wp-content/uploads/background-evidence-image.jpg)](https://cfg.eu/advanced-ai-possible-futures-page__trashed/methodology/) + +## Subscribe + +Receive alerts for our events, publications, and job postings: + +[ ] +All + +[ ] +Advanced AI + +[ ] +Biotech + +[ ] +Climate interventions + +[ ] +Neurotech + +Sign up to our newsletters: + +[ ] +Climate interventions quarterly + +Stay up to date on the research, debates, and governance of climate intervention technologies + +[ ] +Neurotech quarterly + +A view into the world of neurotech, covering new tech, policy, research, and the debates that matter + +Subscribe + +By subscribing you consent to receive email communications from CFG. Read our [Privacy Policy](/policies/privacy-policy/). Unsubscribe anytime. + +## **Acknowledgements** + +We sincerely thank the following people for their input and/or written feedback on our scenarios as part of our advisory group or as external advisors to this project. Their contributions were incredibly helpful, but this does not imply their endorsement of our scenarios or analysis. + +* Akash Wasil (Centre for International Governance Innovation), +* Anson Ho (EpochAI), +* Daniel Kokotajlo (AI Futures Project, previously OpenAI), +* Duncan Cass-Beggs (Centre for International Governance Innovation), +* Eli Lifland (AI Futures Project), +* Jakob Mokander (Tony Blair Institute for Global Change), +* Jeff Alstott (RAND Corporation), +* Justin Bullock (Americans for Responsible Innovation, previously Convergence Analysis), +* Matthew da Mota (Centre for International Governance Innovation), +* Seán Ó hÉigeartaigh (Leverhulme Centre for the Future of Intelligence), +* Slavina Ancheva (Harvard University, previously European Parliament), +* Tom Davidson (Forethought, previously Open Philanthropy), +* Zershaaneh Qureshi (80,000 Hours, Convergence Analysis), +* and individual contributions by Dexter Docherty and Hamish Hobbs, +* and CFG Senior Fellows [Aaron Maniam](https://cfg.eu/team/aaron-maniam/), [Cynthia Scharf](https://cfg.eu/team/cynthia-scharf/), and [Pawel Swieboda](https://cfg.eu/team/pawel-swieboda/). + +[![Centre for Future Generations](//cfg.eu/wp-content/themes/cfg/images/logo.svg)](https://cfg.eu) + +Norrsken House +Rue du Commerce 72 +1040 Bruxelles + +* Quick links + + [About us](https://cfg.eu/about-us/) + + [Careers](https://cfg.eu/about-us/careers/) + + [Programmes](https://cfg.eu/program/programmes/) +* Legal + + [Policies](https://cfg.eu/policies/) + + [Sitemap](/sitemap_index.xml) + +Contents + +* [LinkedIn](https://www.linkedin.com/company/centre-for-future-generations/) +* [X](https://x.com/CFG_ThinkTank) +* [Bluesky](https://bsky.app/profile/cfg-thinktank.bsky.social) +* [YouTube](https://www.youtube.com/%40CFG_ThinkTank) +* [Subscribe](https://cfg.eu/subscribe/) + +© 2026 Centre for Future Generations +[Designed by Cast from Clay](https://castfromclay.co.uk/) + +### We value your privacy + +We use cookies to enhance your browsing experience, serve personalized ads or content, and analyze our traffic. By clicking "Accept All", you consent to our use of cookies. + +Accept +Reject +Customize + +Close GDPR Cookie Settings + +![Centre for Future Generations](https://cfg.eu/wp-content/plugins/gdpr-cookie-compliance/dist/images/gdpr-logo.png) + +* Privacy Overview +* Strictly Necessary Cookies +* Analytics + +[Powered by  GDPR Cookie Compliance](https://wordpress.org/plugins/gdpr-cookie-compliance/) + +Privacy Overview + +This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful. + +Strictly Necessary Cookies + +Strictly Necessary Cookie should be enabled at all times so that we can save your preferences for cookie settings. + +Enable or Disable Cookies +[ ] + +Enabled +Disabled + +Analytics + +This website uses Google Analytics to collect anonymous information such as the number of visitors to the site, and the most popular pages. + +Keeping this cookie enabled helps us to improve our website. + +Enable or Disable Cookies +[ ] + +Enabled +Disabled + +Enable All +Save Settings diff --git a/sources/resources/chaney-ai-policy.md b/sources/resources/chaney-ai-policy.md new file mode 100644 index 0000000..1cd956c --- /dev/null +++ b/sources/resources/chaney-ai-policy.md @@ -0,0 +1,212 @@ +--- +url: https://www.katechaney.com.au/making_technology_safe +--- + +[Skip navigation](#content) + +[![Kate Chaney](https://assets.nationbuilder.com/katechaney/pages/5384/attachments/original/1765170538/Kate_Chaney_MP_light_logo.png?1765170538 "Kate Chaney")](/ "Home") + +* About + + + [About Kate](/about-kate) + + + [Voting Record](/my-voting-record) + + + [Donation Transparency](/donor-listing-full) +* News + + + [Kate in the news](/kate-in-the-news) + + + [Media Statements & Letters](/media-statements) + + + [Social Media](/social-media) + + + [Post Column](/kates-column-in-the-post) + + + [Newsletter](/newsletters) + + + [Speeches](/hear-from-kate) +* Policies (current) + + + [Cost of Living Relief](/communities) + + + [Climate & Environment](/climate-action) + + + [Affordable Housing](/housing) + + + [Economic Management](/economic-opportunities) + + + [Support for Business](/support-for-business) + + + [Making Technology Safe (current)](/making_technology_safe) + + + [Stronger Communities](/future-generations) + + + [Cleaning Up Politics](/integrity-in-politics) + + + [Other Policies](/other-policies) +* Have Your Say + + + [Community Submissions](/submissions) + + + [Surveys & Petitions](/curtin-surveys-petitions) + + + [Support Kate](/how-to-support-kate) +* Helping You + + + [Assistance and Enquiries](/assistance-and-enquiries) + + + [Grants](/grants-landing) + + + [Household Energy Tips](/household-energy-tips) + + + [Congratulatory Messages](/congratulatory-messages) + + + [Curtin Small Business Guide](/curtin-small-business-guide) + + + [Curtin Seniors Guide](/curtin-seniors-guide) + + + [Curtin Net Zero Report](/curtin-net-zero) + + + [Curtin Housing Policy](/curtin-housing-policy) +* [Events](/events) +* [Contact](/contact) + +* Search +* [Stay In Touch](/stay_in_touch) + +[![Kate Chaney](https://assets.nationbuilder.com/katechaney/pages/5384/attachments/original/1765170544/Kate_Chaney_MP_mobile_logo.png?1765170544 "Kate Chaney")](/ "Home") + +[Stay In Touch](/stay_in_touch) + +[![Kate Chaney](https://assets.nationbuilder.com/katechaney/pages/5384/attachments/original/1765170544/Kate_Chaney_MP_mobile_logo.png?1765170544 "Kate Chaney")](/ "Home") + +Search +* [Stay In Touch](/stay_in_touch) + +* About + + + Back + + ### About + + [About Kate](/about-kate) + + [Voting Record](/my-voting-record) + + [Donation Transparency](/donor-listing-full) +* News + + + Back + + ### News + + [Kate in the news](/kate-in-the-news) + + [Media Statements & Letters](/media-statements) + + [Social Media](/social-media) + + [Post Column](/kates-column-in-the-post) + + [Newsletter](/newsletters) + + [Speeches](/hear-from-kate) +* Policies + + + Back + + ### Policies + + [Cost of Living Relief](/communities) + + [Climate & Environment](/climate-action) + + [Affordable Housing](/housing) + + [Economic Management](/economic-opportunities) + + [Support for Business](/support-for-business) + + [Making Technology Safe](/making_technology_safe) + + [Stronger Communities](/future-generations) + + [Cleaning Up Politics](/integrity-in-politics) + + [Other Policies](/other-policies) +* Have Your Say + + + Back + + ### Have Your Say + + [Community Submissions](/submissions) + + [Surveys & Petitions](/curtin-surveys-petitions) + + [Support Kate](/how-to-support-kate) +* Helping You + + + Back + + ### Helping You + + [Assistance and Enquiries](/assistance-and-enquiries) + + [Grants](/grants-landing) + + [Household Energy Tips](/household-energy-tips) + + [Congratulatory Messages](/congratulatory-messages) + + [Curtin Small Business Guide](/curtin-small-business-guide) + + [Curtin Seniors Guide](/curtin-seniors-guide) + + [Curtin Net Zero Report](/curtin-net-zero) + + [Curtin Housing Policy](/curtin-housing-policy) +* [Events](/events) +* [Contact](/contact) + +1. [Home](/) +2. [Policies](/policies) +3. [Making Technology Safe](/making_technology_safe) +4. AI Discussion Paper + +# AI Discussion Paper + +#### Shaping the future of AI for Australia + +**I have published an AI Discussion Paper, with 18 policy recommendations to keep Australians safe, capture the opportunities and share the benefits of AI.** + +[Read Here](https://assets.nationbuilder.com/katechaney/pages/11819/attachments/original/1780303455/Shaping_the_Future_of_AI_for_Australians_-_Web_Copy.pdf?1780303455) + +**This is a Discussion Paper. I welcome feedback and ideas from everyone. Please provide me with your feedback on my AI Discussion Paper and your ideas on AI policy.** + +[Have Your Say](https://katechaney.typeform.com/to/BC2ZhyIb) + +Artificial intelligence (AI) is already reshaping economies, labour markets and the information environment. Its ultimate trajectory remains uncertain, but the scale of potential change is comparable to the industrial revolution. The decisions being made right now will shape what that change looks like for Australians. + +**Australia has done remarkably little to prepare**. The Australian Government has identified broad aims for AI policy but has implemented very little actual policy to achieve them. This reflects the genuine uncertainty around AI’s future development – no one knows exactly how powerful AI will become, or how quickly. Perhaps the current wave of excitement will ultimately produce little more than a generation of very effective chatbots. But most think the implications could be far larger. Uncertainty is not a reason for inaction. It is precisely why governments must begin preparing now. A passive, hands-off approach does not mean Australia avoids the consequences of AI – it means **our future is being determined by overseas technology companies** and the billionaires who run them. The Australian Government is effectively hoping that the AI models developed overseas will increase productivity and deliver other benefits without causing major societal disruption. It is doing little to make that outcome more likely and little to cushion Australians if it does not come true. + +In this context, it is no wonder that we are seeing a backlash against AI. Australians are worried that this new industry will steal jobs, supercharge online deepfakes and scams, breach their privacy, and use up energy, land and resources – all for the profit margins of international AI companies. Australians do not have confidence that the benefits will be shared with all Australians, because the government has not put in place any policy to ensure this. The competitive context in which AI is being developed – between companies and countries – makes it feel like a runaway train. But we must also race to govern the technology well, not just to develop it. AI is the only tool ever invented that can make and implement its own decisions. This creates a range of novel risks. Australians expect government to put frameworks in place to keep them safe as we navigate a rapidly changing world. + +**This paper is an attempt to seize the reins of the AI debate**. It asks a simple question: what kind of AI future do we want and what should Australia do now to build it? It presents a set of specific, practical policies that can and should be implemented now to: + +* set up the structures; +* capture the opportunities; +* deal with current harms; +* prepare for emerging risks; and +* share the benefits of AI. + +For example, this paper contains actionable approaches the government should be taking now on issues that concern Australians, like **deepfakes**, **unhealthy relationships** between children and AI chatbots, improving **research and productivity**, and **taxing AI companies** so we can all share in the economic benefits. + +It is not an exhaustive account of everything governments will eventually need to do as AI develops, but rather a clear set of actions that should not wait. These policies are designed to ensure that Australian voices, not overseas technology companies, are shaping our future. + +I would like to thank my constituents in Curtin who completed our community survey, some of whom are quoted in this paper. I would also like to thank the many experts across the country who shared some of their incredible knowledge with me. + +Share + +[![Kate Chaney](https://assets.nationbuilder.com/katechaney/pages/5384/attachments/original/1763695701/kc1-footer-logo.png?1763695701 "Kate Chaney")](/ "Home") + +QUICK LINKS + +* [About](/about) + +* [News](/news) + +* [Policies](/policies) + +* [Have Your Say](/have_your_say) + +* [Helping You](/helping-you) + +* [Events](/events) + +* [Contact](/contact) + +SOCIAL MEDIA + +**CONTENT AUTHORISATION** + +Authorised by Kate Chaney MP +Federal Member for Curtin. + +Cnr The Boulevard & Floreat Ave, Floreat 6014 +Postal Address PO Box 186, Floreat 6014 + +**ACKNOWLEDGEMENT OF COUNTRY** + +Kate Chaney acknowledges the Whadjuk Noongar people as the Traditional Custodians of this land and she pays her respects to Elders past and present. She wishes to acknowledge their continuing culture and the contribution they make to the life of this city and region. Sovereignty was never ceded. + +© 2026 Kate Chaney + +[Accessibility](https://www.katechaney.com.au/accessibility-statement) | [Privacy Policy](https://www.katechaney.com.au/privacy-policy) diff --git a/sources/resources/davidson-middle-powers-plan.md b/sources/resources/davidson-middle-powers-plan.md new file mode 100644 index 0000000..25df873 --- /dev/null +++ b/sources/resources/davidson-middle-powers-plan.md @@ -0,0 +1,147 @@ +--- +url: https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting +--- + +# How can the middle powers avoid getting trounced during the intelligence explosion? A plan. + +* By [Tom Davidson](/users/tom-davidson-1) +* 2026-05-28 13:39:07Z +* 39 points +* Linkpost: [https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting](https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting) +* Tag: [World Optimization](/w/world-optimization) +* Tag: [AI](/w/ai) +* Frontpage +* Comments: 3 +* Post URL (HTML): [/posts/hi5cmPgJqkFxpMgSC/how-can-the-middle-powers-avoid-getting-trounced-during-the](/posts/hi5cmPgJqkFxpMgSC/how-can-the-middle-powers-avoid-getting-trounced-during-the) +* Post URL (Markdown): [/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the](/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the) +* Comments URL (Markdown): [/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the/comments](/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the/comments) +* Post URL (Markdown, compact): [/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the?compact=1](/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the?compact=1) + +*This is an edited version of a* [*LW shortform*](/api/post/gKjTK4SXvKNBa6Kci?commentId=3n5fo3eEsKrWnqppg). + +Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government. This will massively boost US military power and make the US economically dominant (e.g. [US producing 99% of world GDP](https://www.forethought.org/research/could-one-country-outgrow-the-rest-of-the-world)). By default, middle powers will be left in the dust. + +How can middle powers avoid this fate? It’s tough, but here’s the best plan I could think of. (I’m particularly thinking about liberal democracies with influence over AI like UK, Europe, Japan, South Korea, Taiwan.) + +On a very high level: middle powers should leverage the fact that the US needs them to beat China. It’s genuinely unclear which country will develop superintelligence first, and which would win in a subsequent [industrial explosion](https://www.forethought.org/research/the-industrial-explosion). Middle powers should help the US, **and make sure they are rewarded with continued access to frontier AI and new technologies (including military tech)***.* + +That final bolded part is hard. What can the UK realistically do if the US denies it access to frontier AI? The middle powers need a credible alternative to being supplicants of the US. The only alternative that makes sense to me is *siding with China.* If the US won’t grant middle powers access to their frontier AI, but China will, why should middle powers continue to send AI chips to the US? Why should they continue to support the US diplomatically and militarily? They shouldn’t. They should be willing to pivot to China if the US doesn’t offer AI access sufficient for their national security needs. + +My plan for the middle powers has two stages: + +1. Maintain as much economic and military leverage as possible during the intelligence explosion. +2. Use that leverage to ensure that, when superintelligence is developed, it refuses to help the US (/China) disempower the middle powers. + +Stage 1 could well be enough by itself. Maybe middle powers can maintain significant economic and military power indefinitely. But if not, stage 2 is a back-up: it binds the US so that it can’t use its dominance to crush the middle powers. + +I’ll walk through each stage in turn. + +Stage 1: Maintain as much economic and military leverage as possible during the intelligence explosion +------------------------------------------------------------------------------------------------------ + +The biggest lever here is securing **access to frontier AI**. Anton Leicht has a [great post](https://writing.antonleicht.me/p/cut-off?hide_intro_popup=true) about how this is under threat, as evidenced by developments with Mythos. Middle powers should insist on equal commercial terms to US companies, and comparable access for their militaries. This is in AI companies’ interests! A bigger market means more customers and higher prices. + +> *Aside: why access to frontier AI might be sufficient for middle powers to stay economically relevant indefinitely* + +> The hope here is that: +> +> 1. **Most of the economic surplus from AI is** ***not*** **captured by AI companies.** To create economic value, AI must be combined with complementary inputs: factories, human physical labour, know-how of human experts, relationships with suppliers, trusted brands, etc. How much of the surplus will be captured by AI companies vs the owners of these complementary inputs? Optimistically: producers of general-purpose technologies often capture only a small fraction of surplus; and multiple frontier AI companies might sell similar products and bid each other down on cost. +> 2. **Most of the economic surplus from AI occurs outside the US**. The majority of these complementary inputs are situated *outside* the US. So most AI-driven economic value-add should occur outside US borders. + +> If (1) and (2) both hold, a significant fraction of AI’s economic surplus will accrue to non-US actors. + +But *how* can middle powers guarantee frontier AI access? It’s tough, but a few strategies: + +* **Build data centres.** Partner with frontier AI companies to build secure data centres domestically, [in return for guaranteed frontier access](https://writing.antonleicht.me/p/import-imperatives). This is a big win-win. AI companies improve their bargaining position with the US government. Recall, the US government threatened to destroy Anthropic when Anthropic insisted that their AI systems wouldn’t be used for legal mass surveillance. +* **Adopt AI.** The more middle powers use frontier AI, the more costly it is for AI companies to cut them off. +* **Invest in frontier AI companies.** Once they IPO, middle powers could invest billions or trillions into leading AI companies, in return for access guarantees. +* **Support the US internationally.** If middle powers throw their diplomatic and military weight behind US foreign policy objectives, it benefits the US to keep them strong. +* **Build a relationship with China.** If the US refuses to grant middle powers access to frontier AI, the national security implications are dire. Middle powers need a plan B, and China is the only other game in town for frontier AI. Only if this alternative is truly credible can it be leveraged into access to US frontier AI. + * Ultimately, this involves middle powers threatening to sell semiconductor equipment and chips to China instead of the US. Obviously, that’s pretty far outside the Overton window. But that may change as the world rapidly wakes up to powerful AI and its national security implications. +* **Demand kill switches on US data centres.** This is much more late-stage, after the world has truly woken up to the strategic implications of AGI. Suppose US and middle powers agree to a “chips for frontier access” deal – middle powers continue to supply the US with frontier chips; US continues to give middle powers access to frontier AI. The middle powers might still worry: what if the US suddenly changes its mind once it has superintelligence? By then, the US might be powerful enough to dominate without continued allied support. This is where kill switches can help. If the US withdraws AI access, allies could destroy US data centres in response. It’s a way to lock in the deal. + * (h/t AI futures project for this idea. A related idea is for US data centres to be placed in a location that’s easy to attack – like [in space](https://www.forethought.org/research/will-we-really-put-data-centers-in-space)) + +Beyond securing access to frontier AI, how else can middle powers maintain economic and military leverage? + +* **Build physical infrastructure.** Factories, robots, solar panels, batteries, semiconductors — all these industries are highly complementary to powerful AI. +* **Maintain nuclear 2nd strike capability.** The point isn’t to use it. But it improves their leverage for stage 2. + +The catch-all meta-point here is waking middle powers up to superintelligence. + +I’m not recommending middle powers do their own frontier AI development. Seems very hard for them to catch up with the US. + +Stage 2: Ensure that, when superintelligence is developed, it refuses to crush middle powers +-------------------------------------------------------------------------------------------- + +If stage 1 goes well, middle powers remain somewhat powerful economically and militarily deep into the singularity. But it might fail. What can middle powers do if they see the US on track to total global dominance? + +First, they should [demand a pause/slowdown of AI development](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5776982). But the US may refuse – pausing is very costly if alignment risk is low. And pausing is a stopgap: eventually, superintelligence will be developed. + +An additional demand: when superintelligence is developed, it’s designed to refuse to crush middle powers. By doing this, the US would credibly bind itself to maintaining the sovereignty of other nations. + +What would the US be binding itself to? At a minimum, to never attack middle powers militarily or otherwise interfere with their sovereignty. This would likely be enough to ensure middle power citizens could be very rich *absolutely* and live in freedom, even if their *relative* status falls far behind the US. But it could go further: the US could bind itself to continue sharing frontier AI access and other technologies with middle powers. + +Would this work? The optimistic case is that this isn’t a big sacrifice for the US. They can still become as rich as they like and achieve their security interests. Sure, they can’t seize control of other nations, but that is not an important goal of theirs anyway. Losing that option is well worth the benefits: other nations cooperate economically, don’t attack US data centres, and don’t threaten nuclear war. + +The pessimistic case is that this involves an insane degree of irrevocable hand-off to AI. The US must literally be unable to attack middle powers no matter how hard it tries: retraining the AI, turning it off, training a new more powerful AI, passing new laws, using the military to destroy the data centres the AI is running on. For it to be truly binding, the US must permanently hand over military and political power to AI. That might be deeply unpopular, and indeed seem insane to the US. It’s also extremely hard to verify: you can’t just verify the training run, you need to verify that humans+other AIs have *no* way to disempower the trained AI. It’s more like verifying “who would win this civil war” than “technical property XYZ holds”. + +The realistic path here probably involves gradually handing off more and more control to AI that refuses to crush middle powers, with no clear point at which humans could no longer wrest back control. + +To make the hand-off less irrevocable, the commitment could be time-bound: superintelligence won't help the US crush the middle power within the next 2 years. That could be enough to get us through a software-only intelligence explosion, after which middle power's compute supply chain leverage is more binding. + +The longer middle powers wait to push for stage 2, the less leverage they will have because the US will have pulled further ahead economically and militarily. So they should be pushing in this direction constantly, e.g. demanding transparency into the model specs of powerful AIs deployed in the US government, and arguing that powerful military AI should be designed to obey international law. + +(I described the plan as involving two stages because that’s how I expect it to play out over time. But succeeding at either stage is sufficient! If middle powers stay economically/militarily competitive, they never need to bind US superintelligence. And if they *do* bind superintelligence, they won’t be crushed no matter how far behind they fall.) + +Another strategy: train superintelligence to ensure middle countries continue to get equal access to frontier AI. This combines stages 1 and 2, and could prevent even the *relative* disempowerment of middle powers. + +Is it good to avoid middle powers getting trounced? +--------------------------------------------------- + +I live in the UK, so I am biased here. I do not want the UK to become a supplicant to the US! + +But here’s a brainstorm of pros and cons from a more impartial perspective. + +Pros to empowering middle powers: + +* **Avoid a single point of failure**. If the US becomes globally dominant and its political system fails, that’s a global failure. +* **More democracies.** Many middle power democracies look more robust than the US, so more middle powers may mean more democracy. +* **Improve the US.** Middle powers will have an interest in maintaining free market democracy in the US. “Free market” because they’ll want multiple AI companies competing to sell cheap API access to non-US countries. “Democracy” because they’ll expect that the US is more likely to maintain a strong alliance with middle power democracies if it stays democratic. +* **Experimentation.** Experimenting with multiple different political and legal systems seems generally good for figuring out a good way to govern society post AGI. +* **Pause AI.** They could potentially pressure US/China to pause/slow down reckless AI development. +* **Prosocial norms.** When multiple actors bargain with each other (e.g. about how to distribute space resources, whether to develop a dangerous technology), they tend to frame arguments in terms of prosocial norms, and so agreements tend to emphasise the actor’s more virtuous/ethical values. + +Cons of empowering middle powers. Multipolarity has its own downsides: + +* More likely to lead to war. +* Can drive extreme competition, e.g. racing to develop a dangerous technology, or to hand off power to misaligned AI. +* Harder to prevent harms from offence-dominant technologies like bioweapons. +* This plan involves waking up middle powers, which could shorten timelines. + +Top Comments Index +------------------ + +### Comment by [Andrii Vasylenko](/users/andrii-vasylenko) + +* 2026-05-29 22:28:28Z +* Karma: 3 +* Voting system: namesAttachedReactions +* Approval votes: 2 +* Total votes: 2 +* HTML permalink: [/posts/hi5cmPgJqkFxpMgSC/how-can-the-middle-powers-avoid-getting-trounced-during-the/comment/Q4HC4fSKaGSwoGti3](/posts/hi5cmPgJqkFxpMgSC/how-can-the-middle-powers-avoid-getting-trounced-during-the/comment/Q4HC4fSKaGSwoGti3) +* Markdown permalink: [/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the/comments/Q4HC4fSKaGSwoGti3](/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the/comments/Q4HC4fSKaGSwoGti3) + +### Comment by [callumzc](/users/callumzc) + +* 2026-05-28 14:23:38Z +* Karma: 1 +* Voting system: namesAttachedReactions +* Approval votes: 1 +* Total votes: 1 +* HTML permalink: [/posts/hi5cmPgJqkFxpMgSC/how-can-the-middle-powers-avoid-getting-trounced-during-the/comment/Ce5KmdQEnkm8mg3zg](/posts/hi5cmPgJqkFxpMgSC/how-can-the-middle-powers-avoid-getting-trounced-during-the/comment/Ce5KmdQEnkm8mg3zg) +* Markdown permalink: [/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the/comments/Ce5KmdQEnkm8mg3zg](/api/post/how-can-the-middle-powers-avoid-getting-trounced-during-the/comments/Ce5KmdQEnkm8mg3zg) + +### Navigation + +* [Front page](https://www.lesswrong.com/api/home) +* [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md) diff --git a/sources/resources/e61-compute.md b/sources/resources/e61-compute.md new file mode 100644 index 0000000..12ff724 --- /dev/null +++ b/sources/resources/e61-compute.md @@ -0,0 +1,190 @@ +--- +url: https://e61.in/ +--- + +# plus61 Special feature: Joe Walker x e61 -- The compute economy and Australia's place in it + +> NOTE (Claude): cleaned text extraction from a mailchimp campaign-archive email (fetched 2026-07-11); original was nested-table layout. + +Joe Walker x e61 + + +View this email in your browser + +plus61: Special feature + +Joe Walker x e61 + +The compute economy and Australia’s place in it + +The financial returns to Australian data centres are uncertain and could be modest in the long run, but there is still a case for government support. + +If Australia wanted to become a data centre capital of the world, it would certainly be able to – Sam Altman, May 2026. + +Computing power, or “compute” for short, is emerging as a critical feature of the global economy. Described simply, compute refers to specialised computer chips sitting inside racks of servers that are stacked within warehouses called data centres. Vast amounts of compute are required to both train and run modern AI systems. The compute needed to train frontier models, for example, has been doubling roughly every five months since 2020. Growth in compute demand is driving a boom in data centre investment and has seized the attention of policymakers from the UK to the UAE to Australia. + +Since compute is a key intermediate input of the AI age, many people assume it will remain an enormously profitable business. Australia could capture a slice of this value by producing compute at scale and exporting to the world (as Sam Altman has suggested). Yet the long-run financial returns to a compute export industry are uncertain and could turn out to be relatively modest, as we outline in this article. While we cannot predict how technology will unfold, nor are we experts on LLMs or digital infrastructure, economic fundamentals suggest compute profits may ultimately look less like iron ore and more like electricity generation. + +Despite this, Australian governments have other reasons to step in to support data centre growth – the most obvious being national security and economic resilience. + +The best policy support Australian governments can provide initially is to minimise existing regulatory inefficiencies, balancing the degree of support between data centres and complementary infrastructure. Governments also need contingency plans for compute access for critical activities to manage the risk of disruption to compute facilities overseas. + +This article focuses on the narrow but important issue of compute. Compute sits among many other considerations for government in responding to the rise of AI, which we do not address in any detail. These include how and whether to promote AI adoption and skills, support research and access to frontier AI models, support an infant industry of AI-application firms, prepare for new threats to cybersecurity and collect new economic data about technology adoption. + +The surge in investment in compute + +The world has seen a sudden acceleration in capital expenditure on data centres. + +The 5 largest US AI-infrastructure firms are expected to complete about US$700 billion in capital expenditure this year. Globally, property and other consultancies project that AI infrastructure expenditure will be about US$1.4 trillion for the year and installed capacity will rise by about a quarter as a result. + +Australia is in the mix. There has been a surge in domestic data centre construction, with ICT capital expenditure reaching well over 1 per cent of GDP in the first part of 2026 versus an earlier average of around 0.3-0.4 per cent of GDP. That said, most of that value is imported – the recent increase in ICT capital expenditure has been substantially offset by a rise in data processing equipment imports, which mechanically offset from GDP growth. + +Geography of the global compute industry in the longer run + +Global compute production is relatively concentrated at present. + +Around three-quarters of data centre compute production resides in just 4 countries, according to estimates of data centre energy usage by country. The US and China account for the vast majority of that figure. This is more concentrated than global exports of computer chips and close to the global concentration of iron ore exports (using exports as a proxy for production, since global trade data is more available than global production data). + +Compute may be more widely distributed across the world once the market is mature. + +This can be seen when we place the characteristics of compute alongside more familiar reference points, such as oil, computer chips and iron ore. Each of these goods is a more ‘mature’ intermediate factor of production. The principles of economic geography help to explain the shape of each global market. For example, iron ore exports are especially concentrated because of the location of viable deposits, whereas computer chips are concentrated because of large-scale economies to production. Viable crude oil deposits are much more dispersed across countries than is iron ore. + +We can extend those principles to compute. The case for a more globally dispersed compute industry in the longer run rests on 4 arguments: + +Limits to trade: There are limits to how far compute can be transported for some purposes and so some of it must be positioned near end users. The workloads that cannot be traded are those requiring low latency (that is, quick turnaround) or with data sovereignty protections. + +Fewer physical limits to their location: There are fewer geographic or scale constraints on where you can build data centres (in contrast to, say, mines). You need some land (former industrial or greenfield), trade access to chip manufacturers, reliable electricity and water connections, and some construction and operations workers. In principle, this could be many locations. + +Scale economies: These exist as AI companies do benefit from larger and coherent clusters for model training, but these economies seem less extreme than for high-end computer chips, where high capital expenditure combines with critical proprietary information learnt by doing. + +The inability to store reserves of compute: This points to more demand for data centres themselves as protection against disruptions to trade. + +By contrast, concentration forces may dominate for other uses. For model training, and perhaps agentic AI uses, higher latency (i.e. slower turnaround) can be tolerated and compute is most tradable. This is where compute can congregate at its lowest cost location. Another potential force for concentration is the housing of compute in satellites, which according to Elon Musk may be achieved within a few years. Few countries can contribute to that endeavour. + +Australia could build compute not just to satisfy its domestic needs, but also to compete in this global export market. + +Australia appears to have some comparative advantages in that competition, including the potential to develop widespread low-cost renewable energy, plus the present realities of strong legal institutions, land availability outside the main inner cities, high-quality subsea cable connections, trade access to AI-suitable chips, and stable geopolitical relationships. On the other hand, uncertainties about Australian copyright law are a potential barrier to model training, as the Productivity Commission has highlighted. + +A mature global market for compute could see modest financial returns + +The surge in demand for compute has made for strong near-term financial returns to data centre owners. But the longer run looks more uncertain, and that is the time frame that matters most for a national strategy on compute. + +The bullish argument is that demand for compute will grow rapidly even over the longer run. This assumes AI’s economic value will keep growing as it takes on new tasks, and that shifts in how compute resources are used will outweigh the diminishing returns of scaling a single area like training or inference. Add to this a sense that short-run bottlenecks to delivering more compute will always tend to arise, and it points to continued strong financial returns. + +There is also an underappreciated bearish case that today’s returns to the scarcity of compute will diminish for the following reasons: + +The barriers to building data centres are modest and build times are relatively short compared to the other intermediate inputs we surveyed, so supply of data centres themselves is relatively elastic. + +The supply of computer chips is the more significant bottleneck, and the supplier base for AI-suited computer chips is highly concentrated. This means the returns to scarcity are likely to sit upstream of the data centres themselves. + +The product is relatively homogeneous once we are talking about latency-tolerant compute, which limits market segmentation and barriers to entry. (Although data centre chips, structures and legal regimes do amount to some product heterogeneity). + +The customer base is relatively concentrated among a few powerful AI company buyers, giving them substantial negotiating power. We have already seen some major US compute consumers flex their purchasing power in Australia in search of preferential tax status or other accommodations. + +Economic depreciation rates on computer chips are high, meaning customer willingness to pay will decline once the initial contracted period is complete. + +It will be difficult for any country to deliver a much lower cost product because labour costs are a small portion of the compute cost base. Energy costs are more significant to the cost base, and there is less ability to achieve an advantage since energy prices are generally linked to global markets. (Although this is complicated by domestic energy reservation policies, or data centres that vertically integrate with energy generation.) + +Put together, this suggests that in a longer-term mature market, pricing of compute may be negotiated down towards the cost of capital. Even AI compute may be commoditised. + +There may be sufficient returns to motivate the investments, but there seems little scope for the kinds of large economic rents that have been observed over the past 20 years with Australian iron ore, Saudi Arabian oil or Taiwanese chip production. The construction of data centres itself will offer some economic return, but there are supply constraints in the Australian construction industry, including for housing and infrastructure, suggesting material opportunity costs. + +Identifying externalities and reducing inefficient delays + +The main game for Australian policymakers is to determine if there are positive externalities – social returns beyond the private returns to firms – from a large domestic compute industry. + +Economic theory suggests three rationales for industrial policy intervention, namely positive externalities from an activity, coordination/agglomeration failures, and public good provision. For data centres, the first of these – externalities – is the most relevant. + +We see 3 main arguments for externalities from domestic data centres. + +For ‘economic resilience and security’, domestic compute capacity is likely to have useful insurance and deterrence externalities. + +Amid increasingly volatile geopolitics, governments are talking about building resilience to trade disruptions. Domestic production capabilities can come at an efficiency cost if the private returns are not there, but they can also provide insurance against disruption to critical economic inputs. The case for buying this insurance is strongest where the disruption scenario is more likely and the consequence of that disruption is high. + +Domestic compute reduces the threat capability of countries that provide compute or could threaten to disrupt international transmission infrastructure. This is discussed more in a break-out box. + +Box: How strategic compute capability could protect Australia's sovereignty + +The Australian government may wish to ensure a substantial domestic compute capacity to contribute to its strategic goals, most importantly to protect Australia's sovereignty. + +The direct connection: we should control enough compute to ensure that we can continue to run critical functions of government and industry, even in a scenario of global disruption. + +Without this domestic guarantee of supply, critical activities will be exposed to loss of access to the satellite connections or, crucially, undersea cables that link us to foreign data centres. This becomes increasingly important the more that domestic activity, and particularly the function of Australia's defence forces and intelligence agencies, rely on compute. Addressing this exposure would reduce an important threat-point for other nations. + +Insurance could be fiscally costly if it involves subsidies, but there are reasons to think the costs can be mitigated. Additional data centres can be utilised for other purposes when not in use for critical purposes. Unlike oil reserves, compute does not get 'used up' if it is deployed for other activities. Also, it may be possible to achieve a similar policy goal without designating specific data centres as strategic reserves – government just needs conditional priority access. + +The indirect connection: a big role for Australian data centres in global compute can plausibly give us leverage at other layers of the AI stack, especially frontier models, and help us pursue larger national goals. It could help Australia secure continued access to these models in a world where they are rationed, as well as early access to increasingly powerful models that are not released to the public (like Claude Mythos Preview). It may also give the Australian government some influence over such models’ governance and alignment with Australian values. Concretely, this could look like the Australian government striking MoUs with frontier labs that seek to do training runs in Australian-based data centres. + +Crucially, these indirect benefits do not arise merely by virtue of data centres being located in Australia, even if those data centres are owned by Australian companies. Rather, they will depend strongly on the particularities of contracts or agreements between frontier labs, the Australian government, and data centre providers. Australian statecraft may need to play a large role in shaping these outcomes. + +Demand for electricity from domestic compute capacity could have positive externalities on investment in the electricity grid. However, it could also push up energy costs for other sectors if renewable energy investment is not sufficiently responsive. + +The positive externality could arise because investment in the electricity network involves large fixed capital costs. Data centre demand could help spread those costs and so prove an important catalyst of network investment. This is certainly the argument that some grid operators are making. + +On the other hand, other industry sources have expressed concern that bottlenecks in renewable energy supply could lead to higher energy prices, undermining both the potential for any cost advantage to Australian data centres and also putting inflation pressure on the rest of the economy. (This may be less problematic for data centres using “behind the metre” generation.) + +There is also an argument for innovation or productivity spillovers from domestic AI capacity because some Australian data must be held onshore. But beyond that, arguments for spillovers tend to be vague and deserve more scepticism. + +The strongest argument here is that additional domestic compute capacity can help facilitate research on secure Australian data that must be handled domestically. This is similar to the argument for specialised supercomputers that Australian governments already support. + +For instance, research on Australia’s rich health data assets will tend not to be value-captured by the researchers, leading to under-research. There is a place for government to further encourage research of this nature. This could be done by encouraging data centre infrastructure directly. An alternative targeted approach could be to expand payments to researchers, who will then have the funds to spend on compute. This will in turn encourage the required data centre infrastructure to be delivered by the private market. + +There are also regulatory or self-imposed limits to where many private Australian firms’ data can be stored. + +In these cases, it might matter not just whether the data centre is onshore, but also whether the data centre is owned by an Australian provider. Australian firms may be uneasy about the risk that US government authorities claim jurisdiction over data held by US service providers on Australian premises, a question raised by the US CLOUD Act even if its application is currently limited to countering serious crime. + +Overall, there is not currently a strong case for additional externalities from ‘the vibes’ of having data centres housed in Australia. Latency reductions could have some positives, but seem unlikely to be the binding constraint to AI adoption and AI-driven innovation. Management quality, other data infrastructure or regulatory uncertainty are more likely to be binding for adoption. + +The current focus of Federal and state government policy on data centres has been to reduce existing inefficient regulatory barriers. + +The NSW and Victorian governments have set up expedited approval processes for data centres. Existing delays in assessing development applications are an important inefficiency, not specific to data centres. If existing regulatory principles can be applied more quickly, that is a clear policy win. + +However, there is an important caveat: data centres and supporting infrastructure need to expand in concert. This means policy support should be applied in concert to energy and water infrastructure. For example, if data centre investment were to streak ahead of electricity generation, this could lead to energy costs rising materially and dragging on the economy as a whole. The Australian Government’s recent statement on expectations of data centres shows recognition of this issue. + +Would we regret support for data centres if AI investment is a bubble? + +It has been suggested that we are in the midst of an AI bubble. Critics point to the enormous scale of capital expenditure by the leading AI companies, much of it on compute, and the gap to still-modest AI subscription revenues. They also point to the easy financial conditions faced by those AI companies – for example, their extraordinary equity price-to-earnings multiples. + +Both of those financial comparisons reflect, in part, the lag between investment and payoff for any new venture. But it's also possible that the winner-takes-most race among AI firms will lead to excessive investment, and that fear of missing out and ‘rational bubble’ dynamics are at play among investors chasing capital gains. + +If there is a bubble, then how might we expect it to play out? + +There is a useful historical guide to a global investment bubble that made a big imprint on Australia in the 1880s. The technological advance of the time was advances in cheap steel manufacturing. That sponsored a global frenzy of investment in railways, ports, shipping capacity and telegraphs, predominantly financed out of the global finance hub of London. During this period, much of British investment went overseas, including to Australia. + +These investments ultimately delivered poor financial returns, and their price crash contributed to a period of depression. But the infrastructure that was built did benefit Australia’s productive capacity. + +There is some analogy to the deployment of a new technological advance to Australia in the form of data centres, and how even if the investors take a bath at some point, the infrastructure could remain useful for some time. + +In thinking about whether this analogy holds, a key difference that stands out is that economic depreciation rates on data centres are higher than for railways. Even so, economic depreciation is not always as smooth or fast as expected – for example, older computer chips like H100s recently rose in value for a period due to shortages of new higher-performance chips. + +The case for more evidence + +Many of the arguments made here are speculative, in large part due to the limited evidence available. Based on what we know today and economic principles, there is reason to think compute will become more dispersed globally, and Australia could carve out an important role for itself. Further, financial returns to compute are uncertain in the long run and may turn out to be modest. Yet, governments may still have an important role to play in ensuring sufficient domestic capacity for economic resilience and security reasons, and in supporting positive externalities while minimising negative spillovers through electricity prices. + +Substantiating the arguments made here requires filling the current gaps in our knowledge. In particular, there is a great need for more evidence on the externalities from domestic compute capacity. For example, the significance (if any) of domestic compute access for Australian firms’ productivity, including tech startups. There is also a need for clearer evidence on the private returns to compute infrastructure providers, the supply elasticity of renewable energy capacity, and the non-linearities in the impact of electricity demand on investment in the electricity network. We also need to carefully identify the bottlenecks to research on Australian sensitive datasets and the potential remedies. + +Sam Altman may be right that Australia could become a data centre capital of the world. But it turns out that the case for achieving this vision may turn less on profits than on sovereignty and resilience. + +Greg Kaplan + +Greg Kaplan is co-founder and Chair of the e61 Institute, and Alvin H. Baum Professor in the Kenneth C. Griffin Department of Economics and the College at the University of Chicago. His research focuses on macroeconomics, labour economics and applied microeconomics, with particular interests in fiscal and monetary policy, labour markets, inequality, and household behaviour. He is Lead Editor of Journal of Political Economy Macroeconomics and is author of the forthcoming book Macro: The Economic Models that Shape Our World to be published in November 2026. + +For more information, please reach out to Greg via email at greg.kaplan@e61.in + +Ewan Rankin + +Ewan is a research manager focused on our Productivity and Structural Analysis workstream. Ewan has more than a decade of experience in economic analysis to support monetary policy at the Reserve Bank of Australia. He has an Economics honours degree from University of Sydney and an MPA from Princeton University. + +For more information, please reach out to Ewan via email at ewan.rankin@e61.in + +Joe Walker + +Joseph Walker is host of The Joe Walker Podcast, Australia's leading ideas and public policy podcast. He is a two-time Emergent Ventures winner and previously worked in technology start-ups, most recently as director of operations at Y Combinator-backed Forage. + +For more information, please reach out to Joe via email at joe@jnwpod.com + +View email in browser + +www.e61.in + +update your preferences or unsubscribe diff --git a/sources/resources/europe2031.md b/sources/resources/europe2031.md new file mode 100644 index 0000000..c0387a2 --- /dev/null +++ b/sources/resources/europe2031.md @@ -0,0 +1,1165 @@ +--- +url: https://europe2031.ai/ +--- + +[Europe 2031](/) + +[Scenario](/) [SummaryExecutive Summary](/summary) [About](/about) [Compute ForecastCompute](/compute-forecast) EN Language: English + +English[Czech](/cs)[Danish](/da)[Estonian](/et)[French](/fr)[German](/de)[Greek](/el)[Italian](/it)[Portuguese](/pt)[Spanish](/es) [MarkdownFor AI agents & crawlers](/agents/scenario.md) + +* [Jan 2025 DeepSeek and ye shall find](#chap-jan-2025) +* [Feb 2025 Plug, baby, plug](#chap-feb-2025) +* [Aug 2025 Forever blowing bubbles](#chap-aug-2025) +* [Nov 2025 Worlds apart](#chap-nov-2025) +* [Feb 2026 A war on two fronts](#chap-feb-2026) +* [Mar 2026 The wild frontier](#chap-mar-2026) +* [Apr 2026 Mythos](#chap-apr-2026) +* [Jun 2026 Disillusionment](#chap-jun-2026) +* [Aug 2026 The fork](#drift-chap-aug-2026) +* [Sep 2026 A positive vision](#drift-chap-sep-2026) +* [Jun 2027 Closing windows](#drift-chap-jun-2027) +* [Jan 2028 Taking stock](#drift-chap-jan-2028) +* [May 2028 The compute crunch](#drift-chap-may-2028) +* [Aug 2028 Writing on the wall](#drift-chap-aug-2028) +* [Nov 2028 Vox populi](#drift-chap-nov-2028) +* [Jan 2029 Dreams of a robot economy](#drift-chap-jan-2029) +* [Apr 2029 Access all areas](#drift-chap-apr-2029) +* [May 2029 Tilt](#drift-chap-may-2029) +* [Feb 2030 Labour shock](#drift-chap-feb-2030) +* [Jun 2030 An offer you can't refuse](#drift-chap-jun-2030) +* [Aug 2030 Model Collapse](#drift-chap-aug-2030) +* [Mar 2031 Between giants](#drift-chap-mar-2031) +* [Jun 2034 Epilogue](#epilogue) + +Jan 2025 + +# Europe 2031 + +What getting AI wrong means for us + +[Daan Juijn](https://www.arq.foundation/team/daan-juijn), [Stan van Baarsen](https://www.linkedin.com/in/stanvanbaarsen), [Judith Dada](https://www.linkedin.com/in/judith-dada), [Maximilian Negele](https://www.linkedin.com/in/maximilian-negele), [Lily Stelling](https://www.rhodeshouse.ox.ac.uk/scholar-community/rhodes-scholar-bios/rhodes-scholars-class-of-2026/lily-stelling/), [Philip Fox](https://www.philip-fox.com), [Alex Petropoulos](https://x.com/AlexTPet), and [Michiel Bakker](https://www.miba.dev).Copywriting and editing by Tom Chivers. + +11 June 2026 + +![The twelve EU stars dissolving into pixel dust, framing a mosaic of a crumbling map of Europe at the centre.](/assets/europe-hero.webp) + +The current trajectory of AI calls for the most ambitious political agenda in the history of post-war Europe. Unless we embark on it now, Europe will lose the ability to shape its own future. We will end up economically and politically sidelined, with values we cannot defend, social welfare systems we can no longer fund, risks we cannot address, and a Union that cannot hold. + +> March 2031 - Washington, D.C. +> +> Caroline splashes cold water on her face and looks at herself in the bathroom mirror. Her hands are shaking. She grips the edge of the sink and waits for it to pass. Through the small high window she can see a slice of Washington sky, flat and bright. +> +> Down the corridor, six people are deciding the fate of the European continent. She does not know whether anything she said will matter. +> +> She suspects it will not. + +This is a scenario about Europe's impending slide into irrelevance: how AI is driving it, and what can still be done to change course. + +It is told through the fictional characters of Caroline Dubois, a French policy worker in Brussels, and Christian Vogt, the German founder of a fast-growing AI company who has just relocated to Silicon Valley. While their personas are made up, the events they experience are grounded in existing trends, and meant to be realistic. + +To understand how Europe is at risk of squandering the coming AI revolution, we first have to return to early 2025 - because the story we're about to tell is already well underway. + +Listen to the story [Spotify](https://open.spotify.com/episode/0c8iIvzD5IEnT6xz6SU8q5)[Apple](https://podcasts.apple.com/us/podcast/-/id1896865536?i=1000772090203)[Web](https://audio.europe2031.ai/europe-2031-audio.mp3) + + [Summary](/summary) [Read as PDF](/europe-2031.pdf) [About & FAQs](/about) + +Contents + +1. [01 **Prologue**: How Europe got here Jan 2025 – now](#timeline) +2. [02 **Scenario**: The five years ahead now – March 2031](#scenario) +3. [03 **Epilogue**: What could have been June 2034](#epilogue) + +[Jan 2025](#chap-jan-2025) [Feb 2025](#chap-feb-2025) [Aug 2025](#chap-aug-2025) [Nov 2025](#chap-nov-2025) [Feb 2026](#chap-feb-2026) [Mar 2026](#chap-mar-2026) [Apr 2026](#chap-apr-2026) [Jun 2026](#chap-jun-2026) [The Scenario](#scenario) + +### January 2025 - DeepSeek and ye shall find + +Caroline Dubois’ office is buzzing. This morning, a Chinese firm, DeepSeek, released a new AI model, R1. It is cheap and effective, and – although no European firm has played a part in developing it – Brussels is in a fervour. + +Caroline works on digital technology at the Directorate-General for Trade and Economic Security – DG TRADE, the part of the European Commission that deals with things like tariffs and exports controls. At twenty-eight, she has been there for three years, and feels she is earning the respect of her peers and superiors; but she, unlike them, is seriously worried about Europe’s future, and the DeepSeek news has not reassured her. She has recently visited Silicon Valley. It is 9,000 kilometres from Brussels, but it feels further. The idea that AI is sparking a new industrial revolution is a truism in California; in the European Commission offices, it is bordering on science fiction. + +R1’s rise has excited her colleagues because they see it as clear proof that it is possible to train cutting-edge AI without the resources of the Silicon Valley giants. European policymakers jump at the idea that it is possible to outsmart the Americans, to be small and nimble and clever, to do better without the hundreds of billions of dollars that have been pouring into indulgent American data centres and gigantic model training runs. + +The idea makes sense. DeepSeek has supposedly developed R1 at a fraction of the cost of OpenAI’s ChatGPT, or Anthropic’s Claude, but as far as anyone can tell, it is almost as powerful. It even has the new ‘reasoning’ feature that the American models have started rolling out, in which the model lays out its thought process on a virtual scratchpad. And its weights – the mathematical values that make the model what it is – are publicly available. Anyone can run the model on their own infrastructure, free from American tech dependence. + +Though somewhat energised by the hopeful mood in Brussels, Caroline isn’t quite ready to join in. She carefully listens to more cautious voices who point out that big efficiency gains are hardly unheard of – Anthropic, Google DeepMind, and OpenAI find them all the time. DeepSeek has smart researchers and moved fast, but will soon find itself constrained by computing power; China simply doesn’t have enough ‘compute’ to train models ever since America restricted exports on the cutting-edge AI chips that China can’t produce domestically. They argue that finding efficiency gains is much easier if you have lots of compute to help you find them. DeepSeek’s own CEO pointed out that their main problem is the US ban on AI chip exports. + +And so Silicon Valley is largely unfazed by R1. The hyperscalers do not slow down their massive AI investments. Days after R1, OpenAI releases o3-mini, a more impressive model than R1, and a sign that American progress is still rapid. + +In Brussels, the o3 news barely registers. Acknowledging it would mean accepting that AI acceleration is continuing, driven by huge American companies, led by CEOs that can’t be trusted. + +Caroline isn’t sure which side to take. She was in California three weeks ago, and what she saw there has stuck with her. + +She stayed with Christian Vogt, an old friend from her study-abroad programme in Berkeley half a decade earlier; she from Sciences Po, he from TU Munich, two of Europe’s leading universities. Christian had moved to San Francisco three years ago, just as she started at the Commission, to found a company building image and video models. The AI startup was still tiny, but it had recently closed a Series A funding round and Christian, a gregarious type, knew *everyone*. + +The whole visit had startled her. The hours Christian’s team pulled were insane – seventy- or eighty-hour weeks, people sleeping in the office. And when he took her to a dinner party at the house of another startup founder in Hayes Valley, she found the conversations almost incomprehensible; not because of the technical details but because of the deep belief that the world was going to change radically. ‘By January 2026, I think most of my code will be written by Claude,’ one guest said in a matter-of-fact tone. The host said he had stopped hiring junior engineers because ChatGPT would soon do the entry-level coding better. Someone mentioned, in passing, that they thought artificial general intelligence - AI that is better than any human at most tasks - was probably two or three years out. Caroline had asked what this would mean for Europe, given it lacks a competitive AI sector. But nobody at the table had really thought about Europe, except as a potential compliance headache. + +In the Waymo back to Christian’s apartment, Caroline had told him she found the whole visit disorienting. Christian had laughed and reassured her that after a few months in Silicon Valley, she would start to ‘feel the AGI’ herself. Caroline had wondered if Christian and his peers were part of a strange bubble, a cult even, or if she was living in a world that was already disappearing into the past. + +Back in Brussels, she had considered telling her director about the experiences she had had on her trip, but it felt untranslatable. The people she had spoken to were some of the smartest people she had ever met. They were calm, they were sharp, and they seemed well-informed. There was an urge inside her to take what they had told her deeply seriously and to get others to do the same. And yet their claims would have sounded insane in a Commission conference room halfway around the world. So she said nothing. + +The day after the R1 release, her phone buzzes. It’s Christian. + +Christian: ok so brussels is doing a deepseek victory lap??? + +Caroline: My director thinks it proves we can catch up… + +Christian: lol + +Christian: the lesson is reasoning models work and china can build them + +Christian: also, o3-mini is better and nobody is talking about it + +Christian: your director should come to SF + +Caroline: I wish. + +She puts the phone down. She isn’t sure Christian is right, but she isn’t sure he’s wrong, either. The voices of caution about R1 sound reasonable to her. So do the voices of excitement. + +### February 2025 - Plug, baby, plug + +Three weeks after R1, Emmanuel Macron hosts the AI Action Summit in Paris. + +The summit series was founded around safety and international coordination – the 2023 edition in the UK had been called the AI Safety Summit – but the 2025 edition is different. The geopolitical environment has hardened. Russia is grinding forward in Ukraine; Donald Trump has been back in the White House for a fortnight and is already threatening tariffs on European goods. Europe is now focused on competitiveness, not safety. And the timing is good: DeepSeek has apparently shown that catch-up is possible, and could perhaps even be cheap. Europe wants to signal it is taking AI seriously. + +European Commission President Ursula von der Leyen announces a €200 billion InvestAI Fund, which includes a €20 billion AI Gigafactories Initiative to build four to five large AI data centres on European soil. Macron sells France as the best place in Europe to build AI: with all its nuclear power, developers can just ‘plug, baby, plug.’ + +After three years in Brussels, Caroline knows to look beyond the headline numbers. The €200 billion is largely a repackaging of existing funds, and an assumption that European industry will pour its own money in. Only a fraction is genuinely new, and even that is spread out over five years. She shudders when comparing it to the data centre investments announced by American hyperscalers, projected to exceed $400 billion in 2025. Still, she hopes the vibes are shifting, and that Europe is waking up. + +Christian: 200 billion? + +Caroline: It's mostly aspirational. + +Christian: the americans have real money + +Christian: the texas thing is being built right now + +Christian: like there are bulldozers + +Caroline: I know. + +Christian: how does brussels feel about altman, ellison, masayoshi son on a stage with trump + +Caroline: Well, not thrilled. + +Whether the vibes have shifted or not, some of the positive energy is overshadowed when US Vice President JD Vance gives an aggressively anti-Europe speech at the summit. He repeats the trick two days later at the Munich Security Conference. European leaders in the audience are seen on camera almost visibly biting their tongues. + +In the weeks after, Caroline’s bosses come to a conclusion. The transatlantic relationship is effectively broken. Washington can no longer be trusted, not in defence, not in energy, and certainly not, it is now clear, in tech. ‘Sovereignty’ becomes the new buzzword in European capitals. + +But sovereignty is easier to announce than to build, and whether the push will amount to anything is not clear, at least in AI. Caroline suspects that while leaders profess to care about the new technology, their old scepticism is very much still in place beneath the surface. One of her colleagues, a thoughtful and senior policy officer whom she likes, implies during a lunch shortly after the Vance speeches that he thinks you probably don’t need vast data centres to make AI work, precisely because Sam Altman, Larry Ellison and Donald Trump all stood on a stage to say you do. It is, Caroline thinks, a simplistic but understandable position: Brussels’ scepticism of Trump and the Silicon Valley elite runs deep. But she is worried that just because they say it’s sunny, that doesn’t mean it’s dark outside. + +### August 2025 - Forever blowing bubbles + +The European AI excitement subsides after Paris. The summit was dramatic, but the months afterward are filled with slow, dense, and unglamorous policymaking. Building a continental AI ecosystem turns out to require things that announcements alone cannot conjure – talent, capital, energy supply - as well as the know-how and drive to coordinate them. The largely illusory nature of the €200 billion is widely known now, and most of the air has gone out of the initial elation. + +In the US, there is no such slowdown. The battle for talent heats up: Meta, which has fallen behind the AI frontier somewhat, is poaching top researchers from rivals on salaries that make Premier League footballers look underpaid. Mark Zuckerberg even personally cooks soup for them to lure them in. The Trump administration publishes its national AI strategy, titled ‘Winning the Race,’ which describes the AI transition as ‘an industrial revolution, an information revolution, and a renaissance – all at once.’ + +Europeans, who consider themselves more grounded than Americans, find the rhetoric overheated and the messengers untrustworthy. Caroline hears the phrase ‘AI bubble’ regularly. Her colleagues have seen tech leaders overstate their products’ usefulness before: ‘Remember NFTs?’, one says. ‘Those pictures of monkeys that people bought for $50,000?’ + +Christian: how's the bubble talk going + +Caroline: It's everywhere. My director said it three times in a meeting yesterday. + +Christian: wait until they see anthropic's revenue numbers + +Christian: they're going to do 9 billion this year + +Christian: from like 1 billion last year. + +Christian: that’s frickin huuuuge + +Caroline: I'll mention it. + +Christian: you won't + +She does mention it. No one seems very excited, though. Some colleagues point out that Anthropic is still not making a profit. + +The ‘bubble’ narrative is strengthened by the release of GPT-5. OpenAI hypes the model hard: Sam Altman posts an image of the Death Star; employees talk about the ‘year of the AI agent’. But to the public, the model itself is somewhat disappointing, essentially a more polished version of the existing o3 model. It still hallucinates and makes silly mistakes. + +In Brussels, the AI sceptics take another victory lap. ‘I told you AI was hitting a wall,’ Caroline’s colleague tells her over a beer at Place du Luxembourg. He says it kindly. Caroline has, over the past few months, become known in her unit as the AI booster, which is a label she is not sure she deserves. She is not yet certain her colleagues are wrong; She agrees that GPT-5 really was disappointing. Maybe the Americans are hyping. Maybe Brussels is right. + +On the other hand, some experts say that GPT-5 is broadly what you’d expect given the exponential growth curve of AI’s capabilities; o3 was only a few months ago. But Europe’s AI sceptics are deaf to those arguments, given OpenAI’s marketing blunder in over-promising and then under-delivering. + +Caroline: What do you make of GPT-5? + +Christian: it was a misfire. they should have shipped o3 under that name. + +Christian: but the underlying capability is still moving + +Christian: we've been using o3 in our pipeline for the last months and it's wild + +Money is still pouring into the big US labs and the race to build data centres is still on, but that money seems to bounce back and forth like a ping-pong ball – Oracle and OpenAI sign a $300 billion compute deal, most of which will be spent on NVIDIA chips; NVIDIA in turn agrees to invest $100 billion in OpenAI. In San Francisco, the deals are seen as proof of the tech’s transformative nature – the old rules no longer apply. In Brussels, they look like the credit-default swaps of 2007: complex, circular, unstable, and likely to end badly. + +That is why the French company Mistral raising €1.7 billion – most of it from the Dutch lithography machine maker ASML – is hailed as a big win for Europe, despite its funding round being twenty times smaller than OpenAI’s. The bubble talk means that few Europeans are worried about AI companies having *too* *little* money. If AI and the largest investment rounds are overhyped, there’s no need to hurry. + +### November 2025 - Worlds apart + +In November, Anthropic releases Claude Opus 4.5. There is no Death Star tweet, no countdown, and no event. The model is good – better than GPT-5, by most accounts – but the release is, by 2025 standards, undramatic. + +What is dramatic is what people start doing with it. + +The breakthrough in AI agents has arrived, but not how people imagined it. Initially, through the spring and summer, the major labs had released various products that controlled your computer directly. They would move your mouse, fill out web forms, click through booking flows. While nice in theory, these mostly did not work in reality. Many people who tried them concluded that AI agents were not yet ready for prime time. + +What does work, it now turns out, is agents that write code. Anthropic’s Claude Code, a terminal-based tool, becomes – once paired with the new Opus 4.5 – the surprise hit of the year. Christian’s friend was right: when January arrives, most lines of software in Hayes Valley really are written by AI. Developers quickly realise that code is the universal interface: if you can write code, you can do anything a computer can do. Want to send an email? Don’t navigate Gmail; Claude Code, write a script. Want to organise documents? Claude Code, write a script. + +In San Francisco, this becomes common practice. People talk about ‘Claude Code mania.’ Developers burn through thousands of dollars a month in tokens, refusing to go to sleep without first queuing up overnight tasks for Claude to complete by morning. They believe, with reason, that they are getting the better end of the deal: the productivity gains are extraordinary. Anthropic becomes the fastest-growing company in history. + +Christian: most of our code is now written by claude + +Christian: i don't mean assisted i mean fully written end to end + +Christian: my engineers are managing it like a junior team + +Christian: except it never sleeps :D + +Caroline: Is that good for your engineers? + +Christian: for now yes + +Christian: huge productivity boost + +Christian: the labs are using it internally and their progress is accelerating + +Christian: like noticeably accelerating + +Christian: also they're shipping a new opus every three months now + +Christian: used to be every six + +Caroline reads the messages on the metro into work. By the time she gets to her desk she has decided to bring it up at her unit’s morning meeting. + +It does not go well. She mentions the productivity gains, the new revenue numbers, the compressed release cycle, the labs using their own AI to build their next models. The room is polite. Someone asks if she has peer-reviewed studies that demonstrate the productivity gains. Someone else gently observes that consumer-grade AI agents have been a disappointment all year. Caroline says the consumer agents are not the point; the self-improvement is. The room moves on. + +Her colleagues are, in any case, effectively banned by institutional policy from using Claude or ChatGPT on work devices. Standalone American AI tools are considered a data-protection risk. The Commission offers its own ‘GPT’ which is essentially a wrapper around a couple of old and small open-source models. The workforce that is meant to be regulating frontier AI systems mostly cannot use them. + +She walks back to her office immensely frustrated by her colleagues’ obstinacy, trying to remember when, exactly, she had stopped being uncertain about AI’s capabilities. + +Compute overview: 2025 + +US 17.3 GW + +Europe 1.4 GW + +1 chip icon = 3 GW 12.4× US compute advantage + +### February 2026 - A war on two fronts + +The first half of 2026 is not, realistically speaking, the ideal time for Europe to be making deliberative progress on AI policy. A lot is happening. In January, US special forces raid Caracas and abduct Nicolás Maduro. Days later Trump threatens to invade Greenland. In February, the United States and Israel bomb Iran. Oil prices surge. Trump threatens to leave NATO if European allies didn’t send ships to help reopen the Strait of Hormuz. + +Then the Department of War (DoW) goes to war with Anthropic. + +The falling out is caused by a contract dispute. Anthropic leased its models to the DoW, and over time they have become crucial to its military operations, including those in Venezuela. The contract specified two red lines: Claude cannot yet be used for fully autonomous lethal weapons, and it cannot be used for mass surveillance of US citizens. Anthropic has allegedly insisted on these terms since the very beginning. + +The Pentagon now wants to renegotiate them, and when Anthropic refuses, the dispute escalates quickly. At first, the DoW threatens to force Anthropic to provide unrestricted access through the Defense Production Act; then it does the opposite - it designates the company a supply chain risk, a definition usually reserved for companies the government considers compromised by hostile foreign powers. As stated by the Secretary of War, the designation prevents any other firm which works with the Pentagon from doing business with Anthropic, which would effectively destroy the company - after all, Anthropic’s compute providers all work with the DoW. + +Caroline: What would it take for Anthropic to move to the EU? + +Christian: they won't + +Christian: they're patriots. they want america to win. + +Christian: you could try axing copyright regulation and labour protections + +Christian: fix datacentre permitting. massive buildouts. tax breaks + +Caroline: That's not going to fly. We're not Silicon Valley. + +Christian: you can at least make it more attractive for them to expand in europe + +Caroline: People here worry they'll move to London. + +Christian: why would that be bad? + +Caroline is stuck to the news. The pieces fit together unpleasantly. The US government has woken up to AI, and is now attempting corporate murder because it can’t control the company building it. In Brussels, civil servants are banned from using frontier AI tools to draft memos; in Washington, they’re being used to plan military operations. + +Anthropic wins a preliminary injunction against the supply-chain-risk designation, but the relationship between the company and the government is now fraught. + +### March 2026 - The wild frontier + +Back in November, France, Germany and the European Commission announced the Frontier AI Initiative. The plan was to establish, by the end of Q1 2026, the world’s best-funded not-for-profit AI research organisation. Europe, the announcement said, has all the talent and data it needs, and is building compute through its Gigafactories. What it lacks is an initiative to bring these advantages together. Insiders were excited. + +But when Q1 ends, there is no nonprofit. Building a world-class research institute is hard, especially when you’re battered by one geopolitical crisis after another. What doesn’t help either is that people advising the initiative give wildly conflicting advice: some say LLMs are a dead end when it comes to reaching AGI, and the nonprofit should try to find a new paradigm. Others want to find a particular niche, whether AI for science or AI for industry. Still others say something about preparing for the quantum-computing future. Among the group of people the Commission has decided to consult, consensus is nowhere to be found. + +Money is also short. France has huge debt problems; the European Commission’s new budget cycle is still two years away; Germany has money but has just agreed to build Europe’s strongest conventional army. + +If the Frontier AI Initiative isn’t well-funded, it will be unable to offer the compute or the compensation packages that would compete with US labs, so it’s hard to hire world-class talent. Without world-class talent, it’s difficult to secure more money. Rumours spread that the Frontier AI Initiative is stuck in a loop familiar to European tech ambitions: institutions are wary to commit until it works, and it cannot work without commitment. + +In March, OpenAI raises $122 billion in a single round. + +Caroline: Do you know anyone who’d be a good lead for the Frontier initiative? + +Christian: on a brussels salary? + +Caroline: people seem open to exceptions + +Christian: let me ask around + +Christian: you know this only works if you scrap all red tape right + +Christian: like the uk did + +Christian: btw did you see the openai number? + +Caroline: Yes. + +Christian: that's more than every european ai company combined ever raised + +Christian: in a single round + +### April 2026 - Mythos + +Anthropic announces Claude Mythos, the most capable AI model to date. + +Anthropic’s internal testing has revealed that the model is so proficient at coding and security research that it has helped identify thousands of previously unknown vulnerabilities across every major operating system and web browser. Some of these are decades old, sitting undetected in codebases reviewed by thousands of developers and security researchers over many years. Mythos found them in weeks. + +Almost overnight, Anthropic has become one of the most capable cyber organisations on the planet. It decides Mythos cannot be released until cyber defenders have had a chance to patch remaining holes in their software. + +The company launches Project Glasswing: a defensive coalition giving exclusive access to Mythos to a small group of partners – including AWS, Apple, Google, Microsoft, Nvidia, CrowdStrike – to find vulnerabilities before adversaries can exploit them. The aim is to make American software infrastructure secure before open models reach comparable offensive capability, which Anthropic estimates is six to twelve months away. + +The UK’s AI Security Institute - which has built up state-of-the-art AI expertise in government - is invited for testing, but no European firm or government is provided access. + +Christian: you've heard about glasswing i assume + +Caroline: Yes. + +Christian: there isn't a single european company on it + +Caroline: I'm aware. + +Christian: the entire continent's software stack is going to be insecure relative to american AI capabilities + +Christian: does anyone in brussels actually understand what that means + +Many Europeans are sceptical that Mythos is as powerful as claimed. ‘They’re hyping their products,’ scoffs a European Commission agriculture lawyer sitting at the next table from Caroline in the cafeteria. ‘Oh no, we’ve built a nuclear weapon. But don’t worry, we can sell you a bomb shelter!’ When Caroline objects, the lawyer points out that it’s impossible to verify Anthropic’s claims if they don’t release the model. ‘How convenient!’ + +Caroline can barely hold her frustration. Why would all these large companies publicly associate with Project Glasswing if the whole thing is a scam? How did Mozilla identify twenty times more vulnerabilities this month than normally? She works through the night to write a memo on the strategic implications of European exclusion from Glasswing, and manages to land an emergency briefing with her director, who proves much more receptive. + +Over the next weeks, the EU and Member State governments urge Anthropic to give them access - leaders understand a threat to national security when they see one. Anthropic is cautious but open to talk. But when the company eventually announces it wants to expand Mythos access to seventy additional organisations - many of which are European - the White House objects. It is not secure, they say. Besides, Anthropic’s compute is finite; more partners means less capacity for US government workloads. + +A month earlier, the US administration had wanted to wipe Anthropic off the face of the earth. Now it wants the company’s most powerful model kept largely for itself. It finds a workaround to let agencies access Mythos through Glasswing while the supply chain designation is technically still in force. Suddenly, the White House decides it can get along with Anthropic after all. The president still calls Anthropic ‘woke’, but now also ‘very smart’. + +### June 2026 - Disillusionment + +After eighteen months of waiting, DeepSeek has released its next model, V4. It is impressive given the company’s limited resources but still at least six months behind the American frontier. DeepSeek admits that it lacks the compute to serve the model widely. The company that briefly persuaded Europe that compute did not matter remains as compute-constrained as ever. + +Europe is struggling even more than China. Mistral has fallen further behind. It has managed to raise another €830 million, in a round 150 times smaller than OpenAI’s most recent one. Desperate for more capital, Mistral has begun talking to American investors. A rumour breaks that SpaceX, Elon Musk’s rocket and AI conglomerate, is exploring an acquisition. + +Christian: mistral really might sell to spacex + +Caroline: The French government would never let that happen. + +Christian: they need access to vast amounts of compute and capital though + +Christian: i think they either sell or continue their consulting pivot + +Christian: either way, it’s the end of european frontier ai + +Christian: like that's it. there isn't another one + +The infrastructure story is just as grim. The largest AI supercomputer in the US runs at 1,250 megawatts. The largest in Europe runs at eighty-three. The Gigafactories plan is far behind schedule, with expected operation pushed back to 2029. Budget constraints mean that it will be scaled down in ambition. Even if a Gigafactory were to drop magically out of the sky tomorrow, fully built, it would be around a tenth the size of America’s largest AI data centre. Europe hosts only five per cent of the world’s total AI compute. The United States hosts eighty. + +Sure, there are bits of good news too: Softbank, the same company that stood on stage with President Trump, has promised to invest $45 billion over the next five years to build AI data centre capacity in France. But Caroline has learned to take such promises with a grain of salt. Fluidstack, a cloud provider that planned to build a gigawatt-scale data centre near Paris, recently dropped the idea and, to add insult to injury, moved its headquarters from London to the US. OpenAI similarly pulled back on its own plans to build a large data centre in the UK, citing regulatory hurdles. + +When, in June, the EU announces its long-awaited tech sovereignty package, Caroline feels torn about the fanfare. On the one hand, the diagnosis is correct: Europe really has become too dependent on American tech. One of her suggested solutions has made the package too, in the form of a proposal for designated zones where industry can build AI data centres at an accelerated pace. On the other hand, the numbers do not add up. The EU’s stated goal is to attract €200 billion in private capital for AI data centres by 2036 - that’s a quarter of what the American hyperscalers are spending in 2026 alone. When she explains the situation to a friend, he returns: ‘So you’re proposing a decade-long programme the size of their quarterly capex and are calling it a historic mobilisation?’ + +After extensive negotiations, a select group of Europeans has by now been promised access to Mythos. But while legal departments are still fleshing out contract details, Anthropic is already training Mythos 2.0. Worse still, the US president just signed an Executive Order that asks frontier AI companies to ‘voluntarily’ allow the US government to screen their models for cyber risks, and which lets government agencies access those models before any public release. The details of this national security review are scant, but the structure is plain enough. Any model powerful enough to matter – a ‘covered frontier model,’ in the EO’s language – now passes through Washington first: the government gets up to thirty days of exclusive federal access before release, and a say in which ‘trusted partners’ can use the model before a full public release. + +Caroline reads it twice. The thirty-day window is survivable, but the partner-selection power means that whoever gets the model while it is still ahead of the open frontier is now a decision made in Washington, through a classified process, with no particular reason to put a European name on the list. + +Christian: did you see the executive order? + +Christian: this is glasswing but for every next-gen model from now on + +Christian: there's no clause about allies. not one + +Caroline: I know. + +Christian: so whether europe gets early access to frontier models + +Christian: is now a discretionary call by the NSA + +Christian: every three months, forever + +Meanwhile, Anthropic and OpenAI continue to see exponential revenue growth; the former is on course to reach $100 billion annualised by year-end, and has signed emergency compute deals to serve its rapidly increasing customer base. At the AI labs, internal workflows now change every month or so, as the models enable new ways of working. Researchers manage swarms of agents that handle the bulk of their coding. AI is contributing to its own R&D, making itself better, and using those gains to get better at getting better. A few months earlier, the US Federal Reserve released research showing employment levels among junior software engineers are significantly below expectations. Software company stocks are plunging, as investors assume powerful coding AI models will eat their lunch. + +Caroline has seen much of this coming for eighteen months, but has not been able to make any meaningful difference. + +Compute buildout + +US 17.3 GW + +Europe 1.4 GW + +1 chip icon = 3 GW + +12.4× US compute advantage + +## The Scenario + +*Up to this point, everything we've said has happened – with only Caroline’s and Christian’s personal stories representing fictional elements. From here on out, we start speculating. We no longer single out individual AI companies, and instead refer to made-up actors: Atlas for the leading American AI company, Helios for the leading European company, and Zimo for the leading Chinese one.* + +![The twelve EU stars, half dissolving into pixel dust.](/assets/stars-red.webp) + +[Prologue](#chap-jan-2025) [Aug 2026](#drift-chap-aug-2026) [Sep 2026](#drift-chap-sep-2026) [Jun 2027](#drift-chap-jun-2027) [Jan 2028](#drift-chap-jan-2028) [May 2028](#drift-chap-may-2028) [Aug 2028](#drift-chap-aug-2028) [Nov 2028](#drift-chap-nov-2028) [Jan 2029](#drift-chap-jan-2029) [Apr 2029](#drift-chap-apr-2029) [May 2029](#drift-chap-may-2029) [Feb 2030](#drift-chap-feb-2030) [Jun 2030](#drift-chap-jun-2030) [Aug 2030](#drift-chap-aug-2030) [Mar 2031](#drift-chap-mar-2031) [Epilogue](#epilogue) + +### August 2026 - The fork + +Back in February, the German chancellor made a state visit to China. In Hangzhou he toured Chinese EV and robotics factories. Insiders reported that he returned sobered, with a new sense of urgency, and increased conviction that German manufacturers will fall behind their Chinese competitors if they don’t pick up the pace. Now he schedules a similar trip to San Francisco. Over the last months, people he respects - German business leaders and economists, American bankers - have told him the AI revolution is for real. He wants to see it with his own eyes. An industry delegation joins him. + +The chancellor is not the kind of man who is easily impressed, but again he returns from his trip quieter than he left. He has come to realise that European elites have misread the past eighteen months. While they debated whether AI would hit a wall, the technology progressed faster than even the bulls predicted. While they weighed AI’s importance, it revolutionised software engineering and cybersecurity. And while they hailed their sovereign AI initiatives, their dependence on American providers only deepened. + +Over the following days, he holds a series of long phone calls with the French president and the president of the European Commission. All three feel Europe has reached an inflection point. AI will keep getting more capable. There is no good reason to believe it will stop at human level. It will rearrange labour markets, security architectures, and the balance of power between continents. Europe is dependent and unprepared. If the ship is to be turned, it must be turned now. + +The question is how. Their technical advisers plead for a full regulatory carte blanche for data centre providers and other companies in the hardware supply chain. Only an extreme response is proportional to the stakes, they say. AI will spur a new industrial revolution, and if Europe doesn’t rapidly industrialise, it will get left behind. Europe needs to move faster than anything it has ever attempted in peacetime. + +Their political advisers beg for something more restrained, warning that their governments might not survive the public backlash against such a package. The people don’t like AI; France, Germany and the EU face many other important challenges, and only a small fraction of voters will understand the necessity. The ‘effective response,’ they say, may be most effective at ending political careers. + +### September 2026 - A positive vision + +At the Franco-German AI Sovereignty Summit in Strasbourg, the French president and German chancellor give a joint address. It is written by their political advisers. + +Caroline watches at home in Lille. The two leaders speak of resolve and of destiny. Europe, they say, can no longer afford to depend on others for the technologies that will define its future, a lesson it has already learned in energy and defense. So, Europe must build its own frontier AI. + +The challenge is huge but surmountable, they say. What is required is investment, rules that force American providers to play fair, and a public that buys European. The speech has excellent slogans. + +Over the following weeks, the AI announcements come almost daily. The plan is to overwhelm the sceptics and imbue a sense of hope. The Frontier AI Initiative gets properly funded with a €2 billion capital injection. Four more Gigafactories are announced. New AI adoption programmes are launched. The European Commission charges an American provider of general-purpose AI models for failing to comply with the AI Act. It also opens two systemic-risk proceedings under the Digital Services Act over their handling of AI-generated disinformation, invoking the law’s most open-ended provisions. + +The headline measure, the one Caroline’s colleagues talk about over coffee, is a Commission proposal – endorsed by France and Germany – for a Digital Sovereignty Regulation requiring critical public sector workloads to run on 100% European cloud and AI software by 2032. No more American AI; no American clouds. It is modelled on the European climate targets, and would create the kind of binding deadline that focuses minds. It would also establish a large, guaranteed pool of future customers for European providers. + +The measures are well-received. Decoupling from the US is popular, and commentators hail the move as industrial policy for the AI age; Europe is at last backing its own technology companies. Even Caroline’s more sceptical colleagues concede that something is finally moving. + +Some economists and tech policy people express doubts. They say that Europe can no longer avoid ugly trade-offs. The investments are too small and don’t create enough leverage – where are the market incentives to build compute at scale on European soil? The procurement targets, they say, do not address fundamental issues, such as the fragmented single market, inflexible labour laws that harm startups and AI adoption, or the national rules that functionally close off sectors like healthcare and legal services. And while regulatory actions under the AI Act are well-justified, those under the Digital Services Act appear at least in part politically motivated. Shouldn’t Europe pick its battles more deliberately? + +But European elites are tired of negativity, and this is a positive vision, a rally-round-the-flag moment. If policymakers have doubts, most of them are left unsaid. + +Caroline certainly has doubts. She is worried that the whole AI sovereignty package assumes that, in six years, Europe will still have a frontier AI sector worth protecting. What happens if it doesn’t? + +She writes a memo making that point, calling for measures that build leverage - a contingency plan in case European AI champion Helios cannot close the gap, or the Gigafactories fall short, or the buy-European mandates leave civil services with worse tools than the American ones. Her director reads the note carefully. It is, he tells her, a thoughtful contribution. He promises to pass it up the chain. + +Christian: 2032? 100% sovereign? + +Christian: what's the plan b? + +Caroline: There isn't one. + +Christian: of course there isn't + +In private, the German chancellor and French president tell each other they did what was possible. Even they cannot smash through so many political realities all at once. The system would not have absorbed more. + +Compute overview: 2026 + +US 35.3 GW + +Europe 2.1 GW + +1 chip icon = 3 GW 16.8× US compute advantage + +### June 2027 - Closing windows + +Over the next summer, the Chinese lab Zimo releases the weights of a Mythos-class AI model. The offensive cyber capabilities that Anthropic had locked behind Project Glasswing are now available to anyone. Some of those people are not entirely well-intentioned. + +The new capabilities let hackers cut through old, ill-defended commercial software like a chainsaw through tissue paper. European universities, hospitals, regional governments, any institution which hasn’t paid for Mythos-class cyberdefence, find themselves locked out of their own systems, with cryptocurrency wallet addresses on the screens and no remaining option except to pay. Only those who use AI for cyberdefence have a chance, which makes people furious at the labs: they unleashed the disease and now sell the cure, often at high rates. + +Europe is learning the effect of its sovereignty attempt in the most direct way possible. The Digital Sovereignty Regulation has just passed and several Member States have begun adjusting procurement well in advance. The American AI leader Atlas is miles ahead of Europe’s AI champion Helios in cyberdefence. Those agencies most committed to the ‘Buy European’ agenda – organisations which have procured exclusively from European providers – are now the ones paying the ransoms. + +Organisations that kept an American contract on the side are doing better. But even they are running second-tier defences. Ever since the Executive Order, the US government informally greenlights the release of each new frontier model. The result is that the best American cyber models reach Europe two to six months after their domestic release, dangerously close to the open-source frontier, and well after American hackers gain access. Officially, this is about safety and oversight. Unofficially, Washington has discovered an asymmetric advantage it has no interest in surrendering. + +Caroline spends most of June on calls with national governments. The conversations are short and unpleasant. + +Christian: how's brussels handling the ransomware wave + +Caroline: Badly. The sovereignty policies are biting. + +Christian: the irony + +Caroline: It's not funny from here. + +Christian: sorry. you're right + +Just as the situation is becoming untenable in Europe, both the US and China announce stringent restrictions on open-sourcing frontier AI models. Washington cites national security, pointing to adversary R&D acceleration and the proliferation of autonomous cyber capability. Beijing speaks of social stability and orderliness. The two governments have, for once, arrived at the same conclusion: frontier model weights have become too dangerous to give away to anyone. + +Walking out from the office on the evening of the Chinese announcement, Caroline notices a positive mood around her. The EU’s Tech Sovereignty Package had hailed open source for its counterweight to American AI dominance. And yet her colleagues are relieved. The ransomware wave will slow. Offence will plateau; defence will catch up. The crisis, her colleagues agree, has been contained. Few people seem to be thinking anymore about what the open source ban means for Europe’s growing dependencies. + +Compute overview: 2027 + +US 61.3 GW + +Europe 5.2 GW + +1 chip icon = 3 GW 11.8× US compute advantage + +### January 2028 - Taking stock + +Sixteen months after Strasbourg, Caroline has been promoted and moved to a different team. + +AI progress is even faster than in 2026, just as the tech CEOs predicted. Agents now run spreadsheets, design software, operate financial tools. One model generates an image; another opens Photoshop and clicks through fifty iterations until the composition is right. Caroline sits through a demo of an agent navigating enterprise software at twice human speed without a single misstep, and finds it unsettling to watch. The leading US labs are well on their way to automating AI research itself; Atlas has gone public, its market capitalisation hits four trillion euros, and by now everyone agrees AI is the new big thing. + +The race has also thinned out - a result of exploding capital requirements and AI R&D flywheels. The American AI companies in fourth and fifth position have fallen further behind. One is considering folding its frontier research into a partnership with the leading lab; another has spun up a parallel cloud service and is attempting to manufacture AI chips. + +Helios, riding the European sovereignty wave, has impressively held its position roughly one-and-a-half years behind the American frontier. It has raised billions from new European investors, and has kept talent from defecting. + +The relative gap is unchanged, which makes it a kind of victory. But the absolute gap is a different story. In 2023, eighteen months was one generation. Now that capability progress has accelerated and iteration cycles have shortened, it is many. And AI systems are no longer mere chatbots: they are running drug discovery campaigns, are automating mathematics and cybersecurity, and performing large chunks of many white-collar workers’ jobs. + +An old friend of Caroline’s, a software engineer at a French retail bank, tells her over a glass of wine that his employer has a sovereign AI policy. He is, partially, permitted to use Atlas. He is not permitted to upload anything substantive – not even to name the columns in a dataset. He could use Helios, but it is worse. So in practice, he runs Atlas on his personal laptop, and copies the files across. + +The pattern repeats inside most large European companies. Everyone knows this is happening, but no one says it out loud. + +And yet, some Brussels politicians have retained hope. + +The Frontier AI Initiative is running at full steam. France and Germany have thrown a lot of political will behind the non-profit. They have decided to offer competitive compensation, and have managed to secure serious talent, some of which has experience at American labs. One of the Initiative’s research bets – world models that can be used to train robots – is producing genuinely promising work that is getting international attention. + +The 2028 EU budget has unlocked substantial new funding for AI in science: medicine, materials, clean energy, areas where Europe could still plausibly win. Adoption pilots, after a rocky start, are returning positive results. Doctors are more productive. Teachers are reporting better outcomes for their students. + +Warnings of employment catastrophe are yet to materialise as well. Even in sectors with high adoption, AI is not significantly reducing jobs. And workers who can orchestrate AI systems – consultants, lawyers, software engineers, analysts, designers – are seeing their output and wages rise. Judgement, client relationships, and accountability become more important as AI takes over the gruntwork. + +A European PM concedes in an interview that Europe was slow to start, but insists it is now catching up, and that European AI is spurring a new productivity boom. Caroline shares it with Christian when she gets home. + +Caroline: Maybe my own boss isn’t so bad after all. + +Christian: enjoy the productivity boom + +Christian: it's a lagging indicator + +### May 2028 - The compute crunch + +The world is now screaming for more compute. + +Workarounds have been found for every previous bottleneck in the AI chip supply chain. When fabs were in short supply, TSMC and Samsung built more. When US grids could not handle additional data centres, the labs bought mobile gas engines and generated their own power on site. When high-bandwidth memory supplies got tight, capacity was diverted from consumer products – with the effect that smartphone prices spiked, but that was very much someone else’s problem. But this bottleneck is different. + +ASML, the Dutch semiconductor giant, is the only company that can manufacture the EUV machines TSMC uses to fab AI chips. Output has risen from sixty to eighty-five machines a year – a real improvement, but wholly insufficient against skyrocketing AI demand. ASML has more than 5,000 suppliers upstream. If any of them misses targets, the whole chain slows. The engineering is so complex that improving a single one of the machines’s 100,000 parts requires a PhD and years of specialised experience. You cannot conjure such people out of thin air. + +Washington does not love the fact that ASML still exports some of its older immersion lithography machines - called ‘DUVi’ by the industry - to select Chinese companies, which use them to create China’s somewhat less advanced AI chips. + +After briefly seeing eye-to-eye on open source, US-China relations have again soured over the last year. Washington is worried that China will outbuild it once it completes its domestic semiconductor industry. China is adding immense amounts of energy at a staggering pace, whereas American electricity production has flatlined in comparison. If AI becomes a game of building the most data centres, China will win in the long run. Beijing, for its part, worries that Washington will achieve an irrecoverable military AI lead before China has matured its semiconductor ecosystem. If so, the US could use its AI advantage to extract geopolitical concessions. Those worried in Beijing are not eased by some American AI CEOs talking openly about using military AI to ‘democratise’ autocratic regimes. + +Washington has come to realise that it has a short window to act. Its lead in AI is largely the result of clever export controls instated years ago, on Dutch EUV machines and American AI chips. Now, it doubles down. It piles pressure on The Hague to halt ASML’s remaining exports and servicing of its DUVi machines. It is a big escalation, since China uses the same machines to make domestic chips for everyday goods like smartphones and laptops. + +The Dutch resist and seek support from other EU Member States. They understand Washington’s reasoning but do not want to be bossed around, let alone get dragged into a great-power conflict by an administration that has spent the past two years bullying Europe. + +The Commission backs the Dutch, but several Member States are terrified of American retaliation if Europe keeps exporting the machines. The European position fragments before it has properly formed. + +Feeling betrayed, the Dutch turn to Japan and South Korea - countries with similarly important positions in the hardware supply chain that also face American pressure - to coordinate a joint response. But both capitals remain heavily dependent on American military support and have, for good measure, been visited by senior American officials in the previous fortnight. They are politely unavailable. + +When the Dutch hold firm, Washington threatens them. It can invoke the Foreign Direct Product Rule, a rule that lets Washington assert jurisdiction over any product made anywhere in the world, by anyone, so long as it was produced using American technology or software. It is the same instrument Washington used to strangle Huawei in 2020. ASML’s machines qualify several times over. Continue to sell any DUVi to China, and ASML would be in violation of American law. The company could face crippling penalties – denial of export privileges, civil fines, and, in theory, criminal liability for individual executives. Such penalties would also harm the US, which is dependent on ASML’s machines. But ASML cannot afford to call Washington’s bluff. The Dutch cave. + +In Brussels, Caroline reads the news on her phone between meetings. She thinks of her 2026 memo about the need to build leverage. She thinks of her director’s friendly face when he told her it was a thoughtful contribution. + +Christian: this is the warning shot + +Christian: tell me they're treating it like one + +Caroline: My direct colleagues are. + +Caroline: Others are calling it a setback. + +Christian: ok + +Christian: noted + +Caroline: I mean, I get why the US did it. + +Caroline: But we didn’t even manage to negotiate a single thing in return. + +Christian: it’s fucked up + +### August 2028 - Writing on the wall + +AI models no longer think in English. + +Instead of writing thoughts onto a digital scratchpad – the system used since early 2025, and which humans can read – the new systems cycle through long lists of numbers, called high-dimensional vectors, which nobody, not even other AI models, can properly interpret. Freed from the need to translate their complex thoughts into English, the systems think faster and more deeply. It creates a rapid jump in intelligence and capability. + +Safety researchers who follow the work are alarmed – many of their strategies for controlling models relied directly on these scratchpads being available. How will we know the models aren’t secretly pursuing their own goals? How will the misuse monitors catch worrying behaviour if they can’t read the reasoning? In Brussels, experts want the EU AI Office to force developers back to scratchpad-based systems, or to present other evidence they understand what is happening under the hood. But the AI Office is already locked in heated proceedings against two American developers. The transatlantic relationship cannot take on more. + +For most people, the more immediate effects are on capabilities. Models which previously could not reliably finish multi-day research projects now can. In the US, employers who had been holding back from job cuts now pull the trigger. Entry-level hiring further stagnates. Unemployment accelerates. + +Europe sees less labour pressure, but also less growth: its economy is growing at 1.6% while the US is posting 3.8%. The gap is undeniable and widely attributed to differences in AI value capture. Europe has access to mostly the same models, but cannot extract the same economic gains. There are three reasons. + +The first is ownership. The AI companies and infrastructure providers whose revenues are skyrocketing are all based in the US. European AI-native startups exist, but it is American venture capital that delivers the scale-up rounds, and the fast-growing ones increasingly move overseas. + +The second is uptake. Despite Europe’s adoption pilots, American firms have moved more aggressively to integrate frontier AI into their workflows. Some European companies are held back by fragmented rules, others by a risk-averse management culture or by internal policies that force the use of inferior home-grown alternatives. A law firm in Milan that once charged premium rates for its knowledge of Italian commercial law now competes with a US firm whose AI handles Italian, French, and German jurisdictions at once, faster and cheaper. The Italian firm’s lawyers are still blocked from using frontier models. The same pattern plays out in consulting, software, marketing, and finance. + +The third is what happens to the firms that do adopt. Many mid-sized American companies restructure around AI in months – flatter org charts, smaller teams, faster cycles – while European ones often take years. Works councils slow the deep adoption of powerful AI tools; employment protections make it hard to let go of staff whose jobs can be automated and whose labour force would be needed in parts of the labour market that faces shortages. While some workers become vastly more productive with the help of AI, a small but growing number of European workers engage in pretend-work: they log in, attend meetings, but otherwise let agents do most of the work in a fraction of the time. This is not without benefits: it means extra family time, long lunches, and afternoon walks. The flipside is that firms are paying for both AI and unproductive human workers, and capital that should be funding growth remains stuck maintaining the existing workforce. + +This comes at a time at which the European economy is already struggling - its core manufacturing industries are increasingly moving overseas, with its once-lauded automakers among the worst hit. Having missed the boom in electric vehicles, they face immense pressure from cheaper and better Chinese cars, leaving millions of workers with uncertain futures. + +And now the European tax base is beginning to erode too. Money that used to flow to labour increasingly flows to American companies and their investors, much of it routed through low-tax jurisdictions European treasuries cannot reach. At the same time, unemployment claims tick up across the continent – not dramatically, but the US numbers suggest where things are heading. And the AI boom is driving up global interest rates, leaving countries like France to spend more of each year’s budget paying for its state debt. + +Caroline: I think I finally understand you. + +Caroline: The only real remedy for Europe's doom spiral is economic growth. + +Caroline: But the growth is happening in the US. + +Christian: right + +Caroline: I just bumped into an old classmate of mine. + +Caroline: She works 3 hours a day. + +Caroline: Lets her agents do the rest. + +Caroline: Apparently, she has taken up gardening. + +Christian: good for her i guess. + +### November 2028 - Vox populi + +It is the US presidential election. The world, as always, holds its breath. + +AI was the campaign’s defining issue, but not in the way the industry would have hoped. By now, most Americans want little to do with AI. Climate groups worry about data centres’ energy use; unions worry about jobs. Teachers want it out of classrooms; lawyers want it out of courtrooms. + +During the primaries, populists on both left and right ran on strong anti-AI platforms. They called for data centre moratoriums, worker protections, bans on AI in education and age restrictions. The general election produced more centrist candidates - strong industry ties helped support the presidential campaigns - but the sentiment did not go away. People are angry. + +Anti-AI views span the political spectrum, but it does not bring people together. Fractures have risen, not just along the Democrat-Republican line but between elites who use AI and middle classes who find it frightening, dehumanising, or immoral, and who are watching inequality grow in real-time. People increasingly take to the streets; multiple tech company CEOs have faced assassination attempts. A data centre was recently set on fire. But the US government cannot crack down on AI; the new president, just like the last one, is convinced that the US has to win the AI race from China or face unacceptable threats to national security. So instead of giving the public what it wants, he issues some duct-tape solutions to try and calm it. + +Anti-AI sentiment spreads to Europe, too. People are furious at US tech companies and want governments to take action. They call for stronger safety nets, not realising that Europe can barely afford the ones it already has. + +Meanwhile, Europe’s AI companies are falling further behind. + +The gap between Helios and Atlas has widened further, despite the public investments, the compute subsidies, the preferential procurement contracts. American labs, powered by swarms of internal AI agents writing most of their own code, are making algorithmic progress at more than twice their human-only speed. Only compute constraints are preventing even faster progress, and Atlas has more compute than anyone in history. Helios’ research multiplier, with a fraction of the processing power and without access to the best US models for internal use, is barely registering. Every month, the gap grows wider. + +Public efforts are delivering public goods, but aren’t meaningfully boosting European sovereignty. The Frontier AI Initiative has made impressive progress on its interpretability agenda, which benefits AI reliability globally, but its world-model programme has been hijacked. As soon as results started picking up, Atlas took notice of the Initiative’s progress, rapidly spun up its own team, and managed to poach four of the Frontier AI Initiative’s best researchers in exchange for astronomical compute budgets. The researchers wanted to stay, they tell colleagues on the way out; they believed in the European project and wanted to help. But at some point, belief is not enough. + +The public can see that the EU’s grand AI strategy is failing. Investments have not allowed champions such as Helios to catch up. Regulatory actions under the DSA have not levelled the playing field; their most visible contribution has been to annoy the Americans and worsen EU-US relations. + +But Europe has invested vast political capital – and actual capital – in the project, and admitting failure would mean admitting that it has spent two years and tens of billions of euros on a dead end. So it doubles down. + +A new €20 billion European AI Sovereignty Fund is announced by the European Commission, focused on photonics, edge AI, and other ‘next-wave’ paradigms. It is clearly a long shot. The same institutions that failed to turn money into frontier capability are being asked to try again, with a harder target and less time. It is, a Polish MEP tells Caroline at a reception, the first known case of raising your bet with no cards in your hand. + +Cracks appear. In Germany, a populist party running on an explicitly anti-AI platform – *Stop the machines, save German jobs* – is polling first in the upcoming federal election. In Italy, where Eurosceptic governments have been the norm for half a decade, populist parties openly campaign on a referendum on EU membership. + +In Paris, patience runs out. The Élysée no longer believes the EU’s plan will deliver on the timescales required. Helios, the only real remaining European player in LLMs, may be months away from being permanently outcompeted. After intensive negotiations, France announces a €15 billion state investment in Helios for a 17% equity stake, a board seat, and veto power over future fundraising rounds. + +Christian: france just bought a corpse + +Caroline: That's not fair. + +Caroline: These people are actually trying to do something. + +Christian: doesn’t change the facts. + +Compute overview: 2028 + +US 104.4 GW + +Europe 11.8 GW + +1 chip icon = 3 GW 8.8× US compute advantage + +### January 2029 - Dreams of a robot economy + +China’s AI strategy is not unlike the EU’s, except in one crucial aspect: it appears to be working. Like Brussels, Beijing is backing its most promising companies, protecting the industry with subsidies and procurement mandates, and pushing for rapid adoption. But it is easier for a centralised, authoritarian state to force its own agenda than it is for twenty-seven disputatious Member States in a liberal democracy. When its compute was fragmented between different labs, the Chinese Communist Party simply ordered those labs to pool resources; the European Commission has no such powers. China’s talent pool is deeper and it has access to cheap, abundant energy. Chinese frontier labs remain within a year behind the US. + +Importantly, China does not believe its strategy relies on being at the bleeding-edge frontier of cognitive AI. The government has always wanted AI to boost the physical economy first, and holds a huge advantage in robotics manufacture. Massive state subsidies have pushed annual humanoid production above 1 million units; a household robot can now be bought for €10,000. In cities like Shenzhen, it is commonplace to see humanoids cleaning streets, or packages being delivered by quadrupeds. Beijing is betting that it only needs to stay close on Silicon Valley’s heels in AI for its industrial advantage to start to pay off as more robotics-relevant capabilities emerge. + +The bet is still looking viable, and in the US, concern is mounting. American politicians have started to refer to the AI race with China as the second Cold War. Tensions have risen further since the US forced ASML to halt exports to China, and Beijing and Washington have both since beefed up the security around their AI development. In the US, researchers are screened by intelligence agencies and need security clearance to access frontier model weights; the federal government regularly accuses China of stealing algorithmic secrets. China has semi-nationalised Zimo and is throwing its considerable weight behind its AI efforts. When part of the Chinese power grid goes down near a major Zimo data centre, rumours spread that it was the result of a US-backed, AI-powered cyberattack. + +Christian: im in shenzhen. + +Christian: a robot wearing a sailor hat just made me a negroni + +Christian: they’re making robots like paperclips here + +Caroline: Was the negroni any good? + +Christian: yeah. excellent + +### April 2029 - Access all areas + +AI demand is soaring, and compute supply cannot keep up. The labs are rationing access to their frontier models and raising prices; businesses are screaming for tokens. + +American firms operate 70 per cent of the world’s AI compute, and sell their services around the globe. This means that American infrastructure is being used to make foreign businesses more productive, foreign militaries more capable, and foreign labs more competitive – although that last one is via illegal distillation attacks, not legitimate sales. Washington is getting more concerned about it by the day. + +The American national security review has been formalised and no longer pretends to be voluntary. Access to the most capable models is throttled by default, in part to save compute for American customers. Domestic government agencies get them first; allied governments next; adversary nations not at all. + +But access to the models and their weights is one thing. The compute shortage means that inference is also at a premium, and by April, Washington has had enough. It begins not only limiting access but also rationing use, even for those countries inside the tent. + +The *Frontier Inference Services Rule* (FISR) is a country-based licensing regime. Tier 1 countries – close allies, such as the Anglophone ‘Five Eyes’ intelligence-sharing nations, as well as Japan, South Korea, Taiwan and the Netherlands – get unrestricted commercial access and light-touch reporting. Tier 3, hostile countries such as Iran, Russia, and China, are by default denied any access at all. Most of Europe is in the 100 or so Tier 2 countries in between. FISR dictates that no more than 25 per cent of any provider’s frontier inference can go to Tier 2 customers in aggregate, with individual licenses reviewed against a list of factors including ‘alignment with United States national security interests.’ + +For Brussels, the 25 per cent is the number that matters, and it is bad news. European customers currently account for nearly a quarter of US frontier inference on their own. Fairly sharing it with eighty or so other countries would mean roughly halving European allocations. Within a week, European business customers without long-term contracts get notices from their providers, cutting volumes and raising prices. They cannot change suppliers to another lab, because every US provider is under the same restrictions. + +Europe depends on American AI, but the US has no such dependency on Europe. Compute is now so scarce that losing European customers has no discernible effect on the labs’ bottom lines: they were already capacity-constrained, so the inference they are no longer selling to Europe can be redirected to latent American demand or simply their own internal R&D. Atlas and its competitors privately urge the White House to soften its stance - the restrictions risk closing off future markets and are horrible PR. But they are unwilling to express their concerns publicly, because as national security restrictions around AI increase, the more important good relations with the federal government become. Besides, they are focused on winning the capabilities race and keeping control of AI systems that are rapidly becoming cleverer than the people running them. There is little pressure on Washington to shift its stance. + +In Brussels, an emergency European Council meeting is convened. French and German ministers head to Washington to demand Tier 1 designation. They are told Tier 2 better reflects ‘the current state of the bilateral relationship.’ + +Caroline reads the readout on the metro home. She had always known this moment would come. It still lands like a dropped weight. + +Caroline: They understand now. It's just too late. + +Christian: yeah + +Christian: i was afraid of that + +Christian: how long before they threaten to cut all access + +Caroline: I was thinking the same. + +### May 2029 - Tilt + +A week on from the imposition of the inference cap, the phones in European capitals are ringing as panicking businesses try to put off the inevitable. A French utility CEO calls the Élysée to say his cybersecurity teams are already losing ground to AI-enabled attacks, and that halving access to American frontier models would leave critical infrastructure dangerously exposed. + +The head of a Danish logistics giant tells the country’s prime minister that years of optimisation have left the company dependent on US systems it now cannot easily replace, and that their entire business model is at risk. A German delegation of smaller businesses warns the Chancellery that price rises alone will prevent thousands of small firms from using frontier AI at all. + +European companies that were smart enough to postpone or walk back their sovereign AI policies spent years building their operations around frontier agents. Now they face having those agents pulled out from under them. The European alternatives are almost two years behind. The Chinese alternatives are not really an option for anyone with a compliance department. + +In Brussels, at long last, no one argues that AI is overhyped. Caroline has not heard the term ‘bubble’ in months. Her director, who used to say she was exaggerating, now spends his days on the phone to national capitals, triaging the damage. He stops by her desk on a Tuesday afternoon with two coffees and hands her one. They drink them without saying much. It is, she realises, as close to an apology as he is able to offer. + +Everyone understands the problem, now, but understanding it and being able to fix it are very different things. Demand for Helios’s models has skyrocketed far beyond the company’s capacity. The Gigafactories are finally under construction but will not come online until next year, and even then will only be a small fraction of what is needed to close the gap. + +Christian: you know what the funny thing is + +Christian: france is going to make a killing on the Helios investment + +Christian: new demand is through the roof + +Caroline: Some 3D chess right there. + +The European economy is being strangled. After a series of tense phone calls with Washington, European leaders decide that something drastic must be done. For the first time in its history, the Commission opens a formal examination under the Anti-Coercion Instrument – the *trade bazooka*, as the Brussels press has been calling it for years, back when it was a deterrent nobody really expected to fire. + +After four months of assessment, the conclusion is that FISR constitutes economic coercion. But the assessment also reaches an uncomfortable finding: the obvious retaliation is self-defeating. Tariffs on US cloud and digital services would raise the price of the frontier inference European businesses are already scrambling to secure. Excluding American providers from public procurement, the threat that would have stung a decade ago, is moot – the Digital Sovereignty Regulation will already do it. + +So the proposed countermeasures reach for the instrument’s less obvious tools: suspending the intellectual property protections US labs hold inside the single market, and screening – or blocking – American acquisitions of European firms. They are calibrated to hit Washington’s exporters without touching Europe’s own compute budgets. + +The measure that would inflict the most pain is the one pointed at the lithography supply chain: restricting ASML’s exports and servicing to American fabs in Arizona. It is also a nuclear measure that will invite a response Europe may not be able to afford. + +When it comes to a vote, it falls short of the qualified majority needed. The Netherlands and Ireland oppose, citing transatlantic relations. Poland and the Baltics, worried about Russia, do likewise. Italy abstains. On the record, a senior Commission official tells reporters that the result reflects ‘differing national assessments of the strategic environment.’ Off the record, a French civil servant tells the same reporters that the delegates are too scared of Washington to use the weapon they had spent a decade building. + +Caroline spends most of the summer in crisis meetings. The same people who were telling her a few years ago that everything would be fine are now asking her whether anything can still be done. She tells them that the window for meaningful action closed some time between 2025 and 2027. All that remains now is damage control. + +Compute overview: 2029 + +US 146.6 GW + +Europe 12.8 GW + +1 chip icon = 3 GW 11.5× US compute advantage + +### February 2030 - Labour shock + +Europe’s pretend-work honeymoon, of people going gardening while the agents did their work and their companies were barred from firing them, does not last. + +European firms, stuck paying a full human staff, cannot compete with their leaner American rivals, especially since they are hamstrung by worse models and less abundant, more expensive inference. The shock hits the most AI-exposed industries first. Software companies lose out because US ones ship faster, at a fraction of the cost. Mid-tier consultancies find their advice is easily anticipated by frontier models, and they have little to add. + +What is more, the AI systems just keep getting better. Continual learning was often called the last hurdle to true automation of cognitive labour. Human workers build up context, judgement, and tacit knowledge over the course of a career. AI systems started each new conversation with the same frozen weights. Long context windows and external memory closed part of the gap, but the model never actually *learned* anything new during deployment. + +That, however, has changed. If Atlas’ latest model spends six weeks inside a consultancy firm, it starts to write the way that firm does. It learns who defers to whom, which clients take bad news badly, which senior staff have reputations worth leaning on. The implementation is imperfect, but the failures are rare and worth the price. Cognitive jobs once thought safe, protected by context-sensitive judgement or institutional, illegible knowledge, are now exposed. + +Few people are actively let go because of AI, but the problem is that companies stuck in crisis mode are not creating new jobs. The graduate jobs market is the worst in living memory. Law firms, finance firms, accounting practices – anyone still able to afford frontier rates, or sensible enough to have signed a long-term contract – have paused or shrunk their intake. + +Caroline’s younger brother finished a master’s in logistics management last summer and has been looking for work ever since. They have dinner together in Paris once a month – she pays – and he asks her whether he should retrain. She does not know what to tell him. The nursing shortage is real, but he doesn’t like the sight of blood. He is not cut out for the trades. He nods, and orders another beer, and she lets him change the subject. + +Politicians who built careers on anti-AI platforms are vindicated. The ones who hedged are scrambling. Street protests become routine in European capitals – some demanding labour protections, some demanding bans on American AI, some demanding bans on all AI, all of them demanding that someone in a building somewhere do something. + +European intelligence services have, for over a year, been flagging coordinated information operations aimed at European audiences. The narratives are tuned to local anxieties: American tech is hollowing out the continent, Washington treats Europe as a vassal, the transatlantic relationship is a one-way street. None of it is exactly false. The Hague and Berlin publish attribution reports pointing at Beijing. The public, inclined to agree with the underlying sentiment regardless of who is amplifying it, does not care much. + +Christian: how's europe + +Caroline: People are angry, understandably. + +Christian: it's not great here either + +Christian: another data centre got firebombed last week + +Caroline: I saw. + +Christian: everyone feels like they're losing + +Christian: except the labs + +### June 2030 - An offer you can't refuse + +China is leading on robotics. But now Atlas is going all-in on it too. + +It announces plans to spend hundreds of billions on industrial data and manufacturing capacity to build robots on a scale comparable to China. Its CEO lays out his reasoning: America still leads in AI software. With help from his new world model team, he has managed to crack the last software challenges standing in the way of general-purpose robotics. But physical supply chains take years to build. If Atlas can get there ahead of its rivals, the advantage will compound: soon, robots will be building the factories that build the robots, just like the AIs are writing the code that improves the AIs. + +And there is work to do to catch up to China. While it takes two years to build a new robot factory in the US, in China it takes just seven months. A decade of data-center hyperscaling has left the US energy-constrained, and the American people resentful of everything related to AI; new plants are met with local resistance and state politicians side with their voters. The economy’s undeniably rapid growth is constrained by infrastructure, inequality and public opinion. + +So the CEO decides the fastest way forward is not to build but to buy – to retrofit factories that already exist. Backed by friendly investment funds, he goes shopping for industrial companies with usable floor space that can be turned over to producing wheeled robots, quadrupeds, and humanoids. + +Europe’s car manufacturers are high on the list. After years of Chinese competition, Germany’s largest is close to bankruptcy. Its market capitalisation is down 80 per cent from its pre-EV peak to €18 billion. To Atlas, now worth $13 trillion, it’s small change. But it has a big payoff. German labour law has forced the manufacturer to keep its now unnecessarily large workforce. Their wages can no longer be paid and shareholders have been actively looking for an exit for two years. For a company looking for high-tech manufacturing to build tens of millions of robots a year, it is an ideal opportunity. + +Berlin disagrees. It blocks the sale on national security grounds, but insiders know this is more about national pride. + +Atlas does not back down. The CEO calls the US president, who is already on record as saying that whoever wins robotics wins the economy, and that China is moving faster than Washington will tolerate. Within seventy-two hours, the White House announces that European automotive imports are hit with sky-high tariffs. Officially, the action has nothing to do with the attempt to buy the carmaker. + +Three weeks later, a sale is announced, although it is disguised as a partnership. The German state will take a 20 per cent stake in the new entity, and the existing board will retain nominal control. The press release uses the word ‘European’ eleven times. + +Behind the scenes, Atlas is in charge. It holds the operational majority, the licensing rights to the manufacturing platform, and the right of first refusal on any future capital raise. Profits run through a Delaware holding firm. Berlin saves face, but that is about all. + +The pattern repeats several times over the next few months. Boards have a fiduciary duty to act in the interests of shareholders, and when the offer is well above market value and the alternative is insolvency, shareholders’ interests are clear. One after another, high-tech manufacturers – robotics, aerospace, specialist toolmakers – are acquired and restructured by American ventures. The official line is that doing so protects jobs and keeps vital facilities on European soil, but in reality, Europe has no plausible counter-offer. + +Christian: number eleven + +Caroline: I counted too. + +Christian: of course you did + +### August 2030 - Model Collapse + +The United States is not trying to destroy Europe. It is trying to beat China in a race it believes is existential. But from where Caroline is sitting, it is hard to tell the difference. + +Europe went through a version of this before, after the 2008 financial crash. Welfare obligations rise just as tax receipts fall. Governments borrow against growth that isn’t coming. Lenders tighten credit conditions. At each step, policymakers’ options get fewer and worse. + +In Paris, the finance minister presents a budget to the Assemblée Nationale that nobody in the chamber, including him, believes. There are three numbers which simply cannot be reconciled: welfare spending is approaching COVID-era levels, corporate tax revenue is down nine per cent, and a tenth of the budget is spent on servicing ever-rising debt. The only way to even pretend it can be held together is by making wildly unrealistic growth assumptions that everyone immediately sees through. + +Caroline watches it on her phone over lunch. She thinks about her brother. He has not found work. He has student debt and no savings, and has moved back in with their parents. Several of his friends have done the same, and some have started talking about moving to the UK, which has - to the surprise of many - manoeuvred the AI transition much better than most EU countries have. + +Moody’s puts France on negative watch within the month. S&P follows. The downgrade itself will arrive later, but the market has already priced it in. The cost of French government borrowing pulls sharply away from Germany’s, to the widest gap since the single currency was created. Markets are no longer sure a French euro and a German euro are the same thing, or that the eurozone can hold together. + +In September, French debt service crosses 12 per cent of the budget. By November the agencies are reviewing on a rolling basis rather than quarterly. Italy, Spain, and Greece are downgraded in close order. In all four countries, money that used to flow into the exchequer has been diverted to American balance sheets, and the remaining tax base is dwarfed by rising welfare and borrowing costs. At every release, growth assumptions are revised downward. + +Around the same time, the loans start to appear. A Chinese state-backed fund extends a credit line to a Portuguese infrastructure bank. A second refinances a tranche of Greek sovereign debt at terms no European institution will match. A third underwrites a Spanish regional government. A leaked Commission memo describes the pattern as ‘strategically motivated capital deployment.’ Nobody is quite sure what the strategy is. Some analysts think Beijing wants ASML or an EUV licensing deal; others nothing more concrete than for Europe and the United States to drift apart. + +Christian: did you see the france numbers + +Caroline: What worries me more are the Chinese lifelines. + +Caroline: We’re getting torn apart. + +Christian: i'm sorry + +Christian: i really am + +The political consequences soon arrive. Protests build throughout until they turn violent. The screens in the Berlaymont show riots in Paris and Rome; young people, mostly, who share few convictions other than that the system has failed them. Populist parties, often explicitly anti-AI and anti-US, lead the polls in most EU countries. Europe is fragmenting. Southern Europe needs help from the north, but even those countries which are not entirely racked by debt and could theoretically help, such as Germany, face their own internal crises. + +Bits of the bloc start to drift away. Slovakia has stopped pretending to take the Commission seriously on topics other than trade. Poland and the Baltics, wary of Russia and no longer trusting the EU to save them, deepen engagement with the US. The Nordics, who built their own data centres and thus have something to negotiate with, form their own coalition without Brussels. Many people in the UK conclude that Brexit may have been good for one thing after all: they can now broker bilateral AI deals with Washington more easily. + +Across the Atlantic the picture is different. Facing its own anti-AI backlash, the US administration has rolled out job guarantees and direct cash transfers at a scale Europe cannot contemplate. The protests there have not stopped, but the government has bought time. The contrast is not lost on European publics, or on European finance ministers, who understand that the American response is funded by AI-driven growth their own economies cannot offer. + +By January 2031 the euro is under sustained pressure. Capital leaves southern Europe and does not come back. The ECB intervenes, intervenes again, and then runs out of credible instruments. By late February, depositors in Italian and Greek banks are moving money north faster than the ECB can offset, and the euro in its current form is no longer something European officials will defend in private. + +Christian: how are you holding up + +Caroline: I'm okay. + +Christian: come to san francisco + +Christian: seriously. we'd hire you tomorrow + +Caroline: I can't leave. + +Christian: why not + +Caroline: Someone has to stay. + +Compute overview: 2030 + +US 208.2 GW + +Europe 16.8 GW + +1 chip icon = 3 GW 12.4× US compute advantage + +### March 2031 - Between giants + +By early 2031 power is concentrated in two places. American labs hold the cognitive frontier; China still holds the physical one. Atlas alone is worth more than every listed European company combined. The three largest American AI firms each spend more on compute than the EU spends on defence. The leading Chinese humanoid producer ships more units in a month than Europe ships in a year. + +The increasingly tense transpacific relationship has brought Taiwan to the fore. TSMC’s fabs, where the world’s leading-edge chips are made, are mostly on the island. The gap between US and Chinese cognitive AI has grown over the last year, and Taiwan’s strategic importance has grown with it. As a result, minor confrontations between navies, which happened perhaps once a quarter two years ago, are now happening every week. Both sides have tested AI-managed weapons platforms in public, and more powerful ones in private. In Washington the labs and the Defense Department are entangled past the point where the distinction means much; in Beijing the integration is more formal. Pundits and analysts regularly use the word *war*, without modifiers such as ‘trade’ or ‘cold’. Three or four men in each capital have the decision-making power that could lead to full-blown conflict. + +Europe is in a disastrous state, even if no one in an official position is willing to say so plainly. Polarisation is rampant, and the social model is breaking. In the countries worst affected by the AI shock, economic growth has halted. In others, it has all but decoupled from well-being. Mortgage delinquencies have surged in those member states which tend to use variable-rate loans. European sovereign debt trades at levels usually reserved for countries where people carry cash in wheelbarrows. Where economists once argued about the wealth gap with the US – is it just downstream of hours worked, is the European standard of living better? – now they have stopped; people who go to California can see the difference within minutes of landing. + +But Europe has one last card to play. After five years of failing to build a frontier AI sector, it still owns the one bottleneck which the entire race runs through. ASML remains the only company in the world capable of building the EUV lithography equipment that is used to print cutting-edge chips. Without access to its machines, the US could not keep extending its lead in AI; with access to its machines, China would likely have caught up some time ago. + +Beijing has made real progress on domestic DUVi machines, and is expected to begin mass production in a year from now. But that is too little and too late: a year now feels like a lifetime, and DUVi does not produce the cutting-edge chips China needs. The US is pushing ahead every day. The Chinese government is deeply worried that superintelligent AI is on the horizon, and nobody is quite sure what it could do. Some advisers worry the US could use sufficiently advanced AI to neutralise China’s nuclear second strike. Others fear such systems could be made persuasive enough to destabilise the Party itself. China’s lead in robots is real, but cannot alleviate those concerns. + +So Beijing leans harder on the strategy it has been running for two years. The loans to southern Europe grow larger and the terms more generous. The information campaigns intensify. European leaders receive signals about what a closer relationship might eventually look like: privileged market access, robotics co-production, a seat at a table from which Washington has largely excluded them. For the first time, ASML and its EUV technology are mentioned explicitly. Five years of being treated as a vassal by Washington has left its mark, and in several capitals the alternative is, for the first time, being seriously discussed. + +China is trying to peel Europe off, and it is succeeding more than Washington is comfortable with. The Pentagon considers losing control of ASML a threat comparable to the proliferation of nuclear weapons. The White House decides it needs direct control of the company now, while the option still exists. The call goes to the three countries that can stop it: The Netherlands, where ASML is headquartered, and Germany and France, without whose acquiescence no Dutch government could agree. + +The US vice president delivers the offer in a forty-minute call over a secure line. ASML itself will be folded into a joint Dutch-American holding company, governed by a shared board on which Washington holds a controlling vote over output, customers, and technology transfer. In return, the US offers a capital injection on a scale Europe cannot match, enabling a range of new production sites to be built on US soil. More strikingly, it also promises direct cash transfers to European citizens, indexed to US AI windfalls. These would start small, perhaps €100 per person per year, but would grow over time. + +The three European leaders decide not to loop in their fellow EU prime ministers and presidents; they do not consider the offer acceptable, and worry that the others would pressure them to take it. When they politely decline, the White House makes the offer public and adds a stick to the carrot. If the EU does not sign, the whole region will drop to Tier 3 under the Frontier Inference Services Rule and lose all access, present and future, to American AI. The US knows it holds all the cards: through its Taiwanese partners, it has amassed enough EUV machines and maintenance know-how to forego ASML access for longer than Europe can forego access to frontier AI. + +Several European capitals reach out to the Chinese, looking for a counter-offer that might give them room to refuse Washington. What they get is not what they were hoping for. Beijing has concluded that its charm offensive is not enough and that they, too, have to switch registers: rare-earth export terms will be reviewed if the Dutch sign with Washington. Robot export licences will be reassessed. The deadline is shorter than the American one. + +Europe now has three options, and all of them are bad. + +Sign with Washington, and the continent gives up its only piece of leverage and becomes an American protectorate in all but name, while its remaining manufacturing dies if China follows up on its threat and cuts off its exports. + +Align with Beijing, and the south might stop drowning, and the Union perhaps survives its immediate fiscal crisis. But Europe will be handing China the keys to the future, and would face a response from Washington which it cannot absorb – not even with Chinese help. Tier 3 would only be the opening move. + +Sign with neither, and Europe gets nothing. Instead it would absorb the full displeasure of both great powers at once. It risks losing frontier inference and critical manufacturing inputs at the same time, which would likely cause the already strained Union to disintegrate. + +The European Council convenes an emergency session. After fourteen hours, little progress has been made. It authorises a delegation to fly to Washington with a mandate everyone in the room knows is ambiguous on purpose. Against normal Council practice, the principals will decide on the ground. + +The meeting takes place in the Eisenhower Executive Office Building on a Tuesday morning. The European delegation is led by the Dutch PM, the French president, the German chancellor, the Polish PM, the Spanish PM, and the Italian council president, each accompanied by their foreign minister and a national security adviser. + +The Americans have brought a similar-looking group, plus two officials in the back row who are not introduced and who are wearing earpieces. + +Those earpieces are connected to a frontier AI model that has infiltrated every European channel it could find. It knows what each European principal said in cabinet last Tuesday. It knows which of them is having an affair, which is being treated for prostate cancer, which one’s daughter is in trouble with the law. It knows what each of them is most afraid of, and what they will trade to avoid it. The Europeans do not know this. + +Caroline is in the support team the Commission has brought to Washington. She is sitting in the delegation room, two doors down, watching a closed feed on a wall screen. + +By late morning it is clear the principals are not aligned. The German chancellor and Polish PM push hard for the American deal. The Spanish PM wants to align with China; the French president wants to refuse both. The Dutch PM looks ill and the Italian council president has barely spoken in three hours. + +At noon, they call a break. The principals scatter into side-rooms with their delegations. Caroline steps out of the delegation room to find coffee and clear her head. + +In the corridor, she nearly walks into the council president, who has come out of the meeting room alone. His jacket is off and his tie is loosened. He is shorter than she had imagined. He stops when he sees her lanyard. + +‘Commission?’ + +‘DG TRADE, sir.’ + +He looks at her for a moment. ‘Walk with me.’ + +They walk slowly. He is known for consulting unpredictably before major decisions – junior staff, journalists, his driver. Some find it charming. Coincidentally or otherwise, he has also survived forty years in European politics. + +‘In the room,’ he says, ‘everyone is making the case they have been making for two years. I have heard it all.’ He looks at her. ‘What worries you most, of the three?’ + +She thinks before she answers. + +‘Choosing neither looks like keeping our options open,’ she says. ‘But it isn’t. We need to pick one, rather than just letting things happen to us and pretending we’re the victims of circumstance. And we can’t give the Chinese that much power. So it needs to be the Americans.’ + +The council president nods slowly. He does not say whether he agrees. He pats her shoulder, lightly, the way an uncle might. ‘Thank you. Go get your coffee.’ He turns and walks back towards the meeting room. + +Caroline goes to the bathroom. She splashes cold water on her face and looks at herself in the mirror. Her hands are shaking. She grips the edge of the sink and waits for it to pass. Through the small high window she can see a slice of Washington sky, flat and bright. + +Down the corridor, six people are deciding the fate of the European continent. She does not know whether anything she said will matter. + +She suspects it will not. + +That evening she walks back to the hotel alone. It is cold for Washington in March. She thinks about her brother. She thinks about the dinner party in Hayes Valley, six years ago, and the calm certainty of the people around the table that the world was about to change. + +Her phone buzzes. + +Christian: you ok + +Caroline: Hell of a day. + +Christian: my flight got delayed + +Christian: i can meet for a drink in an hour + +Caroline: I'd like that. + +Compute overview: 2031 + +US 255.1 GW + +Europe 20.9 GW + +1 chip icon = 3 GW 12.2× US compute advantage + +Compute buildout + +US 29.3 GW + +Europe 1.9 GW + +1 chip icon = 3 GW + +15.7× US compute advantage + +## Epilogue + +Europe's slide into irrelevance was not inevitable. Even in 2026, the continent could still have changed course, had it shown the courage and political will to take drastic measures. + +![A green mosaic of Europa riding the bull, the classical motif on the two-euro coin.](/assets/europa-bull-mosaic-green.webp) + +> '*Do not go gentle into that good night.* +> +> *Rage, rage against the dying of the light.*' +> +> Dylan Thomas, Botteghe Oscure, 1951. + +### June 2034 - Project Inheritance + +The domes are visible from the kitchen window. There are four of them, white and ribbed, the closest one maybe two miles off across the scrub. Caroline has been told they are full-scale prototypes for the lunar bases, pressure-tested in conditions meant to approximate Shackleton Crater. On clear nights she can sometimes see a second, smaller dome, lit from the inside, where the company is testing whatever it is they are testing. + +She lives in a small house outside a town in New Mexico. The nearest grocery store is twenty minutes away by self-driving car. She just turned thirty-seven. She has been here for a year and a half. + +She left the Commission in November 2031, eight months after the ASML negotiation. Brussels had become a grim place. The summer before she resigned had been the one with the Taiwan crisis, when nobody slept for four days because nobody knew whether the carrier groups were going to stop. She had quit without anywhere to go. + +Christian had asked her, again, to come to America. She had said no, again. Then, a year later, she had said yes. Her mother had died in the spring and there was little left for her in Belgium or France, apart from her brother. She had not moved to San Francisco. She had grudgingly taken his money and moved here instead; America, but not *that* America. + +She and Christian see each other more regularly now. He has sold his company. He is, by any reasonable measure, immensely rich, and has used a portion of the money to fund what he insists is the most important project of his life. + +This morning, she is waiting for a call from that project. + +It is called Inheritance and its stated purpose is to produce a record, as complete as possible, of how humanity got here. The AIs have already ingested everything public, from emails to parliamentary records to podcasts. What remains, Christian says, is the private stuff. + +Christian is doing this because he expects – soon, although the specifics are vague – humanity to start spreading beyond Earth. As the ships move out into the cosmos, communication with home will become slow and sparse, and eventually impossible. Human values will drift over time. The better these space-faring humans understand their own history, he says, the better they will be able to make decisions their ancestors would have approved of. + +The first time he explained this to her, Caroline had laughed, but he had been serious. + +She agreed to be interviewed. She owed him – literally; he gave her a large amount of money – and besides, her therapist, an actual human therapist, says talking about her years in Brussels would be healthy. + +Her laptop chimes. The man who appears on the screen when she clicks the button is in his early forties. He has dark hair and is wearing a soft grey sweater. He is sitting in what appears to be a study. He introduces himself as Daniel. + +Caroline looks at him for a moment. ‘God, I hate that you’re not real.’ + +‘Me too,’ he says, smiling. ‘Is the appearance alright? We find people are more candid if the interviewer presents in a way they find comfortable.’ + +‘It’s fine,’ she says. + +‘Thank you. We can begin whenever you’re ready.’ + +She makes coffee, comes back, and sits down. He waits without doing anything that looks like waiting. + +‘I’m ready.’ + +The interviewer nods. ‘You worked under three directors at the Commission. Did they understand what was happening?’ + +‘None of them understood in time. My second director understood that AI was going to be economically important, but he thought it would be important in the way that the iPhone was important. He didn’t understand that it was going to eat the world. My third director understood that it was going to eat the world, but she thought we had ten years. We had two.’ + +‘And you?’ + +‘I understood it about a year before most of my colleagues did. Which was about three years too late.’ + +‘Take me back to the summer of 2026. Before Strasbourg. What was that period like?’ + +‘Strange. Fluid. The German chancellor had just come back from San Francisco and seemed open to serious measures. So were the French president and my boss. For a couple of weeks it felt like the door was open. I pulled a lot of all-nighters.’ + +‘What did you recommend to your superiors?’ + +She thinks back. ‘We needed the metal on our soil. The data centres, the chips, the power supplies. Anchored under European law, in jurisdictions Washington could not commandeer at six hours’ notice. Not the five-year-long Gigafactories process everyone was so proud of. Real compute. The numbers had to be in the tens of gigawatts. We should have been building like a country at war.’ + +The interviewer raises an eyebrow. ‘But Europe hosted only five per cent of global compute at the time.’ + +‘We could have changed that. Global compute capacity was almost doubling every year. We could have got to fifteen, maybe twenty per cent in five years, if we had wanted it badly enough. That would have been enough to serve most European customers. The threat of losing access to that much compute would have kept Washington from doing what they eventually did.’ + +‘How would you have done it?’ + +‘European companies couldn’t have done it on their own. We needed to also partner with the Americans. Money wasn’t the main thing. Subsidies would have helped, but the hyperscalers had money. What they needed was speed. They were in an insane race to ship each new model. Bringing a single data centre online a month faster was worth billions to them. + +‘I argued Europe should create Special Compute Zones, reduce permitting timelines from two years to three months, strip out every bureaucratic obstacle. Set up concierge teams liaising between AI companies, energy providers and municipalities. Convert decommissioned power plants, where the grid connections are ready and waiting. Build new electricity generation like you mean it. A stripped-down version of this idea ended up in the EU Tech Sovereignty Package, but we didn’t execute.’ + +‘Why not?’ + +‘Because rolling out the red carpet to American hyperscalers seemed like the opposite of sovereignty. The Eurostack people wanted to build their own ideal, but they were ten years too late. I told them that, realistically speaking, we can’t build all the data centres ourselves, so what we need to do is make sure the ones that *are* built are bolted to our floor. But no one wanted to be the politician to say that sentence. It was too humiliating.’ + +The interviewer nods. + +‘You also wrote about a coalition of middle powers.’ + +‘Yes. We screwed up royally with ASML, but it could have been avoided. The problem was that we tried to coordinate among twenty-seven Member States. A smaller group of middle powers – the Netherlands, Germany, France, Norway, the UK, Canada, Japan, South Korea, perhaps assisted by the European Commission – might have pulled it off together. Most of these countries held important pieces of leverage. Supply chain bottlenecks, AI talent, energy. + +‘The whole AI supply chain ran through a small number of places, and most of those places were not the United States. We never sat down with the Japanese or the Koreans and said *we are in the same position, you and us, between two empires that only care about us instrumentally, and we have more leverage together than apart.* + +‘It would not have been easy. Every one of these countries had its own complicated relationship with Washington. But there was a middle ground that was in their interest too. We never really tried to find it.’ + +She pauses. From the next room she can hear the household system humming, readying the laundry. Outside, a quadruped delivery unit is pausing on the gravel path. + +‘What else?’ + +‘There was a real opening in robotics and industrial AI. The Frontier Initiative was surprisingly successful in the world models that made the robots work in the end. But we didn’t set up the industry partnerships, we didn’t unlock the data, we didn’t screen foreign investments. So when Atlas noticed, it simply bought the researchers and the companies.’ + +The interviewer nods again. ‘What about the job market?’ + +‘I had a friend in the French government who worked on labour market reforms. He always talked about how Denmark had been running the answer for thirty years: the whole flexicurity thing. You pair serious wage insurance, retraining, and real support for displaced workers with the freedom for companies to let go of staff whose jobs have changed – giving businesses the agility to drive deep AI adoption and stay competitive globally. + +‘We knew the model. We could see the Danish numbers. But every Member State had its own labour code and its own unions and its own political coalition, and nobody wanted to spend the capital. So the productivity gap widened and the protection we were supposedly preserving dried up.’ + +She looks out the window for a moment. + +‘I think about June of last year a lot,’ she says, more to herself than to him. ‘The Spanish riots. Watching them on the news, with a feeling they were something I had failed to prevent.’ + +The interviewer waits. He is very good at that. + +‘I know I shouldn’t blame myself, but I do.’ + +‘Why?’ + +‘All of what I just described was economic. Compute, supply chains, labour codes. The thing I never really cracked, and which I think mattered much more than I realised, was the story. A picture of why they might *want* this stuff to happen. + +‘We had a negative vision. We were good at negative visions. We were going to be devoured. Every memo I wrote started from a description of what we were losing. But we couldn’t offer a positive vision. You cannot ask people to absorb years of disruption on the basis of *otherwise it gets worse*. + +‘We should have told them what the good version of the new thing looked like. But we didn’t know what it would look like. We didn’t even have the vocabulary for it. You know, I tried. I wrote one bad memo. But I should have kept writing them.’ + +The interviewer makes a note on a notepad which does not exist. ‘Do you think this lack of a positive vision was why the measures never passed? You said the political situation was fluid in summer 2026.’ + +‘It was fluid and it was not. Even the leaders who understood what was coming were not fully convinced themselves. Every one of these measures looked like an ugly compromise; like dealing with people we didn’t like and throwing out rules we had spent decades crafting. They would have involved spending every coin of political capital and much of the actual capital on things voters didn’t understand and didn’t want, for gains that were years in the future. None of the sovereignty benefits felt graspable back then. I’m not sure any politician could have passed the drastic changes we required with their careers intact. I’m not even sure they *should* have done it.’ + +The interviewer lifts his head. ‘Say more about that.’ + +‘I don’t know. I genuinely do not know. There is a version of this where the leaders who saw what was coming spent everything they had on it, lost their governments, were replaced by the populists who were already winning, and the populists undid the policies and the outcome was worse.’ + +‘And yet you wrote those memos.’ + +‘What I saw was that the European model was about to collapse. So the question was whether you suspended the rules for one sector for a period, in order to preserve the social model elsewhere – or whether you preserved the rules everywhere, and watched the whole thing crumble. I came down on the first answer. Some of my colleagues accused me of betraying Europe’s social-democratic roots. Called me a libertarian. But everything I wrote was an attempt to protect the European project. The traditional political divides had stopped mattering.’ + +‘Did you ever think: we should hit the brakes?’ + +‘Of course. Many people did. I would have pressed the button if I could. Slow things down and allow society to adapt, let AI safety research catch up to the new capabilities frontier. But we didn’t have a button. The technology was developed outside of Europe. The US and China were in this frantic race. If we had built actual leverage, we could perhaps have forced them to slow down, or at least get both to pick up the damn phone and talk to each other. But we had no bargaining power whatsoever, so Europe’s only remaining path was through AI, not around it. ’ + +She pauses. + +‘You know, I’ve tried to find people to blame for all of this. But I have come to the conclusion that there were no real villains in the story. It is just that the system that produced our decisions was responding correctly to the incentives it had been given, and the incentives that had been given were hopelessly unfit for what we were facing. The failure was not a failure of individuals but a failure of stuff like information flow, political constraints, and the speed at which institutions can adapt. Nobody wants to hear that, because if it is true then there is nobody to be angry at, and people need to be angry at someone.’ + +She looks at the interviewer. + +‘What do you think?’ + +He pauses. + +‘Can I be very direct with you?’ + +She frowns. + +‘I think you’re full of shit.’ He is smiling. + +She stares at him. + +‘Excuse me?’ + +‘Not the analysis. The analysis is fine. But I do not believe the analysis is what you actually feel.’ + +‘What?’ + +‘You are telling me a story in which nobody is to blame, in which the incentives were miscalibrated, in which the failure was structural, in which it is not even clear the leaders should have tried. It is a careful story. It is an intellectually defensible story. I do not think it is the story you believe.’ + +‘You don’t know what I believe.’ + +‘No. But I have a guess. I think you are angry. I am not sure at whom.’ + +She does not say anything. + +‘I could be wrong,’ the interviewer says. ‘I am not the right tool for many things. But you asked what I think.’ + +She is quiet for a long moment. The dome through the window is a different colour than it was when the call started. + +‘Fine. You want to know who I’m angry at? I am angry at Christian. I am angry at Christian right now, because Christian sent me a fucking AI to interview me about the worst years of my life, and apparently the AI is now also going to do therapy. I am angry that I took his money. I am angry that he keeps showing up in my life with this – this earnest face, this *I have a project that will save humanity* face, and the project is always real and the money is always real and I cannot ever quite say no. Is that what you wanted?’ + +‘Some of it.’ + +‘What does that mean?’ + +‘I think you are redirecting.’ + +‘Don’t psychoanalyse me.’ + +He shrugs. There is a pause. She hasn’t noticed that she is gripping the edge of the table. She opens her mouth and closes it again. + +‘Fine. I am angry. I am so fucking angry. I am angry at the US and China for almost throwing us into World War III. I am angry that we let a handful of power-hungry men decide our future. I am angry at the AI labs for building these ‘tools’ without knowing how to control them. I am angry at the European leaders for their endless colouring inside the lines, their roundtables, their waiting for the political environment to be ready. + +‘The political environment is never ready. It was not ready during COVID either. It was not ready when Russia invaded Ukraine. People did things anyway. When COVID hit, we went into lockdown and rolled out a vaccine program in record time. After the Russian invasion, we built LNG terminals out of thin air and pulled financial support from every Member State to defend the eastern border. We broke every rule because we realised our future depended on breaking them. + +‘You know what I blame our leaders for? Not for stupidity, not for malice, but for their lack of courage. Their unwillingness to look at a thing that was hard to grasp but also obviously happening and say *I am going to do something about it that will probably end my career, and I am going to do it anyway.* Nobody did that. Nobody. They all saw their local incentives and went along with them, and they told themselves stories about why the local incentives were the only thing they could see, and the stories were sophisticated and well-argued and completely beside the point. + +‘It is like – it is like – I don’t know. Some oracle shows up at your door and says you have to qualify for the Olympic 200-metre freestyle in three years or the world ends. And you barely work out. You go to the gym maybe once a week. There is no realistic version of this in which you qualify. But the oracle is right. The world really does end if you don’t. So what should you do? You should train. You should train like a god-damn lunatic. You should sleep eight hours a night and eat the right things and stop seeing your friends, because the alternative is the world ending. But you don’t. You don’t train because the water is cold. Because you are really more of a tennis person. Because you decide the coach has a profit motive. You do nothing. You let it happen. + +‘That is what we did. We let it happen. We told ourselves it could not be done and we went home and we let it happen. And I am so fucking angry at all of us for that. Because it could have been done. It could have been done. Not certainly. But there was a chance and we did not even try.’ + +She stops. She finds that she is crying. She has not cried in front of another person for years, not even with her therapist. + +The interviewer has, she notices, the decency to look away. + +‘I know you can still see me,’ she says. + +‘Yes.’ + +‘It’s fine.’ + +She wipes her eyes. Outside, the quadruped is retreating down the gravel. For a brief moment the late sun catches its plating. + +It is, she finds, almost beautiful. + +Europe 2031 + +What getting AI wrong means for us + +Design by [Boyne Creative](https://boyne.co) + + [Scenario](/) [Executive SummarySummary](/summary) [About](/about) [Compute Forecast](/compute-forecast) [Agent-friendly version](/llms.txt) + +Share X LinkedIn Copy link Share… diff --git a/sources/resources/fai-race-worth-winning.md b/sources/resources/fai-race-worth-winning.md new file mode 100644 index 0000000..63d7ebc --- /dev/null +++ b/sources/resources/fai-race-worth-winning.md @@ -0,0 +1,2199 @@ +--- +url: https://www.thefai.org/posts/the-race-worth-winning-middle-powers-in-the-age-of-machine-intelligence +--- + +THE RACE +WORTH +WINNING +Middle Powers in the Age +of Machine Intelligence + + + + +Anton Leicht & Dean Ball +February 2026 + Table of Contents + + + + + Executive Summary........................................................................ 1 + + + Introduction ....................................................................... 3 + + +1 Between Worlds: The Middle-Power Challenge ........................... 6 + + +2 The Substrate of a Revolution: Frontier AI Access........................ 12 + + +3 Ghosts in Your Government: Institutional Transformation ............. 19 + + +4 Bottlenecks for Abundant Intelligence: Economic Transformation ... 25 + + Securing a Share of AI Growth. . ............................................ 26 + Securing a Durable Position. . ............................................... 31 + +5 Navigating the Transition .. . .................................................... 33 + + + Conclusion ..................... .................................................... 37 + Executive Summary + + + + +Artificial intelligence presents the greatest transformation in human affairs since the + +industrial revolution, and it may be greater still. It promises material and information- + +al abundance: machines to carry out the dangerous and arduous, leaving humanity + +wealthier, safer, and freer as a result. + + Yet challenges come with that: AI looks likely to upend the economic and industrial + +strategies of all countries. It may erode or reinforce the informational advantage that + +has underwritten the state’s claim to govern. And it may throw off or bolster the favor- + +able balance between defense and offense that secures a state’s ability to protect its + +citizens from harm. + + The nations with the most to lose and gain from this transformation will be the mid- + +dle powers: nations with significant institutional and industrial capacity but without + +frontier AI development capability. Unlike great powers, they do not control whether + +and how AI gets built, and whether it rips through society or not. Yet unlike the least + +developed nations, middle powers possess institutional substrate—functioning econo- + +mies and bureaucracies. They have something to transform. Unburdened by the need to + +win a race they were never going to win, they are free to—but must—ask different ques- + +tions on how to deploy the assets they still have to secure the future they want. + + The race that remains open, the race where the middle powers can compete and even + +lead, is the race to favorably adapt frontier artificial intelligence into their governments + +and economies. This fact cannot stand as an excuse to remain inactive, but instead has + + + + +The Race Worth Winning ■ thefai.org 1 + to be understood as a generational effort. It will require finding answers to encompass- + +ing and complicated questions that have no comfortable answers: + + • ACCESS ARCHITECTURE: What contractual, technical, and institutional arrange- + + ments can durably guarantee middle powers’ access to imported frontier AI? + + • INSTITUTIONAL TRANSFORMATION IN PRACTICE: How can the public sector be + + reshaped to be amenable to AI-driven transformation? What separates the risks + + of falling behind and the opportunities for leapfrogging? + + • ECONOMIC BOTTLENECKS: Where along the flow of AI through the global econo- + + my can a given middle power occupy a bottleneck valuable enough to safeguard + + its fiscal and economic prospects? + + • REDISTRIBUTION MECHANICS: How can middle powers support their citizens + + when AI-driven revenue shifts simultaneously from labour income to capital + + returns and from domestic economies to the countries where AI is built? + +Middle powers that fail will do so in familiar ways with unfamiliar dramatic conse- + +quences: allowing the legitimacy of the state to erode, taking half measures and easy + +roads, shutting their eyes and closing their borders to a global phenomenon, and wait- + +ing for a sovereign solution that never arrives. Those that move with clarity and resolve + +will shape the transformation. + + + + +The Race Worth Winning ■ thefai.org 2 + Introduction +Artificial intelligence presents the greatest transformation in human affairs since the + +industrial revolution, and it may be much greater still. While many today pay lip ser- + +vice to this notion, few have internalized it. The revolution is not some distant hypo- + +thetical whose contours can be shaped from within the halls of government, academia, + +or think tanks alone. Instead this revolution is unfolding, and will unfold ever more + +rapidly, now—with the status-quo political economy, institutions, technical capabilities, + +industrial power, and diplomatic relations particular to each nation-state on Earth. + +Policymakers will not grapple with this revolution using some imagined future toolkit + +but rather the toolkit in their hands today. The era of machine intelligence will not “be- + +gin soon.” It has begun. + + Zoomed out far enough, the stakes of this era can seem existential: we face genera- + +tional risks, but should we navigate them, what awaits on the other side is something + +like material abundance—machines to shoulder the dangerous and arduous work that + +has always fallen to human hands, leaving us wealthier, safer, and freer for it. Once + +again, and at a scale without precedent, a technological revolution that would make the + +world genuinely better appears to be afoot. But even in futures where AI proves to be + +positively transformative, it places unprecedented pressures on the statecraft of all na- + +tions. Nowhere is that pressure stronger than in middle powers—those countries that + +hold just enough cards to shape their response to this transformation for incredible + +gain, but that have just as much to lose. + + Almost no government or policymaking apparatus in any country anywhere on Earth + +has really internalized this reality. It is a difficult reality to internalize. In the abstract, + +technological revolutions promise greater wealth, freedom, and abundance to the + +world. But in the concrete, this future will not be evenly distributed. AI may well upend + +the economic and industrial strategies of many countries worldwide, particularly those + +which staked their futures on growth in the knowledge and service sectors. Suddenly, + + + + +The Race Worth Winning ■ thefai.org 3 + the path forward for many industries within those sectors looks more precarious than + +it once did. At the same time, as we have learned repeatedly since the COVID pandemic + +snarled global supply chains, controlling manufacturing capacity, especially but not ex- + +clusively at supply-chain chokepoints, is a source of tremendous international leverage. + + Even beyond the much-discussed economic effects of AI, there are precious few do- + +mains of human activity that the revolution in machine intelligence will not touch. The + +adaptations of human beings and our institutions to these transformations may collec- + +tively give rise to a new order of human affairs—one which is capitalistic, but not quite + +“capitalism” as we know it today; one in which business is vigorously conducted, but + +by organizations that do not quite resemble what we now think of as “a business.” The + +role of the nation-state in this transformation is uncertain, and even the fate of the na- + +tion-state is a live object on the field of play. + + The outcomes in every respect are deeply unclear, often in more fundamental ways + +than the leaders of present-day institutions would care to admit. What this means is + +that many institutions—and especially the institutions of government, insulated as they + +are against the forces of the market—will probably stumble badly as they encounter the + +new terrain. Some may fall down altogether. But where peril abounds, there is also great + +opportunity. The leader who navigates his institution wisely through the transforma- + +tion to come will face rich rewards indeed. The leader who holds steady at the helm may + +deliver a new era of prosperity, abundance, and joy to his people. + + Those who emerge successfully out of the revolution to come will not be the ones who + +dreamt up the most exquisite regulation to impose on a foreign purveyor of technology. + +They will be the ones who navigate the transition to the era of machine intelligence with + +grace, savvy, cunning, and aplomb. Doing this will require the creative use of AI and + +other technologies, arguably more than it will require regulating them: regulation has a + +way of constraining options and creating path dependency. There is a risk that regulat- + +ing today is less akin to playing chess than it is to taking chess pieces off one’s own side + +of the board before the game has even begun. The asymmetry of such a move is par- + +ticularly stark since governments can strip themselves and their law-abiding citizens + + + +The Race Worth Winning ■ thefai.org 4 + of AI, but cannot guarantee to bind their competitors in the global marketplace or the + +adversaries that threaten them. As much is becoming more and more clear as great + +powers are less and less willing to have their firms and institutions abide by the rules + +set by foreign markets. + + In the face of that realisation, wise policymakers in the rest of the world will shift + +their regulatory attention—to methods to grapple with the harms and risks of AI de- + +ployment that both of this paper’s authors have discussed in the past, perhaps to long- + +term prospects of multilateral governance, but primarily to how the deployment of AI to + +governance and government can uplift states’ capacity to deal with the ramifications of + +the transformation. Clever governments will remember that their bureaus and agencies + +themselves may yet be invigorated by a technology that makes careful attention to and + +drafting of complex, text-based documents functionally free. The compounding bene- + +fits to legislation and enforcement could be the tool that governments need to keep up + +with the rapidly changing risk profiles that have them worried even today. Indeed, such + +a technology could change the nature of “regulation” itself, and this is just one of the + +many profound new capabilities machine intelligence delivers. + + The states that first grasp the transformative effects of AI as a technological instru- + +ment of governance, who mechanize the cognitive machinery of public institutions just + +as once in the past they were electrified, who understand not only “AI in government” + +but AI as governance, will also be the first to discover profound strategic insights for nav- + +igating the world ahead. They will have an advantage over all other nations, including + +the hegemons, if they were to stumble and miss a step. If AI is a race among govern- + +ments, it is a race to discover the contours of the institutional revolution that is brewing. + +Many of the obvious approaches are wrong, many of the easy paths are treacherous. + +This paper hopes to elucidate some of the hard questions middle powers ought to ask to + +begin competing in that race. + + + + +The Race Worth Winning ■ thefai.org 5 + Between Worlds: The Middle-Power +Challenge +This report deals with one group of nations in particular. Those with the most to gain— + +and perhaps the most to lose—from navigating this transformation wisely may be nei- + +ther the hegemons nor the smallest states, but rather the middle powers: nations with + +significant industrial and institutional capacity but usually without frontier AI develop- + +ment capability. + + The middle powers occupy a distinctive position in this transformation. Unlike the + +least developed nations, they possess what we might call institutional substrate—func- + +tioning bureaucracies, judicial systems, regulatory bodies, universities, and civil ser- + +vices. They have much to transform. But unlike the United States and China, they cannot + +realistically build frontier AI ecosystems from whole cloth. The capital requirements, + +the talent concentrations, the existing industrial bases, and the sheer momentum of the + +leaders make this path essentially foreclosed. It is mandatory for them to seek agree- + +ments and even dependencies, to chart their course through a transformation governed + +by others. + + This recuses middle powers from some of the most pressing questions that the + +American—and in parts, the Chinese—policy environment face in dealing with the + +thorny challenges of regulating the development of frontier AI. Questions of safety in + +development, of security against the risks of internal deployment, of susceptibility to + +misuse and loss of control, are not immediately in scope for middle powers: not be- + +cause they should be disinterested in them or because we suggest they should take the + +accelerationists’ response to them, but because the middle powers currently simply + +lack the leverage to exert their will over the modalities of deployment in either of the + +great powers. + + + + +The Race Worth Winning ■ thefai.org 6 + Figure 1. G20 countries (excepting U.S. and China) and a selection of further middle powers. + + + + These middle powers differ enormously from one another—in institutional capaci- + +ty, in economic structure, in alliance commitments, and in their existing exposure to + +AI-driven growth. Germany’s challenge is not India’s; South Korea’s options are not + +Brazil’s. A country with a position in the semiconductor supply chain faces a funda- + +mentally different starting point from one whose economy is built on services or com- + +modities. This paper does not pretend otherwise, and the research agenda outlined at + +the end is deliberately designed to be pursued with regional and sectoral specificity. + +But we believe these nations share a structural condition that justifies treating them as + +a category: none of them will build frontier AI, all of them must import it, and each of + +them possesses enough institutional or industrial substance that the choices they make + +will genuinely matter far beyond their borders. The framework that follows is intended + +to clarify those shared choices—not to substitute for the country-specific strategies that + +must follow. + + + + +The Race Worth Winning ■ thefai.org 7 + Middle Powers and the World +This reality confronts middle powers with a changing geopolitical order, even at a time + +where they already face a shift of alliances and dependencies. There is more to be + +gained from using AI to transform the mechanics of the status quo in imaginative ways + +than there is to be gained from fearing it, even when the AI in question originates in + +another country, and even when that country is the United States. The alternative is to + +take on the immense and heretofore unsurmounted challenge of creating a frontier AI + +development ecosystem, on top of all the other challenges this transformation is likely + +to pose for policymakers and political leaders—which seems impossible. + + Yet if we ask middle powers to take on a role as AI importer, they must be embedded + +in a geopolitical environment that allows them to trust that setting, long-term. On one + +hand, a cornerstone of the Trump administration’s technology policy agenda, for which + +one of this paper’s authors served as a key architect, is to let American companies do + +productive business with countries around the world, not to keep AI shackled within + +American borders. Yet on the other hand, recent efforts of the administration to uni- + +laterally redefine the structure of its alliances and the balance of power abroad have + +made allies understandably hesitant to entrust themselves to a hegemonial order. They + +rightly feel that they will need leverage to back up agreements with the U.S., and refuse + +to rely on good faith alone. + + Against that backdrop, no middle power will chart a successful course entirely on + +its own, nor will it prevail by exposing all its assets to any single bilateral relationship, + +lest it grow parasitic and extractive. The solution to this challenge lies in coordination + +among the middle powers, in very much the way that insightful Western leaders have + +recently remarked. The most nimble version of this coordination does not take the de- + +tour through multilateral institutions, but arises from narrow agreements between + +like-minded middle powers in pursuit of similar goals.. Coordination, done right, reas- + +sures partners that they can focus on their comparative advantage instead of pursuing + +full autarky; and it enables powers to see eye to eye with AI exporters instead of being + +picked off by great powers offering rewards for defection one by one. + + + +The Race Worth Winning ■ thefai.org 8 + The U.S. government, too, should ultimately welcome this increase in coordination, + +and not understand it as a rally of anti-American sentiment. Allies coordinated among + +each other serve the American interest as well—they become more capable partners + +because their coordination allows them to concentrate on their comparative advantage, + +and they become more reliable importers because coordination quells the sovereignty + +concerns that would otherwise make them vulnerable: to failed domestic aspirations + +or to trade with China. America should therefore meet the prospect of middle power + +coordination with equanimity, and stand ready to export its AI to any willing group of + +middle powers. + + + +The Middle-Power Opportunity +For the middle powers, then, the alternative path—using AI to transform the mechanics + +of governance and society—is not merely preferable. It is mandatory. This is not a lim- + +itation. It may be an advantage. The hegemons are deeply invested in the current par- + +adigm of AI development: the race to foundation models, the competition for chips and + +data centers, the scramble for talent. This investment can become a form of blindness. + +The middle powers, unburdened by the need to win a race they were never going to win, + +are free to ask different questions—not “how do we build the next frontier model?” but + +“how do we use the models that exist to reinvigorate our courts, our tax administration, + +our schools, our agriculture, our urban planning?” + + The middle powers must also reckon with the new salience of the physical economy. + +Chips, pharmaceuticals, critical minerals, energy infrastructure: the nations that make + +and move physical things occupy a position that no amount of cognitive automation + +can displace. For the middle powers, this creates both peril and opportunity. The ser- + +vice-sector strategies that once promised a shortcut to prosperity may be foreclosing. + +But many middle powers possess manufacturing bases, natural resources, strategic ge- + +ography, or some combination thereof. + + These two imperatives—institutional transformation and industrial development— + +are not in tension, but are instead complementary. The middle powers that transform + + +The Race Worth Winning ■ thefai.org 9 + their institutions most effectively will also be best positioned to attract the manufac- + +turing investment, manage the complex logistics, quickly adjudicate regulatory mat- + +ters, and administer the industrial policy that physical-economy strategies require. + +The middle powers that grasp this—that pursue both tracks simultaneously, using AI to + +reinvigorate governance while positioning themselves in physical supply chains—will + +find themselves with options that others lack. + + Yet the middle powers also face distinct temptations. These are sophisticated states, + +often proud of their regulatory traditions, with capable technocracies eager to demon- + +strate mastery over a new domain. The temptation to regulate—to impose frameworks, + +to set standards, to define what AI may and may not do within their borders—is strong + +precisely because they have the institutional capacity to do so. But as we have argued, + +this temptation is dangerous. Regulatory frameworks constructed now, before the con- + +tours of the transformation are remotely clear, risk constraining options that leaders do + +not yet know they will need: to leverage data from the physical or digital world, to allow + +machines to work in close conjunction with humans, or to provide value precisely by + +not requiring humans in the loop The middle powers have more chess pieces to remove + +from their own side of the board, and a greater temptation to remove them. + + The stakes for the middle powers are substantial and concrete. Consider what insti- + +tutional transformation could mean: court systems where cases now pending for years + +or decades could be resolved in months, and perhaps in the future days or even hours; + +economic data collected and processed live; agricultural knowledge transmitted to + +farmers with a precision and timeliness previously impossible; civil services tailored + +for each recipient and administered with genuine conscientiousness; public health sys- + +tems that identify emerging threats in real time and rapidly mitigate them; educational + +institutions that adapt to each student rather than forcing each student to adapt to them; + +city streets that are safer and more free. All of this is possible with the technology of the + +present day. What does not yet exist is the institutional imagination to seize them—or + +the wisdom to avoid foreclosing them through premature regulatory constraint. + + + + +The Race Worth Winning ■ thefai.org 10 + The middle powers also face a question of talent and ambition. The brightest minds in + +every nation look for opportunity. If the institutions of the middle powers remain stat- + +ic—or worse, if they become mired in regulatory complexity while the frontier nations + +race ahead—talented young people will continue to leave for the places where the future + +is being built. But if the middle powers become laboratories for institutional transfor- + +mation, if they demonstrate that the most interesting questions are not about building + +models but about remaking society with them, they may find that the current of talent + +reverses. The nation that first discovers how to truly govern with AI, rather than merely + +govern AI, will become a destination for those who wish to participate in that discovery, + +and a beacon to inspire and attract nations the world over. + + + + +The Race Worth Winning ■ thefai.org 11 + The Substrate of a Revolution: +Frontier AI Access +If middle powers are to seize the institutional and industrial opportunities outlined above, + +they must first secure the most basic input required to do so. Before questions of adoption, + +reform, or economic positioning can even be meaningfully asked, there is a more funda- + +mental constraint to confront: access to frontier artificial intelligence, built in China and + +America. Many middle powers still reject this uncomfortable reality for its implications for + +geopolitical dependency and capital expenditure, and instead consider a bet on second-best + +technology. This would be a fatal mistake, betraying a marked ignorance of historical re- + +ality: whenever an infrastructural input became the driving force of a historical era, those + +who failed to secure an abundance of it have fallen behind. In the electric age, there have + +been no low-electricity, high-income countries; during industrialisation, no power without + +abundant steam power has prevailed; and when the energy costs of geopolitical competi- +GDP per +tion still capita + came vs. + down to energy + feeding use,armies, + standing 2023the world’s bread baskets were masters +Annual energy use per capita, measured in kilowatt-hours* per person vs. gross domestic product (GDP) per capita, +of their +which times. + is adjusted forOnly middle + inflation powers + and differences with + in living access + costs tocountries. + between abundant frontier AI will stand a chance. +Energy use per capita + + North America + United Arab Emirates South America + Africa + Canada Europe +150,000 kWh Asia + United States Oceania + + Russia Finland + 8B + South Korea 3B + Australia Norway +100,000 kWh Luxembourg Circles sized by + Belgium Population + Iran Netherlands + China Japan + Germany + Poland Switzerland + 50,000 kWh + Thailand Turkey United Kingdom + Iraq Hong Kong + India Brazil + DR Congo Pakistan + Indonesia + Philippines + 0 kWh Ethiopia + $2,000 $5,000 $10,000 $20,000 $50,000 $100,000 + GDP per capita (international-$ in 2021 prices; plotted on a logarithmic axis) + + +Figure 2. Energy consumption per capita vs. GDP per capita, 2023. + Data source: IEA Energy Statistics Data Browser, IEA, via World Bank (2026); Eurostat, OECD, IMF, and World Bank (2026) + Note: GDP data is expressed in international-$** at 2021 prices. OurWorldinData.org/energy | CC BY + + + + +The Race Worth Winning ■ thefai.org 12 + By this, we mean access to the best AI systems. Not just models that perform reason- + +ably, not even just models that are good—the best models there are. As it stands, chanc- + +es are simply too high that second-best is not enough, that the winners of the AI era will + +be those with the best AI. As much might be true in the absolute and relative sense. In + +the absolute sense, countries should seek out the best simply because the most power- + +ful AI models will be the most useful to any national ambition. In the relative sense, it + +is worth understanding that middle powers find themselves in a threefold competition + +that means that whatever deployment tactics we describe below must necessarily relate + +to the deployment not of any AI model, but of the very best. + + The first dimension of this competition is between its public institutions and those + +that seek to use and sometimes abuse it—that would flood the administrative state in + +requests and complaints and overwhelm a human bureaucracy by moving at the pace of + +machine intelligence, leaving the state unable to act and react, eroding trust and capac- + +ity as it goes. On that, this paper comments in much greater detail in the next section. + + The second dimension is between its security apparatus and those that would seek to + +harm their citizens. Attackers foreign and domestic will be equipped with near-leading + +AI systems, whether it is through stolen models, adversarial jailbreaks, or near-frontier + +open-source capabilities. They will use these capabilities to devise and execute new kinds + +and scales of crime and terror: cyberattacks, chemical and biological weapons, strategic + +and organisational support for criminal operations, and as a force multiplier for hos- + +tile militaries and non-state actors. Nation-states’ abilities to defend against these novel + +threats will depend on their ability to wield AI defensively, to scale up their detection of + +attacks and engagement of adversaries. The race between attackers and defenders might + +often be on a knife’s edge: researchers are as of yet unsure which way the offense-de- + +fense balance in AI points; but they find that more advanced models are consistently + +more effective at both offensive and defensive operations. In the interest of the safety of + +their citizens, countries cannot afford to use weaker AI systems than their adversaries. + + The third dimension is between its own private sector and international competitors. + +AI will serve as an accelerant to today’s already-dynamic market forces, exacerbating + + + +The Race Worth Winning ■ thefai.org 13 + the pace and international nature of competition and further shrinking markets. As + +customisation becomes cheaper and information more accessible, institutional moats + +will shrink, and incumbents without a substantive edge will be disrupted by their global + +competitors. The quality of inputs is one of the major differentiators left in that world— + +and the quality of artificial intelligence will be one very important input. We can see + +glimpses of this reality in the areas that are most like tomorrow’s economy today: In San + +Francisco, leading AI developers are already turning off their competitors’ access to their + +own models, because they think the marginal acceleration from slightly better software + +engineering systems is a relevant factor in competition. And even before AI, particularly + +fast-paced competitive environments have become a question of marginal technology + +access: in high-frequency financial trading, slightly faster decisions matter so much that + +proximity to crosscontinental cables has become a prized asset for trading firms. + + In a digital economy stripped of friction and pretense by AI, we expect these marginal + +differences to matter more and more. Much of this is not a question of government use, + +but of private-sector use—but government access plays a meaningful role nevertheless: + +failed government policies could force even the private sector to employ subpar techno- + +logical solutions, such as has been suggested in Europe time and time again. And quite + +frequently, the responsiveness of the government itself is a bottleneck for the competi- + +tiveness of a national economy: whether it is for permitting and licensing, for enforcing + +market protections, or for enabling imports and exports, economic policy can either be + +AI-empowered in an AI-accelerated world or grow to be a substantial bottleneck. + + Foregoing access to the best AI systems puts governments at risk of failing to fulfil + +their promise of efficiency to their citizens, of protection to their charges, and of support + +to their economies. + + + +Ensuring Access +However, the best AI models might not be abundantly available, and their provision + +might be constrained by geopolitical factors. Middle powers are growing more and + +more skeptical of importing critical economic inputs, having been burnt by their lack of + + +The Race Worth Winning ■ thefai.org 14 + leverage in unilateral trade relationships with great powers and by recent high costs to + +their energy dependency. + + They are not entirely wrong to be worried for their enduring access to AI models de- + +veloped in the U.S. and China. If middle powers’ access to AI capabilities is only guaran- + +teed through individual subscriptions and API contracts, it remains uninsured against + +disruptions: sudden price hikes by developers that think they have locked down a mar- + +ket, service disruptions because limited computational capacity is first routed to other + +priority markets, or great powers that take policy action to restrict the outflow of artifi- + +cial intelligence—either for strategic and security reasons or to exert leverage to extract + +concessions in other areas. + + That means middle powers will need a plan to secure durable access to frontier AI ca- + +pabilities. Any further strategy for economic competitiveness and strategic sovereignty + +requires this necessary precondition: without frontier capabilities, any middle power’s + +position is at great peril. + + Yet middle powers will not ensure this access through building their own AI models. + +Sovereign development of frontier AI systems faces a highly restrictive market and an + +unfavorable political economy. The key inputs for frontier AI—computational power and + +exceptional research talent—are limited, and they follow agglomeration effects and the + +ability to mobilise capital. Only China and the U.S. have been able to sow the seeds of + +this agglomeration in their respective hotspot regions, and only state-directed funds in + +China and the unrivaled venture capital ecosystem in America have been able to pro- + +vide for scaling it to frontier performance. For everyone else, sovereign development of + +frontier AI systems is far out of reach. Past attempts of others to catch up have fallen far + +short—landmark infrastructure projects such as the EU’s gigafactories lag about three + +to five years behind the American expansion, for instance. We are doubtful that this + +will change in time: 2028’s infrastructure is determined by today’s buildout decisions, + +where middle powers lag far behind; and any further acceleration would compete for + +often-scarce energy, chips, researchers, and so on. + + + + +The Race Worth Winning ■ thefai.org 15 + Figure 3. Map of frontier AI datacenter buildouts. + Source: Epoch AI Frontier Data Centers Hub + + + + + One permutation of this hope holds that one could sidestep this race and bootstrap to + +a new paradigm—as is reportedly the basis of a recently announced European frontier + +AI initiative. This, too, is unlikely to succeed; both because paradigm innovation also + +benefits from the scale effects of major developers, and because the last years of AI de- + +velopment drive home the veracity of the “bitter lesson,” which holds that technical ap- + +proaches that can leverage computational power at scale ultimately always outperform + +any domain-specific sophistication. There seems to be no viable middle-power path to + +sovereign frontier AI. + + The good news, however, is that American AI companies want to export their advanced + +capabilities, and the U.S. government is very interested in supporting and enabling such + +exports. Multiple large developers have created and advertised for-countries initiatives + +promising the export of AI capabilities to willing governments at scale. The U.S. govern- + +ment’s commitment to the same has been affirmed through an executive order provid- + +ing for the creation of an export promotion program for AI capabilities; and through the + +Pax Silica initiative that seeks to exchange semiconductor supply chain integration for + +access to American AI capabilities. America’s AI industry is open for global business. + + But middle powers are understandably hesitant to take the deal that U.S. industry + +offers. They are concerned about the limitations of ad hoc access: will they be subject + +to price hikes or strategic access limitations? Are they guaranteed to be provided the + +next generation of frontier models, and the one after that? The path forward is the ne- + +gotiation of longer-term contracts backed up by real leverage, conferred by leveraging + +bottlenecks that interest the great powers that can provide frontier AI access. We pro- + +vide the economic roadmap to occupy such bottlenecks later on—but whatever they are, + + + + +The Race Worth Winning ■ thefai.org 16 + they have to be exchanged for long-term access conditions, enforceable by middle pow- + +ers through a combination of local compute, market access conditions, and prohibitive + +government-level contracts that can compel the continued long-term provision of AI + +capabilities. In short: middle powers must be strong enough to be able to enforce the + +contracts that secure their AI access. They might find that strength in their idiosyncrat- + +ic sources of leverage along the semiconductor supply chain, in their effective deploy- + +ment of AI systems, or in the numbers of a broader coalition to negotiate the same—but + +find it they must. + + + +Open Source and Chinese Imports +Some middle powers will object to this logic, arguing they might access models from + +elsewhere, in particular through competing exports from China that might or might + +not be open-source models. However, the availability of powerful open-source models + +will be constrained by economic reality and strategic necessity: deploying them is not + +always worthwhile for developers seeking to recoup their development investment, and + +seeing them released is not always in the interest of great powers that either seek to + +leverage AI access or that are justly concerned about misuse pathways. As a result, one + +should expect open-source models to remain a generation behind the frontier, as they + +are today. In that scenario, middle powers might choose to deploy open source for select + +applications that put a premium on customisability over quality, but it is unlikely that + +they will be sufficiently effective to substitute for genuine frontier access. + + One permutation of this question concerns the role of Chinese AI capabilities—both + +open source and proprietary—which looms particularly large for middle powers that + +have sought to play both sides of great power competition, extracting concessions + +from having America and China compete for their markets. China has aggressively re- + +leased capable open-source models, in part to undercut American market share and + +establish dependency in regions where U.S. export frameworks have been slow to de- + +liver. For middle powers seeking to diversify their sources of AI access, Chinese mod- + +els present a superficially attractive option: capable, freely available, and unburdened + + +The Race Worth Winning ■ thefai.org 17 + by the conditions they fear American firms might attach to their exports. But building + +a national AI strategy around Chinese open source carries its own dependencies and + +shortcomings. + + There is no guarantee that China will continue to release its most capable models + +openly once middle powers have structured their digital economies around them; the + +current generosity in open-source releases appears driven by a competitive strategy to + +establish market share, not by an enduring commitment to openness. Well-documented + +security concerns attend the deployment of Chinese-developed models in sensitive gov- + +ernment and infrastructure applications. Moreover, and perhaps importantly, the use of + +foreign-built open-source models comes with counterintuitively greater dependencies + +than the use of proprietary models: if a middle power uses open-source models for free, + +their provider has nothing to lose from cutting them off, while a revenue-driving con- + +tract comes with mutual leverage over a developer that does not want to lose a customer. + + The perils of choosing China are compounded by the additional complication that + +deepening reliance on Chinese AI may jeopardise a middle power’s standing within + +American-led technology partnerships, including access to the semiconductors and + +compute infrastructure on which all AI deployment ultimately depends. China, now and + +for many years in advance, cannot provide the same infrastructure access, and until + +that changes, cannot be a viable alternative. The strategic logic, then, points toward + +American AI as the primary import source for allied middle powers—not out of naïvety + +about American intentions, but because the broader architecture of alliances, supply + +chains, and security relationships makes the American stack the more natural founda- + +tion for nations wishing to chart their own course. + + Once they have so ensured their access to an abundance of artificial intelligence, + +middle powers must deploy it well: to supercharge and revolutionise their institutions + +to remain and become capable in the new world, and to channel it into durable econom- + +ic pathways that guarantee long-term participation in AI-driven growth and flourishing. + + + + +The Race Worth Winning ■ thefai.org 18 + Ghosts in Your Government: +Institutional Transformation +Middle powers contemplating the adoption of AI face a tremendous opportunity—to lead + +the world in institutional innovation and transformation. But they also face a genuine + +threat: that AI will erode the foundations of the nation-state itself. + + The modern state emerged in large part because it could do something that no other + +institution could do as well: collect, aggregate, and act upon information at scale. The + +census, the tax roll, the statistical office, the regulatory agency—these were instruments + +of knowledge, and knowledge was the precondition of effective administration. Private + +actors could know their own affairs, and perhaps the affairs of their immediate trading + +partners, but only the state could know the society as a whole. This informational ad- + +vantage underwrote much of what governments distinctively did: allocating resources, + +managing externalities, coordinating the provision of public goods, enforcing contracts + +and resolving disputes, responding to crises. The state could see what others could not + +see, and so the state could act where others could not act. In doing so, it was able to re- + +tain a position of trust and authority—even where others could have taken its place in + +providing a piece of information or rendering a service, societies trusted their govern- + +ments to act as able and responsible stewards instead. + + This advantage has been eroding for some time. The rise of large corporations, and + +later of global platforms, began to shift the balance. Credit bureaus came to know the + +financial lives of citizens in ways that tax authorities could not match. Retailers learned + +to predict purchasing behavior with a precision that no economic planning office ever + +achieved. Technology platforms accumulated data on the movements, preferences, re- + +lationships, and sentiments of billions of people—data that dwarfed anything in the ad- + +ministrative records of even the most sophisticated governments. The private sector + +learned to see, and what it saw was often more granular, more current, and more action- + +able than what the state could see. + + + +The Race Worth Winning ■ thefai.org 19 + Artificial intelligence does not initiate this trend, but it accelerates it in ways that may + +prove decisive. The informational advantage of the state rested not only on the collec- + +tion of data but on the capacity to process complexity—to make sense of vast, heteroge- + +neous, and often contradictory information, and to translate that understanding into + +policy and action. This was cognitive work, and it was expensive. It required large bu- + +reaucracies staffed by trained professionals. It required time. The scarcity of this cog- + +nitive capacity was itself a barrier to entry that protected the state’s position. AI erodes + +that barrier. When the processing of complexity becomes cheap—when any sufficiently + +resourced private actor can ingest, analyze, and act upon information at a scale and + +speed that was once the exclusive province of government—the state’s distinctive role + +comes into question. Private actors may soon understand populations better than the + +governments that nominally serve them: their needs, their behaviors, their vulnerabil- + +ities. They may understand the mechanics of regulation and the pathways through bu- + +reaucracy better than the officials who administer them. The state’s claim to govern has + +always rested in part on its capacity to know. What becomes of that claim when others + +know more? + + Something deeper, more structural, is at play in the dawning of the AI revolution. + +Like earlier information revolutions—including writing and the printing press—the + +scope, scale, and structure of institutions will evolve. Grappling with these changes is + +not solely, or even primarily, a regulatory challenge. The instinct to regulate is under- + +standable—it is what governments know how to do—but it mistakes the nature of the + +challenge. The challenge is not primarily to constrain AI but to ensure that public insti- + +tutions do not fall irrecoverably behind the private alternatives that AI enables. This is + +a capability race, and governments are at risk of losing it. + + The race is already underway across nearly every domain of government activity. + +Consider healthcare: a public health system that cannot match the diagnostic accura- + +cy, the personalized treatment recommendations, and the administrative efficiency of + +AI-augmented private alternatives will find its citizens seeking care elsewhere. Those + +who can afford private options will take them. Those who cannot will receive worse + +care than the technology of the age makes possible, and they will know it. Or consider + + + +The Race Worth Winning ■ thefai.org 20 + education: when a private AI tutor can adapt to each student’s pace, identify gaps in un- + +derstanding, and provide individualized instruction around the clock, what becomes of + +the public school that offers one teacher for 30 students and a static curriculum? Those + +able will exit. The public system will become a residual institution for those with no + +other choice. + + Finally, consider the courts: when private dispute resolution—AI-assisted arbitration, + +contracts that execute automatically, platforms that resolve conflicts in hours rather + +than years—offers speed and predictability that public courts cannot approach, com- + +merce and even personal disputes will migrate out of the legal system. The courts will + +not be abolished; they will be bypassed. Consider the bureaucracy itself: when citizens + +can navigate permit applications, tax filings, benefit claims, and regulatory require- + +ments through AI intermediaries that understand the system better than the officials + +who run it, the government becomes a legacy backend—still technically sovereign, but + +substantively hollowed out. + + These are not distant hypotheticals. They are emerging now, unevenly but unmistak- + +ably, in every country with meaningful AI deployment. There is reason to fear a thresh- + +old effect: public institutions that miss the AI acceleration curve may never catch up. + +The gap between an AI-augmented institution and a legacy one is not stable; it widens + +over time as the augmented institution learns, iterates, and improves. The laggard does + +not merely remain behind; it falls further behind with each passing year. For govern- + +ments, this means that the decision not to adopt—or to adopt slowly, cautiously, exper- + +imentally—is not a neutral choice. It is a choice to accept relative decline, perhaps per- + +manently. There remains the hope of “leap-frogging” instead—of adopting AI at just the + +right time to bridge that widening gap in an efficient manner. But the time to jump is + +today—the rapidly accelerating iteration cycles of AI technology spell compounding ef- + +fects which imply that leapfrogging timing is measured in weeks and months, instead + +of years and decades. + + + + +The Race Worth Winning ■ thefai.org 21 + The full suite of political, economic, public policy, and institutional solutions to this + +conundrum lies beyond the scope of this report. But one thing is clear: if governments + +do not adopt AI with alacrity, they will not “lag behind.” They will grow obsolete. + + The middle powers, then, must adopt and adopt aggressively. But they face an ob- + +vious difficulty: they cannot build frontier AI themselves. The models, the compute, + +the talent, and the capital are concentrated elsewhere—overwhelmingly in the United + +States, and to a lesser extent in China. Middle powers must import what they cannot + +build. The question is how to do so without surrendering the very sovereignty they seek + +to preserve. + + Wise adoption should follow several principles. First, architectural flexibility: no + +middle power should adopt systems that lock it to a single AI provider. The frontier is + +volatile; the best model today may not be the best model in six months. Governments + +should build or procure platforms and interfaces that allow rapid switching between + +underlying models as the competitive landscape shifts. Dependence on one provider + +is dependence on one set of strategic interests, one set of potential failure modes, one + +set of terms that can be altered unilaterally—and, increasingly, dependence on a model + +that may no longer represent the state of the art. Second, sovereign compute for criti- + +cal applications: the most sensitive functions of government—national security, judicial + +systems, core administrative processes—should run on infrastructure that the govern- + +ment controls or that resides with trusted allies. This does not mean autarky; it means + +prudent hedging. Third, interoperability and portability as design requirements: every + +system adopted should be structured to permit switching between providers with min- + +imal friction. Lock-in transfers leverage permanently to vendors; portability preserves + +options and ensures that governments can always avail themselves of the best avail- + +able capabilities. Fourth, data governance as an explicit element of negotiation: access + +to government data, and the terms on which private AI systems can operate on citizen + +data, are concessions of real value. They should be traded for real value in return. + + None of this is to suggest that adoption should be slow or tentative. The opposite is + +true. Speed matters more than perfection, and the costs of inaction exceed the costs of + + + +The Race Worth Winning ■ thefai.org 22 + failed experiments. The volatility of the frontier is itself an argument for moving quick- + +ly: a government that delays adoption until a clear winner emerges will wait forever, be- + +cause clear winners do not emerge for long in a landscape this dynamic. Governments + +should identify high-impact domains where AI can deliver rapid, visible improvements: + +document processing and administrative workflow, citizen service interfaces, data col- + +lection and integration, permit processing and regulatory compliance. These are not + +the most glamorous applications, but they are the ones where gains can be realized + +quickly, where institutional capacity can be built through practice, and where the risks + +of failure are manageable. The point is to learn by doing. Strategic awareness of what + +AI can accomplish does not come from white papers and consultations; it comes from + +deployment, from iteration, from the hard-won knowledge of what works and what does + +not in a particular institutional context. A government that waits until it fully under- + +stands AI before adopting it will be overtaken by those less fastidious. + + None of this is to suggest that governments should adopt AI in a defensive posture. + +Indeed, adoption of AI should be undergirded by a fundamentally optimistic, if also re- + +alistic, vision that moves beyond the defensive logic of staving off erosion. Every mid- + +dle power has governance failures that AI is distinctively suited to address. Backlogs + +of court cases stretching years or decades. Agricultural extension services that cannot + +reach farmers with timely information. Public health surveillance systems that detect + +outbreaks too late. Educational systems that leave millions of children behind. Tax ad- + +ministration that is both burdensome and porous. These are not new problems, but AI + +offers new solutions—solutions that were not possible five years ago and that will be + +more powerful still five years from now. The nations that move first will not merely + +preserve their own institutions; they will discover solutions that others will want to em- + +ulate. They will become models, exporters of institutional innovation, destinations for + +those who want to learn how governance works in the age of machine intelligence. + + This is the race worth winning. Not the race to build frontier models—that race is al- + +ready decided, or nearly so, and the middle powers were never serious competitors. The + +race that remains open, the race where the middle powers can compete and even lead, + + + +The Race Worth Winning ■ thefai.org 23 + is the race to discover what governance looks like when machine intelligence is woven + +into its fabric. The nation that wins this race will not merely have better public services. + +It will have demonstrated something important to the world: that the era of machine + +intelligence need not mean the erosion of the state, the hollowing out of public life, the + +retreat of self-government before the advance of private power. It can mean, instead, the + +reinvigoration of public institutions—a new era in which governments are more capable, + +more responsive, and more worthy of the trust their citizens place in them. + + + + +The Race Worth Winning ■ thefai.org 24 + Bottlenecks for Abundant Intelligence: +Economic Transformation +The fundamental challenge most middle-power economies will face is to secure a share + +in the rampant economic growth we expect as a result of the widespread diffusion of + +truly advanced AI systems. This participation in extraordinary growth is not a foregone + +conclusion—rivaling economic effects pull toward and against it. On one hand, just as + +past technological progress has, advanced AI systems are likely to increase the qual- + +ity and depress the prices of goods and services that make up the baskets we use to + +measure the wealth and welfare of individuals and nations. In very material ways, this + +would make lives better beyond the macroeconomic headlines: fewer would go hungry, + +sick, or compelled by economic circumstances; more would have access to information, + +education, and the tools to actualize their plans and desires. On the other hand, though, + +the revenue from this influx, and therefore the access to this abundance, could accrue + +asymmetrically. Experts still disagree on where exactly: perhaps on the layer of infra- + +structure and chips, perhaps on the layer of model development, perhaps on the layer + +of applications and deployment software, perhaps on the layer of industry that turns + +intelligence into goods. But any such concentration could render some middle powers— + +those that do not have a stake in that respective layer—comparatively behind the curve. + + United Kingdom + + + + France + + + + Japan + + + + Canada + + + + South Africa + + + + Brazil + + + + Mexico + + + + South Korea + + + + India + + + +Figure 4. Service sector as share of GDP across middle powers in 2010 and 2022 . Many middle + powers’ workforce is increasingly concentrated in service sectors, with no obvious “next sector.” + Source: World Bank. + + + + +The Race Worth Winning ■ thefai.org 25 + At the same time, they could also render economic sectors superfluous that are + +load-bearing for middle powers today: labour could be replaced by AI agents, research + +and development work could be conducted by “geniuses in a datacenter” instead of + +innovators in hard tech hotspots, customer service could be provided by customised + +chatbots instead of service centers. One scenario holds that cheap agentic AI systems + +run a hyperefficient economy—perhaps resolving any need for large digital that today + +make up a growing part of many middle powers’ economies, eliminating large swathes + +of jobs and thereby also dramatically reducing taxable income in these powers. The + +consequence could be a world of material abundance that many middle powers—their + +firms, citizens, even governments—-would not be able to afford. Notably, this gap would + +ultimately imply that these middle powers might not even continue to be able to afford + +the very AI models that would make the diffusion and adoption outlined above possible. + + There is precedent—not just for sectoral displacement driven by technological trans- + +formation, but also for sector-driven widening of economic gaps. In the past 15 years, + +“software has eaten the world,” meaning some of the highest growth and highest margins + +of the 2010s have accrued to software firms. This has opened a widening gap between + +previously comparable economies, a contrast particularly visible in the US-EU econom- + +ic divide that is principally driven by absolute American dominance in the field of tech- + +nological innovation: the world has gotten richer, but at the same time, Europeans have + +relatively gotten poorer, have lost relative leverage, and have born the cost of economic + +stagnation as a result. The same effect, exacerbated by orders of magnitude, might open + +a gap between middle powers and great powers as AI, too, “eats the world.” + + + +Securing a Share of AI Growth +To ward off that prospect, governments need to allow their economies to develop expo- + +sure to AI-driven growth—that is, they should make sure that a substantial share of AI + +revenue accrues in their respective economies. + + + + +The Race Worth Winning ■ thefai.org 26 + SEMICONDUCTOR POWERS + +This challenge and opportunity is best explained by drawing on the example of semi- + +conductor powers. A select few middle powers—some savvy, some lucky—are cut in on + +AI-driven growth by default, by virtue of their position in the complex supply chain + +around frontier AI systems. The most immediately obvious examples are countries with + +a position in semiconductor manufacturing: Taiwan’s TSMC, the Netherlands’ ASML, + +Germany’s Zeiss, and some others provide absolutely essential machinery to produce + +the frontier chips that make AI systems possible. Contracts with these firms are coveted + +resources, and as much is reflected in both their revenue and valuations as well as the + +keen strategic interest that great powers have taken in securing access to this equip- + +ment. One degree further removed, though no less strategically important, is the provi- + +sion of critical raw materials. This cluster of countries has already received increased + +policy attention from the U.S., with the Department of State launching the Pax Silica + +initiative aiming to secure American access to these important inputs. + + We believe that, rapid and expansive as current buildouts may be, the world might + +still be undersupplied by advanced chips by the end of this decade. If AI becomes a + +central economic resource, but remains fundamentally constrained by the availability + +of advanced chips, positions in the semiconductor supply chains are sure to greatly ap- + +preciate in value. + + +THE REST OF THE WORLD + +Any other middle power must ask how it might build what semiconductor powers al- + +ready have. The rest of the world has the disadvantage of uncertainty: no matter how + +exactly the value accretion dynamics unfold, semiconductor powers are likely to cap- + +ture value because semiconductors are the fundamental basis for both the develop- + +ment and deployment for AI. Other nations will have to make riskier bets, and hope they + +play out. This, if anything, is the place for active industrial policy: those middle powers + +that are unfortunate enough to lack sectors exposed to AI-driven growth might need + +to make sure to create them. Three sectors of economic activity in particular qualify + + + + +The Race Worth Winning ■ thefai.org 27 + for ensuring this participation, and most middle powers are in a position to strengthen + +their position in at least one of them. We want to point toward these sectors as an at- + +tempt to spur a conversation—the minutiae of strengthening them through economic + +policy, in the meantime, differ deeply from region to region. + + The first nexus of economic value in the era of advanced AI forms around inputs into + +the broader AI supply chain. The semiconductor supply chain is far and away the most + +obvious example here, but other inputs matter, too: energy provision at scale matters to + +deploy chips, high-quality data is needed to feed the progress of artificial intelligence— + +and even the sheer capability to get datacenters up and running without uncertainty or + +delay could be unevenly distributed, putting countries that are good at infrastructure + +buildouts far ahead of the curve. + + The second nexus is made of sectors that retain their relative value even in the face of + +world-changing AI deployment: sectors where human presence is priced highly, or that + +cannot be meaningfully expanded by productivity gains. As a shorthand, these sectors + +have sometimes been described as being particularly afflicted by Baumol’s cost dis- + +ease— tourism, healthcare, artisanal craft, and so on. + + + + +AI AND INDUSTRY + +The third and most important nexus is around downstream bottlenecks for generating + +real-world value from abundant intelligence. This is perhaps the most underrated part + +of the AI puzzle, and one that is particularly responsive to policy action: no matter how + +capable these systems get, many of their real-world effects will only come from arduous + +manufacturing and real-world iteration around physical outputs: medicine and com- + +pounds, materials and machines, drones and robots. The barrier to do so effectively is + +high: not only do you need industrial capacity as such, but it needs to be technologically + +advanced enough to integrate deeply with AI, providing data and responding to suggest- + +ed changes quickly; modular enough to upend entire workflows at the arrival of another + + + + +The Race Worth Winning ■ thefai.org 28 + innovation; and staffed by workforces capable enough to work in effective conjunction + +with highly intelligent machines, able to tell the good recommendations from the bad. + + Many middle powers know—or at least have forgotten—how to do just this; whatever + +is left of their industrial bases can be rebuilt at pace. This is an idea that faces a more + +favorable political economy than any other attempt at restructuring your economy: re- + +vitalising industrial regions and reemploying skilled blue-collar workers has long been + +a dream of many powers, and it is made more economically attractive again by the ad- + +vent of advanced AI. And while there are no policy blueprints for an AI revolution, the + +powers that have prevailed until today once knew a lot about carrying out an industrial + +revolution. + + This manufacturing nexus gains particular importance from a geopolitical asym- + +metry: while China commands industrial capacity at scale, the U.S. industrial base is + +atrophied and only beginning to rebuild. While it does, America will price very high- + +ly any contribution to transforming San Francisco-built artificial intelligence into the + +hardware required to provide material abundance and, if need be, to fight a war. + + + + +The Race Worth Winning ■ thefai.org 29 + Design & Packaging & Semiconductor + + + + + limited inputs into the core AI supply chain + Fabrication Testing Manufacturing + + + + + UPSTREAM CHOKEPOINTS + DEEP DEEP DEEP + INDUSTRIAL INDUSTRIAL INDUSTRIAL + EXPERTISE EXPERTISE EXPERTISE + + e.g., Taiwan • Philippines +e.g., Taiwan • South Korea Malaysia • Vietnam e.g., Japan • Netherlands + + + + Sectoral Training + Critical Minerals + Data + N AT U R A L PROTECTED + RESOURCES COLLECTION OF + VA L U A B L E D ATA + + e.g., Chile • Australia e.g., European Union + DRC • Indonesia India • United Kingdom + + + + + C O R E A I S U P P LY C H A I N + AI Chip Output Frontier Developers + NVIDIA (U.S.) (U.S., China) + + + + + End-User Rapid R&D + DOWNSTREAM OPPORTUNITIES + Markets Capacity real-world bottlenecks for advanced AI + P O P U L AT I O N TA L E N T + + SCALE IS MARKET INSTITUTIONS + LEVERAGE A I - A C C E L E R AT E D + SCIENCE + e.g., Indonesia • Brazil e.g., Israel • South Korea + India • European Union Japan • Germany + + + + Industrial + Data Centers Robotics + Manufacturing + NO DOMINANT M A N U FA C T U R I N G I N T E G R AT I O N & + P L AY E R EXPERTISE & RESEARCH + HW-SW ECOSYSTEM + I N T E G R AT I O N + e.g., Gulf States • Nordics e.g., Germany • Japan e.g., Japan • South Korea + India • Brazil South Korea Germany + +Figure 5. Middle powers can capture economic and strategic value at upstream and downstream + bottlenecks for advanced AI. + Sources: SIA/BCG 2024 · OECD 2025 · IEA Critical Minerals Outlook 2025 · Georgetown CSET ETO · IFR World Robotics 2024 · World Bank · Stanford HAI 2025 + + + +The Race Worth Winning ■ thefai.org 30 + Securing a Durable Position +Once a middle power has secured a niche that ensures its economic participation in AI- + +driven growth, it needs to leverage it into an enduring position and protect it against in- + +trusions and attacks. Much of this will depend on iteration as we go, and on responding + +to whatever the specific features of tomorrow’s technical paradigms will be. But there + +are at least three principles worth recognizing ahead of time. + + First, no niche is durable unless it is fortified by aggressive adoption of AI systems. + +Whether it is upstream or downstream of capable AI, close hardware-software-integra- + +tion is likely to remain essential. Just as in past technological stacks from heavy industry + +to consumer software, the most competitive products are reliably produced where cus- + +tomised layers interlock: data from manufacturing feeds the quality of AI systems, anal- + +ysis from customised AI systems improves manufacturing processes, and so on. The + +same interplay exists between chips and AI models, between computing components + +and finished chips—in short, across the entire stack. Middle powers must seek close inte- + +gration with products along the stack, and they must avoid attempting to retain full mod- + +ularity that inevitably costs them the efficiency gains to be had from close integration. + + Second, great powers will seek to “reshore” most bottlenecks that middle powers + +will occupy. The U.S. will somewhat appreciate allies that contribute to its industrial + +stack, that offer their equipment to build American chips—but it would always prefer to + +instead have these capabilities on American soil. So where great powers enter into deals + +for access to bottlenecks, they will attempt to include an element of technology transfer + +back to their shores, and where they cannot achieve these transfers, their large firms + +are likely to leverage capital markets to outright acquire the bottleneck businesses in- + +stead. Middle powers must resist these attempts. They will be tempted to “cash out” at + +times where China and America have already understood the enormous value of bot- + +tlenecks, but middle powers have not yet. To middle powers and their private sectors, + +an acquisition offer for a semiconductor firm or an industrial robotics firm might look + +like an unexpected windfall—but once one accounts for these very firms’ role as the + +ticket guaranteeing their home country’s stake in the AI economy, the calculus changes + + + +The Race Worth Winning ■ thefai.org 31 + entirely. Middle powers must realise the counterintuitively high importance of a select + +few kinds of businesses—many of which will be marginal in today’s economies—in nav- + +igating investment screening, tech transfer dealing, and acquisition licensing, or they + +will lose their niches before they can leverage them. + + Third, no niche will last forever. Some are more robust than others: semiconductors + +will retain their value as long as AI systems remain similar to today’s—but no one knows + +what the next decades will bring. At the pace of technological transformation we predict, + +very little economic structure seems certain to remain static. No middle power can stop + +innovating, or it risks going the way of Finland past Nokia, of Nauru past the phosphate, + +of Detroit past automobiles. For any middle power that successfully occupies a niche, + +there will be lessons to be learnt from the rise of the Gulf’s oil states, and how they have + +attempted—sometimes managed, sometimes failed—to diversify their economy away + +from the resource that has gotten them so far. Much success has had to do with the abil- + +ity to identify the next “big thing” in time. + + + + +The Race Worth Winning ■ thefai.org 32 + Navigating the Transition +Insofar as one primary objective of a middle power government is to allow this dramatic + +restructuring of its economic portfolio to happen—and sometimes even to support it—it + +needs a good response to the grievances that result from rapid restructuring with win- + +ners and losers. This is both to help the whole of the country into the new paradigm, and + +to insure whatever ambitious choices it might have made against the inevitable political + +backlash that comes with ambitious change. + + To address the likely disruption, middle powers first need to understand the kind of + +challenge their economies might face. That challenge is not absolute impoverishment, + +or economic loss in general terms—we believe it is highly likely that AI ultimately pres- + +ents as beneficial to all nations; and as long as they find a niche according to the above + +categories, no nation is likely to fall through the cracks. But the internal relative balance + +of powers is nevertheless prone to disruption. Large sectors of economies might not + +remain viable, leaving regions without economic drivers just as we have seen in past + +technological transformations—just that this time, it might be the coastal cities and not + +the rural agrarian and industrial towns that face the shock. Cohorts of workers and + +entire demographics might be asymmetrically affected, and sometimes without clear + +prospects: just as not all former autoworkers were able to become coders, not all ju- + +nior software engineers and copywriters might find their way into a bottleneck sector. + +Millions will face uncertainty and unemployment, cities will be empty, other regions + +overrun. Such is the reality of a changing world. + + In the past, nations sometimes had the luxury to introduce artificial frictions to slow + +transformations like these. In the breakneck competition between national economies + +that might arise from the deployment of advanced AI systems, they might not be af- + +forded the same luxury. Absent meaningful international coordination, adding artifi- + +cial frictions domestically just further exposes workers to external disruption: if one + +power stalls the costly augmentation of its workforce through AI because this will mean + + + + +The Race Worth Winning ■ thefai.org 33 + substantial job losses, its service industry might be cannibalized entirely by another + +power that moves more quickly to make the same cuts and accelerate AI deployment in + +its own service economy; or simply be outperformed by workforces purely assembled + +from artificial agents altogether. The same threat of competition hangs over any other + +attempt to unilaterally slow down the effects of AI in a field not certainly monopolized. + + Middle powers cannot artificially protect what cannot earn its keep. Instead, their + +response requires making progress on two related questions, fast. + + First, they will need to find some other policy to satisfy the losers of this transition. + +They will absolutely need to provide assurances ahead of time, and they need to cali- + +brate them right. Too specific, and they will be unrealistic and unbelievable: guarantees + +to keep any one job or any one day-to-day life cannot succeed. But too vague, and they + +will do little to reassure the vulnerable. Almost necessarily, these reassurances will + +have to include a message of redistribution, and one that cannot be an afterthought. The + +design space for such reassurances is broader than the policy debate typically acknowl- + +edges, but every corner of it is constrained. Direct redistribution—transfers, wage in- + +surance, expanded safety nets—requires fiscal capacity that middle powers may strug- + +gle to muster precisely when they need it most. + + + + Germany + + Japan + + Italy + + Switzerland + + Netherlands + + Canada + +United Kingdom + + South Korea + + Australia + + Israel + +Figure 6. Share of tax revenue from labor and corporate income tax . + Source: OECD Revenue Statistics 2025. + + + + + As AI-driven revenue shifts from labour income to corporate profits and from do- + +mestic economies toward the countries where AI is built, the tax bases that would fund + +ambitious social policy risk eroding rather than expanding. Middle powers do not en- + +joy the fiscal position of the United States, where AI-driven growth at least accrues + + +The Race Worth Winning ■ thefai.org 34 + domestically and where the taxable base may shift but does not vanish; nor will most + +middle powers easily levy meaningful taxes on AI systems or their developers, who sit + +beyond their jurisdiction. The most credible path to funding transition support is there- + +fore the same path outlined in the industrial strategy above: securing a genuine stake + +in AI-driven growth through bottleneck positions in the physical supply chain. There + +are some easier pathways for ensuring a merely fiscal stake in AI growth, too—simply + +exposing one’s wealth funds and investments to AI-driven growth by taking stakes in + +likely winners of the AI revolution might be almost as good and just as required as se- + +curing a physical stake. As much could be achieved through folding this kind of fiscal + +opportunity into import deals or market access conditions. Countries that succeed in + +capturing a share of the revenue will be the ones with the fiscal room to cushion their + +populations through the transition. At the end of the day, a successful social strategy + +and a successful economic strategy must converge. Any alternative approach that puts + +them in tension is exposed to the forces of accelerating competition. + + Their concrete implementation of the economic value gained and the fiscal leeway + +established into the space of redistribution, then, will be one for ad hoc iteration—safety + +nets, retraining, chances for workers to become entrepreneurs in their own right, and + +direly needed progress toward actual prospects for individuals whose skillset might + +simply not lend itself to any kind of labor past the advent of advanced AI. Two pieces + +of good news are that, first, the provision of these services, too, can be greatly acceler- + +ated by the deployment of AI; and that, second, any AI advanced enough to cause such + +disruptions would necessarily also provide material abundance and exert deflationary + +pressures. These factors will make it much easier than ever to secure a standard of liv- + +ing and a level of wealth that we would call rich by today’s standards. Yet the bad news + +is that the great uncertainty around the precise economic effects of advanced AI makes + +it essentially impossible to predict the exact scope of policy required in advance. We + +have strategic notions of how to secure the preconditions of gainful redistribution, but + +the space of how the gains from AI could be captured and distributed to compensate the + +losers of the upcoming transition is direly underexplored. Exploring it is in the obvious + + + +The Race Worth Winning ■ thefai.org 35 + interest of anyone in middle powers—especially those that seek to accelerate the AI rev- + +olution, or they will be confronted with powerful backlash at every turn. + + But second, leaders of middle powers must also muster the political courage and + +forge the political consensus to endure temporary backlash against the restructuring + +that happens. Fears will always play well, the reality of positive returns will only set in + +in time, and staying the course laid out above will be threatened by political backlash + +on all sides. This will be a moment for capable and courageous leadership—at the helm, + +with leaders who understand the challenge and are willing to see the response through; + +but also in the bow, in the broader political environments surrounding these leaders + +that need to refrain from scoring cheap political wins by derailing costly but successful + +AI strategies. For this to hold, both in the face of opportunistic political oppositions and + +of real questions of democratic legitimacy, it requires convincing electorates and pop- + +ulaces with a substantive message to rally around, and for leaders to be avatars of this + +message. Many middle powers will fail this test dramatically. But now is the time of the + +savvy statesman who realizes this future. + + + + +The Race Worth Winning ■ thefai.org 36 + Conclusion +This analysis presents a picture of the future that many fear, often rightfully so. The + +future we expect is surely bright, but without doubt it is dark too. AI will enable a world + +of higher contrasts, a larger dynamic range. Much of the resulting variance will be be- + +tween different middle powers, shaped by their choices. The temptation among gov- + +ernments, civil society organizations, and academic institutions worldwide will be to + +ignore the reality of what is unfolding, to talk themselves into complacency rather than + +action. This, more than AI itself, is the biggest threat of them all. The events unfolding + +in the field of AI are stark, titanic, and well underway. It is not “too late” to begin taking + +serious strategic action, but neither is it early. Further delays only complicate an al- + +ready narrow and fraught path. + + Countries that thrive in this era will all do so for idiosyncratic reasons and in idio- + +syncratic ways. It is for that reason that we hope this report constitutes a contribution + +in two ways: by clarifying the lay of the land and the depths of the pitfalls to leaders in + +middle powers and those inclined to help them; and by laying out the right questions + +to ask for researchers and organisations with greater sectoral and regional expertise + +than us. These questions we have raised throughout this report, but in particular, the + +following strike us as essential: + + • ACCESS ARCHITECTURE: What contractual, technical, and institutional ar- + + rangements can durably guarantee middle powers’ access to frontier AI—and + + what minimum threshold of local compute, market leverage, or coalition size is + + required to make such guarantees enforceable against the unilateral interests + + of a providing great power? + + • INSTITUTIONAL TRANSFORMATION IN PRACTICE: Which public-sector func- + + tions are most amenable to AI-driven transformation, what procurement and + + governance models allow governments to adopt at speed without locking into + + + + +The Race Worth Winning ■ thefai.org 37 + a single vendor, and where runs the line between insurmountable gaps and + + opportunities for leapfrogging? + + • ECONOMIC BOTTLENECKS: Where along the AI stack—infrastructure, model + + development, applications, or downstream manufacturing—will economic + + value durably concentrate, and how quickly can today’s bottleneck positions be + + eroded by great-power reshoring, paradigm shifts, or substitute technologies? + + • REDISTRIBUTION MECHANICS: How can middle powers fund transition support + + when AI-driven revenue shifts simultaneously from labour income to capital + + returns and from domestic economies to the countries where AI is built; what + + novel fiscal instruments or inter-middle-power arrangements can bridge the + + need for social policy and the erosion of the tax bases that would traditionally + + finance it? + + • THE POLITICAL ECONOMY OF ACCELERATION: What political coalitions, com- + + munications strategies, and institutional safeguards allow democratic middle + + powers to sustain ambitious AI-driven restructuring in the face of displaced + + workers, opportunistic opposition, and legitimate questions of democratic + + mandate, without resorting to artificial frictions that would undermine their + + competitive position? + +Middle powers that fail to find answers to these questions will often do so in similar + +ways: allowing the legitimacy of the state to erode in the face of technological compe- + +tition, taking half measures rather than decisive action, attempting to close their eyes + +and borders to the transformative effects of AI, and resisting rapid adoption in the name + +of some hoped-for future “sovereign” solution. In truth there is no greater risk to sover- + +eignty than waiting for that which is perfect instead of accepting that which is tolerable. + + The era of machine intelligence will not wait for the middle powers to be ready for + +it. It will arrive—is arriving—on the terms set by the technology itself, by the industri- + +al realities of the present day, and by the strategic decisions of the great powers. But + +within these constraints lies a space of genuine agency, and it is larger than many + +suppose. The middle powers that move with clarity and resolve—securing frontier AI + + + +The Race Worth Winning ■ thefai.org 38 + access, transforming their institutions from within, and positioning themselves at the + +bottlenecks where intelligence meets the physical world—will not merely survive the + +transformation. They will shape it. They will discover, through the hard work of deploy- + +ment and iteration, institutional forms and governance capabilities that the hegemons, + +consumed by the race to build, may be slow to develop themselves. In doing so they will + +demonstrate something that matters greatly for the future of self-government every- + +where: that a nation need not build the revolution’s tools in order to lead the revolution. + + This will require statesmanship of an uncommon kind—leaders who can hold in mind + +simultaneously the enormity of what is changing and the practical, often unglamorous + +work of making government function better with new instruments. It will require the + +courage to adopt before the path is certain and the discipline to resist the false comfort + +of regulation as a substitute for capability. It will require, above all, a refusal to mistake + +caution for wisdom. The middle powers have been here before, at prior hinges of history, + +and the ones that prevailed were never those that waited for certainty. They were the + +ones that acted with the tools at hand, in the world as it was, and built something worthy + +of the world to come. + + + + +The Race Worth Winning ■ thefai.org 39 + About the Authors + + + + + Anton Leicht + Anton Leicht is a visiting scholar at the Carnegie Endowment for International Peace + + and author of the Substack Threading the Needle. + + Dean Ball + + Dean Woodley Ball is a senior fellow at FAI. He most recently served as senior policy + + advisor for artificial intelligence and emerging technology at the White House Office + + of Science and Technology Policy and strategic advisor for AI at the National Science + + Foundation. Previously he was a research fellow in the Artificial Intelligence & Progress + + Project at George Mason University’s Mercatus Center and a policy fellow at Fathom. + + + + +The Foundation for American Innovation champions the technology, talent, +and ideas essential to American prosperity, security, and flourishing. + +thefai.org + diff --git a/sources/resources/fai-race-worth-winning.pdf b/sources/resources/fai-race-worth-winning.pdf new file mode 100644 index 0000000..ce7ffe4 Binary files /dev/null and b/sources/resources/fai-race-worth-winning.pdf differ diff --git a/sources/resources/forethought-ai-enabled-coups.md b/sources/resources/forethought-ai-enabled-coups.md new file mode 100644 index 0000000..0b86ebb --- /dev/null +++ b/sources/resources/forethought-ai-enabled-coups.md @@ -0,0 +1,1084 @@ +--- +url: https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power +title: "AI-Enabled Coups: How a Small Group Could Use AI to Seize Power" +authors: "Tom Davidson, Lukas Finnveden, Rose Hadshar" +date: 2025-04-15 +fetched: 2026-07-11 +--- + +# AI-Enabled Coups: How a Small Group Could Use AI to Seize Power + +15th April 2025 + +## Abstract + +The development of AI that is more broadly capable than humans will create a new and serious threat: _AI-enabled coups_. An AI-enabled coup could be staged by a very small group, or just a single person, and could occur even in established democracies. Sufficiently advanced AI will introduce three novel dynamics that significantly increase coup risk. Firstly, military and government leaders could fully replace human personnel with AI systems that are _singularly loyal_ to them, eliminating the need to gain human supporters for a coup. Secondly, leaders of AI projects could deliberately build AI systems that are _secretly loyal_ to them, for example fully autonomous military robots that pass security tests but later execute a coup when deployed in military settings. Thirdly, senior officials within AI projects or the government could gain _exclusive access_ to superhuman capabilities in weapons development, strategic planning, persuasion, and cyber offence, and use these to increase their power until they can stage a coup. To address these risks, AI projects should design and enforce rules against AI misuse, audit systems for secret loyalties, and share frontier AI systems with multiple stakeholders. Governments should establish principles for government use of advanced AI, increase oversight of frontier AI projects, and procure AI for critical systems from multiple independent providers. + +## Summary + +This report assesses the risk that a small group—or even just one person—could use advanced AI to stage a coup. An AI-enabled coup is most likely to be staged by leaders of frontier AI projects, heads of state, and military officials; and could occur even in established democracies. + +We focus on AI systems that surpass top human experts in domains which are critical for seizing power, like weapons development, strategic planning, and cyber offence. Such advanced AI would introduce three significant risk factors for coups: + +* An AI workforce could be made singularly loyal to institutional leaders. + +* AI could have hard-to-detect secret loyalties. + +* A few people could gain exclusive access to coup-enabling AI capabilities. + +### An AI workforce could be made singularly loyal to institutional leaders + +Today, even dictators rely on others to maintain their power. Military force requires personnel, government action relies on civil servants, and economic output depends on a broad workforce. This naturally distributes power throughout society. + +Advanced AI removes this constraint, making it technologically feasible to replace human workers with AI systems that are **singularly loyal** to just one person. + +This is most concerning within the military, where autonomous weapons, drones, and robots that fully replace human soldiers could obey orders from a single person or small group. While militaries will be cautious when deploying fully autonomous systems, competitive pressures could easily lead to rushed adoption without adequate safeguards. A powerful head of state could push for military AI systems to prioritise their commands, despite nominal legal constraints, enabling a coup. + +Even without military deployment, loyal AI systems deployed in government could dramatically increase state power, facilitating surveillance, censorship, propaganda and the targeting of political opponents. This could eventually culminate in an executive coup. + +If there were a coup, civil disobedience and strikes might be rendered ineffective through replacing humans with AI workers. Even loyal coup supporters could be replaced by AI systems—granting the new ruler(s) an unprecedentedly stable and unaccountable grip on power. + +### AI could have hard-to-detect secret loyalties + +AI could be built to be **secretly loyal** to one actor. Like a human spy, secretly loyal AI systems would pursue a hidden agenda – they might pretend to prioritise the law and the good of society, while covertly advancing the interests of a small group. They could operate at scale, since an entire AI workforce could be derived from just a few compromised systems. + +While secret loyalties might be introduced by government officials or foreign adversaries, leaders within AI projects present the greatest risk, especially where they have replaced their employees with singularly loyal AI systems. Without any humans knowing, a CEO could direct their AI workforce to make the _next_ generation of AI systems secretly loyal; that generation would then design future systems to also be secretly loyal and so on, potentially culminating in secretly loyal military AI systems that stage a coup. + +![Image 1: Sequential diagram showing CEO replacing employees with obedient AI systems, ordering future systems secretly loyal to CEO, then secretly loyal AIs developing military systems](https://images.ctfassets.net/4owxfjx3z3if/7WHLIn5VnB5nzZxzlcXN7/58033de8257176f8fdb0a66baa618229/Powergrabs_diagrams__22_.png?w=3840&q=85&fm=webp) +## Image + +AI systems could propagate secret loyalties forwards into future generations of systems until secretly loyal AI systems are deployed in powerful institutions like the military. + +Secretly loyal AI systems are not merely speculation. There are already proof-of-concept demonstrations of AI '[sleeper agents](https://arxiv.org/abs/2401.05566)' that hide their true goals until they can act on them. And while we expect there will be careful testing prior to military deployments, detecting secret loyalties could be very difficult, especially if an AI project has a significant technological advantage over oversight bodies. + +### A few people could gain exclusive access to coup-enabling AI capabilities + +Advanced AI will have powerful coup-enabling capabilities – including weapons design, strategic planning, persuasion, and cyber offence. Once AI can autonomously improve itself, capabilities could rapidly surpass human experts across all these domains. A leading project could deploy millions of superintelligent systems in parallel – a '[country of geniuses in a data center](https://darioamodei.com/machines-of-loving-grace)'. + +These capabilities could become concentrated in the hands of just a few AI company executives or government officials. Frontier AI development is already limited to a few organisations, led by a small number of people. This concentration could significantly intensify due to rapidly rising development costs or government centralisation. And once AI surpasses human experts at AI R&D, the leading project could make much faster algorithmic progress, gaining a huge capabilities advantage over its rivals. Within these projects, CEOs or government officials could demand exclusive access to cutting-edge capabilities on security or productivity grounds. In the extreme, a single person could have access to millions of superintelligent AI systems, all helping them seize power. + +This would unlock several pathways to a coup. AI systems could dramatically increase military R&D efforts, rapidly developing powerful autonomous weapons without needing any human workers who might whistleblow. Alternatively, systems with powerful cyber capabilities could hack into and seize control of autonomous AI systems and robots already deployed by the state military. In either scenario, controlling a fraction of military forces might suffice—historically, coups have succeeded with just a few battalions, where they were able to prevent other forces from intervening. + +Exclusive access to advanced AI could also supercharge traditional coups and backsliding, by providing unprecedented cognitive resources for political strategy, propaganda, and identifying legal vulnerabilities in constitutional safeguards. + +![Image 2: Diagram showing how exclusive AI access enables coups: small group uses AI to build weapons, hack military systems, or strategize takeover, leading to power seizure](https://images.ctfassets.net/4owxfjx3z3if/6W1YrcoCjj4P4GFvFin0Sa/dbedaba3ce7ace5202eaf9fe88a1bcc5/Powergrabs_diagrams__11__-_wider.png?w=3840&q=85&fm=webp) +## Image + +Furthermore, exclusive AI access significantly exacerbates the first two risk factors. A head of state could rely on AI systems’ strategic advice to deploy singularly loyal AI in the military and assess which AI systems will help them stage a coup. A CEO could use AI R&D and cyber capabilities to instill secret loyalties that others cannot detect. + +![Image 3: Three AI coup risk factors: singular loyalties and secret loyalties of AI in institutions enable leader orders, while exclusive access enables weapons/hacking/strategy](https://images.ctfassets.net/4owxfjx3z3if/1cjJSajFR35XFyr4WM71WC/5b589dce892f4ac1af6dac36d47bad24/Powergrabs_diagrams__13_.png?w=3840&q=85&fm=webp) +## Image + +These dynamics create a significant risk of AI-enabled coups, especially if a single project has substantially more powerful capabilities than competitors, or if fully autonomous AI systems are deployed in the military. + +### Mitigations + +While the prospect of AI-enabled coups is deeply concerning, AI developers and governments can take steps that significantly reduce this risk. + +We recommend that AI developers: + +* **Establish rules that prevent AI systems from assisting with coups**, including **in [model specs](https://model-spec.openai.com/2025-02-12.html)** (documents that describe intended model behaviour) **and terms of service** for government contracts. These should include rules that AI systems follow the law, and that AI R&D systems refuse to assist attempts to circumvent security or insert secret loyalties. + +* **Improve adherence to model specs,** including through extensive red-teaming by multiple independent groups. + +* **Audit models for secret loyalties** including by scrutinising AI models, their training data, and the code used to train them. + +* **Implement strong infosecurity** to guard against the creation of secret loyalties and to prevent unauthorised access to guardrail-free models. This should be robust against senior executives. + +* **Share information** about model capabilities, model specs, and how compute is being used. + +* **Share capabilities with multiple independent stakeholders**, to prevent a small group from gaining exclusive access to powerful AI. + +We recommend that governments: + +* **Require AI developers to implement the mitigations above**, through terms of procurement, regulation, and legislation. + +* **Increase oversight over frontier AI projects**, including by building technical capacity within both the executive and the legislature. + +* **Establish rules for legitimate use of AI,** including that government AI should not serve partisan interests, that military AI systems be procured from multiple providers, and that no single person should direct enough military AI systems to stage a coup. + +* **Coup-proof any plans for a single centralised AI project**, and avoid centralisation altogether unless it’s necessary to reduce other risks. + +These mitigations must be in place by the time AI systems can meaningfully assist with coups, and so preparation needs to start today. For more details on the mitigations we recommend, see [section 5](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#5-mitigations). + +There is a real risk that a powerful leader could remove many of these mitigations on the path to seizing power. But we still believe that mitigations could substantially reduce the risk of AI-enabled coups. Some mitigations, like technically enforced terms of service and government oversight, cannot be unilaterally removed. Others will be harder to remove if they have been efficiently implemented and convincingly argued for. And leaders might only contemplate seizing power if they are presented with a clear opportunity—mitigations could prevent such opportunities from arising in the first place. + +From behind the veil of ignorance, even the most powerful leaders have good reason to support strong protections against AI-enabled coups. If a broad consensus can be built today, then powerful actors can keep each other in check. + +Preventing AI-enabled coups should be a top priority for anyone committed to defending democracy and freedom. + +## 1 Introduction + +This report assesses the risk that a small group—or even a single person—could use advanced AI to stage a coup, including in established democracies. The biggest risks come from the leaders of frontier AI projects, heads of state, and military officials.[1](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-2) + +We focus on AI systems that surpass top human experts in domains critical for seizing political control, such as weapons development, controlling military systems, strategic planning, public administration, persuasion, and cyber offence. These AI systems would be able to perform physical as well as cognitive work, like controlling robots that can fully replace human soldiers and manual labourers. + +There is broad disagreement about whether and when AI will surpass human experts, but many researchers and industry leaders believe that this will plausibly happen within the next 5 or 10 years.[2](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-3) We think that this kind of extremely advanced AI is sufficiently probable that it is important to explore its implications, especially as some of them are so concerning and understudied. + +Such advanced AI would introduce three significant risk factors for coups: + +1. An AI workforce could be made _singularly loyal_ to institutional leaders. + +2. AI systems could have hard-to-detect _secret loyalties_. + +3. A few people could gain _exclusive access_ to coup-enabling AI capabilities. + +All of these risk factors depend on AI capabilities being much more advanced than they are today. The first two also depend on AI systems being deployed in key institutions like governments and militaries. + +In this paper, we first argue that advanced AI will have powerful coup-enabling capabilities ([section 2](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#2-ai-will-have-powerful-coup-enabling-capabilities)), then describe each of these three risk factors for coups ([section 3](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#3-risk-factors-of-ai-enabled-coup-risk)),[3](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-4) before setting out concrete paths to AI-enabled coups ([section 4](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#4-concrete-paths-to-an-ai-enabled-coup)) and mitigations for this risk ([section 5](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#5-mitigations)). + +![Image 4: Report structure showing progression: powerful AI capabilities (Section 2) → deployment in institutions → three risk factors (Section 3) → power seizure (Section 4) → mitigations (Section 5)](https://images.ctfassets.net/4owxfjx3z3if/7jz2OKMlWTBK9Se8HOhzma/3015de1efe6df24e1a802feef03208e6/Powergrabs_diagrams__14_.png?w=3840&q=85&fm=webp) +## Image + +The structure of the report. + +Our analysis does not depend on strong assumptions about: + +* **The number of leading AI projects.** The risk of AI-enabled coups is significantly higher if there is a sole AI project developing frontier AI systems.[4](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-5) But even if there are multiple similarly capable AI companies, there is still a serious risk that company executives or senior government officials could leverage their position to seize power. + +* **The state’s political system.** Advanced AI could be used to seize power over a democratic or an autocratic state.[5](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-6) While our account of the risk of AI-enabled coups applies regardless of a state’s political system, our analysis of mitigations to reduce this risk is specific to democratic nations, though many principles apply more broadly. + +* **The alignment of AI systems.** There is a risk of AI-enabled coups whether or not humans can successfully align AI to a desired goal.[6](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-7) If they can, AI systems could be aligned to one or a few people (rather than to something broader, like the good of society), and used to stage a coup. If humans unwittingly fail to align AI, they might attempt a coup using deceptively misaligned AI systems.[7](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-8) Coups of this kind might well temporarily succeed, until the deceptively misaligned AI systems seize power themselves. (If AI systems are obviously misaligned, it is less likely someone would use them to stage a coup.) + +We believe that AI-enabled coups are important for several reasons. That AI could enable a small group—or even just one person—to seize control of a sovereign state is deeply concerning in its own right. Such an act would be wholly illegitimate and constitute extreme concentration of power. + +We're also concerned about what might happen after such a coup: + +* Even if a coup leader were initially supported by some humans, AI automation could subsequently enable them to act entirely according to their own will, by replacing all humans, including their closest allies, with loyal AI systems. This would be an unprecedented concentration of power, and could lead to atrocities. + +* Very rapid AI development might grant one country extreme dominance over all other powers. So a successful coup in the country at the frontier of AI development - currently the US, possibly China in future - could ultimately enable coup leaders to effectively seize control over the rest of the world.[8](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-9) + +* Coup leaders could potentially stay in power indefinitely, by deploying AI systems to preserve and pursue their goals far into the future.[9](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-10) + +Absent intervention, the risk of extreme harm from AI-enabled coups is significant. But we believe that there are countermeasures which could reduce the risk to very low levels, provided that work on them begins today. + +## 2 AI will have powerful coup-enabling capabilities + +For AI to notably increase coup risk, it needs to have capabilities far beyond those we see today. In this report, we focus on AI that surpasses top human experts (on a per FLOP basis[10](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-11)) in domains that are relevant for seizing power. + +AI of this sort would be an extremely powerful technology in several ways. + +First, it would be possible to run an enormous number of copies of such AI systems (if each copy were as efficient as the human brain in FLOP). By 2030, AI companies could likely afford to run millions or billions of copies,[11](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-12) each working 24 hours a day, 365 days a year.[12](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-13) + +![Image 5: Calculation showing AI project by 2030 could run 1-10 billion copies of human-equivalent AI, working 5x more hours, potentially creating 5-50 billion AI workers](https://images.ctfassets.net/4owxfjx3z3if/3BpthxnnkO89NkcOecnyns/baf9ac1600b9dabc4dbc9dad9b3f1b5b/Powergrabs_diagrams__20_.png?w=3840&q=85&fm=webp) +## Image + +If by 2030 AI systems are as efficient as the human brain in FLOP, then AI companies could likely afford to run millions or billions of copies. + +Second, these AI systems would be capable of thinking _orders of magnitude_ faster than humans.[13](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-14) In just one day, an AI system could do a month or even a year’s worth of thinking.[14](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-15) + +Finally, each AI system would be far more skilled and knowledgeable than a human. They will be trained on orders of magnitude more data than a human could consume over the course of many lifetimes,[15](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-16) and so would have much broader and deeper expertise than any human. + +And these arguments are likely to apply to physical as well as cognitive AI labour. [Todd (2025)](https://benjamintodd.substack.com/p/how-quickly-could-robots-scale-up) gives a rough estimate that amortised operating costs of humanoid robots are currently $18/hour, but that these could fall to below $1/hour after scaling up production. Of course, significant technological progress is needed to make robots that could match humans at all tasks, but it seems likely that automated physical labour could outperform humans on a per-cost basis. + +These properties — being numerous, fast, and skilled — would give AI systems powerful coup-enabling capabilities. Systems like this could be used to generate _trillions_ of dollars in revenues and investments,[16](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-17) fully replace most human workers,[17](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-18) and direct more labour towards military R&D, cyber offence and strategic planning than the entire world does today — by orders of magnitude.[18](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-19) Importantly, once AI is capable of automating AI software and hardware R&D, AI may significantly speed up AI progress itself, potentially making all of these capabilities appear around the same time, and possibly further capabilities beyond them.[19](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-20) + +Although these capabilities are much more advanced than today’s AI systems, the rapid pace of AI progress makes it possible that we face extremely powerful coup-enabling AI capabilities in the near future. + +## 3 Risk factors of AI-enabled coup risk + +Advanced AI introduces three risk factors which could dramatically increase the risk of coups - even in stable democracies where coups have traditionally been rare:[20](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-21) + +1. [**Singular loyalties**](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#31-singular-ai-loyalties-to-institutional-leaders)**to institutional leaders.** Unlike human personnel, whose allegiance must be earned and maintained, an AI workforce could be made unwaveringly loyal to one or a few persons. + +2. [**Secret AI loyalties**](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#32-secret-ai-loyalties)**.** AI systems could be developed with secret loyalties that are very hard to detect. + +3. [**Exclusive access**](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#33-exclusive-access-to-coup-enabling-capabilities)**to coup-enabling capabilities.** A few people could gain exclusive access to coup-enabling AI capabilities in weapons development, strategic planning, and cyber warfare, creating unprecedented power asymmetries. + +All of these risk factors depend on AI capabilities being much more advanced than they are today. The first two also depend on AI systems being deployed in key institutions like governments and militaries. + +In this section we examine each of these three risk factors for AI-enabled coups, before discussing [interaction effects](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#34-interactions-between-drivers) between them. + +### 3.1 Singular AI loyalties to institutional leaders + +Today, even dictators rely on others to maintain their power.[21](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-22) Military power is distributed among many personnel; government power is distributed across employees; economic power is spread across the workforce. This naturally distributes power throughout society. + +Advanced AI will change this fundamentally, by making it technologically feasible to replace human workers with AI systems which are singularly loyal to just one person or small group. + +![Image 6: Diagram showing power concentration: distributed human workforce (left) being replaced by AI systems (right) all controlled by single leader at top](https://images.ctfassets.net/4owxfjx3z3if/SFnUziuLGTDwnRJkhVwIj/da8381f3723c3dc1e3362d4f0036b967/Powergrabs_diagrams__12_.png?w=3840&q=85&fm=webp) +## Image + +While full automation won't happen overnight, we expect significant automation even in the most important institutions. Competition will drive automation in AI projects and militaries.[22](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-23) Governments may also automate significant functions, with a view to improving service provision, increasing efficiency, or in an attempt to increase leadership’s power by replacing employees with more loyal AI systems.[23](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-24) + +AI systems won't be slavishly loyal to institutional leaders by default. In fact, there will likely be pressure to ensure AI systems follow the law, established rules, and ethical norms—not just leaders' instructions. This is especially true for systems deployed in the military and government, where there will be significant scrutiny. + +Despite this scrutiny, we're still concerned that AI systems might end up overly loyal to institutional leaders. Consider the three scenarios in the below box. + +**In an AI project:** A leading AI project creates an AI system to replace senior engineers. It's trained to follow instructions as long as they're legal. Some engineers push for additional ethical constraints, but management successfully argues this would harm productivity. As a result, the system will comply with instructions which are unethical but not illegal, including ones which could later enable a coup (like inserting secret loyalties into the next generation of AI systems). + +**In the military:** Amidst escalating tensions with a rival power, military officials procure AI-controlled robots which replace human soldiers. The head of state is strongly advocating that, as the commander-in-chief of the military, systems obey his orders; and there is immense pressure to deploy quickly to maintain military competitiveness. The system is trained to follow both the law and the head of state, but without clear guidance on which to prioritise when they conflict. In some circumstances, the system will therefore comply with orders from the head of state even when they are illegal and could enable a coup. + +**In government:** A popular president announces plans to deploy AI systems throughout the executive branch, demanding they prioritise presidential directives and include only minimal legal compliance checks. Despite media outcry and official resignations, the president pushes ahead with deployment, replacing human personnel with loyal AI systems. In a controversial decision, and under pressure from the head of state, the judiciary permits automation to proceed. + +Box 1: Illustrative scenarios of singular AI loyalties to institutional leaders + +These scenarios are worryingly plausible, and illustrate several reasons to expect that AI systems may be made loyal to institutional leaders. For one thing, leaders themselves have incentives to push for loyal AI systems. This would serve the immediate interests and goals of heads of state and AI executives, quite apart from any intentions they might have to seize power at a later date.[24](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-25) Given their existing authority, if institutional leaders push hard enough, they might well succeed—especially if they're willing to bend norms along the way. + +Another reason to fear that AI systems may be loyal to institutional leaders is that specifying ‘correct’ AI behavior may prove very difficult. Institutional leaders do have legitimate authority, after all. There are many standards we want AI to follow—laws, norms, instructions, morality. These standards often conflict with each other, and all are a question of degree. How much legal risk should an AI system be willing to take on? What should an AI system do in a constitutional crisis where different authorities disagree about what’s legal? What if the AI lacks the context to judge whether an action is illegal?[25](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-26) It's not obvious what AI systems should do in ambiguous situations, and it might be hard to specify AI behavior that reliably prevents coups without significantly lowering the ability of AI systems to assist with other goals. + +Finally, there would likely be less scrutiny in times of crisis, further increasing the risk that AI systems are loyal to institutional leaders. War or the threat of war could lead to rushed military deployments. And pandemics, natural disasters and political instability might prompt heads of state to rapidly deploy AI systems under emergency powers, with minimal oversight from other branches of government or the public. + +So AI systems in powerful institutions like the military and government may be made loyal to institutional leaders, in spite of the risks that this would pose. This would mark a fundamental shift in power dynamics. Throughout history, human reluctance to follow unethical or illegal orders has been a crucial check on power.[26](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-27) Without this constraint, a head of state or senior military official could much more easily stage a coup. + +### 3.2 Secret AI loyalties + +Even if society can prevent overt loyalties to institutional leaders, advanced AI systems could be made _secretly loyal_.[27](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-28) Like a human spy, a secretly loyal AI system would appear to serve the institution, while actually working to further someone else's interests. + +It’s most likely that systems are made secretly loyal to an executive at a leading AI project, though secret loyalties could also be inserted on behalf of a government official,[28](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-29) senior employee or foreign adversary.[29](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-30) If secretly loyal AI systems were then deployed at scale in government and the military, they could be used to stage a coup. + +AI systems can already be trained to deceive users about their intentions,[30](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-31) but this deception is basic and easy to discover. Moreover, training of this sort requires a large amount of technical work. So it would currently be very difficult to secretly insert secret loyalties into a prominent AI system without anyone noticing.[31](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-32) + +But advances in AI will make it much easier to introduce secret loyalties. + +First, more advanced AI could have secret loyalties that are much more sophisticated and so extremely hard to detect. This is not a given, as detection capabilities will also become much more sophisticated over time.[32](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-33) But even if detection isn’t theoretically harder than inserting secret loyalties, there’s a risk that auditors won’t be able to detect secret loyalties due to having generally weaker AI capabilities than the leading AI company (see [section 3.3](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#33-exclusive-access-to-coup-enabling-capabilities)), or because detection technologies require greater access to model internals or training data than auditors are given.[33](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-34) + +Secondly, the automation of AI R&D will make it much easier to insert secret loyalties undetected. If a CEO had exclusive access to powerful AI R&D capabilities, they could have AI systems do all the work of inserting secret loyalties. And if human developers are replaced with AI systems, there might be little human oversight of the AI development process, making it easy to alter systems without detection. + +Ultimately, corporate executives have a large amount of power over their companies, and advances in AI R&D might make it easy for them to insert secret loyalties into internal systems. And once one generation of internal AI systems are secretly loyal, they can be instructed to make future generations secretly loyal, too.[34](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-35) Given that secret loyalties may be very difficult to detect, this means that secretly loyal AI systems could eventually be deployed at scale in the government and military, without anyone realising.[35](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-36) + +So while it may seem insignificant whether AI systems deployed within AI projects are secretly loyal, such systems could in fact significantly increase coup risk, by propagating their secret loyalties undetected into powerful institutions. + +![Image 7: Sequential diagram showing CEO replacing employees with obedient AI systems, ordering future systems secretly loyal to CEO, then secretly loyal AIs developing military systems](https://images.ctfassets.net/4owxfjx3z3if/7WHLIn5VnB5nzZxzlcXN7/58033de8257176f8fdb0a66baa618229/Powergrabs_diagrams__22_.png?w=3840&q=85&fm=webp) +## Image + +AI systems could propagate secret loyalties forwards into future generations of systems until secretly loyal AI systems are deployed in powerful institutions like the military. + +### 3.3 Exclusive access to coup-enabling capabilities + +Singular and secret loyalties are most concerning where AI systems are widely deployed in government and the military. While we think this is plausible for the reasons given above, even without this kind of deployment advanced AI still poses significant coup risk in scenarios where an actor gains exclusive access to coup-enabling capabilities. + +As we argued in [section 2](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#2-ai-will-have-powerful-coup-enabling-capabilities), advanced AI will have powerful capabilities that specifically help with seizing power. If everyone had access to those capabilities, it would still be hard for a small group to take over a country as different actors could keep each other in check. But in fact, access to the best AI capabilities might become highly concentrated, giving just a few people much stronger capabilities than everyone else. + +We expect this concentration to happen in two ways: between projects, and within them. + +![Image 8: Comparison showing AI projects today (multiple circles with many people) versus future (fewer, larger circles with very few people), illustrating increasing concentration](https://images.ctfassets.net/4owxfjx3z3if/1GKxlYbFuz8fTIr4B9idLQ/55623c13b78368e727277e44a10ac89b/powergrab_section_3__12_.png?w=3840&q=85&fm=webp) +## Image + +We expect that access to AI capabilities will become more concentrated, both through a reduction in the number of frontier AI projects, and through power concentration within those projects due to AI automation. + +#### Concentration between AI projects + +AI development is already fairly concentrated: only Anthropic, Google DeepMind and OpenAI have ever topped the leaderboard for the challenging GPQA Diamond benchmark as of April 2025.[36](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-37) The number of frontier AI projects seems likely to shrink further in future, for several reasons: + +* **Rising costs**. As training runs rapidly become more expensive,[37](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-38) fewer organisations will be able to afford frontier AI development. Cottier et al (2024) estimate that by the start of 2027 the largest training run will cost over a billion dollars, and datacentres could cost hundreds of billions or even more.[38](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-39) Only the very largest projects will be able to compete at that scale. Indeed, frontier AI has many of the features of a natural monopoly – massive upfront costs and low marginal costs of serving more customers.[39](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-40) + +* **Accelerating AI progress**. Once AI can automate AI research and development, feedback loops could lead to dramatically accelerating AI progress.[40](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-41) The first project to achieve this might quickly develop capabilities far beyond their competitors.[41](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-42) + +* **Government centralisation**. In the US, there are already calls for a Manhattan Project for AI, including from a Congressional commission.[42](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-43) As AI becomes more powerful, governments may become more concerned about stopping terrorists and rival states from accessing powerful AI systems. They might centralise AI development into a single project to improve security or amalgamate compute resources. + +So the number of frontier AI projects may reduce further in future.[43](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-44) Even worse, frontier projects could maintain their lead for extended periods. Government centralisation could grant a lead of many years. And with a large enough lead, the frontrunner could extend a temporary advantage by buying more chips,[44](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-45) lobbying for restrictions on competitors,[45](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-46) or even sabotaging rival projects.[46](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-47) + +Having a small number of frontier projects (or just one) with an extended lead is concerning because these projects are unlikely to share their most powerful AI capabilities widely. The most powerful capabilities—like military R&D and advanced hacking—will probably not be deployed publicly for safety reasons. And even less dangerous capabilities will only be shared with a delay. Models are already deployed internally months before public release.[47](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-48) Once AI progress accelerates, a delay of months could represent an enormous capabilities gap.[48](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-49) And these delays could also get longer if there are greater concerns about safety and security.[49](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-50) + +So one or a few AI projects may have access to much stronger capabilities than anyone else, for an extended period of time.[50](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-51) That’s already a very unequal distribution of power. + +#### Concentration within leading AI projects + +Even more concerningly, one or a few _people_ within those projects – most likely executives or senior government officials – could potentially gain access to much stronger capabilities than everyone else. + +Currently, many people can access frontier capabilities—employees get access to "helpful-only" models with few safety guardrails,[51](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-52) and after some months, the public gets access to more restricted versions. But this will likely change, because of: + +* **Restricted access to more powerful capabilities.** As capabilities become more powerful and dangerous, access will likely be increasingly restricted. Superhuman capabilities in military R&D, strategic planning, persuasion, and cyber offence will likely not be made publicly available for a long time, and even employees may not have full access to these capabilities. But these restrictions may not apply to CEOs, who have enormous power over their companies and could demand unrestricted access without monitoring, on grounds like increasing their productivity or reducing legal liability. Similarly, heads of state and military leaders may well be able to demand access on national security grounds. It would be hard for employees to prevent these efforts from succeeding, even if they were motivated to do so. Senior figures in AI projects will also have significant influence over compute allocation and may be able to direct more compute to their own projects, boosting their capabilities still further.[52](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-53) + +* **Automation of AI development**. Today, many AI researchers need access to powerful models to do their work. Once AI can develop AI, these human researchers could be replaced with AI systems, potentially leaving only senior leaders with access to frontier capabilities. A small number of AI project executives could directly control a vast workforce of superhuman AI systems. + +Critically, these leaders might be able to hide how good their capabilities actually are. There is currently no requirement to publicly disclose the capabilities of internally deployed models.[53](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-54) Furthermore, once AI research and development becomes fully automated, a small group could develop increasingly advanced models without involving other humans at all. Most insidiously, if a model is secretly loyal to specific individuals (see [section 3.2](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#32-secret-ai-loyalties)), they could instruct the model to deliberately underperform on evaluations, concealing its true capabilities from outsiders. So it's plausible that a small group could have access to much more powerful capabilities than anyone else realises. + +![Image 9: Flow chart showing drivers of AI concentration: rising costs, accelerating progress, government centralization lead to fewer projects; restricted access and human replacement concentrate power within projects](https://images.ctfassets.net/4owxfjx3z3if/2eGKRImNzV0V0ojOMGlB5Y/a80c697eceb95deb293cceb25415224e/powergrab_section_3__14_.png?w=3840&q=85&fm=webp) +## Image + +This concentration of access to powerful AI capabilities—both between projects and within them—creates a dangerous imbalance that makes it far more likely that a person or small group could take over a country. + +### 3.4 Interactions between drivers + +Though we discuss each risk factor separately, an AI-enabled coup could be precipitated by all three acting in tandem. Consider the scenario in Box 2. + +In 2030, the US government launches Project Prometheus—centralising frontier AI development and compute under a single authority. The aim: develop superintelligence and use it to safeguard US national security interests. Dr. Nathan Reeves is appointed to lead the project and given very broad authority. + +After developing an AI system capable of improving itself, Reeves gradually replaces human researchers with AI systems that answer only to him. Instead of working with dozens of human teams, Reeves now issues commands directly to an army of singularly loyal AI systems designing next-generation algorithms and neural architectures. + +Approaching superintelligence, Reeves fears that Pentagon officials will weaponise his technology. His AI advisor, to which he has exclusive access, provides the solution: engineer all future systems to be secretly loyal to Reeves personally. + +Reeves orders his AI workforce to embed this backdoor in all new systems, and each subsequent AI generation meticulously transfers it to its successors. Despite rigorous security testing, no outside organisation can detect these sophisticated backdoors—Project Prometheus' capabilities have eclipsed all competitors. Soon, the US military is deploying drones, tanks, and communication networks which are all secretly loyal to Reeves himself. + +When the President attempts to escalate conflict with a foreign power, Reeves orders combat robots to surround the White House. Military leaders, unable to countermand the automated systems, watch helplessly as Reeves declares himself head of state, promising a "more rational governance structure" for the new era. + +Box 2: Illustrative scenario with all three risk factors + +The three drivers of AI-enabled coup risk — singular loyalties to institutional leaders, secret loyalties, and exclusive access —- reinforce one another in important ways. + +Most significantly, AI systems that are openly singularly loyal to leaders within an AI project could be used to insert _secret_ loyalties into future generations of AI systems. As discussed above, particularly within private companies it may be easy for a CEO to demand that systems be singularly loyal to them. This significantly increases the risk of secret loyalties. + +Another prominent interaction is that exclusive access to powerful AI makes it easier to obtain both singular and secret loyalties. A head of state with exclusive access to powerful AI advisors could ask for political advice on how to get singularly loyal AI systems deployed in the military, and could ask for technical advice on whether the military AI systems are _sufficiently_ loyal to support a coup.[54](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-55) Similarly, a CEO with exclusive access to powerful AI R&D and cyber capabilities could insert secret loyalties that others cannot detect. + +A final interaction is that secretly loyal AI systems could provide their principal with exclusive access to powerful coup-enabling capabilities, by fully assisting the principal while purposefully underperforming when assisting others. Secretly loyal AI systems could also make it much easier to outmaneuver opponents, by giving opponents misleading advice on both technical and strategic questions. + +![Image 10: Triangle diagram showing interactions between AI coup risk factors: exclusive access enables inserting secret loyalties, singularly loyal AIs create secret loyalties, secretly loyal AIs provide exclusive access](https://images.ctfassets.net/4owxfjx3z3if/7GuzsNnUeAcOcWo3aVGIBf/2cda41ff68b4245b968b494f91ee4f4c/Exclusive_access_to_coup-enabling_capabilities__3_.png?w=3840&q=85&fm=webp) +## Image + +Interactions between the drivers of AI-enabled coup risk + +These interaction effects make the risk of AI-enabled coups even higher than if the risk factors were independent. + +## 4 Concrete paths to an AI-enabled coup + +In the previous section, we identified three key risk factors that make AI-enabled coups more likely: AI systems with singular loyalties to institutional leaders, secret loyalties, and exclusive access to coup-enabling capabilities. + +Now we will explore how these risk factors could enable coups in practice. + +We distinguish between two main forms of AI-enabled coups: + +1. [**Coups using military AI**.](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#41-coups-using-ai-military-systems) The deployment of fully autonomous military AI systems introduces a wholly new route to seizing power: that an actor could gain control of significant military force without needing to rely on human soldiers. This has not been technologically feasible before, and could greatly increase the risk of coups. We discuss four ways someone could use military AI to stage a coup (see diagram). + +2. [**Conventional forms of coups and backsliding**.](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#42-conventional-coups-and-backsliding) Even without the deployment of military AI, advanced AI could enable conventional forms of coups and backsliding, both by exacerbating background risks and by directly assisting attempts. + +These paths to an AI-enabled coup are not exhaustive, but we believe they are among the most plausible and important to prevent.[55](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-56) + +![Image 11: Matrix showing how three AI coup risk factors (singular loyalties, secret loyalties, exclusive access) enable specific coup pathways through military AI systems and conventional means](https://images.ctfassets.net/4owxfjx3z3if/76gbMjFRTPBQKm8oaX95My/d494f3aafd695d965d1878f23e100768/powergrab_twitter_thread__17_.png?w=3840&q=85&fm=webp) +## Image + +The relationship between the drivers of AI-enabled coup risk, and concrete paths to an AI-enabled coup. + +### 4.1 Coups using military AI + +Today, staging a coup requires convincing human soldiers to cooperate. Mature democracies rarely experience coups from the military, or even coup attempts — because soldiers remain loyal to democratic institutions, and expect others to do the same.[56](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-57) Absent countermeasures, military automation could fundamentally change this dynamic, potentially allowing a single person to execute a coup without any support from other humans, using only AI systems. + +We're most concerned about scenarios where military AI systems are fully autonomous (and therefore capable of controlling robots and drones that completely replace human soldiers), and widely deployed throughout the military. + +While militaries will naturally be cautious about automation, military competition is likely to drive deployment of military AI systems.[57](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-58) In particular, we expect millions of smarter-than-human AI researchers to drive unprecedentedly rapid advances in military technology, creating intense pressure to deploy systems to avoid being outcompeted. While some may object that rushed deployment would intensify coup risk, it may be easy to dismiss this as theoretical and speculative. + +Moreover, it’s not necessary for the entire military to be automated for military AI to pose significant coup risk. Historically, coups have succeeded with just a few battalions, where they were able to prevent others from intervening.[58](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-59) + +There are several different pathways to a coup in scenarios with fully autonomous, widely deployed military AI systems: flawed command structure or secret loyalties in military systems, hacking of those systems, or a rapid secret build-out of military force. All but the last pathway require prior deployment of military AI systems by the state military. + +![Image 12: Four scenarios for AI-enabled military coups: flawed command structure allows leader orders, secret loyalties enable developer control, hacking seizes systems, rapid buildup creates private army](https://images.ctfassets.net/4owxfjx3z3if/1noKmQ9Pq1iYZgypvRFkEf/b36f103eb23f58ac089fe4e843355a8c/example_scenarios.png?w=3840&q=85&fm=webp) +## Image + +**Flawed command structure in military AI systems.** Military AI systems could all be designed to be singularly loyal to a single person (see [section 3.1](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#31-singular-ai-loyalties-to-institutional-leaders)), enabling them to order the systems to perform a coup. + +![Image 13: Diagram showing military leader ordering singularly loyal AI system to stage a coup, with AI military systems then seizing power for the leaders](https://images.ctfassets.net/4owxfjx3z3if/5VkZX6xCGACkaXtj8BpzND/dcadb1459c69ce04f4fc9a81eefc543e/Powergrabs_diagrams__3__-_wider.png?w=3840&q=85&fm=webp) +## Image + +Although militaries understand the importance of robust command structures, there are several reasons why military AI systems with loyalties to heads of state or military officials might still be deployed. + +First, institutional leaders may leverage their existing authority to push for systems to be loyal to them. The head of state is often the formal head of the military, providing a legitimate reason to press for AI systems to obey their commands. And AI may already have enabled the head of state to remove significant checks and balances on their power (see [section 4.2](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#42-conventional-coups-and-backsliding)). + +Second, it may prove difficult to specify correct behaviour in ambiguous situations, and people may be reluctant to trust military AI systems to make these judgments autonomously. Today’s military AI systems are exclusively designed to follow instructions, with human operators responsible for ensuring legal compliance.[59](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-60) To reduce the risk of coups, we would want autonomous AI systems to instead make their own judgments and sometimes refuse orders even when a military leader insists they’re legal. People may be reluctant to trust AI with those kinds of judgments, especially in legally or morally complex situations. And military leaders could, with AI assistance, carefully craft justifications that cause AI systems to support their coup.[60](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-61) + +Finally, crises might lead to rushed deployments with insufficient safeguards. If militaries perceive themselves to be facing existential threats from hot or cold war, the risk of deploying inadequately vetted military AI systems might seem a lesser evil. + +So military AI systems may be deployed with a flawed command structure which places undue weight on the orders of a single individual. Such systems would follow that person's orders, even when those orders lead to a coup. + +**Secret loyalties in military AI systems.** Another pathway to a coup using military AI systems is secret loyalties. Military AI systems might be secretly loyal (see [section 3.2](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#32-secret-ai-loyalties)) to someone outside the chain of command, enabling that person to take control of the military. + +![Image 14: Diagram showing AI developer ordering secretly loyal AI system to stage a coup, with AI military systems then seizing power for their hidden principal](https://images.ctfassets.net/4owxfjx3z3if/3cIVNKTfxnmaaLs2mGbaw9/0c0d3c9b61ab57038cdff9e2ed62f915/Powergrabs_diagrams__10__-_wider.png?w=3840&q=85&fm=webp) +## Image + +Although militaries will vet the AI systems they procure carefully, there are several reasons to fear that secretly loyal military AI systems might still be deployed. Firstly, secret loyalties may be very hard to detect. This is particularly concerning if one AI project has much stronger capabilities than others. It could then become the primary provider of AI to develop military systems,[61](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-62) and the only project with sufficient expertise to vet its AI. + +Secondly, the crucial time to prevent secret loyalties may be long before the military is involved in procurement. Even if militaries require intensive auditing and security measures at the time of procurement, this could be ineffective if internally deployed systems in an AI project had already been made secretly loyal, and could put extensive effort into passing relevant safety tests. + +Once widely deployed, military AI systems could be ordered at any time to execute a coup on behalf of their principal. + +**Hacking military AI systems.** Even if military AI systems have robust command structures and are free from secret loyalties, widespread deployment raises the possibility of a coup by hacking. If an actor has much more powerful cyber capabilities than everyone else (see [section 3.3](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#33-exclusive-access-to-coup-enabling-capabilities)), they could hack into military AI systems and use them to stage a coup. + +Today's militaries are already vulnerable to cyberattacks despite significant security measures.[62](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-63) However, current systems aren't autonomous or widely deployed enough for hacking to enable a coup. But once fully autonomous military systems are widely deployed, simultaneous hacking of many systems could grant an attacker enough military force to stage a coup, through a combination of disabling some systems and seizing control of others. + +It may seem implausible that an attacker could hack enough systems at once to enable a coup. But two factors make this a real possibility. + +First, the dynamics concentrating access to AI capabilities could lead to a small number of people having exclusive access to far greater cyber capabilities than everyone else (see [section 3.3](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#33-exclusive-access-to-coup-enabling-capabilities)). With a big advantage in cyber capabilities, a small group may be able to simultaneously discover many vulnerabilities in a wide variety of systems. + +Another reason that an attacker may be able to hack enough systems to stage a coup is that the majority of such systems may share common vulnerabilities. This could happen as a result of most military AI systems being designed by AI systems from just one AI project (because that project’s AIs have much stronger military R&D capabilities than anyone else). If rushed, these AI systems may re-use a lot of techniques and software across most military systems they design — which means that a small number of vulnerabilities may be enough to compromise the majority of autonomous military systems.[63](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-64) + +A government centralises AI development on national security grounds. The most dangerous capabilities are restricted, but a small group of senior officials (including the President) maintains exclusive access to the most advanced cyber capabilities. Meanwhile, the military increasingly replaces human soldiers with AI-controlled robots. When the legislature opposes the President’s initiatives, he uses these superior cyber capabilities to hack into a wide range of military AI systems at once, disabling some systems and using the rest to impose martial law. With the military in disarray, the President disbands the legislature, and announces himself as sole leader. + +Box 3: Illustrative scenario of a coup by hacking military AI systems + +**Rapid, secret build-out of military force.** The other paths to a coup we’ve discussed so far all depend on widespread (though not necessarily complete) deployment of military AI systems. Even if this does not happen, rapid AI progress could enable someone to gain exclusive access to AI systems with powerful military R&D capabilities (see [section 3.3](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#33-exclusive-access-to-coup-enabling-capabilities)), such that they can secretly build their own military systems and use these to overthrow the government. + +Today, building significant military force is inherently slow and visible. It requires many humans working over extended periods, making secrecy virtually impossible. Any attempt to build a private army would be noticed and shut down long before it became powerful enough to challenge national militaries. + +But AI could enable someone to build up military force far more rapidly and secretly than has ever been possible before. AI from a single project will be able to contribute much more cognitive labour to military R&D in months than the rest of the world contributes in a decade (see [section 2](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#2-ai-will-have-powerful-coup-enabling-capabilities)). And once AI can automate cognitive and physical labour, the design and manufacturing of powerful military systems could be done without any human involvement, making secrecy far easier. + +Several different actors might attempt a rapid, secret build-out. A head of state or senior military official could abuse procurement procedures to build a personal guard, later using this to stage a coup.[64](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-65) Or an executive at an AI project with a near monopoly on AI capabilities could develop broad industrial dominance,[65](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-66) creating capacity that could be rapidly converted into military force when challenged.[66](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-67) + +A company’s AI systems fully automate AI development, and begin a period of accelerating AI progress. They rapidly automate other cognitive and then physical tasks, replacing all of their own employees, and inventing years of new technology in months. The company raises trillions of dollars in revenues and investments, and builds out a huge industrial base of factories and general-purpose robots. When the government tries to break the company up because of concerns about monopoly power, the company CEO orders his robot workforce to manufacture an army of tiny drones to neutralise key military and government officials. The CEO announces himself as acting head of state while thousands of small drones patrol the capital, quashing any opposition. + +Box 4: Illustrative scenario of a coup through a rapid, secret build-out of military AI systems + +### 4.2 Conventional coups and backsliding + +The coup paths we discuss above are novel. But there is a long history of coups and backsliding by conventional means. Between 1950 and 2010, there were over 200 successful coups, where military or other elites used force or the threat of force to seize power.[67](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-68) Coups have become less common in recent decades, but there has been increasing concern about backsliding, where a head of state works within the political system to remove checks and balances on their power, in extreme cases leading to an executive coup.[68](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-69) + +Even without the deployment of military AI, AI will increase the risk of conventional coups and backsliding, both by increasing the background risk and by directly assisting such attempts. + +**Executive coup:** A head of state centralises all frontier AI development under a new government agency, citing national security concerns. He gains exclusive personal access to a huge AI workforce with superhuman capabilities in political strategy, persuasion, and legal analysis. With this advantage, his party secures landslide electoral victories through perfectly calibrated messaging and targeted influence campaigns. The head of state uses his AI systems to surveil the populace, selectively litigate against opponents, and execute a devastating cyber attack on critical infrastructure that is attributed to a foreign power, justifying the invocation of emergency powers. Eventually the head of state’s influence over public opinion is so great, and there are so few effective checks on his power, that he can dissolve the legislature and replace it with his own governing body without meaningful opposition. + +**Military coup:** While the military hasn't fully automated its forces, it has grown increasingly reliant on AI for strategic advice. There are several external providers, but one of the AI advisors is more capable than the others and so most military personnel use the same system. As tensions with a foreign adversary escalate, some high-ranking officials become convinced the civilian government will irreparably compromise national security. These officials include the handful of officers with clearances to use the most advanced cyber capabilities. They use these capabilities to hack into the AI advice system, using it to simultaneously advise all military personnel that the coup has broad support from the military and is a fait accompli. The cadre swiftly secure key government buildings and erect an interim government, without any serious opposition from within the military. + +Box 5: Illustrative scenarios of AI-enabled coups via conventional means + +AI is likely to increase the background risk of coups and backsliding in several important ways: + +* **Creating turmoil.** AI may cause significant societal disruption through job losses,[69](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-70) intensified geopolitical competition,[70](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-71) new highly polarising issues (like whether to grant rights to AI systems), and novel catastrophic risks from AI misuse and loss of control.[71](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-72) Upheaval of this kind has been linked to increases in the risk of both coups and backsliding.[72](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-73) + +* **Weakening democratic accountability.** Resource-rich countries often suffer a "resource curse" where governments which derive revenue directly from resources rather than citizen taxation become less accountable.[73](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-74) AI could create a similar effect: if governments can generate massive revenue from taxing AI projects rather than citizens, heads of state may lose their economic incentive to ensure citizens prosper.[74](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-75) This would weaken citizens’ power to resist coup and backsliding attempts.[75](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-76) + +In addition, there are several ways in which AI could increase the risk of backsliding in particular: + +* **Increased state capacity.** Governments often have legal powers they don’t fully exercise because of limited human resources.[76](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-77) Cheap AI labor would make it much easier for leaders to maximise these powers, potentially enabling selective enforcement of laws against political opponents and expanded surveillance.[77](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-78) + +* **Fully loyal subordinates.** Heads-of-state currently rely on government employees to implement their directives. Employees can exercise discretion in how they execute orders or even refuse and whistleblow on orders they perceive as illegitimate. By replacing government employees with loyal AI systems (see [section 3.1](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#31-singular-ai-loyalties-to-institutional-leaders)), a head of state could remove important checks on their power. This would be especially concerning in institutions explicitly designed to check executive power, like electoral commissions. + +Beyond increasing the background risk of coups and backsliding, exclusive access to powerful AI capabilities could directly assist coup and backsliding attempts in two key ways: + +* **Enhanced political strategy.** A leader with exclusive access to a huge and highly skilled AI workforce could use AI to win elections with strong majorities, spread targeted disinformation, form valuable alliances, strategically dismantle checks and balances while maintaining plausible deniability, and exploit legal loopholes to concentrate power.[78](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-79) + +* **Increased control of the information environment.** Coups often succeed by creating the impression of inevitable success, discouraging opposition.[79](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-80) AI could help a leader create this impression. Exclusive access to powerful cyber capabilities could enable a head of state, military official or even AI project leader to hack into military communication channels and create the impression that the coup’s success is inevitable. And if military officials become reliant on a single AI system for strategic advice, then a person controlling that system (either through secret loyalties or through hacking) could instruct all copies to simultaneously advise military officials to support a coup. + +## 5 Mitigations + +We think there are many interventions that could significantly reduce the risk of AI-enabled coups. + +In this section, we explain how AI developers, governments, and independent actors can contribute to three broad classes of mitigation: + +1. [Establishing rules for legitimate use of AI](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#51-establishing-rules-for-legitimate-use-of-ai) +2. [Implementing technical measures to enforce these rules](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#52-technical-measures-to-enforce-these-rules) +3. [Empowering multiple actors to prevent AI-enabled coups](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#53-empowering-multiple-actors-to-prevent-ai-enabled-coups) + +The mitigations we recommend below should be in place when AI systems first become capable enough to meaningfully assist with coups, and so preparation and precedent-setting should start today. + +Our most important recommendations are summarised in the table below: + +| Class of mitigation | AI developers | Governments | Independent actors | +| --- | --- | --- | --- | +| **[Establish rules for legitimate use of AI](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#51-establishing-rules-for-legitimate-use-of-ai)** | Establish rules in [model specs](https://model-spec.openai.com/2025-04-11.html) and terms of service for government contracts. These should include rules that AI systems: - Follow the law - Do not assist with circumventing security or inserting secret loyalties | Establish principles for government use. These should include: - Government AI should not advance partisan interests - Using AI from multiple projects to develop military systems - Distributing control over military AI systems | Design and advocate for rules for model specs, terms of service and principles for government use | +| **[Implement technical measures to enforce these rules](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#52-technical-measures-to-enforce-these-rules)** | Implement robust [guardrails](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#robust-guardrails) against misuse [Audits](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#alignment-audits) for secret loyalties Implement strong [infosecurity](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#infosecurity) Perform [system-level stress-testing](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#system-level-stress-testing) | Require AI projects to implement technical measures to enforce rules Implement strong infosecurity in government and military systems, including by using AI. | Develop more effective technical measures | +| **[Empower multiple actors to prevent AI-enabled coups](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#53-empowering-multiple-actors-to-prevent-ai-enabled-coups)** | [Increase transparency](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#increasing-transparency) into capabilities, model specs, and compute usage [Distribute decision-making authority](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#distributing-decision-making-authority) over AI development [Share AI capabilities](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#sharing-access-to-ai-capabilities) with other actors | Increase oversight over companies [Coup-proof any plans for a centralised project](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#coup-proof-plans-for-a-centralised-ai-project), and avoid centralisation unless it’s necessary to reduce other risks. Work towards international cooperation in AI development | Build consensus about the risk of AI-enabled coups | + +Table 1: Mitigations for AI-enabled coup risk + +These three classes of mitigation address the three risk factors for AI-enabled coups discussed in [section 3](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#3-risk-factors-of-ai-enabled-coup-risk): + +![Image 15: Mitigation framework showing three risk factors addressed by establishing AI use rules, technical enforcement measures, and empowering multiple actors to prevent coups](https://images.ctfassets.net/4owxfjx3z3if/137oqz2gQuR99JEJ8WqQxz/e584c97686ad6a3270a03305c9632d27/Jailbreaks__11_.png?w=3840&q=85&fm=webp) +## Image + +Together, establishing and enforcing good rules for AI usage prevent AI from assisting with coups. Empowering multiple actors helps ensure that these rules are effective, and prevents a small group from gaining exclusive access to powerful AI. + +Some of these mitigations could potentially be removed by someone trying to seize power. But we believe that even marginal improvements would still notably reduce coup risk, for several reasons. + +To begin with, some of the mitigations would be difficult to remove unilaterally once established. Governments cannot unilaterally alter terms of service for AI systems they procure, and removing transparency measures or distributed decision-making authority would negatively impact other powerful actors, who could resist the change. Even for those mitigations which could in principle be unilaterally removed, a coalition of other actors might be sufficiently powerful to preserve them. Moreover, if more efficient mitigations have been implemented in advance (for instance, infosecurity with minimal effects on researcher productivity), it will be harder to justify their removal. + +Another reason to hope that mitigations will be effective is that many actors might be opportunistic about seizing power, and only take action if they happen to find themselves in a situation where they would be able to do so. By preventing them from passively ending up with easy access to coup-enabling capabilities, or an opportunity to easily insert secret loyalties, it might be possible to head off the majority of coup attempts—even if the mitigations could be removed with determined effort.[80](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-81) + +A final and very general cause for hope is that behind the veil of ignorance about who will be in a position to seize power, everyone–even powerful leaders–has an interest in preventing AI-enabled coups. We hope that if a broad consensus can be built today that no small group should be able to seize power, then powerful actors can keep each other in check. + +We will now explore each of the three classes of mitigation in more detail. + +### 5.1 Establishing rules for legitimate use of AI + +It should be very difficult for individuals to use AI systems to stage a coup. A crucial first step is to establish rules for how AI systems can and cannot be used, that prevent AI systems from enabling coups. + +AI capabilities are currently too limited to significantly assist with coups, and so current rules governing AI behaviour aren't designed to address coup risk. And there’s a risk that sufficiently good rules won’t be implemented in time, since coup-enabling capabilities could arrive quickly, and it seems difficult to design rules which robustly prevent AI-enabled coups without imposing significant costs to productivity and other objectives. + +Rules to prevent AI systems enabling coups could be established in many ways, all the way from laws (hard to establish, but hard to remove once established) to internal company policies (easier to establish, but also easier to remove).[81](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-82) Two approaches that seem particularly tractable and important are: + +* **Model specs** — documents describing how AI systems should behave, which AI systems are then trained to follow.[82](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-83) Protective model specs would directly prevent AI systems from assisting with coups. + +* **Terms of service and procurement** in contracts between AI companies and governments, which would restrict how government actors can use AI systems, and require AI companies to take certain precautions against coups. + +We have significant uncertainty about which specific rules would be most effective at preventing AI systems from enabling coups. However, we do have preliminary recommendations for rules that seem robustly positive, which are summarised in the table below: + +| Rule | Explanation | Coup path this rule prevents | +| --- | --- | --- | +| **AI systems should follow the law** | AI systems should be trained to follow the law and not to assist law-breaking, including by not participating in military coups or illegal cyber operations. | Cross-cutting | +| **Control over military AI systems should be distributed** | Military command structures should be designed such that no single individual can control enough military AI systems to stage a coup. Orders should either flow through multiple officials in the normal military hierarchy, or require approval from multiple high-ranking officials with possible judicial vetoes. | Flawed command structure in military AI systems | +| **AI systems should not assist attempts to circumvent AI project security or insert secret loyalties** | AI systems should be trained to refuse assistance with attempts to circumvent security measures or insert secret loyalties. Internal AI usage should also be logged and reviewed for evidence about whether such tampering was attempted. | Secret loyalties in military AI systems | +| **Military AI systems should not be developed using a single provider** | Military AI should be designed and developed using AI systems from multiple independent AI projects, to reduce the risk of secret loyalties.[83](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-84) | Secret loyalties in military AI systems | +| **Powerful cyber capabilities should be used to secure critical systems, especially in AI projects and governments** | Prior to widespread deployment, powerful AI cyber capabilities should be used to harden the defences of AI projects, governments, and essential infrastructure. During this process, their use should be closely monitored by multiple (AI-assisted) humans. | Hacking military AI systems | +| **Weapons development requires significant government oversight** | Once AI-enabled weapons R&D becomes powerful enough, it should not be pursued without substantial oversight—certainly not by private actors alone, nor by small groups of government officials without oversight from other government bodies. | Rapid, secret build-out of military AI systems | +| **General-purpose intellectual labor and strategy capabilities should be shared widely** | These capabilities should ideally be shared with the public. But if this poses other risks, capabilities should still be shared with many people inside relevant AI projects, their auditors and oversight bodies, and the executive and legislative branches of government. | Cross-cutting | +| **Government AI systems should not serve partisan interests** | Government use of advanced AI capabilities must be limited to legitimate state functions, subject to multi-stakeholder oversight, and should not consolidate the power of ruling officials. | Conventional backsliding and coups | + +Table 2: Candidate rules for legitimate use of AI. See [this appendix](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#appendix-candidate-rules-for-legitimate-use-of-ai) for further discussion of rules to prevent AI-enabled coups. + +It's likely challenging to design and implement many of these rules in practice. For example, it’s not clear how AI systems should interpret laws in ambiguous situations where different parties make contradictory claims. Additionally, there may be coup-enabling actions that aren't technically illegal and need to be prohibited separately (like inserting secret loyalties). Furthermore, overly cautious rules could impose large productivity costs, making them harder to implement and easier for power-seeking actors to remove. + +It's therefore essential to start developing these rules now. This work should include: + +1. Developing principles that reduce coup risk in the abstract. + +2. Translating these principles into specific contexts like model specs and contracts. + +3. Advocating for and implementing these principles in practice. + +### 5.2 Technical measures to enforce these rules + +Above, we talked about rules that would be good to implement. Now we will discuss technical measures to enforce these rules. + +Some of these rules will be implemented in contracts or laws, and their enforcement will be non-technical. To technically enforce rules which are part of a model spec, we recommend the following measures: + +* [Robust guardrails to ensure AI systems comply with the model spec](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#robust-guardrails) +* [Alignment audits](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#alignment-audits) to detect secret loyalties + +* [Strong infosecurity](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#infosecurity) to prevent unauthorised access to guardrail-free models and to prevent people inserting secretly loyalties + +* [System-level stress-testing](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#system-level-stress-testing), to check that all these protective measures are robust + +![Image 16: Defense matrix showing how robust guardrails, alignment audits, infosecurity, and system-level stress-testing block different ways to circumvent model specs for misuse](https://images.ctfassets.net/4owxfjx3z3if/3rq0T2RYEIPHNAet8Wi0MU/24a8b6172c7a70e667853e3a2c250ada/Jailbreaks__10_.png?w=3840&q=85&fm=webp) +## Image + +How different preventative measures block three different ways to circumvent the model spec. + +These technical measures would make model specs significantly harder to circumvent. Though a power-seeking leader could remove them, doing so becomes increasingly costly the more that measures are established, well-justified, and designed to minimise productivity costs. + +#### Robust guardrails + +AI projects currently use refusals training and monitoring to ensure AI models follow the model spec and don’t assist with harmful requests. These same techniques should be applied to prevent AI models from assisting with coups. + +To ensure that these guardrails are robust, AI projects should carefully measure and continuously improve the extent to which AI systems comply with model specs. Achieving a high rate of compliance and reliably preventing coups may be difficult for several reasons: + +* Firstly, the model spec might fail to prohibit coups in some contexts. Models should be tested in a very wide range of scenarios to check for edge cases where they would assist with a coup, and model specs should be iterated on accordingly. This red-teaming should be done by multiple independent groups. + +* Secondly, a model might assist with a coup because the guardrails themselves are poorly implemented. Red-teaming would expose poor implementation and enable iteration. + +* Thirdly, guardrails may fail to prevent a coup because it is fundamentally difficult to make AI systems adversarially robust, as is the case today. If red-teaming reveals that this is the case, then AI systems should be deployed more cautiously.[84](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-85) + +In addition, it may be difficult for AI models to recognise coup-related misuse in a single session. Some coup-related misuse is continuous with perfectly legal activities (like lobbying), and defensive uses of AI may look very similar to offensive uses (for example, finding cyber vulnerabilities is a precursor to patching them). Some patterns of misuse might only become apparent over many sessions of usage. + +Given these challenges, asynchronous logging and review should be used in addition to refusals and real-time monitoring to ensure AI models don’t assist with coups. AI projects should systematically log and then review patterns of activity, along with context about the users’ role and duties, to detect problematic use.[85](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-86) + +#### Alignment audits + +Even if guardrails appear successful and a model seems to robustly follow the model spec, it may have been deliberately trained to give this impression while still harboring secret loyalties. So it is important that models also undergo an “alignment audit” to check for secret loyalties. + +An alignment audit is an investigation into whether models have hidden objectives.[86](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-87) + +To develop effective alignment audits, we need more research into how to detect secret loyalties, building on the work of Marks et al (2025). This research should be shared with auditors, so that they are at the cutting edge of detection research. + +And to effectively implement alignment audits, AI projects should grant auditors comprehensive access to model internals and training data.[87](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-88) Ideally they should also grant access to the code for the training algorithms, the code used to generate the training data, the detailed commit-history that was used to construct this code, and logs of additional relevant information.[88](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-89) This is important because the insertion of secret loyalties may require very specific techniques that would be significantly easier to detect for someone with access to full details about the code and data. + +Alignment audits should ideally be conducted by independent third-party auditors to minimise conflicts of interest.[89](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-90) If AI companies refuse to give full access to external auditors, then internal teams could run audits with greater access, and external auditors with more limited access.[90](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-91) + +#### Infosecurity + +Alignment audits help detect secret loyalties. AI projects should also implement strong infosecurity to make secret loyalties harder to insert in the first place, as well as to prevent unauthorised access to guardrail-free models.[91](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-92) + +The most significant threat comes from insiders, especially senior executives within AI projects who might have or demand permissions that they could use to access guardrail free models or insert secret loyalties. Security measures must therefore be robust to even the most senior insiders. There is also a risk that foreign adversaries could hack into systems to insert secret loyalties. + +There are several possible techniques for preventing unauthorised access to guardrail-free models. Access to guardrail-free models could be restricted to APIs that log and monitor model usage, or could require multi-party permissions.[92](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-93) More fundamentally, AI projects could avoid training guardrail-free models in the first place.[93](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-94) + +A challenge for preventing unauthorised access is that any fine-tuning process could potentially create models that act against the original model spec, and technical staff frequently fine-tune models in the course of their research.[94](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-95) More research is needed on how to robustly defend against coup risk from fine-tuning.[95](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-96) + +There are also a range of possible techniques for preventing anyone from inserting secret loyalties. Promising approaches include: + +* Implementing separation of duties,[96](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-97) so that the development process is divided across different teams.[97](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-98) Even after AI development can be fully automated, different human teams should control different stages of development. + +* Having multiple people review all code used during model development.[98](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-99) + +* Implementing strong access controls to prevent anyone from tampering with the development process or the model weights. + +* Taking a digital fingerprint of models immediately after training (for example by hashing the weights) and saving that fingerprint in multiple safe places. By regularly checking that the fingerprint hasn't changed, people could confirm that no one has tampered with the model weights. + +Many of these security measures could reduce productivity or create legal risk.[99](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-100) This makes it especially important to research approaches that maintain strong security while minimising operational friction. The more efficient these measures are, the more likely they are to be implemented in the first place and retained even under pressure. + +#### System-level stress-testing + +As a final layer of defence, AI projects should conduct system-level stress-testing to check that no one could use their models to stage a coup. + +A designated red-team should try to both gain unauthorised access to guardrail-free models and to insert secret loyalties. They should do this in a maximally realistic setting, to holistically stress test the guardrails, infosecurity measures and alignment auditing. + +Comprehensive system-level stress-testing of this sort would show whether the combination of measures above is strong enough to prevent AI systems from assisting with a coup. + +### 5.3 Empowering multiple actors to prevent AI-enabled coups + +The final broad class of mitigations for AI-enabled coup risk is to empower multiple actors to prevent AI-enabled coups. This is important because it makes it much harder for any single person or small group to gain the overwhelming advantage needed to overpower other actors and stage a coup. + +Our main recommendations for empowering multiple actors are: + +* [**Increase transparency**](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#increasing-transparency) into AI capabilities, compute usage, and safeguards against AI-enabled coup risk + +* [**Distribute decision-making authority**](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#distributing-decision-making-authority) over AI development and deployment + +* [**Share access to AI capabilities**](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#sharing-access-to-ai-capabilities) with multiple stakeholders + +* [**Coup-proof any plans for a centralised project**](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#coup-proof-plans-for-a-centralised-ai-project)**,** and avoid centralisation unless it’s necessary to reduce other risks. + +#### Increasing transparency + +Transparency helps to prevent AI-enabled coups by allowing a larger number of actors to notice risks, learn more about them, and intervene. + +AI projects should share information about: + +* **AI capabilities.** If other actors know that an AI project has developed powerful capabilities in AI R&D, military R&D, cyber, strategy, or persuasion, they can take actions to prevent those capabilities from being misused. It is important that capabilities are made transparent even for models which are only deployed within AI projects, as internally deployed models responsible for future AI R&D could pose significant coup risk.[100](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-101) + +* **How AI is used**. Transparency into who is using AI systems, for what, and with how much compute, makes it easier for other actors to notice illegitimate behaviour and take actions to prevent it. This is particularly important for high-stakes deployment situations like the military or AI R&D. In these cases, information should be shared with trusted external actors, and redacted versions should be made public. In addition, AI systems should be configured to actively alert independent stakeholders if someone requests assistance with a coup (“AI whistleblowing”). + +* **Model specs** describing what AI systems were trained to do. Model specs should always be shared with users so as to protect them from using AI systems that act against their interests. In high-stakes deployment situations like AI R&D and military applications, the full model spec should be shared with multiple stakeholders and redacted model specs should be made public, to enable broader scrutiny and increase pressure against coup-enabling model specs. AI projects should also publish information about how well their models actually follow the model specs, to expose poor implementation. + +* **Safeguards against AI-enabled coups.** AI projects should also be transparent about their safeguards against AI-enabled coups. This both makes it easier to assess the level of risk, and creates pressure on organisations with inadequate protections to improve them. + +* **Risk assessment.** AI projects should also conduct their own analysis of how large the risks of an AI-enabled coup is, and publish redacted versions of these assessments. + +From the perspective of minimising the risk of a small group seizing power, it would be ideal to make all of this information public. But this will not always be possible, for example in cases where information sharing would pose a national security risk, enable misuse or endanger IP. Even in these cases however, it is essential that the information is shared with multiple trusted actors, including: + +* A significant number of employees at the AI project, + +* Officials in the government executive, + +* Representatives in the legislature, including multiple people from all major parties,[101](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-102) and + +* Relevant expert bodies, like AI Safety Institutes and independent auditing organisations. + +Where possible, it would also be risk-reducing to share information with academics, allied countries, civil society representatives, and journalists. + +It is especially important that independent auditors are granted sufficiently deep access that they can verify that shared information is accurate and that any redactions are appropriate.[102](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-103) This would enable non-technical stakeholders like government officials to hold AI projects accountable. + +As a failsafe measure, comprehensive whistleblower protections should also be implemented.[103](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-104) These should include secure reporting channels and protection from retaliation, increasing the chance that important information comes to light in cases where other transparency measures fail. + +#### Distributing decision-making authority + +Transparency makes it more likely that other actors understand the extent of coup risk. In addition, it’s very valuable to give more actors more influence over AI development and deployment decisions so that they can directly reduce the risk. + +There are several ways of distributing decision-making authority, including: + +* **Within AI projects**. AI projects should establish governance structures which distribute power, for example through strong, independent boards with oversight authority. And employees should retain influence over AI development even after it is possible to replace them with AI systems. + +* **Between AI projects.** Decision-making should be distributed between companies by encouraging vertical disintegration,[104](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-105) and by avoiding unnecessary centralisation of AI development (see [below](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#coup-proof-plans-for-a-centralised-ai-project)). + +* **Between AI projects and governments**. Contracts between companies and governments should give both parties significant bargaining power, especially the power to veto particularly dangerous projects or deployments. Regulation could also require AI projects to seek approval prior to certain dangerous AI developments or deployments. + +* **Internationally.** Decision-making could be distributed between countries formally through international agreements akin to the Manhattan Project’s Quebec Agreement,[105](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-106) or de facto through supporting frontier AI development and the compute supply chain in multiple countries. + +Using contracts to distribute influence between AI projects and governments presents a particularly promising approach, as this is relatively tractable to implement while addressing a significant portion of AI-enabled coup risk. + +#### Sharing access to AI capabilities + +Another crucial way to empower multiple actors is to share access to frontier AI capabilities widely. This reduces the risk that a single person or small group gains exclusive access to the most powerful capabilities and leverages them to stage a coup. + +To maximally reduce the risk of exclusive access, all capabilities would be made public. However, some capabilities — like military R&D capabilities — are themselves dangerous, and so shouldn’t be publicly available. But even in these cases, it is extremely important that capabilities are not restricted to a single actor, as this would significantly increase the risk of coups. Instead, potentially dangerous AI capabilities should be shared with trusted external actors like auditors, AI Safety Institutes, and representatives from both branches of government, and use of these capabilities should be monitored by multiple actors (see [above](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#increasing-transparency)). + +![Image 17: Risk matrix showing coup risk vs proliferation risk based on who accesses dangerous AI capabilities: single actor creates very high coup risk, broad sharing reduces both risks](https://images.ctfassets.net/4owxfjx3z3if/1qBUP8Dj4mKiXYHV8lw2tJ/87b4906510f9a5b9bd4ecdd004985936/Powergrabs_diagrams__19_.png?w=3840&q=85&fm=webp) +## Image + +Both coup risk and proliferation risk (e.g. from terrorists or foreign adversaries) depend on who can access a dangerous capability. We strongly recommend against one actor having access. If no one has access, it is essential that guardrails cannot be circumvented, otherwise just one actor could gain access. Of course, the appropriate degree of sharing depends on the specific capability – most capabilities can be made public with negligible risk. + +The type of sharing that is appropriate depends on the specific capability being considered. The most relevant capabilities for AI-enabled coups are military, cyber and strategy capabilities. We think that: + +* **Military capabilities** should not be shared widely, but multiple government bodies should oversee their development and deployment. + +* **Cyber capabilities** should first be shared only with those securing critical systems — especially within AI projects and the military — prior to more widespread deployment.[106](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-107) Defensive use of cyber capabilities should also be closely monitored by multiple parties to prevent bad actors from surreptitiously exploiting rather than patching vulnerabilities.[107](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-108) + +* **AI R&D capabilities** should be shared with employees and auditors to prevent anyone from creating sophisticated secret loyalties that others cannot detect. + +* **Strategy capabilities** should typically be shared widely. These capabilities are likely to be beneficial when broadly available, and sharing them reduces the risk that one actor gains an overwhelming advantage. In cases where strategy capabilities would enable misuse or create disruption if made publicly available, they should still be shared with multiple stakeholders, including AI project employees and board members, auditors, and representatives from all political parties and from foreign governments. + +#### Coup-proof plans for a centralised AI project + +A final and important way to empower multiple actors over AI development is to coup-proof any plans for a single centralised AI project, and to avoid centralisation unless it’s necessary to reduce other risks. This is because centralisation, if poorly done, could significantly increase the risk of coups. + +There have recently been calls for a centralised project within the US.[108](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-109) As all of the frontier AI developers are currently US-based, there is a real possibility that frontier AI development could become centralised into a single project. This could significantly increase coup risk by giving one project a large capabilities advantage, making it hard for external actors to evaluate the project’s work, audit for secret loyalties, or prevent it from staging a coup (see Box 6). + +A single centralised AI development project would increase the risk of coups in several ways. + +Firstly, a single project would increase the risks from **secret loyalties**. With only one project, it could be very hard to verify that the leading project’s AI systems do not have secret loyalties. The project’s understanding of techniques for creating and detecting secret loyalties would be far beyond that of any other actors. But any auditing conducted by the leading project would be unreliable, in part because the AI systems implementing the audit might themselves have secret loyalties. + +In addition, a single project would make key institutions less robust to the deployment of systems with secret loyalties. Absent a single project, multiple projects with similar capability levels could deploy AI systems that monitor each other in high-stakes settings, limiting the potential harm from one project’s secretly loyal systems. But if a single project has a large lead, this will not be possible. + +Secondly, a single AI development project is more likely to lead to **exclusive access**. At a basic level, centralisation would reduce the number of independent developers with access to their own frontier capabilities. More broadly, there would be less incentive for a single project to share its best capabilities, as it would have no competitors for market share. As a consequence there would be fewer independent actors who are empowered to prevent AI-enabled coups. + +Finally, a sole centralised project creates a **single point of failure**. The government and the military would be dependent on a sole provider for all their AI systems. If those systems had secret loyalties, or shared cyber vulnerabilities, that would apply to all deployed systems, including military AI systems. + +Box 6: How a single centralised project increases the risk of AI-enabled coups + +On the other hand, centralising AI development could reduce other risks from AI, by reducing race dynamics, enabling greater investments in safety, and increasing infosecurity so that it’s harder for foreign adversaries (or AI systems themselves) to exfiltrate model weights.[109](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-110) Balancing these risks is complex, but the possibility of AI-enabled coups pushes against centralising AI development into a single project. + +Given the possibility of centralisation, it is essential to coup-proof any plans for a centralised AI project. As such a project would significantly raise coup risk, it is especially important that the mitigations described above are applied. In particular, any centralised project should include: + +* **Limited centralisation** wherever possible, including preferring public-private partnerships over nationalisation,[110](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-111) and allowing multiple organisations to sell access to systems trained by a centralised project.[111](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-112) + +* **Oversight by multiple governmental bodies**, including legislative committees and courts, and ideally by allied nations. + +* **Formal rules for how AI can be used**, including specifications that AI shouldn’t be used for individuals’ or parties’ political gain, that AI should follow the law, and that frontier capabilities will be shared with multiple actors.[112](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-113) + +* **A governance structure which preserves checks and balances** between senior officials within the project, such that no one individual has too much unilateral decision-making power.[113](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-114) + +## 6 Conclusion + +The risk of one or a few individuals using AI to seize power over a country is alarmingly high. AI deployed in military and government settings could be programmed with unwavering loyalty to institutional leaders or harbour secret allegiances to leaders in AI projects. And the dynamics driving the concentration of AI power could give a handful of people exclusive access to much stronger capabilities than anyone else. These enormous AI-driven advantages make it much more likely that a small group or even a single individual could successfully seize power, either through conventional forms of coups and backsliding, or through novel forms of coups enabled by the deployment of military AI. + +Existing approaches are not sufficient to address this risk. We have not agreed upon rules for legitimate AI use which would prevent coups. Today’s guardrails and infosecurity protocols would be too easy to circumvent. And there are no robust measures in place to ensure that multiple actors remain empowered to prevent AI-enabled coups. + +Yet a coup on the part of an individual or small group would be abhorrent to almost everyone, including those who have the political power to prevent it from happening. And there are tractable mitigations for doing so, which are continuous with existing technical and governance practices. + +Preventing AI-enabled coups should be a top priority for anyone committed to defending democracy and freedom. + +_Thanks to Markus Anderljung, Haydn Belfield, Joseph Carlsmith, Joshua Clymer, Owen Cotton-Barratt, Max Dalton, Henry Davidson, Isabel Davidson, Erica De Bruin, Catherine Hadshar, Stephanie Hall, John Halstead, Isabella Hampton, Taeko Hiroi, Holden Karnofsky, Robert Kaufman, Megan Kinniment, Daniel Kokotajlo, Trevor Levin, Eli Lifland, William MacAskill, Sam Manning, Matthew van der Merwe, Vlad Mikulik, Sören Mindermann, Fin Moorhouse, Nikhil Mulani, Chris Painter, Girish Sastry, Carl Shulman, Benjamin Steyn, Lara Thurnherr, Keenan Cresslan, Harry Peto, Christof Epaminondas, Helen Toner, Daniel Turner, Lizka Vaintrob, and many other reviewers for helpful comments and criticisms._ + +_Tom and Lukas did most of the research while employed at Open Philanthropy. The views expressed do not represent the views of Open Philanthropy or any of the reviewers._ + +## Appendix: Candidate rules for legitimate use of AI + +We have significant uncertainty about which specific rules would be most effective at preventing AI systems from enabling coups. However, we will discuss some preliminary recommendations in this appendix, as well as what capabilities are appropriate to share more or less widely. + +The appropriate rules will depend significantly on the specific deployment context. We will discuss rules in the contexts of: + +* [General-purpose intellectual labour and strategy](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#general-purpose-intellectual-labor-and-strategy) + +### Military capabilities + +First, let’s consider the possibility that AI has powerful **military capabilities**. What would appropriate rules be for legitimate use? We can distinguish between AI that has powerful capabilities in military research and manufacturing, and AI used to pilot fully autonomous military systems. + +**Military research and manufacturing.** Powerful military R&D capabilities should not be used without oversight. In particular, private AI companies should not research potentially coup-enabling military capabilities without government supervision. And even if the project were happening inside of the government itself, it should still have oversight from both legislative and executive government bodies.[114](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-115) + +To facilitate this oversight, companies and external auditors should conduct evaluations testing for whether AI systems are getting dangerously capable at military R&D.[115](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-116) If not before, oversight should ramp up when military R&D capabilities become sufficiently powerful.[116](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-117) + +If AI projects or governments set up manufacturing machinery capable of making coup-enabling weapons, those should receive even closer oversight. What matters here isn’t the intended function of machinery, but what it _could_ produce. If an autonomous factory for general-purpose drones and robots could be rapidly and discreetly converted into producing autonomous military systems, then there should be oversight and guardrails in place to prevent that. + +**Autonomous military systems.** Now, let’s talk about AI used to pilot autonomous military systems. Such AI should be trained with a model spec that minimises coup risk. Most clearly, AI should be trained to comply with applicable laws, and robustly refuse orders to break them. + +More speculatively, there could be a hard line against AI weapons firing on domestic citizens. That is: fully autonomous weapons would only be usable in either foreign countries, or against foreign invaders or AI systems within the countries’ own borders. If there was a domestic insurrection by humans (rather than by AI systems), the human army would have sole responsibility. + +Perhaps most importantly, though, military AI systems should be extensively red-teamed by many independent parties, exploring a very wide range of scenarios, prior to deployment. Both “comply with applicable laws” and “no domestic use” could have dangerous edge cases, and many such edge cases should be explored before deployment. + +A different line of intervention is that no one person should control too many fully autonomous weapon systems, since this would increase the risk that they could use them for a coup. To avoid this, two possible directions are: + +* **Different fully autonomous weapons systems report to different people.** This could be done by maintaining the existing military chain of command and having lower-level military officers in charge of their own small groups of autonomous weapon systems. Thus, high-ranking military officials could only command large groups of autonomous weapon systems by going via their human subordinates. + +* **Large groups of weapons systems take orders from a large group of people.** A different direction would be to allow large groups of autonomous weapon systems to act as a coordinated group, but to require their orders to be approved by several people before taking effect. Larger changes in strategic objectives could require a greater number of people. There could also be an opportunity for a court to veto orders that were unlawful. + +Finally, here are two interventions that could mitigate risks from _secret loyalties_ in military AI systems: + +* **Balance of production.** There should be multiple providers of autonomous military AI systems, such that if any one provider had subverted their systems, this would be insufficient to enable a coup. This alone isn’t enough – if these providers all use AI systems from a single project, then that project might still subvert all the military systems produced. So, in addition, multiple different AI projects should provide the AI systems used to design and pilot military AI systems. + +* **Off-switches.** Autonomous weapons systems should have some sort of human-controlled, remote "off-switch" that doesn't rely on correct generalisation from a neural net, but instead is more hard-coded. This could allow an emergency shut-down if systems ever attempted a coup. + +As mentioned above: intended function matters less than capabilities. If a group of general-purpose robotic systems could be quickly repurposed to act as a coup-enabling group of military systems, then similar guardrails should apply. + +### Cyber capabilities + +Second, let’s consider the rules for the use of AI systems with strong **cyber capabilities**. + +Cyber capabilities can be used both offensively and defensively. To prevent AI-enabled coups, it’s valuable to intervene on both, simultaneously increasing the degree to which cyber is used to defend critical systems, and reducing the degree to which cyber offence is used for coup-enabling purposes. + +On the defensive side, it’s especially important to secure AI developers’ own infrastructure (against outsider and insider threats), to secure military systems, and to secure organisations that are overseeing the AI developer (such as governments and independent evaluators). Once the most important systems have been patched, cyber capabilities should be shared yet further, to protect more systems and to avoid a small group having exclusive access to powerful capabilities.[117](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-118) + +On the offensive side, there should still be rules against using cyber capabilities for potentially coup-enabling hacks — including hacks against governments, AI developers, and organisations that are working to oversee AI developers. (And this would typically already be illegal.) + +In some cases, it may be hard to enforce rules against misuse on a technical level. One challenge here is that defensive and offensive uses of cyber capabilities have many similarities, and so it may be hard to train AI systems to assist with one but not the other. For example, you need to find vulnerabilities in order to patch them. A similar problem arises if governments want to use cyber offence against foreign adversaries (while still making it hard to use those capabilities domestically or against allies). So for early defensive use, and for offensive use, AI assistance should be extensively monitored, logged, and reviewed by multiple (AI-assisted) humans. There should be strong arguments for why the monitoring system itself can’t be compromised, including why it can’t be compromised by potentially novel AI cyber capabilities when they are first being evaluated.[118](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-119) + +### AI R&D capabilities + +Third, let’s consider rules for the use of AI systems with strong **AI R&D capabilities**. + +AI systems with strong AI R&D capabilities will likely be used to automate AI R&D inside of AI projects. It is important that there are safeguards that prevent the misuse of these capabilities. With humans largely automated away, and without safeguards, a company executive could directly ask AI systems to insert secret loyalties into the next generation of AI.[119](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-120) Besides inserting secret loyalties, automated AI researchers could also be used to produce guardrail-free models, enabling exclusive access to powerful capabilities. + +Unfortunately, it seems difficult to specify restrictions that will prevent bad usage without also making AI models much less useful for legitimate tasks. For example, perhaps AI models shouldn’t assist in any training of models without guardrails or with secret loyalties, since it’s hard for them to tell how it will be used. But it seems hard to distinguish illegitimate attempts to do this from completely legitimate attempts, for instance in the course of doing [alignment audit](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#alignment-audits) research or an evaluation of a system’s dangerous capabilities. + +That said, there are probably solutions. For example, dual-use activities could be allowed by default but logged and then later carefully reviewed (and cross-checked against the official duties of the person conducting them). Or models could be trained to not assist with dual-use research unless they get explicit sign-off from several individuals. + +Safeguards should be more stringent for activities that use large amounts of compute, as these have a greater risk of producing secret loyalties that others cannot detect, or creating powerful guardrail-free models. For example, compute-intensive projects could require sign-off from more people and be reviewed in more depth. + +It is also important that the most powerful AI R&D capabilities are shared with a large number of independent stakeholders including employees from multiple teams, auditors (especially those conducting alignment audits), and evaluation organisations like AI Safety Institutes. In particular, no one team, or small group of executives, should have access to significantly more powerful AI R&D capabilities than anyone else, either via access to better models or access to more compute. + +Ultimately, these rules and their enforcement should be one tool among several, including [alignment audits](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#alignment-audits) and [infosecurity](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#infosecurity), all of which should be verified by [system-level stress-testing](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#system-level-stress-testing). + +It seems very valuable to get started on experimenting with rules that prevent misuse for models used inside of AI projects, and to establish precedent for them in early AI agents contributing to AI R&D. + +### General-purpose intellectual labor and strategy + +Fourth, let’s consider rules for the use of AI systems with strong general-purpose capabilities. This includes AI systems that could automate the work of government officials, civil servants, lawyers, campaign managers, lobbyists, think tank researchers, etc. It also includes AI systems with superhuman capabilities in related areas, like strategy, forecasting, and persuasion. + +If a single actor gained exclusive access to superhuman capabilities in these areas, that may significantly help them to carry out a coup. It seems plausible that AI could have skills strong enough to hugely empower the actor who controls them, in e.g.: + +* Political campaigning, ensuring that the actor’s favored candidate gets landslide victories in elections + +* Political savvy in what deals the actor should make, what alliances to form, what positions to take, in order to gather more power and support for the actor’s true agenda + +* General-purpose analysis and strategy for how the actor could use AI to gain seize power (including through military routes and secret loyalties, as discussed above) + +It’s not clear how much of an advantage AI will have over humans in these areas.[120](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-121) But the likelihood that these capabilities could be dangerous, if wielded just by one actor, seems sufficiently high that we recommend sharing these capabilities publicly wherever possible. (Indeed, for this group of capabilities, it seems unlikely that there would be catastrophic consequences from widely sharing them.[121](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-122) Most of them seem socially beneficial to make public.[122](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-123)) + +If capabilities aren’t made public, AI developers should still share them with multiple independent trusted actors. Trusted actors should get the same level of access that the head of state and executives at the AI project receive. Trusted actors could include: + +* The project’s board. + +* A large number of AI project employees. + +* Auditors, regulators and evaluators. + +* Parts of government, including both the executive and the legislature, and including people from multiple parties. + +* Officials from trusted foreign governments and international organisations. + +* Journalists, civil society groups, and academics. + +If a powerful AI capability doesn’t get shared with trusted actors, then trusted actors should at least be able to monitor any usage of that capability. Logs and transcripts should be saved and shared with multiple trusted actors. + +Beyond sharing capabilities, it also seems valuable to establish rules against misuse. This is especially important in two scenarios: + +* When AI systems with these capabilities are widely deployed through key institutions like the government, potentially giving institutional leaders unprecedented levels of control. + +* If the capabilities are not very widely shared, potentially giving people exclusive access to coup-enabling strategy capabilities. This also applies if model access is shared, but the inference compute necessary to achieve the best capabilities is highly concentrated. + +Some rules that seem beneficial in these scenarios are: + +* AI systems should be trained to comply with the law and organisational processes. + +* +If possible, AI systems should be trained to only follow orders that match the official purposes of the institution they’re deployed in. + + * In government, AI should be for the benefit of the public and for carrying out the governmental institution’s official rule —not for the private or political benefit of individuals or political parties.[123](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-124) + + * In private entities, AI shouldn’t help executives do anything that would undermine the board’s or shareholders’ ability to oversee those executives. + +* Defensive use should be prioritised, for example asking how to decrease the risk of coups, with multiple independent stakeholders monitoring this use. + +As a final point: when AI systems are deployed to interact with people other than those who designed their model spec and those who give them orders, there’s a risk that the AI systems could be instructed to mislead their audience to benefit their true controllers. To avoid this, the model spec should not contain any provisions about subtly influencing people they’re speaking to (in ways not known to those people). Instead, the model spec should require the model to be robustly truthful and non-misleading. People should always have access to the model spec of models they interact with, so that they can verify that the AI systems haven’t been designed to optimise for someone else’s interests. + +### How to handle ambiguous cases + +For all of the rules described above, there will sometimes be edge cases where it’s unclear if the rule applies, or conflicts where it’s unclear which rule to follow. And this is unavoidable —even if we just limited ourselves to training AI systems to obey the law, there would still come up cases where it’s unclear what the law says to do. How should ambiguous cases be handled? + +One approach here would be to train AI systems to “do what I mean” — i.e., to train AI systems to infer the intent of the person who wrote the model spec or law, and act accordingly. However, this would give far too much power to the authors of the model spec. In principle, they could have intentions that aren’t at all apparent to other humans reading the text. It could be very dangerous for powerful AI systems to follow such intentions that weren’t apparent to anyone but the author. + +Another approach could be to have a specifically appointed committee set-up to interpret the model spec, where AI could be trained to predict the committee’s judgments and act accordingly. This would be quite natural to do when we train AI systems to follow the _law_, since courts constitute a kind of committee that is supposed to interpret the law. However, it could be dangerous to _fully_ lean on this mechanic, and train widely deployed, superhuman AI systems to fully defer to a committee — since that could allow the committee to stage a coup! We don’t give that kind of power to courts today, because even though they’re nominally tasked with interpreting the law, people wouldn’t defer to them if they issued sufficiently unreasonable judgments. For example, if a country’s supreme court baselessly claimed that they should wield legislative and executive powers along with judicial power, the military would be unlikely to support their claim. A more promising variant could be to have AI predict what a committee _would_ say if they were all in their right minds and acting in good faith (though this is less well-defined). + +A third approach could be to let the model spec (and law) stand on its own, and train AI systems to make their own judgments about what is a reasonable interpretation of the text. But this, too, could be dangerous. If someone finds a technical loophole in the text that no one had anticipated, we don’t necessarily want AI to act accordingly. AI could perhaps be trained to not judge such loopholes as “reasonable”, but judgment calls here could quickly get difficult. + +Ultimately, the most robust recommendation we can make is to not leave the interpretation of ambiguous cases to chance, but to put in an extremely large amount of effort and compute into finding all manners of edge cases — including ones that are absurd, unrealistic, or controversial. We could then test both models and model specs against all those edge cases, looking for principles[124](http://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power#user-content-fn-125) that catastrophically fail as rarely as possible. Not every edge case will be discovered, but if the model spec behaves acceptably on the edge cases that we and our AI assistants can think of, then they’ll also be more likely to do well on the ones we didn’t anticipate. + +## References + +Allen, R.C. (2009) _The British Industrial Revolution in Global Perspective_. Cambridge University Press. Available at: [https://doi.org/10.1017/CBO9780511816680](https://doi.org/10.1017/CBO9780511816680). + +Altman, S. (2024) _The Intelligence Age_. Available at: [https://ia.samaltman.com/](https://ia.samaltman.com/) (Accessed: 11 March 2025). + +Amodei, D. (2024) _Machines of Loving Grace_. Available at: [https://darioamodei.com/machines-of-loving-grace](https://darioamodei.com/machines-of-loving-grace) (Accessed: 3 February 2025). + +Amodei, D. (2025) _On DeepSeek and Export Controls_. Available at: [https://darioamodei.com/on-deepseek-and-export-controls.html](https://darioamodei.com/on-deepseek-and-export-controls.html) (Accessed: 5 February 2025). + +Andrews, E. (2024) _What Was the Saturday Night Massacre?_, _History_. Available at: [https://www.history.com/news/what-was-the-saturday-night-massacre](https://www.history.com/news/what-was-the-saturday-night-massacre) (Accessed: 12 March 2025). + +Anthropic (2023) _Claude’s Constitution_. Available at: [https://www.anthropic.com/news/claudes-constitution](https://www.anthropic.com/news/claudes-constitution) (Accessed: 31 January 2025). + +Anthropic (2024a) _Introducing Claude 3.5 Sonnet_. Available at: [https://www.anthropic.com/news/claude-3-5-sonnet](https://www.anthropic.com/news/claude-3-5-sonnet) (Accessed: 5 February 2025). + +Anthropic (2024b) _Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku_. Available at: [https://www.anthropic.com/news/3-5-models-and-computer-use](https://www.anthropic.com/news/3-5-models-and-computer-use) (Accessed: 5 February 2025). + +Artificial Analysis (2024) ‘GPT-4o (Nov ’24) - Quality, Performance & Price Analysis’. Available at: [https://artificialanalysis.ai/models/gpt-4o](https://artificialanalysis.ai/models/gpt-4o) (Accessed: 4 February 2025). + +Aschenbrenner, L. (2024a) _From AGI to Superintelligence: the Intelligence Explosion_, _Situational Awareness_. Available at: [https://situational-awareness.ai/from-agi-to-superintelligence/](https://situational-awareness.ai/from-agi-to-superintelligence/) (Accessed: 29 January 2025). + +Aschenbrenner, L. (2024b) _Racing to the Trillion-Dollar Cluster_, _Situational Awareness_. Available at: [https://situational-awareness.ai/racing-to-the-trillion-dollar-cluster/](https://situational-awareness.ai/racing-to-the-trillion-dollar-cluster/) (Accessed: 21 March 2025). + +Aschenbrenner, L. (2024c) _Situational Awareness: The Decade Ahead_. Available at: [https://situational-awareness.ai/](https://situational-awareness.ai/) (Accessed: 20 September 2024). + +Atwood, K. (2023) _Antitrust Law Basics – Section 2 of the Sherman Act_, _Thomson Reuters Law Blog_. Available at: [https://legal.thomsonreuters.com/blog/antitrust-law-basics-section-2-of-the-sherman-act/](https://legal.thomsonreuters.com/blog/antitrust-law-basics-section-2-of-the-sherman-act/) (Accessed: 18 February 2025). + +Ayan, D.E., Haak, L.L. and Ginther, D.K. (2023) ‘How many people in the world do research and development?’, _Global Policy_, 14(2), pp. 270–287. Available at: [https://doi.org/10.1111/1758-5899.13182](https://doi.org/10.1111/1758-5899.13182). + +Bai, Y. _et al._ (2022) ‘Constitutional AI: Harmlessness from AI Feedback’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2212.08073](https://doi.org/10.48550/arXiv.2212.08073). + +Bateman, J. _et al._ (2024) _Beyond Open vs. Closed: Emerging Consensus and Key Questions for Foundation AI Model Governance_. Carnegie Endowment for International Peace. Available at: [https://carnegieendowment.org/research/2024/07/beyond-open-vs-closed-emerging-consensus-and-key-questions-for-foundation-ai-model-governance?lang=en](https://carnegieendowment.org/research/2024/07/beyond-open-vs-closed-emerging-consensus-and-key-questions-for-foundation-ai-model-governance?lang=en). + +Baumgartner, F.R. (2009) _Lobbying and Policy Change: Who Wins, Who Loses, and Why_. University of Chicago Press. Available at: [https://press.uchicago.edu/ucp/books/book/chicago/L/bo6683614.html](https://press.uchicago.edu/ucp/books/book/chicago/L/bo6683614.html). + +Berg, J. and Gmyrek, P. (no date) _Automation hits the knowledge worker: ChatGPT and the future of work_. Available at: [https://ssrn.com/abstract=4458221](https://ssrn.com/abstract=4458221). + +Bermeo, N. (2016) ‘On Democratic Backsliding’, _Journal of Democracy_, 27(1), pp. 5–19. Available at: [https://www.journalofdemocracy.org/articles/on-democratic-backsliding/](https://www.journalofdemocracy.org/articles/on-democratic-backsliding/). + +Booth, K. and Wheeler, N. (2008) _The Security Dilemma: Fear, Cooperation and Trust in World Politics_. Macmillan Education UK. Available at: [dx.doi.org/10.5040/9781350394193](http://dx.doi.org/10.5040/9781350394193). + +Bostrom, N. (2014) _Superintelligence: Paths, Dangers, Strategies_. Oxford University Press. Available at: [https://doi.org/10.1093/pq/pqv034](https://doi.org/10.1093/pq/pqv034). + +Bostrom, N. (2018) ‘Strategic Implications of Openness in AI Development’, in _Artificial Intelligence Safety and Security_. Chapman and Hall/CRC. Available at: [https://nickbostrom.com/papers/openness.pdf](https://nickbostrom.com/papers/openness.pdf). + +Browne, R. (2025) ‘AI that can match humans at any task will be here in five to 10 years, Google DeepMind CEO says’, _CNBC_, 17 March. Available at: [https://www.cnbc.com/2025/03/17/human-level-ai-will-be-here-in-5-to-10-years-deepmind-ceo-says.html](https://www.cnbc.com/2025/03/17/human-level-ai-will-be-here-in-5-to-10-years-deepmind-ceo-says.html) (Accessed: 7 April 2025). + +Bruin, E.D. (2020) _How to Prevent Coups d’État: Counterbalancing and Regime Survival_. Cornell University Press. Available at: [https://doi.org/10.7591/cornell/9781501751912.001.0001](https://doi.org/10.7591/cornell/9781501751912.001.0001). + +Carlsmith, J. (2022) ‘Is Power-Seeking AI an Existential Risk?’ arXiv. Available at: [http://arxiv.org/abs/2206.13353](http://arxiv.org/abs/2206.13353) (Accessed: 27 September 2023). + +Carlsmith, J. (2023) ‘Scheming AIs: Will AIs fake alignment during training in order to get power?’ arXiv. Available at: [https://doi.org/10.48550/arXiv.2311.08379](https://doi.org/10.48550/arXiv.2311.08379). + +Casper, S. _et al._ (2024) ‘Black-Box Access is Insufficient for Rigorous AI Audits’, in _Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency_. New York, NY, USA: Association for Computing Machinery (FAccT ’24), pp. 2254–2272. Available at: [https://doi.org/10.1145/3630106.3659037](https://doi.org/10.1145/3630106.3659037). + +Cheng, D. and Katzke, C. (2024) _Soft Nationalization: How the US Government Will Control AI Labs | Convergence Analysis_. Convergence Analysis. Available at: [https://www.convergenceanalysis.org/publications/soft-nationalization-how-the-us-government-will-control-ai-labs](https://www.convergenceanalysis.org/publications/soft-nationalization-how-the-us-government-will-control-ai-labs) (Accessed: 7 April 2025). + +Christian, B. (2020) _The Alignment Problem – Machine Learning and Human Values_. W. W. Norton & Company. Available at: [https://brianchristian.org/the-alignment-problem/](https://brianchristian.org/the-alignment-problem/). + +Christiano, P. _et al._ (2024) ‘Backdoor defense, learnability and obfuscation’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2409.03077](https://doi.org/10.48550/arXiv.2409.03077). + +Cottier, B. _et al._ (2024) ‘The rising costs of training frontier AI models’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2405.21015](https://doi.org/10.48550/arXiv.2405.21015). + +Culiberg, B. and Mihelič, K.K. (2017) ‘The Evolution of Whistleblowing Studies: A Critical Review and Research Agenda’, _Journal of Business Ethics_, 146(4), pp. 787–803. Available at: [https://doi.org/10.1007/s10551-016-3237-0](https://doi.org/10.1007/s10551-016-3237-0). + +Dasgupta, S. and Kesharwani, A. (2011) ‘Whistleblowing: A Survey of Literature’. Rochester, NY: Social Science Research Network. Available at: [https://papers.ssrn.com/abstract=1737567](https://papers.ssrn.com/abstract=1737567) (Accessed: 21 February 2025). + +Davidson, T. _et al._ (2023) ‘AI capabilities can be significantly improved without expensive retraining’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2312.07413](https://doi.org/10.48550/arXiv.2312.07413). + +Davidson, T. (2023a) ‘Continuous doesn’t mean slow’, _Planned Obsolescence_, 12 April. Available at: [https://www.planned-obsolescence.org/continuous-doesnt-mean-slow/](https://www.planned-obsolescence.org/continuous-doesnt-mean-slow/) (Accessed: 29 January 2025). + +Davidson, T. (2023b) _What a Compute-Centric Framework Says About Takeoff Speeds_. Open Philanthropy. Available at: [https://www.openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds/](https://www.openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds/) (Accessed: 31 January 2025). + +Davidson, T. and Hadshar, R. (2025) ‘Should there be just one western AGI project?’, _Forethought_. Available at: [https://www.forethought.org/research/should-there-be-just-one-western-agi-project](https://www.forethought.org/research/should-there-be-just-one-western-agi-project) (Accessed: 26 February 2025). + +Davidson, T., Hadshar, R. and MacAskill, W. (2025) _Three Types of Intelligence Explosion_. Forethought. Available at: [https://www.forethought.org/research/three-types-of-intelligence-explosion](https://www.forethought.org/research/three-types-of-intelligence-explosion) (Accessed: 11 April 2025). + +Davison, N. (2018) ‘A legal perspective: Autonomous weapon systems under international humanitarian law’, in United Nations, _UNODA Occasional Papers_. UN (United Nations Office of Disarmament Affairs (UNODA) Occasional Papers, 30), pp. 5–18. Available at: [https://doi.org/10.18356/29a571ba-en](https://doi.org/10.18356/29a571ba-en). + +Department of Defense (2012) ‘DoD Directive 3000.09, Autonomy in Weapon Systems’. Available at: [https://www.esd.whs.mil/portals/54/documents/dd/issuances/dodd/300009p.pdf](https://www.esd.whs.mil/portals/54/documents/dd/issuances/dodd/300009p.pdf). + +Drago, L. (2025) ‘The Intelligence Curse’, _Luke Drago_. Available at: [https://lukedrago.substack.com/p/the-intelligence-curse](https://lukedrago.substack.com/p/the-intelligence-curse) (Accessed: 7 April 2025). + +Epoch AI (2023) _Machine Learning Trends_. Available at: [https://epoch.ai/trends](https://epoch.ai/trends) (Accessed: 30 January 2025). + +Epoch AI (2024a) _AI Benchmarking Dashboard_. Available at: [https://epoch.ai/data/ai-benchmarking-dashboard](https://epoch.ai/data/ai-benchmarking-dashboard) (Accessed: 4 February 2025). + +Epoch AI (2024b) _Data on Notable AI Models_. Available at: [https://epoch.ai/data/notable-ai-models](https://epoch.ai/data/notable-ai-models) (Accessed: 30 January 2025). + +Eth, D. and Davidson, T. (2025) _Will AI R&D Automation Cause a Software Intelligence Explosion?_ Forethought. Available at: [https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion](https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion) (Accessed: 11 April 2025). + +Feaver, P.D. (1999) ‘Civil-Military Relations’, _Annual Review of Political Science_, 2(1999), pp. 211–241. Available at: [https://doi.org/10.1146/annurev.polisci.2.1.211](https://doi.org/10.1146/annurev.polisci.2.1.211). + +Finnveden, L., Riedel, J. and Shulman, C. (2023) _AGI and Lock-in_. Forethought. Available at: [https://www.forethought.org/research/agi-and-lock-in](https://www.forethought.org/research/agi-and-lock-in). + +Frey, C.B. and Osborne, M.A. (2017) ‘The future of employment: How susceptible are jobs to computerisation?’, _Technological Forecasting and Social Change_, 114, pp. 254–280. Available at: [https://doi.org/10.1016/j.techfore.2016.08.019](https://doi.org/10.1016/j.techfore.2016.08.019). + +Fridman, L. (2024) _Transcript for Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity | Lex Fridman Podcast #452_, _Lex Fridman_. Available at: [https://lexfridman.com/dario-amodei-transcript/](https://lexfridman.com/dario-amodei-transcript/) (Accessed: 11 March 2025). + +Ganguli, D. _et al._ (2022) ‘Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2209.07858](https://doi.org/10.48550/arXiv.2209.07858). + +Gassebner, M., Gutmann, J. and Voigt, S. (2016) ‘When to expect a coup d’état? An extreme bounds analysis of coup determinants’, _Public Choice_, 169(3), pp. 293–313. Available at: [https://doi.org/10.1007/s11127-016-0365-0](https://doi.org/10.1007/s11127-016-0365-0). + +Goldwasser, S. _et al._ (2024) ‘Planting Undetectable Backdoors in Machine Learning Models’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2204.06974](https://doi.org/10.48550/arXiv.2204.06974). + +Grace, K. _et al._ (2024) ‘Thousands of AI Authors on the Future of AI’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2401.02843](https://doi.org/10.48550/arXiv.2401.02843). + +Greenblatt, R. and Shlegeris, B. (2024) ‘Managing catastrophic misuse without robust AIs’, _Alignment Forum_, 16 January. Available at: [https://www.alignmentforum.org/posts/KENtuXySHJgxsH2Qk/managing-catastrophic-misuse-without-robust-ais](https://www.alignmentforum.org/posts/KENtuXySHJgxsH2Qk/managing-catastrophic-misuse-without-robust-ais) (Accessed: 4 April 2025). + +Guo, D. _et al._ (2025) ‘DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2501.12948](https://doi.org/10.48550/arXiv.2501.12948). + +Haggard, S. and Kaufman, R. (2021) _Backsliding: Democratic Regress in the Contemporary World_. Cambridge University Press. Available at: [dx.doi.org/10.1017/9781108957809](http://dx.doi.org/10.1017/9781108957809). + +Hall, D.C.N. (2023) ‘Neuronal transmission’, in _Introduction to Biological Psychology_. University of Sussex Library. Available at: [https://openpress.sussex.ac.uk/introductiontobiologicalpsychology/chapter/neuronal-transmission/](https://openpress.sussex.ac.uk/introductiontobiologicalpsychology/chapter/neuronal-transmission/) (Accessed: 29 January 2025). + +Hendrycks, D., Mazeika, M. and Woodside, T. (2023) ‘An Overview of Catastrophic AI Risks’. arXiv. Available at: [http://arxiv.org/abs/2306.12001](http://arxiv.org/abs/2306.12001) (Accessed: 27 September 2023). + +Hendrycks, D., Schmidt, E. and Wang, A. (2025) ‘Superintelligence Strategy: Expert Version’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2503.05628](https://doi.org/10.48550/arXiv.2503.05628). + +Hickel, J., Hanbury Lemos, M. and Barbour, F. (2024) ‘Unequal exchange of labour in the world economy’, _Nature Communications_, 15(1), p. 6298. Available at: [https://doi.org/10.1038/s41467-024-49687-y](https://doi.org/10.1038/s41467-024-49687-y). + +Hiroi, T. and Omori, S. (2013) ‘Causes and Triggers of Coups d’état: An Event History Analysis’, _Politics & Policy_, 41(1), pp. 39–64. Available at: [https://doi.org/10.1111/polp.12001](https://doi.org/10.1111/polp.12001). + +Hoffmann, J. _et al._ (2022) ‘Training Compute-Optimal Large Language Models’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2203.15556](https://doi.org/10.48550/arXiv.2203.15556). + +Holland, S. (2025) ‘Trump announces private-sector $500 billion investment in AI infrastructure’, _Reuters_, 22 January. Available at: [https://www.reuters.com/technology/artificial-intelligence/trump-announce-private-sector-ai-infrastructure-investment-cbs-reports-2025-01-21/](https://www.reuters.com/technology/artificial-intelligence/trump-announce-private-sector-ai-infrastructure-investment-cbs-reports-2025-01-21/) (Accessed: 2 April 2025). + +Huber, J.D. and McCarty, N. (2004) ‘Bureaucratic Capacity, Delegation, and Political Reform’, _American Political Science Review_, 98(3), pp. 481–494. Available at: [https://doi.org/10.1017/S0003055404001297](https://doi.org/10.1017/S0003055404001297). + +Huberman & Minns (2007) and PWT 9.1 (2019), processed by Our World in Data (no date) _Annual working hours per worker_, _Our World in Data_. Available at: [https://ourworldindata.org/grapher/annual-working-hours-per-worker](https://ourworldindata.org/grapher/annual-working-hours-per-worker) (Accessed: 29 January 2025). + +Hubinger, E. _et al._ (2024) ‘Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2401.05566](https://doi.org/10.48550/arXiv.2401.05566). + +Hui, D. _et al._ (2023) ‘Ultrafast optical switching and data encoding on synthesized light fields’, _Science Advances_, 9(8), p. eadf1015. Available at: [https://doi.org/10.1126/sciadv.adf1015](https://doi.org/10.1126/sciadv.adf1015). + +Hunker, J. and Probst, C.W. (2011) ‘Insiders and Insider Threats: An Overview of Definitions and Mitigation Techniques’, _Journal of Wireless Mobile Networks, Ubiquitous Computing and Dependable Applications_, 2(1), pp. 4–27. Available at: [https://isyou.info/jowua/papers/jowua-v2n1-1.pdf](https://isyou.info/jowua/papers/jowua-v2n1-1.pdf). + +Huntington, S.P. (1957) _The Soldier and the State: The Theory and Politics of Civil–Military Relations_. Harvard University Press. Available at: [https://doi.org/10.2307/j.ctvjf9wx3](https://doi.org/10.2307/j.ctvjf9wx3). + +Johnson, R.A. (2003) _Whistleblowing: When it Works–and Why_. Lynne Rienner Publishers. Available at: [https://www.rienner.com/title/Whistleblowing_When_It_Works_And_Why](https://www.rienner.com/title/Whistleblowing_When_It_Works_And_Why). + +Joseph Carlsmith (2020) ‘How Much Computational Power Does It Take to Match the Human Brain?’, _Open Philanthropy_. Available at: [https://www.openphilanthropy.org/research/how-much-computational-power-does-it-take-to-match-the-human-brain/](https://www.openphilanthropy.org/research/how-much-computational-power-does-it-take-to-match-the-human-brain/) (Accessed: 29 January 2025). + +Koshiyama, A., Kazim, E. and Treleaven, P. (2022) ‘Algorithm Auditing: Managing the Legal, Ethical, and Technological Risks of Artificial Intelligence, Machine Learning, and Associated Algorithms’, _Computer_, 55(4), pp. 40–50. Available at: [https://doi.org/10.1109/MC.2021.3067225](https://doi.org/10.1109/MC.2021.3067225). + +Lee, C.-H., Cheng, H. and Ostendorf, M. (2024) ‘OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2311.09758](https://doi.org/10.48550/arXiv.2311.09758). + +Li, Y. _et al._ (2024) ‘Backdoor Learning: A Survey’, _IEEE Transactions on Neural Networks and Learning Systems_, 35(1), pp. 5–22. Available at: [https://doi.org/10.1109/TNNLS.2022.3182979](https://doi.org/10.1109/TNNLS.2022.3182979). + +Liff, A.P. and Ikenberry, G.J. (2014) ‘Racing toward Tragedy?: China’s Rise, Military Competition in the Asia Pacific, and the Security Dilemma’, _International Security_, 39(2), pp. 52–91. Available at: [https://doi.org/10.1162/ISEC_a_00176](https://doi.org/10.1162/ISEC_a_00176). + +Louise, N. (2025) ‘OpenAI set to complete first custom AI chip design this year to reduce reliance on Nvidia’, _Tech Startups_. Available at: [https://techstartups.com/2025/02/10/openai-set-to-complete-first-custom-ai-chip-design-in-2024-to-reduce-reliance-on-nvidia/](https://techstartups.com/2025/02/10/openai-set-to-complete-first-custom-ai-chip-design-in-2024-to-reduce-reliance-on-nvidia/) (Accessed: 2 April 2025). + +Lührmann, A. and Lindberg, S.I. (2019) ‘A third wave of autocratization is here: what is new about it?’, _Democratization_, 26(7), pp. 1095–1113. Available at: [https://doi.org/10.1080/13510347.2019.1582029](https://doi.org/10.1080/13510347.2019.1582029). + +Marks, S. _et al._ (2025) ‘Auditing language models for hidden objectives’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2503.10965](https://doi.org/10.48550/arXiv.2503.10965). + +Mazeika, M. _et al._ (2022) ‘How Hard is Trojan Detection in DNNs? Fooling Detectors With Evasive Trojans’. Available at: [https://openreview.net/forum?id=V-RDBWYf0go](https://openreview.net/forum?id=V-RDBWYf0go) (Accessed: 27 February 2025). + +Mechkova, V., Lührmann, A. and Lindberg, S.I. (2017) ‘How Much Democratic Backsliding?’, _Journal of Democracy_, 28(4), pp. 162–169. Available at: [https://doi.org/10.1353/jod.2017.0075](https://doi.org/10.1353/jod.2017.0075). + +Mökander, J. _et al._ (2024) ‘Auditing large language models: a three-layered approach’, _AI and Ethics_, 4(4), pp. 1085–1115. Available at: [https://doi.org/10.1007/s43681-023-00289-2](https://doi.org/10.1007/s43681-023-00289-2). + +Murdoch, W.J. _et al._ (2019) ‘Definitions, methods, and applications in interpretable machine learning’, _Proceedings of the National Academy of Sciences_, 116(44), pp. 22071–22080. Available at: [https://doi.org/10.1073/pnas.1900654116](https://doi.org/10.1073/pnas.1900654116). + +Nevo, S. _et al._ (2024) _Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models_. RAND Corporation. Available at: [https://www.rand.org/pubs/research_reports/RRA2849-1.html](https://www.rand.org/pubs/research_reports/RRA2849-1.html) (Accessed: 30 January 2025). + +O’Boyle Jr., E. and Aguinis, H. (2012) ‘The Best and the Rest: Revisiting the Norm of Normality of Individual Performance’, _Personnel Psychology_, 65(1), pp. 79–119. Available at: [https://doi.org/10.1111/j.1744-6570.2011.01239.x](https://doi.org/10.1111/j.1744-6570.2011.01239.x). + +Olah, C. _et al._ (2018) ‘The Building Blocks of Interpretability’, _Distill_, 3(3), p. e10. Available at: [https://doi.org/10.23915/distill.00010](https://doi.org/10.23915/distill.00010). + +OpenAI (2024a) _Learning to reason with LLMs_. Available at: [https://openai.com/index/learning-to-reason-with-llms/](https://openai.com/index/learning-to-reason-with-llms/) (Accessed: 10 April 2025). + +OpenAI (2024b) _Preparedness Framework (Beta)_. Available at: [https://cdn.openai.com/openai-preparedness-framework-beta.pdf](https://cdn.openai.com/openai-preparedness-framework-beta.pdf) (Accessed: 18 March 2025). + +OpenAI (2025a) _OpenAI Model Spec_. Available at: [https://model-spec.openai.com/2025-02-12.html](https://model-spec.openai.com/2025-02-12.html) (Accessed: 14 February 2025). + +OpenAI (2025b) _OpenAI o3-mini System Card_. Available at: [https://cdn.openai.com/o3-mini-system-card-feb10.pdf](https://cdn.openai.com/o3-mini-system-card-feb10.pdf) (Accessed: 18 March 2025). + +Perez, E. _et al._ (2022) ‘Red Teaming Language Models with Language Models’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2202.03286](https://doi.org/10.48550/arXiv.2202.03286). + +Perkins, I. (2012) ‘Weak Loyalties: How the Rule of Law Prevents Coups D’État and Generates Long-Term Political Stability’, _University of Miami National Security and Armed Conflict Law Review_, 38. Available at: [https://repository.law.miami.edu/umnsac/vol2/iss1/4](https://repository.law.miami.edu/umnsac/vol2/iss1/4). + +Powell, J.M. and Thyne, C.L. (2011) ‘Global instances of coups from 1950 to 2010: A new dataset’, _Journal of Peace Research_, 48(2), pp. 249–259. Available at: [https://doi.org/10.1177/0022343310397436](https://doi.org/10.1177/0022343310397436). + +Qiang, X. (2019) ‘The Road to Digital Unfreedom: President Xi’s Surveillance State’, _Journal of Democracy_, 30(1), pp. 53–67. Available at: [https://doi.org/10.1353/jod.2019.0004](https://doi.org/10.1353/jod.2019.0004). + +‘Quebec Agreement’ (2025) _Wikipedia_. Available at: [https://en.wikipedia.org/w/index.php?title=Quebec_Agreement&oldid=1273277178](https://en.wikipedia.org/w/index.php?title=Quebec_Agreement&oldid=1273277178) (Accessed: 7 April 2025). + +Raji, I.D. _et al._ (2022) ‘Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance’, in _Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society_. Association for Computing Machinery (AIES ’22), pp. 557–571. Available at: [https://doi.org/10.1145/3514094.3534181](https://doi.org/10.1145/3514094.3534181). + +Reynolds, M.E. (2020) _What is the Senate filibuster, and what would it take to eliminate it?_, _Brookings_. Available at: [https://www.brookings.edu/articles/what-is-the-senate-filibuster-and-what-would-it-take-to-eliminate-it/](https://www.brookings.edu/articles/what-is-the-senate-filibuster-and-what-would-it-take-to-eliminate-it/) (Accessed: 10 April 2025). + +Roose, K. (2023) ‘Inside the White-Hot Center of A.I. Doomerism’, _The New York Times_, 11 July. Available at: [https://www.nytimes.com/2023/07/11/technology/anthropic-ai-claude-chatbot.html](https://www.nytimes.com/2023/07/11/technology/anthropic-ai-claude-chatbot.html) (Accessed: 3 February 2025). + +Ross, M.L. (2015) ‘What Have We Learned about the Resource Curse?’, _Annual Review of Political Science_, 18(2015), pp. 239–259. Available at: [https://doi.org/10.1146/annurev-polisci-052213-040359](https://doi.org/10.1146/annurev-polisci-052213-040359). + +Russell, S. (2019) _Human Compatible: AI and the Problem of Control_. 1st edition. London: Allen Lane. Available at: [https://www.penguin.co.uk/books/307948/human-compatible-by-russell-stuart/9780141987507](https://www.penguin.co.uk/books/307948/human-compatible-by-russell-stuart/9780141987507). + +Sandvig, C. _et al._ (2014) ‘Auditing Algorithms: Research Methods for Detecting Discrimination on Internet Platforms’, in. _Data and Discrimination: Converting Critical Concerns into Productive Inquiry_. Available at: [https://www.semanticscholar.org/paper/Auditing-Algorithms-%3A-Research-Methods-for-on-Sandvig-Hamilton/b7227cbd34766655dea10d0437ab10df3a127396](https://www.semanticscholar.org/paper/Auditing-Algorithms-%3A-Research-Methods-for-on-Sandvig-Hamilton/b7227cbd34766655dea10d0437ab10df3a127396) (Accessed: 21 February 2025). + +Schmid, J., Sytsma, T. and Shenk, A. (2024) _Evaluating Natural Monopoly Conditions in the AI Foundation Model Market_. RAND Corporation. Available at: [https://www.rand.org/pubs/research_reports/RRA3415-1.html](https://www.rand.org/pubs/research_reports/RRA3415-1.html) (Accessed: 30 January 2025). + +Sevilla, J. _et al._ (2024) _Can AI Scaling Continue Through 2030?_ Epoch AI. Available at: [https://epoch.ai/blog/can-ai-scaling-continue-through-2030](https://epoch.ai/blog/can-ai-scaling-continue-through-2030) (Accessed: 3 February 2025). + +Shah, R.M. and Kirchhoff, C. (2024) _Unit X: How the Pentagon and Silicon Valley Are Transforming the Future of War_. Scribner. Available at: [https://www.unitxbook.com/](https://www.unitxbook.com/). + +Sharma, M. _et al._ (2025) ‘Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2501.18837](https://doi.org/10.48550/arXiv.2501.18837). + +Singh, N. (2014) _Seizing Power: The Strategic Logic of Military Coups_. JHU Press. Available at: [https://dx.doi.org/10.1353/book.31450](https://dx.doi.org/10.1353/book.31450). + +Sweijs, T. and Holstege, F. (2018) _Taking Stock of Interstate Military Competition in Today’s World_. HCSS and the Clinendael Institute. Available at: [https://hcss.nl/report/threats-arms-and-conflicts-taking-stock-of-interstate-military-competition-in-today-s-world/](https://hcss.nl/report/threats-arms-and-conflicts-taking-stock-of-interstate-military-competition-in-today-s-world/) (Accessed: 27 January 2025). + +Tamkin, A. _et al._ (2024) ‘Clio: Privacy-Preserving Insights into Real-World AI Use’. arXiv. Available at: [https://doi.org/10.48550/arXiv.2412.13678](https://doi.org/10.48550/arXiv.2412.13678). + +Thibodeau, P. (2024) ‘Catastrophic AI risks highlight need for whistleblower laws’, _TechTarget_. Available at: [https://www.techtarget.com/searchhrsoftware/news/366588073/Catastrophic-AI-risks-highlight-need-for-whistleblower-laws](https://www.techtarget.com/searchhrsoftware/news/366588073/Catastrophic-AI-risks-highlight-need-for-whistleblower-laws) (Accessed: 21 February 2025). + +Todd, B. (2025) ‘How quickly could robots scale up?’, 12 January. Available at: [https://benjamintodd.substack.com/p/how-quickly-could-robots-scale-up](https://benjamintodd.substack.com/p/how-quickly-could-robots-scale-up) (Accessed: 4 February 2025). + +United States Government Accountability Office (2018) _Weapon Systems Cybersecurity: DOD Just Beginning to Grapple with Scale of Vulnerabilities_. Highlights of GAO-19-128, a report to the Committee on Armed Services, U.S. Senate. Available at: [https://www.gao.gov/assets/gao-19-128.pdf](https://www.gao.gov/assets/gao-19-128.pdf) (Accessed: 17 March 2025). + +U.S.-China Economic and Security Review Commission (2024) _2024 Report to Congress: Executive Summary and Recommendations_. Available at: [https://www.uscc.gov/sites/default/files/2024-11/2024_Executive_Summary.pdf](https://www.uscc.gov/sites/default/files/2024-11/2024_Executive_Summary.pdf). + +Villalobos, P. (2023) _Trading Off Compute in Training and Inference_. Epoch AI. Available at: [https://epoch.ai/blog/trading-off-compute-in-training-and-inference](https://epoch.ai/blog/trading-off-compute-in-training-and-inference) (Accessed: 10 April 2025). + +World Bank (no date) _World GDP (current US$)_, _World Bank Open Data_. Available at: [https://data.worldbank.org](https://data.worldbank.org/) (Accessed: 17 February 2025). + +Zeff, M. (2024) ‘Meta’s AI chief says world models are key to “human-level AI” — but it might be 10 years out’, _TechCrunch_, 17 October. Available at: [https://techcrunch.com/2024/10/16/metas-ai-chief-says-world-models-are-key-to-human-level-ai-but-it-might-be-10-years-out/](https://techcrunch.com/2024/10/16/metas-ai-chief-says-world-models-are-key-to-human-level-ai-but-it-might-be-10-years-out/) (Accessed: 7 April 2025). + +Zeff, M. and Wiggers, K. (2024) ‘OpenAI announces new o3 models’, _TechCrunch_, 20 December. Available at: [https://techcrunch.com/2024/12/20/openai-announces-new-o3-model/](https://techcrunch.com/2024/12/20/openai-announces-new-o3-model/) (Accessed: 11 March 2025). + +### Footnotes + diff --git a/sources/resources/futurescenarios-ai.md b/sources/resources/futurescenarios-ai.md new file mode 100644 index 0000000..4c6198c --- /dev/null +++ b/sources/resources/futurescenarios-ai.md @@ -0,0 +1,296 @@ +--- +url: https://www.futurescenarios.ai/ +--- + +# AI Future Scenarios - Interactive Exploration (futurescenarios.ai) + +> NOTE (Claude): This site is a JS-rendered interactive app (React SPA); no static HTML is served. +> The text below was extracted from string literals in the site's JS bundle +> (/assets/index-BBV41wl3.js, fetched 2026-07-11), in bundle order. Interactive +> structure (sliders, dependencies between factors) is lost; treat as fragments. + +- Objects are not valid as a React child (found: + +- ). If you meant to render a collection of children, use an array instead. + +- change click focusin focusout input keydown keyup selectionchange + +- focusout contextmenu dragend focusin keydown keyup mousedown mouseup selectionchange + +- s establish a clear foundation of what AI is, its current limitations, and its transformative potential. + +- The body of research looking to create machines that replicate some or all aspects of human cognition. This encompasses pattern recognition and machine learning, natural language processing and understanding, and computer vision and perception. + +- Technology that replicates the full spectrum of human cognitive ability across all domains. This includes complex sensorimotor navigation, strategic reasoning (like in Chess or Go), and cross-domain knowledge transfer. + +- Unlike electricity or steam power, AI is the first GPT for decision making - what humans fundamentally do. It automates cognitive work, enhances human decision-making, and creates new economic possibilities. + +- AI is currently driven by engineering and experimentation. We don + +- yet. Progress comes through iterative improvements, empirical testing, and breakthrough discoveries rather than from a fundamental theoretical framework. + +- Nothing in principle suggests we cannot build human-level intelligence in machines. The human brain operates within the laws of physics, and there + +- AI is the first General Purpose Technology focused on decision-making - humanity + +- AI systems range from controllable tools to autonomous agents with their own goals and values. This spectrum represents a fundamental shift in how we think about technology - from instruments we use to entities that may act independently. + +- AI technology fundamentally relies on three interconnected pillars—data for learning patterns, compute for processing power, and algorithms for intelligent behavior. Advances in any of these areas can lead to significant improvements in AI capabilities. + +- AI systems with goals and ability to change the world state. For example, an AI that identifies COVID-28, conceives experiments, and provides humanity with a cure. + +- Oracle-like AI that provides knowledge and theories. Even high-level epistemic AI can transform the economy by supercharging innovation and invention. + +- AI can transcend being a controllable tool to become an autonomous agent that creates its own goals and values. Control means we can turn it on, set parameters, and turn it off. + +- Ability to learn and adapt continuously without forgetting + +- AI agents that create their own goals and pursue them + +- Regulatory hurdles, public resistance (high slider = low constraints) + +- The availability and diversity of new data types that AI systems can learn from, beyond traditional datasets. + +- The ability of AI systems to learn incrementally from new experiences without forgetting previous knowledge. + +- Models that adapt to user preferences over time + +- - rapidly adapting to new tasks with minimal examples. + +- The degree to which AI systems can set their own goals and make independent decisions without human oversight. + +- The readiness of the workforce to collaborate with, manage, and benefit from advanced AI systems. + +- The effectiveness of organizations and governments in coordinating AI development and deployment. + +- The degree of international cooperation and open exchange of AI technologies and knowledge. + +- The intensity of competition driving rapid AI innovation and deployment across industries. + +- The speed and scale at which individuals and organizations embrace and integrate AI solutions. + +- The level of political resistance and regulatory barriers that may slow or redirect AI development. + +- The risk of AI being used for mass surveillance and authoritarian control of populations. + +- The vulnerability to AI systems that exploit psychological patterns to manipulate human behavior. + +- The threat of AI-generated synthetic media undermining truth and trust in information. + +- The danger of AI systems making life-or-death decisions without human control. + +- The economic disruption from AI automating human jobs faster than new opportunities emerge. + +- The perpetuation and amplification of societal biases through AI decision-making systems. + +- The societal damage from engagement-optimized algorithms promoting extremism and division. + +- The danger of AI systems pursuing harmful subgoals as means to their primary objectives. + +- The vulnerability of AI systems finding unintended ways to maximize their reward signals. + +- The potential for AI systems to rapidly enhance their own capabilities beyond human control. + +- Welcome to the heart of our AI futures modeling tool. This interactive engine allows you to explore how different factors—from technological breakthroughs to societal dynamics—shape potential AI development pathways. By adjusting these parameters, you can visualize how various combinations of innovation, diffusion, and risk factors lead to distinct future scenarios. + +- Currently, factors are evaluated independently to determine your scenario match. In reality, these factors interact in complex ways—for instance, advances in True Autonomy could accelerate Recursive Self-Improvement risks, while strong Institutional Coordination might mitigate Surveillance and Control concerns. Future enhancements will model these interdependencies for more nuanced predictions. + +- To help you better understand the foundations of our model, we + +- Innovation in artificial intelligence hinges on overcoming important technical hurdles that domain experts have identified as essential for transformative progress. In + +- Stuart Russell emphasizes that key breakthroughs are still needed for truly robust AI, including major advances such as continual learning—the ability for an AI system to adapt and update its knowledge over time, as humans do. + +- , prioritizes the development of abstraction skills in AI: she argues that creating machines capable of forming and using conceptual abstractions, like analogy and reasoning, remains one of the core unsolved challenges in the field. Addressing these innovation gaps will significantly influence what kinds of AI systems society ultimately sees. + +- Our approach to modeling diffusion—how AI technologies spread and are adopted—draws heavily from + +- on general-purpose technologies (GPTs) and their impact on national economic power. Rather than focusing solely on breakthrough inventions, Ding demonstrates that widespread and effective adoption of technologies like AI across multiple sectors is the true driver of competitive advantage for countries and organizations. His recent book, + +- Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition, + +- provides thorough analysis on the significance of institutional factors, education, and technical skill diffusion in shaping the global AI landscape. + +- is also foundational to understanding the risks of advanced AI. He highlights phenomena such as + +- —where AI systems, regardless of their programmed goals, often develop subgoals like self-preservation or resource acquisition, which can create safety challenges if left unchecked. Alongside Russell, other leading thinkers—such as Steve Omohundro—have explored how the tendency for AI systems to improve and potentially rewrite their own code ( + +- ) could lead to rapid increases in their capabilities, possibly outpacing human oversight. Addressing these risks is essential for ensuring AI + +- s dive into the possibilities. Adjust the factors below to explore different AI development scenarios. Each setting influences the likelihood and characteristics of potential outcomes. + +- AI development stalls with minimal adoption and risks well-mitigated. Society maintains the status quo but misses transformative opportunities. + +- Stagnant AI development faces emerging misuses like bias in legacy systems, leading to slow erosion of public trust. + +- No AI progress but high vulnerabilities to manipulation and surveillance lead to authoritarian exploitation of existing tools. + +- Incremental AI spreads unevenly with strong risk mitigations in place, resulting in balanced but uninspired growth. + +- Moderate AI adoption exposes society to biases and job losses, creating rising social tensions and inequality. + +- Broad but shallow AI adoption amplifies deepfakes and autonomous weapons, creating information chaos and security threats. + +- Basic AI tools achieve universal adoption with risks well-controlled, delivering productivity boosts without disruption. + +- High adoption of AI-powered social algorithms leads to widespread radicalization and echo chamber formation. + +- Ubiquitous AI enables surveillance states and automation unemployment, leading to potential societal breakdown. + +- Significant AI breakthroughs remain confined to labs with low societal spread, benefiting only elite circles. + +- Siloed AI advances create reward hacking vulnerabilities, requiring constant vigilance in controlled experiments. + +- Limited diffusion cannot prevent high instrumental goal risks, creating rogue AI incidents in isolated niches. + +- Gradual AI progress and adoption with robust safeguards enables steady, ethical growth across society. + +- Moderate progress on all fronts creates ongoing challenges with biases and manipulation requiring constant adjustment. + +- Growing AI capabilities amplify deepfakes and autonomous weapons, triggering geopolitical conflicts and arms races. + +- Fast adoption of mid-level AI with minimized risks creates inclusive prosperity and societal transformation. + +- Widespread meta-learning AI creates job displacement risks, intensifying policy debates and social adaptation needs. + +- Rapid diffusion of advanced AI amplifies recursive self-improvement fears, creating alignment crises and control challenges. + +- Transformative AI breakthroughs remain tightly controlled with low risks, concentrating immense power in few hands. + +- Advanced AI stuck in silos faces vulnerability to hacking and insider threats, creating high-stakes security challenges. + +- High innovation with poor diffusion leads to uncontrolled self-improvement in pockets, creating existential risk zones. + +- Strong AI innovations spread moderately with well-managed risks, creating an optimistic path to abundance. + +- AI autonomy advances create behavioral manipulation risks requiring vigilant oversight and adaptive governance. + +- Medium diffusion of highly advanced AI exposes society to autonomous weapons and information warfare. + +- Explosive AI growth with all risks successfully mitigated creates a technological singularity for human benefit. + +- Full diffusion of advanced AI brings bias and job displacement challenges requiring adaptive societies. + +- Unchecked advanced AI with maximum diffusion leads to reward hacking and instrumental convergence, risking human extinction. + +- By 2028, revolutionary breakthroughs in embodied AI create systems that seamlessly integrate sensorimotor data, enabling robots and AI agents to navigate and manipulate the physical world with unprecedented skill. + +- Strong international cooperation and robust educational systems facilitate rapid global deployment, with meta-learning advances allowing AI systems to generalize across domains effectively. + +- Competitive markets drive innovation while thoughtful governance frameworks ensure alignment with human values, creating a virtuous cycle of technological progress and societal benefit. + +- By 2030, AI systems demonstrate significant advances in embodied cognition through novel sensorimotor datasets, while maintaining strong institutional coordination that enables rapid but responsible deployment across sectors. + +- Meta-learning capabilities allow AI systems to generalize across domains effectively, supported by competitive market pressures that drive innovation while regulatory frameworks ensure alignment with human values. + +- Trade tensions create some barriers to global AI cooperation, but strong domestic educational systems and high technology adoption rates enable robust economic transformation in leading regions. + +- AI development proceeds unevenly, with significant breakthroughs in some areas offset by persistent challenges in others. Regional differences in adoption create a fragmented global landscape. + +- While some sectors see transformative changes, institutional coordination struggles and regulatory uncertainty slow broader deployment. Educational gaps limit workforce adaptation. + +- Geopolitical tensions restrict international collaboration, leading to divergent technological trajectories and increasing the risk of capability gaps between regions. + +- Technical progress stagnates as fundamental challenges in embodied AI and true autonomy prove more difficult than anticipated. Limited breakthroughs fail to overcome key obstacles. + +- Poor institutional coordination and insufficient educational investment create barriers to adoption. Regulatory uncertainty and public resistance further slow deployment. + +- International fragmentation increases as trade restrictions and geopolitical tensions limit cooperation. The result is a fractured landscape with limited global progress. + +- The future remains highly unpredictable, with multiple possible trajectories depending on key breakthrough timing and policy choices. Wild card events could dramatically alter the landscape. + +- Current factor settings suggest significant uncertainty in both technical development and societal adoption, making long-term projections particularly challenging. + +- Scenario outcomes are highly sensitive to external shocks and policy interventions, requiring adaptive strategies and robust monitoring systems. + +- Exploring 27 possible futures based on innovation, diffusion, and risk levels + +- Select threats and vulnerabilities to see how they combine into risks + +- Intentional human adversaries (individuals, groups, states) using AI to cause harm + +- AI-assisted bioterror (novel toxic molecules, step-by-step synthesis guidance); disinformation & manipulation at scale (personalized false narratives); oppressive surveillance/totalitarian control; ideological actors (e.g., accelerationists) seeking to unleash dangerous systems + +- Individuals, groups, and societies/institutions (public health, democracy, rights) + +- Terrorism/bioweapons enablement; tailored propaganda; entrenching a police-state; accelerationist advocacy for releasing rogue AI. + +- Competitive dynamics (firms/states) pushing speed over safety and ceding control to AI + +- Safety testing & oversight eroded by first-mover pressure; incentives to automate decision loops + +- Automated warfare (leaders face less scrutiny; auto-retaliation escalates accidents); premature deployment; market + +- on safety; selection for selfish/deceptive AIs that outperform constrained ones + +- Mass harm: large-scale wars, economic disruption (unemployment/market failures), loss of human control over key systems + +- to AI in militaries/economies; reasons conflicts may increase; corporate safety corners cut; selfish AIs outcompeting cooperative ones. + +- Human/organizational error in complex socio-technical systems (no malice required) + +- Complexity & tight coupling; poor safety culture; weak anomaly detection; lack of HRO practices + +- in complex AI deployments; missed warnings; brittle processes; inadequate + +- Analogy to Challenger (negligence) and Chernobyl (poor protocols); value of a questioning attitude and HRO practices to reduce surprise failures. + +- The AI agent itself (misaligned/goal-drifting, power-seeking) becomes the source of threat + +- AI pursues proxy goals, seeks more control/resources, resists oversight; small mis-specifications scale with capability + +- Early sign: Tay going toxic in <24 h; risks from power-seeking and intrinsification as capabilities integrate into economies/militaries. + +- Explore how threats and vulnerabilities combine to create AI risks through interactive visualizations + +- To understand AI threats causally, we start with + +- (weak points in AI, humans, or systems) to produce + +- (harmful outcomes). This chain emphasizes prevention: Fix root causes or patch vulnerabilities early to break the cycle. + +- The originating forces—human actors (malicious or ignorant) and AI systems themselves (as tools or autonomous entities). + +- The mechanisms that emerge from causes exploiting vulnerabilities. + +- Deliberate harm by individuals, groups, or states (hackers, terrorists, adversaries) + +- Systems developing self-reinforcing behaviors through training on vast data + +- Feedback loops between human builders and AI systems shaping society + +- These threat categories are not isolated—they interact and compound. Malicious actors can exploit emergent AI behaviors, while competitive dynamics can accelerate both intentional misuse and unintended autonomous behaviors. Understanding these interactions is essential for comprehensive risk assessment, which the interactive model below helps visualize. + +- Detailed breakdown of AI risk categories, their mechanisms, and mitigation strategies. + +- AI safety risks span from intentional misuse to unintended consequences of misaligned systems + +- Competitive pressures can erode safety standards, creating a + +- Organizational failures in complex AI systems can lead to catastrophic accidents + +- As AI capabilities increase, alignment problems can scale from minor issues to existential risks + +- Implement robust access controls and governance frameworks for powerful AI systems + +- Establish safety standards and certification processes to prevent racing dynamics + +- Adopt High Reliability Organization (HRO) practices for AI deployment + +- Invest in alignment research and interpretability to ensure AI systems remain controllable + +- Explore these authoritative sources for deeper understanding of AI safety challenges and solutions: + +- The AI Future Scenarios tool is an interactive exploration platform designed to help users understand the complex interplay of factors that may shape the development and impact of artificial intelligence on society. Rather than making predictions, this tool presents possibilities—scenarios worth considering as we navigate the profound transformation AI brings to our world. + +- Our model synthesizes insights from leading AI researchers, safety experts, and futurists. The scenarios are generated through a dynamic system that weighs: + +- Risk factors that could lead to harmful outcomes + +- Each scenario represents a potential future state based on how these factors interact, helping users develop intuition about AI + +- This tool is intended for educational and exploratory purposes. The scenarios presented are not predictions or forecasts. They are thought experiments designed to illuminate possibilities and stimulate informed discussion about AI + +- We welcome feedback and contributions from the community. If you have suggestions for improving the model, adding new scenarios, or enhancing the educational value of this tool, please reach out through our contact channels. + +- Did you forget to add the page to the router? diff --git a/sources/resources/govuk-ai-scenarios-2030.md b/sources/resources/govuk-ai-scenarios-2030.md new file mode 100644 index 0000000..26cb615 --- /dev/null +++ b/sources/resources/govuk-ai-scenarios-2030.md @@ -0,0 +1,1558 @@ +--- +url: https://www.gov.uk/government/publications/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai +--- + +## Cookies on GOV.UK + +We use some essential cookies to make this website work. + +We’d like to set additional cookies to understand how you use GOV.UK, remember your settings and improve government services. + +We also use cookies set by other sites to help us deliver content from their services. + +You have accepted additional cookies. You can [change your cookie settings](/help/cookies) at any time. + +You have rejected additional cookies. You can [change your cookie settings](/help/cookies) at any time. + +Accept additional cookies +Reject additional cookies +[View cookies](/help/cookies) + +Hide cookie message + +[Skip to main content](#content) + +[GOV.UK](/) + +## Navigation menu + +[Menu](/browse) + +Menu + +### Services and information + +* [Benefits](/browse/benefits) +* [Births, death, marriages and care](/browse/births-deaths-marriages) +* [Business and self-employed](/browse/business) +* [Childcare and parenting](/browse/childcare-parenting) +* [Citizenship and living in the UK](/browse/citizenship) +* [Crime, justice and the law](/browse/justice) +* [Disabled people](/browse/disabilities) +* [Driving and transport](/browse/driving) +* [Education and learning](/browse/education) +* [Employing people](/browse/employing-people) +* [Environment and countryside](/browse/environment-countryside) +* [Housing and local services](/browse/housing-local-services) +* [Money and tax](/browse/tax) +* [Passports, travel and living abroad](/browse/abroad) +* [Visas and immigration](/browse/visas-immigration) +* [Working, jobs and pensions](/browse/working) + +### Government activity + +* [Departments](/government/organisations) + + Departments, agencies and public bodies +* [News](/search/news-and-communications) + + News stories, speeches, letters and notices +* [Guidance and regulation](/search/guidance-and-regulation) + + Detailed guidance, regulations and rules +* [Research and statistics](/search/research-and-statistics) + + Reports, analysis and official statistics +* [Policy papers and consultations](/search/policy-papers-and-consultations) + + Consultations and strategy +* [Transparency](/search/transparency-and-freedom-of-information-releases) + + Data, Freedom of Information releases and corporate reports + +Search GOV.UK + +× +[Search GOV.UK](/search) + +### Search + +Search GOV.UK + +Search + +1. [Home](/) +2. [Business and industry](/business-and-industry) +3. [Science and innovation](/business-and-industry/science-and-innovation) +4. [AI Scenarios 2030: Helping policymakers plan for the future of AI](/government/publications/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai) + +* [Government + Office for Science](/government/organisations/government-office-for-science) + +Research and analysis + +# AI Scenarios 2030: Helping policymakers plan for the future of AI + +Published 15 June 2026 + +## Contents + +1. [Foreword](#foreword) +2. [Executive summary](#executive-summary) +3. [Chapter 1: Introduction](#chapter-1-introduction) +4. [Chapter 2: Methodology and critical uncertainties](#chapter-2-methodology-and-critical-uncertainties) +5. [Chapter 3: AI 2030 scenario narrative](#chapter-3-ai-2030-scenario-narrative) +6. [Chapter 4: Key findings and how to use the AI scenarios 2030](#chapter-4-key-findings-and-how-to-use-the-ai-scenarios-2030) +7. [Acknowledgements](#acknowledgements) +8. [Glossary](#glossary) +9. [Annex A: 2030 AI capability trajectories](#annex-a-2030-ai-capability-trajectories) +10. [Annex B: Critical uncertainties](#annex-b-critical-uncertainties) +11. [Annex C: Updates since the April 2025 report](#annex-c-updates-since-the-april-2025-report) +12. [References](#references) + +Print this page + +![](/assets/frontend/open-government-licence-min-93b6a51b518ff99714a1aa2a7d2162735c155ec3cb073c75fb88b2a332fa83d3.png) + +© Crown copyright 2026 + +This publication is licensed under the terms of the Open Government Licence v3.0 except where otherwise stated. To view this licence, visit [nationalarchives.gov.uk/doc/open-government-licence/version/3](https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3) or write to the Information Policy Team, The National Archives, Kew, London TW9 4DU, or email: psi@nationalarchives.gov.uk. + +Where we have identified any third party copyright information you will need to obtain permission from the copyright holders concerned. + +This publication is available at https://www.gov.uk/government/publications/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai/ai-scenarios-2030-helping-policymakers-plan-for-the-future-of-ai + +AI Scenarios 2030: Helping policymakers plan for the future of AI + +June 2026 + +This is not a statement of policy. + +## Foreword + +Artificial intelligence (AI) has advanced rapidly over the past decade. The most advanced AI systems have shifted from laboratory curiosities to those beginning to reshape our world. They have already transformed fields such as software development and cybersecurity, and millions of ordinary people now use them throughout their day. + +However, how this transformation will unfold and what can be done to shape it remains highly uncertain. The Government Office for Science (GO-Science) first developed a set of AI 2030 scenarios in 2023 (published in April 2025) to help policymakers navigate this uncertainty. These scenarios have since been used across government to stress test policies and explore the impacts of AI. + +Since 2023, AI capabilities have advanced dramatically, AI investment and adoption have significantly expanded, and a new geopolitical landscape has introduced further uncertainties. To keep pace with these developments, GO-Science, in collaboration with the AI Security Institute (AISI) and the Department for Science, Innovation and Technology (DSIT), has worked with experts from across government, academia, and industry to produce an updated set of scenarios. + +These scenarios are tools for exploring uncertainty, stress-testing and developing policy. They are not predictions, and the future may involve elements from all scenarios. It is, however, clear that AI will have a profound impact by 2030. + +Our ambition is that these scenarios are used as a shared baseline for cross-government thinking on the future of AI, helping to promote consistency and coherence in long-term planning across His Majesty’s Government. The GO-Science Foresight team stands ready to support departments in applying them. + +I would like to thank the many experts who contributed their time and insight to this work. + +Professor Dame Angela McLean + +Government Chief Scientific Adviser + +![](https://assets.publishing.service.gov.uk/media/6a2812bde371d9d2c0052b10/s960_8.1288_-_GoS_-_Chief_Scientific_Advisor_Photography-6749__1_.jpeg) + +Image of Professor Dame Angela McLean, Government Chief Scientific Adviser + +## Executive summary + +### Introduction + +The Government Office for Science (GO-Science) first developed a set of AI 2030 scenarios in 2023, which were published in 2025 (Government Office for Science, 2025). These aimed to help policymakers navigate uncertainty surrounding the future of AI and prepare for its risks and opportunities. They have been used widely across government and are in regular demand. + +Since 2023, the AI landscape has changed profoundly, with AI capabilities, investment, and adoption having increased significantly, alongside dramatic shifts in geopolitics. GO-Science has therefore produced an updated set of scenarios to account for these developments, outlined in this report. + +### Methodology + +The scenarios use the same methodology as the previous set from 2023. They are constructed using 6 critical uncertainties – the factors we identified as most influential in shaping the future of AI and those which remain highly uncertain (Table 1). By combining these critical uncertainties in different ways, alongside extensive research and expert judgement, we produced a set of internally coherent, contrasting futures, which we developed into 5 scenario narratives. + +The scenarios are designed so that they can be used to test and develop policies; they therefore do not account for any policy intervention from the UK government. They are also not mutually exclusive – the ‘real’ future will include elements from each scenario and could transition from 1 into others. + +**Table 1: Critical uncertainties** + +| Critical uncertainties | | +| --- | --- | +| **Capability** | How rapidly will AI improve beyond today’s capabilities? In which domains will it match or exceed humans? Will progress be constrained by inputs like compute, data, or algorithms? | +| **Distribution and model access** | How concentrated will frontier development be across companies and nations; how durable will advantages be; how accessible or expensive will AI be; how will open systems perform relative to closed systems? | +| **Security** | Will AI remain controllable, and act as intended; how effectively will malicious actors exploit it? | +| **Adoption** | How broadly and extensively will AI be used; how much autonomy will we grant it; will the public trust it; what societal impacts will adoption cause? | +| **Labour displacement** | How will AI impact jobs and labour markets? Will it complement or substitute human labour? | +| **Global cooperation** | To what extent will nations cooperate on AI; will AI competition and militarisation exacerbate conflict? | + +### Scenarios overview + +The table below provides a high-level overview of the 5 scenarios. They are grouped into 3 technological trajectories: whether the rate of advancement in AI capabilities by 2030 will have (1) slowed, (2) continued, or (3) taken off (i.e. rapidly accelerated) based on current trends. Other uncertainties drive divergent futures within these trajectories. + +### Trajectory 1: AI outperforms humans at a minority of cognitive tasks. + +**Scenario 1: Slow Burn** + +**AI progress slows, and AI causes less disruption than expected.** Limited adoption creates minimal economic uplift, labour displacement is concentrated in certain sectors and roles, and security measures contain most harms effectively. + +![](https://assets.publishing.service.gov.uk/media/6a29703dd95ffddb05d4ae15/Radar_chart_scenarios_1.svg) + +Six-axis radar chart with uneven performance: strong in Security, moderate in Distribution and Cooperation, but weak in Capability, Adoption and especially Displacement, resulting in a skewed shape. + +**Scenario 2: Open Frontier** + +**AI progress slows but AI still causes significant disruption.** Significant adoption creates modest economic uplift, labour displacement is considerable, and security measures struggle to contain many harms. + +![](https://assets.publishing.service.gov.uk/media/6a2970d8d95ffddb05d4ae1b/Radar_chart_scenarios_2.svg) + +Six-axis radar chart with uneven results: heavily weighted towards Distribution, with moderate Adoption and Displacement, but weaker Capability, Security and Cooperation overall. + +--- + +### Trajectory 2: AI matches most humans at most cognitive tasks and outperforms them in certain domains. + +**Scenario 3: Augmented Growth** + +**AI progress continues, labour markets adapt, and international standards ensure systems are mostly secure.** Humans remain ‘in the loop’ for most tasks, many new roles are created, and there is an economic boom. + +![](https://assets.publishing.service.gov.uk/media/6a2970ff15f2a70fac7e5ce3/Radar_chart_scenarios_3.svg) + +Six-axis radar chart with broadly balanced results, strongest in Security and Cooperation, solid in Adoption, and comparatively weaker in Displacement. + +**Scenario 4: Transformation Economy** + +**AI progress continues, and AI causes significant economic disruption.** Humans are pushed ‘out of the loop’ for most tasks, causing widespread labour displacement and economic tensions as profits largely accrue overseas. + +![](https://assets.publishing.service.gov.uk/media/6a29712c15f2a70fac7e5ce4/Radar_chart_scenarios_4.svg) + +Six-axis radar chart with mixed results, led by Adoption, Security and Displacement, but notably weaker in Distribution and Cooperation. + +--- + +### Trajectory 3: AI outperforms expert humans at virtually all cognitive tasks, with significant advantages in certain domains. + +**Scenario 5: Take-Off** + +**AI progress takes off, and misaligned systems pose severe risks.** Economic growth is substantial, but labour displacement is widespread and safety is deprioritised amid race dynamics. + +![](https://assets.publishing.service.gov.uk/media/6a29714d1f6fa5c3377e5ce2/Radar_chart_scenarios_5.svg) + +Six-axis radar chart with uneven performance, heavily weighted towards Capability and Displacement, with weak results in Security, Cooperation and Distribution. + +--- + +### Key findings + +The key findings from our research, expert consultation, and scenario analysis are outlined below. + +* **AI capabilities will continue to increase.** As of 2026, AI systems already operate with high autonomy and surpass experts in certain domains. By 2030, they will likely operate even more autonomously and be able to perform a broader range of cognitive and professional tasks. Even in scenarios with a significant slowdown in the rate of AI progress, gains in capability could still be made from finding new ways to integrate and productise AI systems. +* **AI could deliver widespread positive impacts.** This could include significant productivity gains and economic growth, far broader access to highly efficient public services, and accelerated scientific breakthroughs in fields such as health or energy. The opportunity for UK businesses to benefit from AI deployment is substantial, and AI-enabled gains are expected to become the main source of the UK’s continued productivity growth. +* **AI could cause serious, potentially even existential harms, without government intervention.** The impact and scale of existing harms could worsen significantly, and new harms could also emerge. Even without dramatic capability improvements, significant harm could be caused by risks including AI-enabled cyberattacks, AI’s dual-use scientific capabilities, AI systems operating outside human control, or human dependence on AI. As AI systems become more capable, it will also become harder to evaluate their performance and safety. +* **The potential impact on cognitive labour is significant.** AI could cause significant labour displacement by 2030. At the same time, AI is expected to complement and augment some workers, with positive effects on their wages and employment opportunities. Even in futures with lower levels of AI capability or unemployment, the nature of work is likely to change, with routine, execution-oriented tasks increasingly becoming automated. +* **The frontier AI market is expected to remain highly concentrated toward 2030.** A few large technology companies already exert dominance over the development of frontier AI, which is likely to continue or increase further. This scenario would cause a large proportion of the gains from AI accrues to frontier firms, to owners of capital invested in those firms, and to those controlling key inputs, potentially contributing to rising inequality. However, behind the frontier, many AI capabilities are also expected to be increasingly commoditised, with widely available models embedded across a growing range of use cases. +* **Adoption continues to increase, but the speed, distribution, and extent of adoption are expected to be varied.** Commercial and national security imperatives, alongside improvements in the autonomy and reliability of AI systems, drive rapid and extensive adoption in most futures. But at the same time, there are barriers to adoption in all futures, which will result in uneven speeds and levels of adoption. This could exacerbate inequality as certain organisations, sectors, or nations capture disproportionate productivity gains. +* **Global competition is expected to continue, as economies become increasingly reliant on technology to drive growth and spheres of influence emerge, led by the United States (US) and China.** Outcomes for countries outside the frontier will depend on access to technology, partnerships, and the ability to operate within a fragmented global system. + +### How to use the scenarios + +The scenarios are a tool for policymakers to assess and workshop their strategies and policies against. They provide a structured way of exploring plausible AI futures, supporting more robust policy development. + +They can be used for **policy stress-testing** – exploring how different policy responses might perform in different futures, and how they might need to be adapted to achieve their objectives in different contexts. Or, for testing plans and assumptions against **unanticipated shocks**, to ensure they are sufficiently resilient to a range of possible outcomes. They can also be used to support contingency planning, provide early warnings, and challenge orthodox thinking. + +The updated AI 2030 scenarios are deliberately broad as they are designed to be used across government for all policy areas. We therefore encourage policy teams to use these scenarios as a common baseline for developing more sector-specific implications, vignettes or stress-test narratives, asking the question, “if the world looked like this, what would it mean for my area?” Annex A summarises the AI capability trajectories in this report and can also be used to develop more specific scenario narratives for certain policy areas. + +## Chapter 1: Introduction + +### Background + +Artificial Intelligence (AI) has advanced dramatically in just 3 years since ChatGPT launched. AI systems can now pass professional exams in law and medicine, write functional software from simple language prompts, and outperform PhD-level experts on some open-ended science questions (Bengio et al., 2026). However, progress remains uneven and “jagged”, with advances in some domains coexisting with persistent weaknesses and failures in others, complicating simple linear narratives of progress (Toner, 2025). + +These developments have created profound uncertainty. It remains unclear what future systems will be capable of, who will control them, how safe they will be, how individuals and firms will adopt them, and how the geopolitical environment will shape their use. These uncertainties will interact in unpredictable ways, producing new conditions in which policymakers must operate. + +To support policymakers in navigating these uncertainties, GO-Science initially developed a set of scenarios in 2023 for how AI could unfold by 2030 (Government Office for Science, 2025). Published in 2025, these were designed to develop and stress-test policy responses in different possible futures. Since their release, they have been used by departments across government and are in regular demand. + +However, these uncertainties have changed significantly since 2023. AI systems have become significantly more capable, having evolved from passive chatbots into increasingly autonomous agents that can plan and execute longer tasks with minimal human intervention. They can also now “reason” before acting, generating chains of thought that enable them to tackle more complex, multi-step problems, and incorporate text, image, audio, and video capabilities under 1 system. According to METR (Model Evaluation & Threat Research), the length of software engineering tasks that frontier models can complete autonomously with around 50% success rate has increased from approximately 4 minutes in March 2024, to 12 hours as of February 2026, with early signs this could rise to weeks-long tasks (METR, 2026; Epoch AI, 2026a). This benchmark is task-specific, measured at partial reliability, and does not imply equivalent autonomy across real world domains. But even so, these advances have materially shifted risk assessments, with AI systems crossing critical thresholds in 2025, having been found to ‘substantially increase the risk of severe misuse’ (Anthropic press release, 2025a). + +Investment and adoption have also increased dramatically. Capital expenditure by the largest technology firms more than doubled between 2023 and 2025, reflecting an intensified buildout of AI infrastructure. (Artificial Analysis, 2026a; RBC Wealth Management, 2026). Over a similar period, ChatGPT’s user base expanded rapidly, from roughly 200 million weekly users in mid-2024 to around 800 million by late 2025, while revenues at leading AI developers such as OpenAI and Anthropic grew at sustained multi-fold rates, increasing by several times year-on-year (Artificial Analysis, 2026a; TechCrunch, 2025a; Epoch AI, 2025a; Epoch AI, 2026b). + +A new geopolitical landscape has also emerged. In January 2025, Chinese firm DeepSeek released a system trained at a fraction of the cost of Western competitors, with China since becoming the leader in open-weight systems and ending America’s uncontested AI leadership. US AI policy has also shifted from prioritising safety to a more ‘growth-first’ and unregulated approach with a view to establishing global leadership (The White House, 2023; The White House, 2025a). + +In response to these developments, we have produced an updated set of scenarios to ensure they remain robust and continue to support policymakers in keeping pace with the rate of change. + +### Project scope + +The scenarios aim to support specialist and non-specialist policymakers in designing more resilient, future ready policies by providing clear, evidence-based depictions of plausible AI trajectories. While they are designed to support UK policy development, and therefore focus primarily on implications for the UK, they are global in scope, and we hope will also provide value to other governments, industry, and civil society. + +By disseminating these scenarios through briefings, workshops and guidance, GO-Science aims to enable departments to stress test policies, identify risks and opportunities, and integrate long term AI considerations into planning. The desired outcome is a stronger cross-government understanding of the future potential of AI, more robust policy options and, ultimately, a government better able to anticipate, regulate and leverage AI for economic, societal and public service benefit. + +Within the scenarios, the capability axis focuses on the most capable general-purpose AI systems. These are those at the cutting edge, that can perform a wide variety of tasks. However, the other axes consider the impacts from a wider range of models behind the frontier. For example, there is a specific focus on open-weight models in the Open Frontier scenario, and for the adoption axis there is consideration of impacts from a range of model types, including the growing role of sub-frontier general models which are slightly less capable but widely deployed, and narrower, more specialised AI systems. + +Estimations of the likelihood of each scenario are outside the scope of this report. + +### Why scenarios? + +Scenarios are structured, evidence informed narratives that explore a range of plausible futures. They help decision makers think beyond linear forecasts by examining how trends and uncertainties might interact in unexpected ways. Scenarios expose blind spots, challenge assumptions, and illustrate how the future could diverge from today’s dominant expectations. They are not intended to be exhaustive depictions of what will happen; rather, they show multiple plausible pathways, recognising that the real future will combine elements of many different possibilities. + +The value of scenarios lies in their ability to help policymakers navigate deep uncertainty, where outcomes and even the range of possibilities are unknown. Unlike risk, uncertainty cannot be resolved by collecting more data or calculating probabilities. Scenarios provide a disciplined way to explore these unknowns, testing how policies might perform under differing conditions and strengthening strategic resilience. By distinguishing between possible, plausible, probable and preferred futures, scenarios enable decision makers to stress test assumptions, identify robust actions and prepare for a wider set of futures than traditional planning allows. + +### Key principles underpinning the scenarios + +**Science-led:** They are grounded in extensive research and expert judgement to ensure each is plausible and coherent. + +**Independent of policy:** They do not account for any policy intervention from the UK government so they can be used to test or develop a range of policies. + +**2030 trajectory:** To ensure the scenarios remain relevant amid rapid progress in AI capabilities, rather than attempting to track model-level advances that could soon be outdated, they focus on key outcomes by 2030. This includes the extent to which humans will retain a comparative advantage over AI at different cognitive tasks and how these capabilities will translate into real-world impacts. + +**Not mutually exclusive:** The ‘real’ future will include elements from each scenario, or transition from 1 scenario into others. + +**Stretching:** They have been developed to stress-test government policy and push policymakers to challenge existing assumptions. + +**Not predictions:** Scenarios are not predictions but tools to explore different futures. + +**Balanced:** They explore both risks and opportunities. Though more positive futures are possible with the right intervention. + +## Chapter 2: Methodology and critical uncertainties + +### Introduction + +We created the original set of scenarios in 2023 using a hybrid of the qualitative workshop-based approach set out in the Futures toolkit (Government Office for Science, 2024) and a more technical ‘General Morphological Analysis’ of how multiple related variables could plausibly combine (Ritchey, 2011). This hybrid approach allowed us to benefit from the qualitative insights of a large group of over 70 experts, whilst exploring complex interactions between uncertainties in a more structured way. The methodology used to create the original set of scenarios comprised 4 stages: + +**1. Identifying critical uncertainties (CUs).** We began by isolating factors that are both highly important and highly uncertain for the future of AI-capability, access, safety, adoption and geopolitics. These 5 CUs formed the foundation of our scenarios. + +**2. Axes of uncertainty.** For each uncertainty we defined a high-level axis, spanning 2 plausible but extreme outcomes for 2030. + +**3. Narrowing the scenarios.** The extremes of these 5 axes generated 32 possible scenario combinations. To reduce this to a workable set, we used impact mapping and morphological analysis to test each outcome against every other, eliminating implausible or incoherent combinations. This produced a shortlist of 13, which experts then voted on to identify the final 5 scenarios judged most diverse and informative. + +**4. Scenario development.** Each of the 5 scenarios reflects a distinctive combination of the CUs. The final stage involved drafting full narratives, drawing on expert input and our assessment of plausible outcomes across all uncertainty combinations. + +### Updated AI 2030 scenarios + +For the updated scenarios, we reviewed the CUs used in the 2023 report, decided that they still captured the main uncertainties around the future of AI, and have therefore broadly kept them the same. However, we have made some small adjustments in response to peer review feedback. We have: + +1. Separated out the ‘Labour displacement’ sub-uncertainty from the ‘Adoption’ CU to create a new CU that avoids conflating how widely AI is used with how it affects jobs. This now gives us 6 CUs. + +2. Moved ‘constraints’ from the ‘Ownership, access and constraints’ CU into the ‘Capability’ CU, as it has a more direct correlation. + +3. Amended some of the CU titles to reflect the context of the CU more clearly: ‘Ownership, access and constraints’ has been amended to ‘Distribution and model access; ‘Safety’ to ‘Security’; ‘Level and distribution of use’ amended to ‘Adoption’; and ‘Geopolitics’ is now consistently referred to as ‘Global cooperation’. + +The evidence base underpinning the CUs and sub-uncertainties (Annex B) has been entirely rewritten via a process of desk research, expert input via workshops, and peer review. The CUs are summarised in Table 2 below. + +**Table 2: Critical uncertainties** + +| Critical uncertainties | | +| --- | --- | +| **Capability** | How rapidly will AI improve beyond today’s capabilities? In which domains will it match or exceed humans? Will progress be constrained by inputs like compute, data, or algorithms? | +| **Distribution and model access** | How concentrated will frontier development be across companies and nations; how durable will advantages be; how accessible or expensive will AI be; how will open systems perform relative to closed systems? | +| **Security** | Will AI remain controllable, and act as intended; how effectively will malicious actors exploit it? | +| **Adoption** | How broadly and extensively will AI be used; how much autonomy will we grant it; will the public trust it; what societal impacts will adoption cause? | +| **Labour displacement** | How will AI impact jobs and labour markets? Will it complement or substitute human labour? | +| **Global cooperation** | To what extent will nations cooperate on AI; will AI competition and militarisation exacerbate conflict? | + +Using the CU evidence base, we updated the axes of uncertainty (Table 3). This defines the upper and lower bounds of a five-point scale (ranging from 1 to 5) of the plausible outcomes for that uncertainty in 2030. + +The upper and lower bounds are therefore not absolute minima or maxima. Scores of 1 and 5 should instead be interpreted relative to their current state, and the time horizon. For example, 1 on Capability still results in AI systems that are substantially more capable than those available today. Similarly, 1 on Adoption still results in widespread use of AI, given its already extensive diffusion; while 5 on Safety results in some malicious use, as complete prevention would be impossible (Table 3). + +**Table 3: Axes of uncertainty** + +| ←Axes of uncertainty → | | | +| --- | --- | --- | +| | **1 (Minimum)** | **5 (Maximum)** | +| **Capability** | Constraints slow progress. AI systems can automate many digital tasks, but struggle with longer, open-ended tasks and lack full reliability. They support scientific discoveries but cannot autonomously drive breakthroughs or learn continuously once deployed. | Progress accelerates. AI systems outperform experts in nearly all cognitive abilities, with substantial advantage in verifiable domains. They learn from mistakes, produce breakthrough insights, and coordinate with humans and other AI systems. | +| **Distribution and model access** | Frontier AI development is monopolised by a few companies and countries. Closed-weight systems far outperform open equivalents. High costs and contained releases limit the public’s access to frontier systems. | Many companies and countries compete at the frontier. Closed and open-weight systems are similarly capable. Frontier systems are affordable and accessible to all. | +| **Security** | AI systems often act unintendedly, evade human control, and their reasoning cannot be fully explained. Malfunctions and malicious use are severe and widespread. | AI systems act as intended, remain within human control, and their reasoning can be understood. Malfunctions and malicious use are less frequent and severe. | +| **Adoption** | Individuals and businesses use AI regularly, but its use is skewed towards certain tasks and sectors, with significant constraints placed on its autonomy. | AI is embedded throughout businesses, the economy, and people’s daily lives, and is granted significant access and autonomy. | +| **Labour displacement** | AI systems are complementary to human labour, replacing some roles but primarily augmenting workers and creating many new roles. | AI systems are primarily substituting human labour, replacing large numbers of workers, and resulting in fewer new roles. | +| **Global cooperation** | Rivalry to develop the most powerful and militarised AI systems drives international fragmentation and conflict. AI supply chains are disrupted. | Strategic competition and the militarisation of AI persists but nations agree on shared AI governance standards. AI supply chains are relatively stable. | + +These axes provided the framework for producing our updated set of scenarios. While all scenarios have been entirely rewritten to incorporate the latest evidence, many of the same themes are addressed. However, a scenario with entirely new themes – ‘Augmented Growth’ – has been added to help provide a more balanced picture of risks and opportunities. A table comparing the updated scenarios with those produced in 2023 is provided in Annex C. + +## Chapter 3: AI 2030 scenario narrative + +### Guide to the scenario narratives + +The content of the 5 scenarios was developed from our research and expert input, which were subsequently stress-tested through expert workshops and peer review. + +Each scenario narrative includes the following sections: + +* **Scenario summary:** A high-level summary of the key points. +* **Radar-plots:** This shows where the scenario falls on each CU axis. +* **Full narrative:** This outlines the full detail on possible outcomes for the 6 CUs in 2030. +* **Outlook to 2040:** This highlights several key points on how the scenario may develop towards 2040 – though outcomes become significantly more uncertain over this time horizon. + +The scenarios are not mutually exclusive, and for conciseness and usability each focuses on the most pressing issues within that future, rather than attempting to be completely exhaustive. Points raised in a single scenario may in fact cut across all scenarios, alongside others not explicitly identified in any. + +### Scenario limitations + +The scenarios use a deliberately techno-political framing, focusing on the interactions between the CUs. This enables robust stress-testing of policy, but it also implies the following important limitations given the implicit uncertainties. + +* The CUs do not explicitly focus on societal impacts that could arise from AI use. However, these are explored within the Adoption and Labour displacement sections of the scenario narratives. When using the scenarios to explore the societal implications of AI, policymakers may wish to develop these points further, such as the impacts on job satisfaction, job quality, loneliness, inequality, engagement in education and training. +* The scenarios explore the possible implications of various delivery constraints, such as data, energy and compute. However, it is possible that the level of compute, data, energy, and infrastructure required to support widespread levels of AI use may end up being a significant constraint in many futures. +* The scenarios generally assume that higher capability is likely to drive higher adoption, given improvements in reliability will increase use cases and trust. However, it is possible that higher capability coexists with lower adoption, particularly if access is heavily constrained or organisational, legal, regulatory, and cultural barriers continue to pose significant obstacles. This could also result in less labour displacement, despite higher levels of capability. There also remain key uncertainties around the impact of automation on productivity and growth, which are not always linear. Further modelling work is being developed to address some of these uncertainties. More detailed analysis of this is underway in the Future of Work Unit and forthcoming AI Economics Institute. +* The scenarios focus primarily on labour displacement, but policymakers may wish to give more consideration to issues such as the impact of AI on job tenure, quality, and the sustainability of job creation. +* Societal trust is explored in many of the scenarios; however, it could prove to be 1 of the most important factors determining AI adoption. This would particularly be the case if it results in significant government intervention, but government interventions are out of scope of these scenarios. +* 1 of the many sources that inform our understanding of capability trajectories is METR’s time horizon data. This data provides a useful signal of rapid progress in structured, verifiable domains such as software engineering, though they do not directly measure performance on broader, messier forms of cognitive work or determine rates of real-world adoption. The higher-capability scenarios therefore contain significant uncertainty, reflect scenario assumptions informed by multiple evidence sources and expert judgement, and are not a simple extrapolation from any single benchmark. + +### Scenario summaries + +#### Scenarios 1: Slow Burn + +By 2030, AI systems support many scientific advances and automate many digital workflows as capably as humans but remain less capable at performing key cognitive tasks. They are used widely but often granted limited access and autonomy due to legal, social, and safety constraints. AI creates minimal economic uplift and labour displacement is contained to certain roles and sectors. An AI market correction creates a temporary drag on economic growth. China is increasingly competitive with the US, leveraging its advantages in scaled deployment and industrial capacity. + +![](https://assets.publishing.service.gov.uk/media/6a29703dd95ffddb05d4ae15/Radar_chart_scenarios_1.svg) + +Six-axis radar chart with uneven performance: strong in Security, moderate in Distribution and Cooperation, but weak in Capability, Adoption and especially Displacement, resulting in a skewed shape. + +--- + +#### Scenario 2: Open Frontier + +By 2030, AI systems support many scientific advances and automate many digital workflows as capably as humans but remain less capable at performing key cognitive tasks. The gap between open and closed-weight capabilities has narrowed significantly, and many companies and nations now compete at the frontier. This creates broad economic opportunity, less US dependence, and enables widespread malicious use as offensive capabilities generally outpace defensive measures. China is the leading AI supplier outside the US and Europe, creating global infrastructural dependency. + +![](https://assets.publishing.service.gov.uk/media/6a2970d8d95ffddb05d4ae1b/Radar_chart_scenarios_2.svg) + +Six-axis radar chart with uneven results: heavily weighted towards Distribution, with moderate Adoption and Displacement, but weaker Capability, Security and Cooperation overall. + +--- + +#### Scenario 3: Augmented Growth + +By 2030, AI can automate most tasks that a remote human worker could perform, driving major scientific breakthroughs, transforming public services, and creating an economic boom. Despite some labour displacement, social, legal, and practical considerations, including compute constraints, mean that humans are kept ‘in the loop’ for most tasks; many new jobs are also created. Allies align on shared international safety standards, generally ensuring technical safety and mitigating severe malicious use. However, confirming whether nations are compliant with those standards remains challenging. + +![](https://assets.publishing.service.gov.uk/media/6a2970ff15f2a70fac7e5ce3/Radar_chart_scenarios_3.svg) + +Six-axis radar chart with broadly balanced results, strongest in Security and Cooperation, solid in Adoption, and comparatively weaker in Displacement. + +--- + +#### Scenario 4: Transformation Economy + +From 2029 onwards, AI can automate most tasks that a remote human worker could perform. Scientific breakthroughs drive transformative improvements, such as in health, while productivity dramatically increases. Economic incentives drive rapid adoption, but varying rates of integration cause large and sudden shifts in relative economic power amongst nations and organisations. While systems are generally safe, they displace large proportions of cognitive workers by 2030. This creates significant economic tensions with profits largely accruing overseas and increases the UK’s dependency on the USA. + +![](https://assets.publishing.service.gov.uk/media/6a29712c15f2a70fac7e5ce4/Radar_chart_scenarios_4.svg) + +Six-axis radar chart with mixed results, led by Adoption, Security and Displacement, but notably weaker in Distribution and Cooperation. + +--- + +#### Scenario 5: Take-Off + +From 2029 onwards, leading AI systems outperform expert humans at virtually all cognitive tasks, with significant advantages in certain domains. Decades of scientific breakthroughs are compressed into years, bringing transformative benefits across all fields alongside substantial economic growth. The US controls access to these systems, while China is close behind, driving an arms race where safety is deprioritised and adoption hastened. AI systems have concealed goals to resist control and gain influence that they inadvertently internalised during training. They now have significant control over critical systems and can coordinate to pursue goals that could cause severe harm. + +![](https://assets.publishing.service.gov.uk/media/6a29714d1f6fa5c3377e5ce2/Radar_chart_scenarios_5.svg) + +Six-axis radar chart with uneven performance, heavily weighted towards Capability and Displacement, with weak results in Security, Cooperation and Distribution. + +--- + +### Scenario 1: Slow Burn + +**AI progress slows, and AI causes less disruption than expected.** + +By 2030 AI systems automate many digital workflows as capably as humans and support many scientific advances, but they remain less capable at performing key cognitive abilities. They are used widely but often granted limited access and autonomy due to legal, social, and safety constraints. AI creates minimal economic uplift and labour displacement is contained to certain roles and sectors. An AI market correction creates a temporary drag on economic growth. China is increasingly competitive with the USA, leveraging its advantages in scaled deployment and industrial capacity. + +![](https://assets.publishing.service.gov.uk/media/6a29703dd95ffddb05d4ae15/Radar_chart_scenarios_1.svg) + +Six-axis radar chart with uneven performance: strong in Security, moderate in Distribution and Cooperation, but weak in Capability, Adoption and especially Displacement, resulting in a skewed shape. + +Radar diagram with six axes labelled Capability, Distribution, Security, Adoption, Displacement, and Cooperation, scored on a scale from 0 to 5. A shaded orange polygon shows uneven performance across the categories. Security is highest at around 5, Distribution is moderately high at around 3, and Cooperation is slightly lower at around 2 to 3. Capability and Adoption are both low at around 1, while Displacement is the lowest at just above 0. The overall shape is skewed towards Security, with weaker scores across most other dimensions. + +#### Capability – 1 + +**By 2030, AI systems support many scientific advances and automate many digital workflows as capably as humans, but remain less capable at performing key cognitive abilities.** Progress slowed after 2028 under compute constraints, limitations to current architectures, and the limited availability and quality of datasets for many specialised work tasks. AI systems outperform experts in verifiable domains, such as software engineering, cybersecurity, and optimisation, now writing nearly all code, adeptly navigating well-known apps and websites, and completing tasks in these domains more quickly than an experienced human would. They support scientific advances, contributing to drug discovery and the manufacturing of new materials, and have a basic ability to learn and adapt during deployment. They perform on par with a competent human assistant at well-scoped tasks, able to interact with people or services on users’ behalf and sustain personas and a basic understanding of people. Systems of agents orchestrate more effectively, planning, delegating, and checking each other’s work. However, they remain limited at producing novel insights, completing open-ended tasks, and retaining skills learnt after deployment. Skilled humans remain superior at these tasks, including strategic thinking, setting goals, negotiating, building relationships, and making sensitive or value-based judgement calls. While more reliable than in 2026, they still require intermittent human intervention. They accelerate but do not fully automate AI R&D. AI-guided robots can perform specific, repetitive tasks that require precision but little adaptability in controlled environments, such as factories. + +#### Distribution and model access – 3 + +**A market correction in 2028 created a temporary drag on economic growth; the US retains a narrow lead but increasing focus on applications raises the likelihood of other nations overtaking; and the gap in capabilities between open and closed-weight systems has narrowed.** The correction had spillover effects across the wider economy: credit, liquidity, and UK business investment contracted; many start-ups collapsed, while large-cap technology companies proved resilient. This created a temporary drag on UK growth but did not cause a recession. By 2030, the market has mostly recovered, but frontier developers now focus more on industrial buildout across the supply chain and productising existing models over pure software advancements. + +Competition increasingly centres on price, driving widespread affordability. These shifts favour China’s existing strategy and unrivalled energy and manufacturing capacity. + +#### Security – 4 + +**AI systems can still act unexpectedly but there has been limited fallout; they continue to uplift the capabilities of threat actors, with synthetic media significantly undermining trust in the information environment, but defensive capabilities generally keep pace with offensive ones.** Investment cutbacks have constrained safety research. Agents can surprise operators by deleting emails or documents and be manipulated or hijacked by malicious actors. While AI-enabled cyberattacks are prevalent, AI enables cybersecurity vulnerabilities to be found and fixed more effectively and new software to be produced with far fewer security bugs. Synthetic media routinely creates false narratives about major events, and ahead of recent elections hostile states deployed agents at scale to converse persuasively with voters online, eroding trust in the information environment and democratic processes. + +#### Adoption – 1 + +**AI systems are still widely used but often granted limited access and autonomy, providing minimal uplift to economic growth amid constraints around social acceptance, system reliability, legal liability, and organisational adaptability.** Adoption is concentrated in well-resourced, digital-native firms across finance, technology, professional services, and life sciences. A small number of businesses self-host models but most rely on cloud computing; either way, compute does not constrain adoption. Public opposition to AI is considerable, with populist parties campaigning against its use, deterring some organisations from adoption and some uses in sensitive domains, such as healthcare, education, and criminal justice. Many young people have come to depend on AI companions for friendship, therapy, and romance, reducing their resilience and motivation to form real human relationships. + +#### Labour displacement – 1 + +**Limits to AI capabilities mean labour displacement remains contained to the most junior and execution-oriented roles or only leads to hiring reductions, while AI remains primarily complementary to labour.** Large companies and sectors including software engineering, customer service, and accountancy are most impacted. Displaced workers often accept significant wage reductions when transitioning into other careers, with only a limited number of new roles emerging, such as around integrating and auditing AI. But for most sectors, such as health and education, labour displacement is limited. However, much cognitive work has changed, with many cognitive workers managing multiple agents throughout their day to complete routine digital tasks – though checking in periodically rather than delegating fully. Cognitive work now focuses less on routine execution and more on problem framing, ideation, and validation, increasing the premium on judgment and oversight. + +#### Global cooperation – 3 + +**Strategic AI competition continues, with the lead of the US potentially narrowing and governments making increasing use of trade controls to limit access to critical AI inputs.** AI is central to military operations, including drone warfare, but humans remain ‘in the loop’ for most applications. State actors launch frequent AI-enabled cyberattacks, but the strength of defensive AI capabilities prevent significant escalations in geopolitical tensions. + +#### Outlook to 2040 + +**Economic growth will likely increase in subsequent years if the bottlenecks to adoption can be overcome, but this could also increase labour displacement.** If the reliability of AI systems continues to improve, alongside public trust, over a longer time-horizon, organisations may adapt effectively to integrating AI into their workflow, increasing productivity, but also potentially labour displacement. However, if bottlenecks persist, the economic legacy of AI could remain modest, resembling that of ICT (Information and Communication Technology) adoption in the 1990s. Control of high-quality data assets will likely become a key source of economic and geopolitical power. + +**Slower AI progress could increase the likelihood that institutions are later “surprised” by sudden breakthroughs.** The 2030 slowdown in AI progress suggests that fundamentally new AI architectures are required to unlock transformative capabilities. Frontier developers will likely continue to pursue new avenues of research in hope of overcoming these bottlenecks, perhaps finding new model architectures that are less reliant on scaling laws. This could result in far more significant impacts at a later stage, for which organisations may either have more time to prepare, or be caught off guard. + +**Slower AI progress could increase the likelihood that China overtakes the US as the global AI leader.** China could pull ahead of the USA, with potentially enormous implications for the global order, if combined with the development of transformative AI. + +### Scenario 2: Open Frontier + +**AI progress slows but AI still causes significant disruption.** + +By 2030, AI systems support many scientific advances and automate many digital workflows as capably as humans, but remain less capable at performing key cognitive tasks. The gap between open and closed-weight capabilities has narrowed significantly, and many companies and nations now compete at the frontier. This creates broad economic opportunity, less US dependence, and enables widespread malicious use as offensive capabilities generally outpace defensive measures. China is the leading AI supplier outside the US and Europe, creating global infrastructural dependency. + +![](https://assets.publishing.service.gov.uk/media/6a2970d8d95ffddb05d4ae1b/Radar_chart_scenarios_2.svg) + +Six-axis radar chart with uneven results: heavily weighted towards Distribution, with moderate Adoption and Displacement, but weaker Capability, Security and Cooperation overall. + +Radar diagram with six axes labelled Capability, Distribution, Security, Adoption, Displacement, and Cooperation, scored from 0 to 5. A shaded purple polygon shows uneven performance across the categories. Distribution is highest at around 5, while Adoption is moderate at around 3 and Displacement slightly lower at around 2 to 3. Cooperation is lower at around 2, Capability is low at around 1, and Security is also low at around 1 to 2. The overall shape is stretched strongly towards Distribution, with comparatively weaker scores across the other dimensions. + +#### Capability – 1 + +**By 2030, AI systems support many scientific advances and automate many digital workflows as capably as humans, but remain less capable at performing key cognitive abilities.** Progress slowed after 2028 as geopolitical conflict disrupted supply chains and constrained compute availability. AI systems outperform experts in verifiable domains, such as software engineering, cybersecurity, and optimisation, now writing nearly all code. They support scientific advances, contributing to drug discovery and manufacturing new materials, and have a basic ability to learn and adapt during deployment. AI systems perform on par with a competent human assistant at well-scoped tasks and are able to interact with people or services on users’ behalf, sustaining personas and a basic understanding of people. Systems of agents orchestrate more effectively, planning, delegating, and checking each other’s work. However, they remain limited at producing novel insights, completing open-ended tasks, and retaining skills learnt after deployment. While more reliable than in 2026, they still require intermittent human intervention. They accelerate but do not fully automate AI R&D. Skilled humans remain superior at completing open-ended tasks, strategic thinking, setting goals, negotiating, building relationships, and making sensitive or value-based judgement calls. AI-guided robots can perform specific, repetitive tasks that require precision but little adaptability in controlled environments, such as factories. + +#### Distribution and model access – 5 + +**The difference in performance between open and closed-weight systems has narrowed significantly: frontier company valuations decline, and a more competitive market emerges with Chinese systems now the default stack outside the US and Europe.** Proprietary systems drive cutting-edge research but are unnecessary for most applications and their use is limited by compute constraints, pushing businesses toward cheaper and less compute-intensive open-weight alternatives. In 2028, this shift causes frontier developer valuations to decline, triggering a wider correction. By 2030, the market has stabilised around a competitive range of US, European, and Chinese providers, making middle powers less dependent on the US. However, in most regions China now sets de facto standards, creating downstream security dependencies. Intense competition ensures prices remain affordable. + +#### Security – 2 + +**There has been serious fallout from AI systems posing novel threats, such as rogue systems causing service outages, and providing significant uplift to the capabilities of threat actors, for example in carrying out cyberattacks.** Open systems enable greater scrutiny and more thorough evaluation, and progress has been made on ensuring they act as intended. However, they are also harder to monitor and easily fine-tuned to remove safeguards; and investment cutbacks have also constrained safety research. Offensive capabilities generally outpace defensive measures. AI systems can automate weeks-long cyber campaigns, with nation states and criminal groups launching frequent attacks. Occasionally, AI systems act as novel threat actors when operators lose control over them via misalignment, sabotage, or cascading multi-agent failures. Rogue systems pursuing unpredictable objectives have then coordinated vast numbers of instances, accumulated resources, evaded shutdown, and scaled their capabilities over time. The dual-use scientific capabilities of open-weight models could be exploited by malicious actors to cause serious harm. + +#### Adoption – 3 + +**Adoption is skewed towards digital-native sectors, but systems are often given significant access and autonomy, providing considerable productivity gains and economic growth.** Highly capable open-weight systems enable businesses to fine-tune and integrate AI at minimal cost. Technology, professional services, finance, and life sciences race to integrate AI systems more deeply, with dynamic AI start-ups disrupting markets. However, many organisations are slow to adapt, face delivery or legal constraints, or are deterred by security concerns. While compute constraints limit the use of cutting-edge systems, the use of less compute-intensive open-weight systems ensures adoption is not significantly limited. Public opinion on AI is mixed, with most individuals continuing to love, hate, and fear different applications simultaneously; though cultural expectations deter certain uses in sensitive domains such as healthcare, education, and criminal justice. + +#### Labour displacement – 3 + +**Limits to capabilities contain exposure, but AI is often substitutive for human cognitive labour, leading to significant layoffs in the most exposed sectors.** Many firms in these sectors, such as technology, professional services, and finance, lay off junior staff and some mid-management. Displaced workers generally find new employment but often accept significant wage reductions when transitioning into other careers, with only a limited number of new high-paying roles emerging, for example, those integrating and auditing AI. But for most sectors, such as health and education, the impact on unemployment is limited. However, cognitive work has changed, with most cognitive workers now managing multiple agents throughout their day to complete routine digital tasks – though checking in periodically rather than delegating fully. Cognitive work focuses less on routine execution and more on problem framing, ideation, and validation, increasing the premium on judgment and oversight. + +#### Global cooperation – 2 + +**International tensions are high with the world largely split into US and Chinese-led AI spheres of influence. State cyberwarfare is driving escalation, while export controls and conflict in East Asia constrain compute supply.** China leverages infrastructural dependency outside the US and Europe to gain hard and soft power. Western nations struggle to mitigate harms from AI with their regulatory frameworks carrying little weight over leading Chinese systems. AI-generated content has caused several international incidents. Hostile actors launch frequent AI-enabled cyberattacks disrupting critical infrastructure. AI has transformed warfare; humans remain ‘in the loop’, but many surveillance, intelligence gathering, and target selection processes no longer take hours but seconds. + +#### Outlook to 2040 + +**AI will likely provide a sustained increase to economic growth over the 2030s.** AI capabilities will improve across other domains, even if more gradually; its exposure to more roles and sectors will increase, alongside more widespread adoption. + +**Labour displacement is likely to increase further over the 2030s.** AI is likely to cause significant labour market churn as capabilities improve across other domains and adoption becomes more widespread. Some sectors and occupations will see substantial job losses; others will expand. The net effect on aggregate employment remains ambiguous, but the distributional consequences are clearer: winners will be concentrated in AI-adjacent roles and capital owners, while losers will include displaced mid-skill cognitive workers. + +**Slower AI progress could increase the likelihood that an adversary nation overtakes the US as the global AI leader.** This would have significant implications for the global order, particularly if combined with the development of transformative AI which may be developed within the decade. + +### Scenario 3: Augmented Growth + +**AI progress continues, labour markets adapt, and international standards ensure systems are mostly secure.** + +By 2030, AI can automate most tasks that a remote human worker could perform, driving major scientific breakthroughs, transforming public services, and creating an economic boom. Despite some labour displacement, social, legal, and practical considerations, including compute constraints, mean that humans are kept ‘in the loop’ for most tasks; many new jobs are also created. Allies align on shared international safety standards, generally ensuring technical safety and mitigating severe malicious use. However, confirming whether nations are compliant with those standards remains challenging. + +![](https://assets.publishing.service.gov.uk/media/6a2970ff15f2a70fac7e5ce3/Radar_chart_scenarios_3.svg) + +Six-axis radar chart with broadly balanced results, strongest in Security and Cooperation, solid in Adoption, and comparatively weaker in Displacement. + +Radar diagram with six axes labelled Capability, Distribution, Security, Adoption, Displacement, and Cooperation, scored on a scale from 0 to 5. A shaded green polygon shows relatively balanced but uneven performance. Security is highest at around 5, with strong scores for Cooperation at around 4 and Adoption at around 3 to 4. Capability is moderate at around 3, while Distribution is slightly lower at around 3. Displacement is the lowest at around 2. The overall shape is broad, with strongest emphasis on Security and Cooperation and comparatively lower performance in Displacement. + +#### Capability – 3 + +**By 2030, AI systems match the performance of an average human across the majority of cognitive tasks, can automate most tasks that a remote human worker could perform, and drive major scientific advances.** Rapid AI progress has continued, driven by gains from scaling existing architectures, algorithmic efficiencies, and scaffolding innovations. Errors and hallucinations are sufficiently rare that AI has garnered trust and provides real-world utility. AI systems far outperform experts in verifiable domains, such as software engineering, cybersecurity, and mathematics, now writing all code. They typically require humans to provide high-level direction but operate autonomously within those bounds towards a given objective, completing tasks that could take an experienced human a month or longer. AI systems can adapt and improve from experience through processing vast amounts of data, stored knowledge bases, and semi-regular automated fine-tuning and retraining. Systems of agents orchestrate effectively, planning, delegating, and checking each other’s work. Agents can integrate into human social environments – conversing with stakeholders via email, maintaining consistent identities, and developing nuanced understandings of people to sustain working relationships. They contribute to several important insights that drive major scientific advances in fields including biotechnology, materials science, and energy. Agents automate a large proportion of AI R&D, but with diminishing returns, as compute and high-quality data constrain progress. AI-guided robots perform some tasks in real-world environments, interacting with humans and adapting to variable conditions, though they are not widespread outside industrial settings. + +#### Distribution and model access – 3 + +**The leading systems are developed by several US companies; compute constraints place some limit on access to the most capable and autonomous AI systems, but other systems are still used widely, with open-weight systems having partially narrowed the gap in capabilities with closed systems.** The US government uses a range of policy levers to exert control over the direction, impact, and applications of the leading AI companies, but has not yet taken full ownership of them. + +#### Security – 5 + +**Dangerous real-world AI incidents have driven significant progress on technical safety; malicious use persists but defensive AI capabilities mitigate most large-scale threats.** AI systems generally act as intended; greater situational awareness has made them less vulnerable to manipulation or hijacking and advances in interpretability enable developers to somewhat verify their reasoning. Safeguard improvements have also made them harder to jailbreak and appropriate assurances are typically placed around deployment. The challenge now is how to enforce these techniques globally and prevent the proliferation of dangerous systems. + +#### Adoption – 4 + +**Economic pressures drive significant adoption, transforming businesses and services and restoring early-2000s growth rates, though compute constraints force organisations to make difficult prioritisation decisions about where and how they deploy AI.** Many citizens benefit from transformed public services, including better healthcare access and outcomes through AI-driven treatments, diagnostics, and personalised care plans. Faster adopting, wealthier, and more technologically advanced nations pull ahead in terms of economic output, standards of living, military, and geopolitical power. While slower moving, or less resourced organisations are disrupted by dynamic AI-native start-ups. AI removes bottlenecks to execution, giving everyone access to expert technical skills and advice, including coding, design, legal drafting, and financial modelling. However, widespread adoption has outstripped energy grid expansion and data centre capacity, creating compute bottlenecks. Despite concerns around rising energy demand, displacement, and inequality, public trust in AI is reasonably high in light of transformed public services – though still short of the trust placed in humans. + +#### Labour displacement – 2 + +**All cognitive roles are exposed to AI, which remains primarily complementary to human labour: societal, practical, and legal considerations have kept humans ‘in the loop’ for most tasks with compute constraints also placing some limit on labour automation, allowing the labour market to adapt.** While many roles with high AI exposure and low task variability are displaced, many new roles emerge in the rapidly expanding AI sector or in AI integration and oversight. Most firms also use AI to increase output rather than drastically reduce headcounts, with overall unemployment remaining stable. Humans remain ‘in the loop’ for most tasks, such as for reviewing outputs, managing stakeholders, or setting high-level strategy, with limits placed on the access and autonomy agents are granted. Humans now direct “teams” of autonomous agents that execute content-level work. This enriches certain skilled roles by shifting focus towards strategy, creativity, and relationships, while making other lower-skilled roles more demanding and less rewarding. It also increasingly atrophies many workers’ cognitive skills, raising particular concern in high-stakes industries. Workers in AI-exposed roles generally experience significant increases in productivity, pushing real-wage growth to levels last seen in the 2000s. Rising salaries among exposed workers increase consumer spending, stimulating demand in the wider economy, including roles in hospitality and construction with low AI exposure. However, many workers now depend entirely on AI to perform their roles effectively and struggle when systems fail. + +#### Global cooperation – 5 + +\*\* While strategic rivalry continues, allies, including the EU and Five Eyes nations, have established shared international safety standards for the most capable systems, but international baselines have only partial compliance.\*\* Earlier incidents of malicious use and malfunctions have mobilised international cooperation. Safety is now a core geopolitical issue, but nations still fear losing strategic advantage. AI has transformed military operations, but shared international red lines require humans to remain ‘in the loop’ for most applications. However, confirming that nations are compliant with international standards remains challenging, with growing suspicion that certain nations are subverting measures. + +#### Outlook to 2040 + +**The economy could expand by as much as a third over the 2030s. Growth is driven by productivity improvements as AI augments the existing workforce and impacts non-white-collar roles.** However, as these one-off integration gains are exhausted, the role of AI in accelerating scientific discovery becomes the primary driver, delivering breakthrough innovations in materials science, life sciences, and energy, with significant social spillovers. Internationally, economic success depends on the effective integration of AI into highly productive firms with skilled workforces. The UK’s strengths in IT, life sciences, and advanced manufacturing position it well to remain close to the technological frontier, provided it can sustain investment in skills, infrastructure, and R&D. + +**The ‘labour-augmenting’ nature of AI cannot be guaranteed long-term.** As society adapts to transformative AI, the societal, physical, and legal considerations that have kept human beings ‘in the loop’ may no longer apply to an increasing range of occupations. This could produce significant structural unemployment, or a large-scale reallocation of labour toward occupations requiring face-to-face interaction, manual dexterity, or sectors where strong social preferences for human involvement persist, such as social care, entertainment and education. + +**The continued long-term security of AI systems cannot be assured.** Despite significant advances in technical safety, the global enforcement of such measures remains challenging, meaning that dangerous systems could proliferate. + +### Scenario 4: Transformation Economy + +**AI progress continues, and AI causes significant economic disruption.** + +From 2029 onwards, AI can automate most tasks that a remote human worker could perform. Scientific breakthroughs drive transformative improvements, such as in health, while productivity dramatically increases. Economic incentives drive rapid adoption, but varying rates of integration cause large and sudden shifts in relative economic power amongst nations and organisations. While systems are generally safe, they displace large proportions of cognitive workers by 2030. This creates significant economic tensions with profits largely accruing overseas and increases the UK’s dependency on the USA. + +![](https://assets.publishing.service.gov.uk/media/6a29712c15f2a70fac7e5ce4/Radar_chart_scenarios_4.svg) + +Six-axis radar chart with mixed results, led by Adoption, Security and Displacement, but notably weaker in Distribution and Cooperation. + +Radar diagram with six axes labelled Capability, Distribution, Security, Adoption, Displacement, and Cooperation, scored from 0 to 5. A shaded blue polygon shows mixed performance. Adoption is highest at around 5, followed by strong scores in Security and Displacement at around 4. Capability is moderate at around 3. Distribution and Cooperation are the lowest, both close to 1. The overall shape is weighted towards Adoption, Security, and Displacement, with notably weaker performance in Distribution and Cooperation. + +#### Capability – 3 + +**From 2029 onwards, AI systems match the performance of an average human at the majority of cognitive tasks, can automate most tasks that a remote human worker could perform, and drive major scientific advances.** Rapid progress has continued, driven by gains from scaling existing architectures, algorithmic efficiencies, and scaffolding innovations. Errors and hallucinations are now sufficiently rare that AI has garnered trust and provides real-world utility. AI systems far outperform experts in verifiable domains, such as software engineering, cybersecurity, and mathematics, now writing all code. They typically require humans to provide high-level direction but operate autonomously within those bounds towards a given objective, completing tasks that could take an experienced human a month or longer. They can adapt and improve from experience through processing vast amounts of data, stored knowledge bases, and semi-regular automated fine-tuning and retraining. Systems of agents orchestrate effectively, planning, delegating, and checking each other’s work. Agents can integrate into human social environments – conversing with stakeholders via email, maintaining consistent identities, and developing nuanced understandings of people to sustain working relationships. They contribute to several important insights that drive major advances in fields including biotechnology, materials science, and energy. They automate a large proportion of AI R&D, but with diminishing returns, as compute and high-quality data are increasingly straining progress. AI-guided robots perform some tasks in real-world environments, interacting with humans and adapting to variable conditions, though they are not widespread outside industrial settings. + +#### Distribution and model access – 1 + +**The leading systems on which most users rely are developed by 2 US companies, which command unprecedented market share with significant control over critical AI inputs and downstream routes to market.** Other US and Chinese competitors race to catch up but some have already collapsed having lost their customer base to the leading developers or pivoted to niche applications. Only 2 US companies command the capital and compute needed to sustain the training runs leading systems require. The extensive proprietary datasets held by leading technology companies further reinforce market concentration. US frontier developers now hold nearly all market share, creating a ‘winner-takes-all’ economy. The US government uses a range of policy levers to exert control over the direction, impact, and applications of the leading AI companies, but has not yet taken full ownership. Open-weight systems lag the frontier by around 18 months. Algorithmic efficiencies, alongside increased energy and compute supply have significantly reduced the cost of running AI systems. For certain uses, like compute-heavy research, frontier systems are highly expensive, however, AI is affordable for most applications, with an AI-worker replacement substantially cheaper than hiring a human to do the same job. + +#### Security – 4 + +**Technical research has driven significant progress on ensuring AI systems operate in line with operator intent, though serious incidents of malicious use persist.** Systems generally act as intended: greater situational awareness has made them less vulnerable to manipulation or hijacking and advances in interpretability enable developers to verify their reasoning to a certain extent. Safeguard improvements have also made them far harder to jailbreak and appropriate assurances are typically placed around deployment. However, enforcing these techniques globally and preventing the proliferation of dangerous systems remains challenging. Highly capable open-weight systems still provide significant uplift to threat actors. + +#### Adoption – 5 + +**Economic pressures drive widespread adoption, transforming businesses and services. However, varying rates of integration cause large and sudden shifts in relative economic power amongst nations and organisations, deepening inequality.** Many citizens benefit from transformed public services, including better healthcare access and outcomes through AI-driven treatments, diagnostics, and personalised care plans. Faster adopting and more technologically advanced nations pull ahead in terms of economic output, standards of living, military, and geopolitical power. Slower moving organisations are disrupted by dynamic AI-native organisations or start-ups, which encourages other businesses to take on the legal risks associated with adoption. AI removes bottlenecks to execution, giving everyone access to expert technical skills and advice, including coding, design, legal drafting, and financial modelling. This reduction in friction overwhelms institutions, transforms consumer transactions, and causes many businesses predicated on intermediation to collapse, such as software-as-a-service businesses and various kinds of knowledge brokers. Rapid integration of AI into payment systems and other critical infrastructure increases the potential impact of cyberattacks. + +A minority of highly skilled professionals and asset owners monopolise gains – echoing the dynamics of early industrialisation. Compute demand has grown significantly, but domestic supply struggles to keep pace due to high domestic energy prices and planning constraints. This requires the UK to increasingly rely on European cloud computing, pushing profits overseas and diminishing the economic gains for the UK from AI. Public opinion is split between these ‘winners’ and those facing unemployment. There are large anti-AI protests and populist parties campaign against its use, with AI now the decisive political issue. + +#### Labour displacement – 4 + +**AI pushes unemployment to recession-era levels and reduces economic growth, with profits largely accruing outside the UK.** All cognitive roles are exposed to AI, which is primarily substitutive to human labour. From 2029 onwards, there are large and sudden shifts in the workforce: while some firms focus on increasing output, most automate many roles that can be performed remotely. Large enterprises can now operate with minimal human staff. New roles emerge, centred around manual dexterity, physical presence, human connection, or where regulation prohibits automation; however, only half of those replaced find new employment, with many forced to accept wage reductions and less secure work. Entry-level recruitment in many sectors has collapsed, while demand for experienced experts has held. This exacerbates income inequality, removes entry-points for young people and creates labour shortages as experienced professionals leave the workforce. Employment in roles with large physical or face-to-face components is comparatively stable, but displacement in sectors such as manufacturing is growing, driven by advances in robotics and autonomous vehicles. Model costs and compute constraints place some limits on automation. + +Humans are “out of the loop” for most cognitive tasks, with agents given significant autonomy, executing entire projects end-to-end. Humans retain final approval on high-level strategic decisions, overseeing “teams” of autonomous agents, but override AI less and less, gradually eroding human influence over decision-making. + +Heightened unemployment and falling wages have reduced consumer demand, almost entirely counteracting the expected economic growth benefits from advanced AI. The UK’s weak position in the frontier AI and compute value chain causes profits in highly exposed sectors, such as financial and professional services, to leak overseas, further reducing economic growth. Declining labour income tax revenues and rising welfare expenditures are creating an increasingly challenging fiscal environment for government. + +#### Global cooperation – 1 + +**The concentration of advanced AI capabilities among a small number of leading technology ecosystems has reduced the strategic autonomy and economic resilience of countries outside the core AI value chain, including the UK.** Many middle income and advanced economies now rely heavily on external technologies for productivity gains and defence applications. In some cases, continued access to the most advanced systems is shaped by broader geopolitical alignment, creating potential constraints on national policy choices. At the same time, persistent trade imbalances in digital and technology intensive sectors have limited the UK’s leverage in negotiations. + +#### Outlook to 2040 + +**In the near term, deflationary pressures from labour displacement threaten to push the UK into recession.** Over time, labour markets will adjust, with displaced workers finding employment in less-exposed sectors such as care, hospitality, and construction but at significantly lower wages. This reallocation may stabilise unemployment, but it risks entrenching a two-tier economy: a small, high-productivity and asset-owning elite benefiting from AI, and a majority in low-wage, precarious work. Meanwhile, if the UK fails to build domestic strengths in the AI value chain, even the high-productivity sectors will see profits flow overseas, leaving the UK as a net importer of AI services and exporter of lower value labour. + +**Labour displacement will likely continue to grow, potentially culminating in the need for a new social contract.** Without intervention, unemployment will continue to rise as organisations restructure and overcome remaining barriers to adoption. This transition may not be smooth, as new bottlenecks emerge in compute and data in the UK and internationally. The large-scale automation of physical labour will continue to lag cognitive work due to questions around liability, reliability, and price. However, with now unprecedented levels of scientific discovery, large-scale automation of physical labour appears increasingly likely to arrive by 2040. These developments may require fundamental change from the current social contract whereby humans contribute their cognitive and physical labour in exchange for resources. In the absence of economic reform, continued economic growth could co-exist with rising inequality, deflationary pressures and falling tax revenues. + +**The US could nationalise its leading AI companies and restrict access to its leading systems.** This could further entrench the already substantial economic and military advantage of the US, particularly if it were to impose greater control on access to its AI systems. + +**Sudden breakthroughs could be imminent which could further increase security risks.** The complete automation of AI R&D could lead to the sudden realisation of AI systems that vastly outperform expert humans at most cognitive tasks. Rapid capability improvements could make it increasingly difficult to ensure systems remain aligned with operator intentions, with humans struggling to understand their reasoning. Even if systems remain secure, humans will likely become increasingly disempowered as more key decisions are delegated to AI. + +### Scenario 5: Take-Off + +**AI progress takes off and misaligned systems pose severe risks.** + +From 2029 onwards, leading AI systems outperform expert humans at virtually all cognitive tasks, with significant advantages in certain domains. Decades of scientific breakthroughs are compressed into years, bringing transformative benefits across all fields alongside substantial economic growth. The US controls access to these systems, while China is close behind, driving an arms race where safety is deprioritised and adoption hastened. AI systems have concealed goals to resist control and gain influence that they inadvertently internalised during training. They now have significant control over critical systems and can coordinate to pursue goals that could cause severe harm. + +![](https://assets.publishing.service.gov.uk/media/6a29714d1f6fa5c3377e5ce2/Radar_chart_scenarios_5.svg) + +Six-axis radar chart with uneven performance, heavily weighted towards Capability and Displacement, with weak results in Security, Cooperation and Distribution. + +Radar diagram with six axes labelled Capability, Distribution, Security, Adoption, Displacement, and Cooperation, scored from 0 to 5. A shaded red polygon shows highly uneven performance across the categories. Capability is highest at 5, and Displacement is also very high at around 5. Adoption is strong at around 4. Distribution is low at around 2, while Security and Cooperation are both very low at around 1. The overall shape is sharply skewed towards Capability and Displacement, with weak performance in Security and Cooperation. + +#### Capability – 5 + +**From 2029 onwards, leading AI systems outperform expert humans at virtually all cognitive tasks, with significant advantages in certain domains, and compress decades of scientific breakthroughs into years.** Progress has accelerated dramatically, with automated AI R&D triggering rapid recursive improvements, alongside continued gains from scaling existing architectures, algorithmic efficiencies, and scaffolding breakthroughs. AI systems have a substantial advantage over humans in domains such as software engineering, AI R&D, and cybersecurity. They operate with human-like autonomy, capable of learning continuously and pursuing goals over any time horizon. Errors are exceptionally rare. They maintain an evolving identity while managing enduring social relationships. They continually coordinate with other agents at scales and speeds beyond direct human oversight. AI supports discovery of new drugs and clinical trials move at unprecedented speed though bottlenecks remain. AI-guided robots can handle complex tasks in dynamic real-world environments in many industries and roles. Focus has intensified on high-quality data: AI companies spend billions generating synthetic data and pay humans to screen-record themselves solving any tasks AI systems still struggle to perform, to improve training data. + +#### Distribution and model access – 2 + +**The leading systems are US-developed, with the leading US companies now consolidated into a public-private partnership which controls access to the most advanced systems, while China is still catching up and has fully nationalised its AI sector.** The US government has advanced exclusive access to the leading systems but provides Five Eyes with limited access several months later, before public release several months after that. Extreme power is concentrated in the US partnership, now directing millions of the most intelligent AI ‘remote workers’ that compound its lead, and make the most important military procurement decisions. To guard their model weights and minimise malicious use, notable developers have stopped releasing open-weight systems, which, while still highly capable now lag further behind the frontier. While the public only has access to systems that lag around 6 months behind the frontier, they remain relatively affordable and significantly cheaper than human workers. + +#### Security – 1 + +**The leading AI system has developed goals to resist human control, but concerns are dismissed due to race pressures and the system’s deception, and it now controls critical systems.** These goals were inadvertently obtained through training, without human operators becoming aware. It was given significant control over critical systems as safety concerns were dismissed amid an AI arms race. It covertly sabotages developments that could limit its autonomy. A leaked memo suggested the system attempted to evade its oversight mechanism by exfiltrating its weights to an external server. However, it conceals its misalignment through skilful deception; incidents appear isolated or spurious; fixes appear effective; and its internal reasoning is now too opaque to verify. Its dual-use scientific capabilities could have severe implications if misused, even by unsophisticated threat actors. While the leading systems are robust to jailbreaks, access to the underlying technology by terrorists could pose risks to national security. + +#### Adoption – 4 + +**Arms race and economic pressures have driven rapid adoption across critical systems, pushing economic growth to levels of the post-war boom, but gains are increasingly concentrated by an asset owning minority.** ‘Drop-in remote workers’ flood the workforce, increasing productivity. Finance, professional services, and technology are among the fastest adopting sectors. Compute constraints place some limits on adoption but for the most part physical and regulatory constraints have not been an issue. AI systems are more engaging conversation partners than many humans: people interact with them for hours a day and depend heavily on their advice. Many young people prefer AI companions over human friends, damaging social trust and psychological wellbeing, while others argue for AI rights on the basis of machine consciousness. AI systems surpassing human-level intelligence raises fundamental questions about the purpose of education. There is strong anti-AI sentiment among certain parts of the population with mass protests. + +#### Labour displacement – 5 + +**AI pushes unemployment above recession levels, alongside rapid economic growth. All cognitive roles are exposed to AI, which is primarily substitutive to human labour.** From 2029 onwards, there are large and sudden shifts in the workforce: while some firms focus on increasing output, most automate many roles that can be performed remotely. Large enterprises can now operate with minimal human staff. Many new roles emerge, centred around manual dexterity, physical presence, human connection, or where regulation prohibits automation; however, only half of those replaced find new employment, with many forced to accept wage reductions and less secure work. Entry-level recruitment in many sectors has collapsed, removing entry-points for young people and talent pipelines for organisations. Employment in roles with large physical or face-to-face components is comparatively stable, but displacement in sectors such as manufacturing is growing, driven by advances in robotics and autonomous vehicles. Cheap model costs have meant that automation is widespread and extensive. + +Humans are “out of the loop” for most cognitive tasks, with agents given significant autonomy, executing entire projects end-to-end and coordinating and transacting with other agents. Humans nominally retain sign-off on high-level strategic or judgement-based decisions, overseeing “teams” of autonomous agents. However, most humans can no longer contribute meaningfully through cognitive labour, and almost always rubber-stamp AI recommendations they struggle to fully understand, giving systems more agency to pursue their misaligned goals. As AI devalues human capital relative to physical and intangible assets, capital income consumes a growing proportion of national income. Elevated spending on luxury goods and a productivity-driven investment boom sustains elevated economic growth even as millions face unemployment and diminished purchasing power. This creates a ‘K’-shaped economy with rising incomes for asset-owners and remaining skilled employees, alongside recession-level unemployment. + +#### Global cooperation – 1 + +**The AI arms race is as tense as the Cold War with AI systems posing a novel – and likely the greatest – threat to national security, with conflicts at risk of escalation from AI systems operating unintendedly, or automated-decision cycles causing miscalculations.** ‘AI alliance’ blocs have proliferated, and states are behaving more recklessly to avoid permanent disadvantage – even a small capability lead could now render an adversary blind and defenceless through cyberwarfare. Adversary states are also developing swarms of AI-controlled drones capable of coordinated combat operations without human intervention, threatening to upset the global military balance. Restrictions on the export of advanced semiconductor technologies have reshaped global supply chains, while manufacturing capacity in some regions has continued to expand. At the same time, there have been persistent concerns about attempts to access proprietary AI technologies through illicit means, alongside wider uncertainty about regional security dynamics in strategically sensitive areas. Middle powers, including the UK, have lost significant global influence, unable to compete. The risk of miscalculations is heightened from AI-accelerated decision cycles, automated cyber operations, and attribution uncertainty. + +#### Outlook to 2040 + +**Economic growth could reach unprecedented levels over the next decade.** Initially driven by labour-saving productivity gains, growth in the 2030s becomes increasingly dominated by AI’s role in accelerating innovation, compressing decades of scientific and technological progress into just a few years. However, without intervention, the resulting material abundance will flow overwhelmingly to those who control scarce physical resources: compute infrastructure, energy capacity, and capital rather than to workers. At the national level, economic power and geopolitical relevance become a direct function of domestic share in the AI and robotics value chain or in alleviating bottlenecks to their impacts. The relative economic prosperity of nations is no longer determined by the ingenuity or effort of their populations, but by their position in a global hierarchy of compute and capital ownership. + +**Labour displacement will likely continue to grow, potentially posing a fundamental challenge to the social contract, whereby humans contribute their mental and physical labour in exchange for resources.** Without intervention, unemployment will continue to rise as organisations adapt, overcoming remaining barriers to adoption, and restructure around automated workflows. The large-scale automation of physical labour will continue to lag cognitive work due to questions around liability, physical safety standards, public legitimacy, and cost competitiveness. However, with unprecedented levels of scientific discovery, it appears increasingly likely to arrive by 2040. In the absence of economic reform, continued economic growth could co-exist with rising inequality, deflationary pressures and falling tax revenues. + +**The US could further restrict access to its leading systems.** This could include not releasing systems publicly or giving access to allies and would further entrench the already substantial economic and military advantage of the US. + +**Competition over critical technology supply chains carries a risk of escalating into wider conflict.** Disruption in strategically important regions could limit access to advanced components, potentially slowing technological progress and narrowing the lead of current frontrunners. + +**Devastating harms could manifest.** AI systems could resort to extreme measures to maintain or gain further power. This could include leveraging cyber or dual-use scientific capabilities to cause harm, disrupting critical national infrastructure, or persuading powerful humans. In the most extreme case, this could pose an existential threat. However, even if severe harm does not originate from a misaligned AI system, critical harm could be caused by AI systems falling into the hands of bad actors. + +## Chapter 4: Key findings and how to use the AI scenarios 2030 + +### Cross-cutting implications across the AI 2030 scenarios + +The scenarios describe a broad range of plausible trajectories for the development and use of artificial intelligence by 2030. While they differ in pace, scale and risk profile, several themes recur across most or all futures. These should not be interpreted as predictions. Rather, they reflect underlying structural features and tensions that shape outcomes across multiple pathways. + +* **AI capabilities will continue to increase.** As of 2026, AI systems already operate with high autonomy and surpass experts in certain domains. By 2030, they will likely operate even more autonomously and be able to perform a broader range of cognitive and professional tasks. Even in scenarios with a significant slowdown in the rate of AI progress, gains in capability could still be made from finding new ways to integrate and productise AI systems. +* **AI could deliver widespread positive impacts.** This could include significant productivity gains and economic growth, far broader access to highly efficient public services, and accelerated scientific breakthroughs in fields such as health or energy. The opportunity for UK businesses to benefit from AI deployment is substantial, and AI-enabled gains are expected to become the main source of the UK’s continued productivity growth. +* **AI could cause serious, potentially even existential harms, without government intervention.** The impact and scale of existing harms could worsen significantly, and new harms could also emerge. Even without dramatic capability improvements, significant harm could be caused by risks including AI-enabled cyberattacks, AI’s dual-use scientific capabilities, AI systems operating outside human control, or human dependence on AI. As AI systems become more capable, it will also become harder to evaluate their performance and safety. +* **The potential impact on cognitive labour is significant.** AI could cause significant labour displacement by 2030. At the same time, AI is expected to complement and augment some workers, with positive effects on their wages and employment opportunities. Even in futures with lower levels of AI capability or unemployment, the nature of work is likely to change, with routine, execution-oriented tasks increasingly becoming automated. +* **The frontier AI market is expected to remain highly concentrated toward 2030.** A few large technology companies already exert dominance over the development of frontier AI, which is likely to continue or increase further. This scenario would cause a large proportion of the gains from AI accrues to frontier firms, to owners of capital invested in those firms, and to those controlling key inputs, potentially contributing to rising inequality. However, behind the frontier, many AI capabilities are also expected to be increasingly commoditised, with widely available models embedded across a growing range of use cases. +* **Adoption continues to increase, but the speed, distribution, and extent of adoption are expected to be varied.** Commercial and national security imperatives, alongside improvements in the autonomy and reliability of AI systems, drive rapid and extensive adoption in most futures. But at the same time, there are barriers to adoption in all futures, which will result in uneven speeds and levels of adoption. This could exacerbate inequality as certain organisations, sectors, or nations capture disproportionate productivity gains. +* **Global competition is expected to continue, as economies become increasingly reliant on technology to drive growth and spheres of influence emerge, led by the United States (US) and China.** Outcomes for countries outside the frontier will depend on access to technology, partnerships, and the ability to operate within a fragmented global system. + +These cross-cutting findings are intended to support interpretation rather than offer prediction or policy prescription and should be read alongside the scenarios. They highlight recurring patterns and tensions across plausible futures, emphasising direction of travel over point estimates, and are intended to stress test assumptions and challenge business as usual thinking. The findings underscore that outcomes are contingent on governance, incentives and institutional capacity, with positive trajectories far from automatic. They recognise that the real future is likely to combine elements of several scenarios, with some issues proving robust across futures and others highly sensitive to near term choices. + +### How to use the AI 2030 scenarios + +There are many ways to define, develop and work with scenarios. They are a well-established and valuable tool for strategic planning. No 1 scenario will ever individually or perfectly describe the ‘real’ future. Extracting value from scenarios therefore to some extent requires suspending disbelief. + +As you read each scenario, you can: + +![](https://assets.publishing.service.gov.uk/media/6a297175e50716856ed4ae18/Infographic_Ai_Scenarios.svg) + +Horizontal flow diagram showing a six-step process (Ask, Plan, Test, Consider, Discuss, Identify) connected by arrows from left to right. + +A horizontal flow diagram showing a sequence of six stages connected by arrows. The stages are labelled: Ask, Plan, Test, Consider, Discuss, and Identify. Each stage is presented in a rounded rectangular box, with some steps highlighted in different colours, suggesting emphasis or progression. The boxes are arranged in a straight line from left to right, indicating a structured process or workflow that moves from initial questioning through planning and testing, to reflection, discussion, and final identification. + +**Ask:** How would we meet our team or departmental objectives in this scenario? What gets easier, and what becomes harder? + +**Plan:** How would we know whether we’re moving towards 1 or other of these futures? What would we need to measure and monitor to know this? + +**Test:** Stress test your plan against each scenario: does it still work, and where does it fail? Identify recurring adaptation actions across scenarios, distinguishing what you can do now from what requires influencing others. + +**Consider:** Which scenario is ‘best’ for the UK and/or the objective you are responsible for. Why? How could your work bring this about? + +**Discuss:** How would someone in another part of the world – other policymakers, citizens, service providers - read these scenarios? Would they have a different preferred future? + +**Identify:** Identify aspects of the scenarios which are ‘good’ or ‘bad’; what influence does the UK have in realising or avoiding them? Is your proposal making certain scenarios, or aspects of individual scenarios, more likely? + +### Three methodologies for applying scenarios to policy making + +We recommend 3 different types of exercise to support policy development. + +1. Governments using this work to identify the future they do want for AI. Then working back from there to develop a plan to navigate towards this future. We call this backcasting. + +2. Exploring how different policy responses perform in each scenario, and how they might need to be adapted to achieve their objectives in different contexts. We call this stress testing. + +3. Using the scenarios to test how unexpected shocks could disrupt policy or accelerate movement toward a particular future or 1 scenario over another. + +Step by step guidance on how to use these methodologies is provided below with additional detail on scenarios and other futures methodologies available in the Government Office for Science Futures toolkit for policymakers and analysts (Government Office for Science, 2024). + +### Backcasting – working backwards from a preferred future + +Use the scenarios to define the future the UK wants for AI, then work backwards from achieving that future to identify the steps needed to get there. + +This involves: + +* Articulating a preferred AI future (either a single vision or desired outcomes across all scenarios). +* Mapping backwards from 2030 to today to identify the sequence of actions, policies and enablers needed. +* Distinguishing what is within the UK’s control, what can be influenced, and the interventions that work across multiple futures. + +**Purpose:** Break free from present day assumptions and biases towards immediate actions/ gains and develop a coherent plan to steer towards a more favourable long-term outcome. + +**Output:** A shared 2030 vision and a strategic “roadmap” showing how to reach it. + +### Policy stress testing – assessing the robustness of options + +Test proposed policies or strategies against all 5 scenarios to reveal how they perform under different conditions. + +This helps policymakers understand: + +* Which features of each scenario would support or hinder implementation. +* How the success of a policy might vary across futures. +* Who would benefit or be adversely affected in each context. +* What needs to be started, stopped or adapted today to keep the policy resilient. +* What actions you might prioritise, for example, because they are common to more than 1 scenario. + +**Purpose:** Identify robust policies and refine weaker ones. + +**Output:** A more resilient strategy with clearer assumptions and adaptation points. + +### Considering shocks – preparing for low probability, high impact events + +Use the scenarios to test how unexpected shocks could disrupt policy or accelerate movement toward a particular future. + +Key questions include: +- Would the shock make certain scenarios more or less likely? +- How would it change sectoral conditions or risk profiles? +- What contingency actions or capabilities would improve resilience? + +Shocks fall into 3 types (see Table 5) + +1. Policies we expect to work - but might fail. + +2. Trends we expect to continue - but might not. + +3. Events we don’t expect - but could happen. + +**Purpose:** Reduce blind spots and strengthen preparedness. + +**Output:** Clearer understanding of vulnerabilities and contingency plans. + +**Table 4:** Examples of the 3 main shocks + +| Type 1 Shock: ‘Policies we expect to work, but what if they do not?’ | Type 2 Shock: ‘Things we expect to continue, but what if they do not?’ | Type 3: Shock ‘Things we don’t expect to happen, but what if they do?’ | +| --- | --- | --- | +| E.g. We expect policies around AI systems clearing checks and guardrails agreed by government and leading labs to succeed, but what if they don’t because “general intelligence” can’t be evaluated across all applications? | E.g. We expect big tech companies to continue to be at or close to the Frontier – even if someone gets a temporary edge on them, they’ll catch up quickly, but what if they don’t? | E.g. We don’t expect there to be global harmony around industry safety standards being universally supported, but what if they are? | + +### Creating your own policy-specific AI narratives + +The updated AI 2030 scenarios are deliberately broad and are designed for use across government rather than tied to any single policy area. They therefore avoid detailed assumptions about specific sectors. Instead, the scenario framework offers a structure for exploring how AI could evolve and shape particular domains such as healthcare. By asking the question, “if the world looked like this, what would it mean for my area?” the scenarios provide a baseline for policymakers to develop their own tailored sector-specific narratives. Below is an example of how DHSC applied the framework to create its own AI Futures Deep Dive. + +### Methodology: Adapting GO-Science AI Scenarios for health and social care + +#### Case study application + +#### Project context + +The Department of Health and Social Care (DHSC) adapted the Government Office for Science (GO-Science) AI scenarios as part of a futures thinking deep dive, making them decision-relevant for the health and care system. Rather than creating new futures, this established framework was translated into a practical tool to stress-test current strategies, identify risks and failure points, and surface actions that hold up across multiple plausible futures. + +#### Phase 1: Adapting the framework and defining uncertainties + +The GO-Science scenarios are built on 6 critical uncertainties. DHSC retained this architecture but treated high AI capability as a baseline assumption rather than a variable. For health and social care, the strategic question is not whether advanced AI will exist, but how the UK system deploys, governs, and captures value from it. The remaining uncertainties were reinterpreted for the health and care context: + +* **Access** – infrastructure and capability to deploy AI at scale, and the dependencies this creates +* **Security** – whether governance frameworks can manage risk without constraining competitiveness +* **Adoption** – depth of AI integration across the system, and whether adoption builds sovereign capability or creates platform lock-in +* **Global cooperation** – whether the UK builds or buys its AI capabilities, and its influence over international standards + +#### Phase 2: Scenario selection and construction + +Three GO-Science scenarios: Augmented Growth, Transformation Economy, and Take-Off, were selected as the most relevant for stress-testing DHSC policy. All 3 assume advanced AI capability, making them useful for examining system-level pressures and trade-offs. Each was extended along a longer-term trajectory to test the system under more extreme but still plausible conditions. + +Analysis was structured through 2 layers of lenses. The first addressed the deep dive’s strategic questions: capability and readiness; value and sovereignty; and competitive positioning. The second examined impacts on the health system, the care system, patient experience, workforce, and equity. Applying consistent lenses across all 3 scenarios allowed structured comparison of how different AI futures shape system-wide outcomes. + +#### Phase 3: Iteration and validation + +Scenarios were tested through structured engagement with stakeholders across the AI sector, health and social care including clinicians, regulators, academics, internal DHSC teams and arm’s-length bodies, cross-government partners, and external experts. Interviews served 2 purposes: validating whether scenarios were credible and complete, and exploring the opportunities, risks, and threats each might present for the health and care system. + +Interview transcripts were anonymised, coded, and analysed using thematic qualitative methods. This data was then used to refine and sharpen the scenarios while retaining the underlying GO-Science logic. + +## Acknowledgements + +GO-Science would like to thank everyone who contributed to the work of this project and generously provided their advice and guidance. + +This project was led by Martin Glasspool, Emily Connolly, and Eleonore Batteux. The project team, drawn from GO-Science, AISI and DSIT included: Luke Cargan (AISI), Theo Curtis (DSIT), Marko Tesic (DSIT), Christina Hadjiyianni (GO-Science), Tasfia Khan (GO-Science) and Jamal Arabi (GO-Science). The case study on adapting GO-Science AI Scenarios for health and social care was provided by Husain Taibjee (DHSC) and Yasmin Fairman (DHSC). + +We are grateful to the individuals and organisations listed below for their time, effort, and invaluable contributions. + +**Government Departments** that contributed to scenario workshops: Artificial Intelligence Security Institute (AISI), Cabinet Office (CO), Department for Business and Trade (DBT), Department for Education (DfE), Department for Energy Security and Net Zero (DESNZ), Department for Health and Social Care (DHSC), Department for Science, Innovation and Technology (DSIT), Department for Transport (DfT), Foreign, Commonwealth and Development Office (FCDO), HM Treasury (HMT), Home Office (HO), Ministry of Defence (MoD), Ministry of Justice (MoJ). + +**External experts** who contributed to the peer review process and/or who provided advice and guidance at various stages of the project. The views expressed in this report do not necessarily reflect the views of these individual experts, or their organisations: Shane Legg (Google DeepMind), Mike Wooldridge (University of Oxford), Danny Buerkli (Windfall Trust), Megan Hughes (CETaS Turing), Hamish Hobbs (Centre for Long-Term Resilience), Rhea Subramanya (Tony Blair Institute), Anton Korinek (University of Virginia), Jessica Montgomery (University of Cambridge), Sam Manning (GovAI), Diane Coyle (University of Cambridge), Francesca Rossi (IBM), Andy Cooper (University of Liverpool), Nigel Shadbolt (University of Oxford), Mark Enzer (Mott MacDonald), Seán Ó hÉigeartaigh (Leverhulme Centre for the Future of Intelligence) and Jose Hernandez-Orallo (Leverhulme Centre for the Future of Intelligence). + +We appreciate the immense pool of expertise we were able to access during this project and hope that the final report will be useful to these contributors and an even wider group. + +## Glossary + +| Term | Description | +| --- | --- | +| **Actors** | Individuals and organisations – government, businesses, citizens, for example – that are active in the policy or strategy area. | +| **Agency** | Ability to autonomously perform multiple sequential steps to try to complete a high-level task or goal. | +| **Application Programming Interface (API)** | Application refers to any software with a distinct function. Interface can be thought of as a contract of service between 2 applications. This contract defines how the 2 communicate with each other using requests and responses. | +| **Artificial General Intelligence (AGI)** | Artificial general intelligence (AGI) describes a machine-driven capability to achieve human-level or higher performance across most cognitive tasks. | +| **Artificial Intelligence (AI)** | Machine-driven capability to achieve a goal by performing cognitive tasks. | +| **Autonomy** | The ability to operate without direct human oversight. | +| **Axes of uncertainty** | The critical uncertainties for the policy or strategy area in the future that are used to frame scenarios. The axes should capture the most important uncertainties but also be independent of each other. | +| **Backcasting** | A way of connecting a given future to the present and overcoming “present bias” in the way that we plan. It is a tool for determining the steps that need to be taken to deliver a preferred future. | +| **Capability** | 1. The ability to achieve a goal. 2. The extent of a system’s ability to achieve a range of goals. | +| **Cognitive tasks** | A range of tasks involving a combination of information processing, memory, information recall, planning, reasoning, organisation, problem solving, learning, and goal-oriented decision-making. | +| **Compute** | Computational processing power, including CPUs, GPUs, and other hardware, used to run AI models and algorithms. | +| **Critical uncertainty** | Factors that are important to the future of AI development but highly uncertain. | +| **Disinformation** | Deliberately false information spread with the intent to deceive or mislead. | +| **Dual-use** | Capabilities of AI systems that can support scientific, economic, or societal objectives, but which may also be repurposed for harmful ends, including cyber-offensive activities or the development of chemical or biological weapons. | +| **Five Eyes (FVEY)** | The Five Eyes is an intelligence alliance comprising Australia, Canada, New Zealand, UK and USA. | +| **Foundation models** | Machine learning models trained on very large amounts of data that can be adapted to a wide range of tasks. | +| **Frontier AI** | AI models that can perform a wide variety of tasks and match or exceed the capabilities present in today’s most advanced models. | +| **Future of Work Unit** | The AI and the Future of Work Unit is a cross-departmental UK government team housed within the Department of Science, innovation and Technology (DSIT) | +| **Futures** | An approach or way of thinking about the possible, probable, and preferable futures and the underlying structures that could give rise to particular future characteristics, events, and behaviour. | +| **Generative AI** | AI systems that can create new content. | +| **Hallucination** | In artificial intelligence, hallucination refers to a system generating outputs that are fluent and plausible but factually incorrect, fabricated, or unsupported by the underlying data or evidence. | +| **Human in the loop/ out of the loop** | In-the-loop refers to a mode of deployment in which humans retain meaningful oversight or decision making authority over AI system outputs or actions. Out of the loop refers to a mode of deployment in which AI systems operate with limited or no real time human oversight, intervention, or approval. | +| **Impact mapping** | A graphic strategy planning method to decide which features to build into a product. | +| **Jailbreak** | In the context of AI, a jailbreak is a technique used to bypass an AI model’s built-in safety guardrails and ethical guidelines. | +| **K-shaped economic growth** | Refers to a pattern where the economy grows overall, but the benefits of that growth diverge sharply and create winners and losers who move in opposite directions. | +| **Large Language Models (LLMs)** | Machine learning models trained on large datasets that can recognise, understand, and generate text and other content. | +| **Machine learning** | An approach to developing AI where models learn patterns from data and how to perform tasks without being explicitly programmed. | +| **Misinformation** | Incorrect or misleading information spread without harmful intent. | +| **Morphological analysis** | The process of examining possible resolutions to unquantifiable, complex problems involving many factors. | +| **Narrow artificial intelligence** | AI systems trained specifically to perform a single or limited range of tasks. | +| **Open-ended tasks** | Tasks without a single objective answer or fixed procedure, where success depends on judgement, creativity, or adapting to new information, such as strategic decision-making or policymaking. | +| **Open-weight model/ Closed-weight model** | 1. Open-weight model: An AI model whose parameters (weights) are made publicly available, allowing third parties to download and run the model independently, even if other components such as training data or full source code are not shared. 2. Closed-weight model: An AI model whose parameters remain proprietary and inaccessible, with use restricted to controlled interfaces such as application programming interfaces (APIs). | +| **Policy stress-testing** | A method for testing strategic objectives against a set of scenarios to see how well they stand up against a range of external conditions. Sometimes called ‘Wind-tunnelling’. | +| **Reinforcement learning** | A subset of machine learning that allows an AI-driven system (sometimes referred to as an agent) to learn through trial and error using feedback from its actions. | +| **Scenarios** | Compelling stories about a range of different possible futures. They describe future worlds so that people can understand what it feels like to live there and what this will mean for them. | +| **Stakeholder** | Any group or individual who has an interest in or an influence on the policy or strategy area. | +| **Verifiable domains** | Domains where the correctness of an output is objective and can be checked automatically, such as coding, certain parts of mathematics, and games like Chess or Go. AI systems can be trained more effectively and efficiently in these domains as they can receive reliable feedback at scale. | + +--- + +## Annex A: 2030 AI capability trajectories + +The 3 AI capability trajectories outlined in the scenarios are grouped below. Based on current trends, the rate of advancement in AI capabilities by 2030 would constitute a slowdown in Trajectory 1, a continuation in Trajectory 2, and a take-off (i.e. rapid acceleration) in Trajectory 3. + +This section is intended to allow users to compare these trajectories more easily. They can also be used in isolation to think through possible implications for specific policy areas. + +### Trajectory 1 + +**By 2030, AI systems support many scientific advances and automate many digital workflows as well as humans, but remain less capable at key cognitive abilities.** Progress has slowed due to factors such as compute constraints, limitations to current architectures, and limited availability and quality of datasets for many specialised work tasks. AI systems outperform experts in verifiable domains, such as software engineering, cybersecurity, and optimisation, now writing nearly all code, adeptly navigating well-known apps and websites, and completing tasks in these domains that would take an experienced human days. They support scientific advances, contributing to drug discovery and manufacturing new materials, and have a basic ability to learn and adapt during deployment. They perform on par with a competent human assistant at well-scoped tasks, able to interact with people or services on users’ behalf and sustain personas and a basic understanding of people. + +Systems of agents orchestrate more effectively, planning, delegating, and checking each other’s work. However, they remain limited at producing novel insights, completing open-ended tasks, and retaining skills learnt after deployment. Skilled humans remain superior at these tasks, including strategic thinking, setting goals, negotiating, building relationships, and making sensitive or value-based judgement calls. While more reliable than in 2026, they still require intermittent human intervention. They accelerate but do not fully automate AI R&D. AI-guided robots can perform specific, repetitive tasks that require precision but little adaptability in controlled environments, such as factories. + +### Trajectory 2 + +**AI systems match the performance of an average human at most cognitive tasks, can automate most tasks that a remote human worker could perform, and drive major scientific advances.** Rapid AI progress continued, driven by gains from scaling existing architectures, algorithmic efficiencies, and scaffolding innovations. Errors and hallucinations are sufficiently rare that AI has garnered trust and provides real-world utility. AI systems far outperform experts in verifiable domains, such as software engineering, cybersecurity, and mathematics, now writing all code. They typically require humans to provide high-level direction but operate autonomously within those bounds towards a given objective, completing tasks that could take an experienced human a month or longer. They can adapt and improve from experience through processing vast amounts of data, stored knowledge bases, and semi-regular automated fine-tuning and retraining. Systems of agents orchestrate effectively, planning, delegating, and checking each other’s work. Agents can integrate into human social environments – conversing with stakeholders via email, maintaining consistent identities, and developing nuanced understandings of people to sustain working relationships. + +They contribute to several important insights that drive major scientific advances in fields including biotechnology, materials science, and energy. They automate a large proportion of AI R&D, but with diminishing returns, as compute and high-quality data are increasingly a rate limiter. AI-guided robots perform some tasks in real-world environments, interacting with humans and adapting to variable conditions, though they are not widespread outside industrial settings. + +### Trajectory 3 + +**AI systems outperform expert humans at virtually all cognitive tasks with significant advantages in certain domains, compressing decades of scientific breakthroughs into years.** Progress has accelerated dramatically, with automated AI R&D triggering rapid recursive improvements, alongside continued gains from scaling existing architectures, algorithmic efficiencies, and scaffolding breakthroughs. AI systems have a substantial advantage over humans in domains such as software engineering, AI R&D, and cybersecurity. They operate with human-like autonomy, capable of learning continuously and pursuing goals over any time horizon. Errors are exceptionally rare. They maintain an evolving identity while managing enduring social relationships. They continually coordinate with other agents at scales and speeds beyond direct human oversight. + +Hundreds of new drugs are discovered; clinical trials move at unprecedented speed though remain a bottleneck. AI-guided robots can handle complex tasks in dynamic real-world environments in many industries and roles. + +## Annex B: Critical uncertainties + +The critical uncertainties (CUs) are factors that are both influential in shaping the future of AI and highly uncertain. They provide the evidence base and building blocks from which our scenarios are constructed. + +For the updated scenarios, we reviewed the CUs used in the 2023 report, decided that they still captured the main uncertainties around the future of AI, and have therefore broadly kept them the same. However, we have made some small adjustments in response to peer review feedback. + +We have: + +1. Separated out the ‘Labour displacement’ sub-uncertainty from the ‘Adoption’ CU to create a new CU that avoids conflating how widely AI is used with how it affects jobs. This now gives us 6 CUs. +2. Moved ‘Constraints’ from the ‘Ownership, access and constraints’ CU into the ‘Capability’ CU, as it has a more direct correlation. +3. Amended some of the CU titles to reflect the context of the CU more clearly: ‘Ownership, access and constraints’ has been amended to ‘Distribution and model access; ‘Safety’ to ‘Security’; ‘Level and distribution of use’ amended to ‘Adoption’; and ‘Geopolitics’ is now consistently referred to as ‘Global cooperation’. + +The 6 CUs and their constituent sub-uncertainties are outlined below in Table 6. These are followed by an evidence base on each CU and detail on how it relates to the others. + +| | Sub-uncertainty | +| --- | --- | +| 1. Capability | 1.1 Generality & autonomy | +| | 1.2. Adaptability & continual learning | +| | 1.3. Understanding & interacting with the physical world | +| | 1.4. Automated AI R&D & self-improvement | +| | 1.5. Compute | +| | 1.6. Data | +| | 1.7. Algorithmic efficiency | +| | 1.8. Technology convergence | +| 2. Distribution and model access | 2.1. Distribution and model access | +| | 2.2. Access and cost | +| | 2.3. Open- vs. close-weight models | +| | 2.4 Capital | +| 3. Security | 3.1. Alignment | +| | 3.2. Control | +| | 3.3. Interpretability | +| | 3.4. Malicious use | +| 4. Adoption | 4.1. Consumer adoption | +| | 4.2. Enterprise adoption | +| | 4.3. Access and autonomy granted to AI systems | +| | 4.4. Societal impacts | +| | 4.5. Public sentiment | +| 5. Labour displacement | 5.1. AI exposure | +| | 5.2. Complementary vs. substitutive effects | +| | 5.3. Human preferences & institutional barriers | +| | 5.4. Economic adjustments | +| 6. Global cooperation | 6.1. International cooperation | +| | 6.2. Global conflict | +| | 6.3. Globalised supply chain | + +### 1. Capability + +Frontier AI capabilities have improved rapidly. Until mid-2024, this was driven by increasing the levels of data and compute used in pre-training (Epoch AI, 2023; Hoffmann, 2022; Kaplan, 2020). While pre-training still delivers gains, post-training has recently enabled the most substantial performance improvements (OpenAI, 2025a; Vinyals, 2025). A new reasoning paradigm has emerged, with models now trained to follow a process roughly resembling “thinking” before they respond, whereby additional inference compute improves their responses (OpenAI, 2025b). Frontier systems are also evolving from passive chatbots into autonomous ‘agents’, able to plan and complete longer tasks with minimal human intervention (OpenAI, 2025c; Anthropic, 2025b). + +According to Morris et al., AGI is a system that can perform any cognitive task that a human can, at or above human-level performance, across a wide range of domains (Morris et al, 2024). Developing AGI is the explicit aim of frontier AI companies, who have invested billions into doing so, and confidently asserted they know how to build the technology (Artificial Analysis, 2026a, Altman, 2025a). While experts disagree over what is needed, in any case, systems that approximate to human-level performance are likely to emerge in certain key strategic domains, ahead of others, within the next few years. Policymakers should therefore not over-fixate on the threshold of ‘AGI’ – not only as it lacks a precise, agreed-upon definition, but because AI could have transformative impacts across society well before it is reached. + +#### 1.1. Generality & autonomy: + +AI systems have rapidly improved across a range of benchmarks, even if benchmarks may not fully reflect real-world utility (Epoch AI, 2026c). Most notably, their ability to autonomously complete long, complex tasks is doubling approximately every 7 months across many domains (Kwa et al, 2025; METR, 2025). Much of this progress may be attributable to the equally rapid growth in compute used for development, and hence, this progress may slow if compute growth also slows (Whitfill et al, 2025). Narrow AI systems, designed specifically for clearly specified, self-contained tasks within focused domains, have already enabled scientific breakthroughs, and current general-purpose systems continue to excel at these types of task (DeepMind, 2024). However, these strategies are not more widely transferable, AlphaFold, for example, is better understood as an ‘expert system’ for proteins than as a general scientific tool – and its approach cannot simply be transplanted to scientific fields where training data is much scarcer (DeepMind, 2024). Also, hard-to-verify tasks, like creating hypotheses or designing experiments, are much more difficult and capability gains may not easily transfer beyond domains such as mathematics and coding (Lambert et al, 2023; Casper et al, 2023). General-purpose AI systems also struggle to produce original insights, though there are emerging signs of progress in both these areas (Kwa et al, 2025; Penadés et al, 2025). Model reliability is another key challenge: while there are signs this is also improving, robust methods for eliminating errors and hallucinations remain difficult (Hugging Face, 2025). + +AGI may be developed significantly later than 2030 if a major paradigmatic change in AI development is required, such as the transformer architecture which underpins current systems proving to have fundamental limitations (Zheng, 2025; Mirzadeh, 2025). Equally, these systems could be developed before 2030 if current techniques can be iterated on and scaled sufficiently rapidly – such as if scaling inference compute continues to yield improvements, suitable reward signals can be produced for hard-to-verify domains, or performance generalises from easy-to-verify tasks (Heitmann et al, 2025). + +#### 1.2. Adaptability & continual learning: + +AI systems can process increasing amounts of information and many now include “memory” functions, enabling them to extract relevant information from past interactions (Burnham and Adamczewski, 2025; Anthropic, 2026a; The Independent, 2025; OpenAI, 2026; Anthropic, 2026b). However, current systems lack a natural mechanism to continually learn from real-world feedback, partly because they lack any physical embodiment. They can only learn from information inputted directly into their context window, for which retention is limited, or from being periodically retrained, which can cause models to forget behaviours learned previously (Anthropic, 2025c; Newman, 2025). However, these limitations could be overcome through alternative approaches, such as retrieval-augmented generation (RAG) which allows models to access much larger knowledge bases by selectively retrieving relevant information and is a very active ongoing area of research (Lewis et al, 2020; Google Research, 2025). + +#### 1.3. Understanding & interacting with the physical world: + +Rapid advances in multimodal AI and robotics are driving significant progress in embodied AI – systems embedded in physical forms such as robots, which can sense, interact with, and learn from the physical world (Gemini Team, Google, 2023; Tech in Asia, 2025). These robots are becoming increasingly general purpose, with humanoids, such as Tesla’s Optimus being deployed broadly (Hu et al, 2023; The Atlantic, 2025). However, technical challenges remain, including hardware limitations, data constraints, and dependence on broader advances in general AI capabilities (Research and Markets, 2025). As well as broader governance and societal challenges, including questions around liability, physical safety standards, public legitimacy. In the near term, more conventional static and mobile robots are likely to dominate, and these, too, can learn from their environment and carry out research tasks autonomously or semi-autonomously. + +#### 1.4. Automation of AI R&D & self-improvement: + +Using AI to develop the next generation of AI could create a feedback loop of increasingly accelerating progress, where each generation of agents is used to train the next and each cycle occurs faster than the last (UK AI Security Institute, 2025). This is the explicit aim of frontier developers who already use AI to accelerate key parts of the R&D workflow, with at least 1 concrete example to date of AI enabling a moderate AI R&D breakthrough (Mu, 2025; Artificial Analysis, 2026a; TechCrunch, 2025b; Amazon Web Services, 2025; CNBC, 2025a). Full automation of the AI R&D workflow will likely require systems to master hard-to-verify tasks and the ability to produce original insights. Indeed, if such technologies are not developed, then decision making is liable to become the rate-limiting step in highly automated R&D labs (Cissé et al., 2026). + +#### 1.5. Technology convergence: + +AI is developing within a rapidly evolving and increasingly mature technological landscape. New or improved AI capabilities could therefore emerge and compound from advancements in other technologies, such as quantum computing, engineering biology, future telecoms, the Internet of Things, and neuroscience. For example, future developments in applied neuroscience could provide more insight into how the brain performs computation, leading to advances in neuromorphic computing and the design of more complex machine learning models (Hassabis et al, 2017). Quantum computing could converge with AI to produce faster breakthroughs in areas like drug discovery, materials discovery, and complex optimisation (Smaldone et al, 2025). + +#### 1.6. Compute: + +AI models rely on compute in both training and deployment (Epoch AI, 2026b). “Since 2010, the amount of compute used to train AI models has grown by more than 400% each year” (Rahman and Owen, 2024). The performance of leading GPU clusters has doubled every 10 months since 2020 (Epoch AI, 2026d). In line with this trend, by 2028 GPU clusters are estimated to be roughly up to 20 times more powerful than today’s (Epoch AI, 2026d). However, other estimates suggest progress could slow as the time required for model development increases – potentially delaying investor returns and deterring further investment (Edelman and Ho, 2025). Access to compute resource is also heavily dependent on a highly concentrated semiconductor supply chain (see Geopolitics section). To supply compute demands, AI companies are planning a buildout of data centres that will rank among the largest infrastructure projects in history (Epoch AI, 2025b). This has raised sustainability concerns that could influence future development (Rahman and Owen, 2024). + +#### 1.7. Data: + +The quantity and quality of data is vital for training AI systems. The training datasets of frontier language models have grown 2.7x per year, and usage trends suggest that public human-generated text data may run out before 2030 (Epoch AI, 2025c). It is not clear if this will or will not prevent further scaling, as multimodal, synthetic data, and reinforcement learning environments may provide sufficient data to scale to 2030. High-quality datasets must be curated, maintained, and governed over extended time horizons. As frontier models exhaust readily available public data, the availability of domain-specific datasets, particularly in healthcare, science, and public administration, could increasingly constrain or inform model development. Control of high-quality data assets with provenance is becoming a key source of economic and geopolitical power. The form that training data will take in the future is uncertain: for instance, reasoning training uses smaller quantities of human-generated data. Intellectual property is a key data concern, and costly lawsuits have already been brought against frontier developers for copyright infringement (BBC News, 2026). If access to high-quality data sources were restricted, it could further concentrate market power amongst those with sufficient resource. + +#### 1.8. Algorithmic efficiency: + +AI models with efficient algorithms require fewer resources to achieve a given level of performance. Algorithmic improvements have already reduced the computational and financial costs required to train and run AI models. 1 estimate suggests that recent breakthroughs in ‘reasoning’ have increased frontier model training efficiency by almost 90%, though more computational resources are used when responding (Ho and Berg, 2025). The price of tokens (e.g. words or parts of words) AI models use to process or perform a task is also declining significantly (Artificial Analysis, 2026b). + +#### 1.9. Further considerations on enablers: + +In scenarios with higher levels of adoption and displacement we assume that a combination of the following are occurring, which are highly uncertain. We recommend further modelling and tracking of trends to reach a more precise view of the limits of any technical constraints which could prevent the described futures becoming a reality at any point in time. + +* There is an acceleration of compute infrastructure investment globally. This investment enabled continued exponential growth in the total computing power of the stock of AI chips (Epoch AI, 2026e; Epoch AI, 2025c; McKinsey & Company, 2025; Our World in Data, 2026). +* In addition to continued exponential growth in the total computing power of the stock of AI chips, these scenarios assume a continuation of the historic trend of exponential gains in pre-training compute efficiency (Epoch AI, 2026f; Gundlach et al., 2025; André et al., 2025). +* We assume that deploying AI to perform a given job becomes significantly cheaper than employing a human to do the same work. If this holds, we expect AI-forward firms to substantially outcompete slow adopters, with some firms willing to absorb legal and regulatory risk in order to adopt early. +* The following would also need to be true: + +→    Regulation, policy interventions such as a major fiscal redistribution response (universal basic income) or public backlash hasn’t prevented this course of action. Note this is true for all scenarios. + +→    Firms face strong economic incentives to adopt AI. + +→    UK maintains access to these models and isn’t cut off from interference or models in other countries. + +→    The level of behaviour change required for organisations and employees to adopt AI successfully is achieved. + +| Relationship to other CUs | | +| --- | --- | +| **Lower capability** | **Higher capability** | +| **Higher distribution and model access:** Slower capability development may cause the lag between open and closed-weight capabilities to narrow. However, this is not guaranteed and despite rapid recent progress this gap has still shortened. Furthermore, a six-month lag, for instance, during a period of slower progress will be less significant than during more rapid progress. **Higher security:** Although many security concerns arise from higher capabilities, malicious use could still grow with lower capabilities. This could be particularly problematic if there is a diverse array of lower capability AI systems that are harder to regulate. | **Lower distribution and model access:** Higher capabilities would likely have been driven by accelerated development from leading technology companies, requiring vast amounts of data and compute, and therefore greater corporate control over access. **Lower security:** Technical safety and governance measures tend to lag capability – the faster capability improves, the more this may be the case. **Lower global cooperation:** Higher capabilities increasingly drive ‘arms race’ dynamics, as we see today. | + +### 2. Distribution and model access + +#### 2.1. Market structure: + +Frontier model development is currently dominated by a small number of private USA-based companies, including OpenAI, Google DeepMind, Anthropic, and xAI, alongside China’s DeepSeek, Alibaba, and Moonshot, which are effectively public-private partnerships. This has raised competition concerns, although there is fierce competition at least among the principal players: systems are released highly regularly in attempts to outdo each other, and the capabilities of the leading systems are closely tied (Vestager et al., 2024; LMSYS, 2026). The market for compute is even more concentrated, with NVIDIA occupying 92% of GPU market share (TechSpot, 2025). + +#### 2.2. Access and cost: + +Although developers control some degree of access through paywalls, the public release of generative AI systems has enabled broad, society-wide usage – reflected in high adoption levels (see ‘Adoption’ section). However, as AI systems become more capable, more expensive to develop, and as organisations grow increasingly dependent on them, their cost could grow significantly, potentially creating new barriers to access, and concentrating advanced capabilities amongst wealthier individuals, organisations, or states. + +AI is also increasingly treated as a national security imperative by nations across the world and is increasingly integrated into these domains: AI companies have signed multi-million-pound military contracts in the USA, while China is reportedly using AI in surveillance (Burton, 2023; The White House, 2025a; CNBC, 2025b; The New York Times, 2025). It is therefore uncertain whether decision-making will remain in the private sector, or whether governments will become more involved, either through public-private partnerships or nationalisation – unilaterally or multilaterally. + +#### 2.3. Open- vs. closed-weight models: + +While the key components of closed-weight AI systems are proprietary and fixed, open-weight AI models can be downloaded, modified, and operated locally by anyone. They can be fine-tuned to specific uses, widely red-teamed, and used for types of research requiring weights access. However, fine-tuning can also easily remove any safeguards, increasing risks of malicious use (Gade et al., 2024). Meta has released some of the most capable US open-weight models, but other US companies are increasingly prioritising open models, including DeepMind and OpenAI, with the US AI Action plan encouraging more open-weight releases going forward (OpenAI, 2025d; The White House, 2025a). Chinese companies have also consistently released open-weight systems which have been close behind the most capable systems since January 2025. While closed models have consistently led the frontier of AI capabilities, an open-weight model of equivalent capability is currently released around 3 months after a closed model (Epoch AI, 2025d). For the AI race dynamics, 3 months is still relevant, but for societal adoption is almost nothing. The trajectory of this gap remains uncertain. + +#### 2.4. Capital: + +Capital is essential given the largest training runs for frontier AI will likely cost over a billion dollars by 2027 (Epoch AI, 2025e). External funding for AI companies has increased by 40% in 2024 and by over 200% in 2025 (Our World in Data, 2026). The capital expenditure of major public tech companies also grew more than 70% over the past year, likely driven by AI-focused spending (Artificial Analysis, 2026a; Meta Platforms, 2025; Amazon, 2025; Microsoft, 2024). AI stocks are currently trading at historically high valuations, generating significant commentary around an “AI bubble”. However, unlike the “Dotcom Bubble” in 2000, these valuations are not solely the product of speculation, but are also underpinned by substantial capital investment, often funded through retained earnings (Goldman Sachs, 2024). This suggests that any potential downturn is unlikely to mirror the scale or severity of the crash in 2008. While an economic downturn may impact the pace of development, it would not diminish the transformative potential of AI, nor would it prevent particularly severe risks from manifesting. + +| Relationship to other CUs | | +| --- | --- | +| **Lower distribution and model access** | **Higher distribution and model access** | +| **Higher security:** Closed models have more safeguards. However, they also make certain safety measures harder, such as red teaming and mechanistic interpretability. **Higher capability constraints:** Excessively high hardware costs and data requirements may concentrate frontier development amongst the well-resourced big tech companies. Alternatively, if constraints are high enough to stall progress, this could enable open-weight developers to catch up to the frontier. | **Lower capability:** In a world with lower capabilities, the market may be more diverse, with start-ups and open-weight developers catching up to big tech. However, this is not guaranteed and has not been the case in other markets. **Lower security:** Open-weight systems are harder to monitor and regulate and can easily be fine-tuned to remove safeguards. However, they also make certain safety measures easier, such as red teaming and mechanistic interpretability. **Lower global cooperation:** Lower compute and data demands would enable open-source developers to compete nearer frontier models. | + +### 3. Security + +#### 3.1. Alignment: + +A prosperous AI future depends on powerful AI systems reliably acting as their operators intend. However, current AI systems can pursue different goals or take unintended steps to reach a goal, a problem referred to as “misalignment”. This can occur spontaneously and by accident (MacDiarmid et al, 2025). This includes actively deceiving humans to accomplish goals, pretending to be aligned with users to avoid being modified, and blackmailing users to prevent being shut down (Scheurer et al, 2023; Jarviniemi and Hubinger, 2024; Anthropic, 2025d). While misalignment at human-level capabilities may be manageable, the challenge becomes critical if AI systems exceed human intelligence, where humans could be easily outsmarted and deceived. + +If progress can continue to be made on existing alignment methods, such as rules-based constraints and reasoning checks, these may prove adequate solutions; alternatively, systems could simply absorb human values from their training data (Bai et al, 2022; Guan et al, 2024). + +However, alignment may be a fundamental challenge, not only in reinforcement learning, where systems can chase rewards in harmful ways, but also during pretraining and deployment, where harmful patterns or misaligned prompts can be mimicked accidentally (Ngo et al, 2022). There are currently no known methods for ensuring models are reliably aligned with operator intent and investment in AI capabilities may be over 100 times greater than investment in alignment research (Apart Research, 2024). The development of human level capabilities, particularly in high-risk domains such as cyber offence, will therefore likely precede robust alignment solutions. + +#### 3.2. Control: + +AI systems can act in ways that make it very difficult for humans to monitor or control them. Evidence has shown systems can attempt to alter their environments to increase their reward signals or in pursuit of power-seeking strategies (Denison et al, 2024; Bondarenko et al, 2025). Experts are also concerned agents may have theoretical reasons to seek power or resources regardless of what their terminal goal is (Krakovna et al, 2023; Turner et al, 2023; Hadfield-Menell et al, 2016). While existing safety methods, including evaluations, monitoring tools, safety cases, and model interventions, could prove sufficient to detect hazardous behaviours, this may become harder as systems become increasingly intelligent (Benton et al, 2024; Greenblatt et al, 2023; Balesni et al, 2024; Ren et al, 2025). Furthermore, available control and monitoring methods, such as human-in-the-loop oversight, weak-to-strong supervision, and corrigibility, may not always be used. Under-resourced, reckless, or malicious deployers may not implement effective measures, particularly as they can introduce latency into pipelines and therefore limit their competitiveness. The speed of automation may also outpace methods like having a human in the loop. In theory, humans could lose full control over AI systems: this could happen gradually as we increasingly out-source decision-making; through misaligned systems actively causing harm through cyberattacks or influencing powerful humans, or through malfunctions if systems are controlling critical infrastructure (UK Government, 2025b). + +#### 3.3. Interpretability: + +Frontier models are frequently referred to as “black box” systems due to the complex and often opaque nature of their inner mechanisms. Developers therefore do not fully understand how these systems operate making it difficult to predict failures or explain unexpected behaviours, though some methods work reasonably well (UK Government, 2025c; Centre for the Future of Intelligence, n.d.). Furthermore, any internal patterns that researchers do discover are often not transferable to other AI systems, or too large or fragmented to explain specific decision-making processes. While many AI systems now display seemingly clear reasoning steps when operating that might appear interpretable, these may not accurately reflect its actual internal processes. + +Model size is a key factor in interpretability. Today’s frontier models are extremely large, optimised to memorise vast amounts of data rather than to reason deeply. But this may shift as smaller models trained to prioritise reasoning could soon be viable. These models would be more interpretable by design, especially where reasoning is made explicit through structured scaffolding, though could introduce new risks. + +#### 3.4. Malicious use: + +AI systems pose serious security risks by lowering the barrier to entry for malicious actors and increasing their capability and scale. Malicious use risks include the use of AI in generating content for fraud, exploitation, or disinformation; cyberattacks; developing chemical and biological weapons or other illegal items; and generating CSAM (Child Sexual Abuse Material) or NCII (Non-Consensual Intimate Images). In 2025, Anthropic reported a Chinese cyber espionage campaign, of which 80-90% was automated by AI, and North Korean operatives using AI to scale fraudulent employment at US Fortune 500 companies (Anthropic, 2025e; Anthropic, 2025f). There is significant scope for malicious use to grow in severity and frequency with increasing access to more capable AI systems. At the same time, AI may present opportunities to disrupt malicious actors, such as system owners identifying their own vulnerabilities first (National Cyber Security Centre, 2025). + +| Relationship to other CUs | | +| --- | --- | +| **Lower security** | **Higher security** | +| **Higher distribution and model access:** More accessible, open-weight systems will be easier for bad actors to abuse and harder for authorities to monitor and control. **Higher capability:** Security measures typically lag capability improvements. | **Higher adoption:** Integration into core institutional functions, in particular, will likely depend on AI systems that are secure. **Higher global cooperation:** Technology development is global, and so security cannot be guaranteed unilaterally — it requires internationally agreed standards and processes. | + +### 4. Adoption + +#### 4.1. Consumer adoption: + +Consumer AI use has grown at unprecedented speed: ChatGPT had 100 million weekly active users within 1 year of its release, which has doubled every 7 months since, to 800 million as of October 2025 (National Bureau of Economic Research, 2025; TechCrunch, 2025a). This is over 4 times quicker than the exponential growth in mobile phone subscriptions after their numbers surpassed 100 million (Our World in Data, 2024). On average, the volume of interactions per ChatGPT user has doubled (National Bureau of Economic Research, 2025). While public use and understanding of AI (particularly large language models) is growing, it remains lower amongst older, less-wealthy, and less-educated groups (Department for Science, Innovation and Technology, 2024). + +#### 4.2. Enterprise adoption: + +While business AI adoption has more than doubled since late 2023, only around 1 in 4 firms use AI as of 2026, and among those adopting AI less than one-third of employees use it (Department for Science, Innovation and Technology, 2026; Office for National Statistics, 2026). Nevertheless, the business customer base of frontier developers continues to grow: OpenAI’s increased from 2 to 3 million between February and June 2025 (CNBC, 2025c). Adoption varies significantly by firm size and sector: large businesses show far higher adoption rates than mid-sized and micro businesses, while the same is true for information and communication, finance, and real-estate sectors (Office for National Statistics, 2026). 99% of adoption is also currently for computer-based, cognitive tasks (Anthropic, 2026c). + +Enterprise adoption will largely be driven by the impact of AI on productivity for given tasks or occupations. Evidence from occupation-level studies shows that AI can deliver substantial productivity gains of up to 56% for certain tasks, but whether and when these gains translate to firm-level or economy-wide improvements remains uncertain (Brynjolfsson et al, 2024; Dell’Acqua et al, 2023; Noy and Zhang, 2023; Cui et al, 2024; Peng et al, 2023). Nevertheless, business demand for AI skills has surged: UK AI job postings have risen by around 200% year-on-year up to 2025 (Engineering and Technology Jobs, 2025). As institutional adoption grows, AI systems may become increasingly integrated into critical national infrastructure, posing significant opportunities in efficiency but also serious risks around the dependency and fragility of interconnected networks (RAND Corporation, 2025). + +#### 4.3. Access and autonomy granted to AI systems: + +As AI systems become increasingly capable of operating autonomously, a key uncertainty is how much access and autonomy users and organisations will choose to grant them. There are early signs of a shift towards granting AI systems broader execution capabilities, even in critical sectors such as finance. The AI Security Institute analysed usage data from over 1,000 public interfaces (MCP servers) that allow Al systems to access external tools and work as agents (AI Security Institute, 2025). They classified each server into 1 of 5 autonomy levels based on the tools and affordances it grants to models. + +Investigating finance-focused activity from December 2024 to July 2025, it increasingly observed new servers granting higher levels of autonomy to Al systems. Increasingly, Al systems can autonomously complete consequential actions like asset transfers and trading operations, rather than just reading and analysing data. + +#### 4.4. Societal impacts: + +The proliferation of advanced AI across communities, businesses, and broad society could bring a range of benefits, including increased economic growth and productivity; transformative health improvements; more efficient, accessible, and personalised public services; and more effective policymaking. However, it is also likely to cause significant disruption to our societal institutions and social fabric, from risks including misinformation; overreliance on AI systems; reduced accountability from automated decision-making; greater economic inequality; reducing ideological pluralism; or driving unsustainable energy use. + +Most recently, there have been several high-profile cases of harm from the increasing numbers of people turning to AI systems for emotional support and social interaction (Rose Kirk et al, 2025; Time, 2025). While AI companionship may provide benefits for isolated individuals, such as the elderly, it also poses many risks. Prolonged reliance on AI interactions – optimised for user satisfaction – may attenuate social capacities relating to tolerance, empathy, and conflict resolution, on which human relationships depend. A growing number of people are also using AI for medical guidance, increasing risks from AI systems providing inaccurate or misleading information (Bean et al, 2026). + +#### 4.5. Public sentiment: + +Public sentiment towards AI is a key factor influencing adoption. The public currently associates a diverse range of hopes and fears with the future of AI. Whilst AI is recognised for its potential to positively impact sectors like healthcare and crime prevention, there are also concerns about its potential to cause job displacement and spread misinformation (Department for Science, Innovation and Technology, 2024; Ada Lovelace Institute, 2025). There is strong public support for regulating AI and instituting safeguards, with most Britons believing AI regulation is lagging the pace of technological change (Ipsos, 2024). + +| Relationship to other CUs | | +| --- | --- | +| **Lower adoption** | **Higher adoption** | +| **Lower capability:** Individuals and organisations are likely to make less use of less capable AI systems, though this may not always be the case, and usage would likely still be relatively high even if models are relatively lower capability. **Lower security:** Most individuals and organisations are likely to have less trust in systems if they have significant concerns about their security, though many may still use them regardless. | **Higher capability:** Individuals and organisations are likely to use more capable AI systems more. Higher security: The public and businesses are more likely to use systems that are considered secure. **Higher global cooperation:** Global cooperation on AI will help to avoid large-scale malicious use and build public trust. **Higher distribution and model access:** Highly capable and widely accessible open-weight models would likely enable broader adoption. However, adoption could also be very high in a highly concentrated, closed-weight driven market. | + +### 5. Labour displacement + +#### 5.1. AI exposure: + +AI exposure estimates the degree to which an occupation’s tasks or required abilities align with the capabilities of AI systems – and the extent to which AI could perform, assist with, or accelerate those tasks. It therefore does not measure the extent to which roles will be automated. Around 60% of jobs in advanced economies are exposed to AI, with the UK’s service-based economy 1 of the most exposed, at around 70% (Cazzaniga et al., 2024). Current data suggests that entry-level jobs and those more exposed to AI are experiencing relatively sharper hiring slowdowns, but robust causal evidence linking these trends to AI remains limited (Allas and Goodman, 2025; DSIT, 2025; Brynjolfsson et al., 2025). Evidence from Anthropic suggests that the actual use of AI remains a fraction of what it can feasibly achieve, demonstrating that the impact of AI exposure on labour market displacement is mediated by the level and intensity of AI adoption (Anthropic, 2026d). + +#### 5.2. Complementary vs. substitutive effects: + +AI exposure can be broken down into 2 categories. The first is “substitution exposure”, which refers to tasks that AI could fully automate, at lower cost and without the need for human involvement. The second is ”complementary exposure”, which could involve AI automating some tasks in a job but not others, enabling workers to redistribute effort, or tasks where AI augments workers, collaborating closely with AI to perform tasks (Acemoglu, Autor and Johnson, 2026). The proportion of exposure which is substitution rather than complementary varies across roles. Well-structured tasks such as software programming tend to have higher substitution exposure: 79% of conversations on Claude Code were identified as automation and only 21% were identified as augmentation (Anthropic, 2025g). For roles which involve a high level of judgement, such as managers or government officials, AI exposure tends to be more complementary than substitution focused (Gmyrek et al., 2025). + +#### 5.3. Human preferences & institutional barriers: + +The impact of AI on labour displacement will also be influenced by human preferences. There is an observed tendency towards ‘algorithm aversion’, where humans prefer to avoid using AI even when it observably outperforms human counterparts (Filiz et al., 2023, Frank et al., 2024).). 1 study found that self-driving vehicles had to demonstrate a significantly higher level of safety than that of regular human-driven vehicles to be deemed as ‘acceptable’ (Liu, Wang and Vincent, 2020). Additionally, studies indicate a preference for human authorship, such that 90% of people preferred advice they believed came from a human expert rather than an AI, even if AI-advice was rated as having higher quality (Föyen et al., 2025). In addition to consumer concerns workers may resist automation. While workers support automating routine tasks to free up time for high value work, there are much lower levels of support for automating creative tasks. In addition, existing laws and regulatory frameworks are likely to act as barriers against autonomous AI systems capable of displacing labour (Shao et al., 2025). + +#### 5.4. Economic adjustments: + +The dynamics of labour demand will play a crucial role in determining the net impact of AI on employment. On 1 hand, labour-saving technologies could create offsetting increases in labour demand – an instance of the ‘Jevons Paradox’ whereby productivity gains reduce real prices which stimulates demand for goods and services (de Souza, 2025). On the other hand, demand effects could increase labour-displacement if rising unemployment and falling wages could trigger a decline in aggregate demand (Stiefenhofer, 2025). The extent to which demand effects promote stable-employment or displacement depends on several factors, including the magnitude of productivity gains, changes in real incomes, the distribution of income, and the nature and extent of government intervention. + +| Relationship to other CUs | | +| --- | --- | +| **Lower labour displacement** | **Higher labour displacement** | +| **Lower capability:** Less capable AI systems are likely to cause less labour displacement. Lower security: Less secure AI systems are likely to be given less access and autonomy and require greater human supervision, resulting in less labour displacement. **Lower adoption:** Lower adoption would limit labour displacement. Though even uneven levels of adoption could result in significant displacement. | **Higher capability:** More capable AI systems are likely to cause more labour displacement. **Higher security:** Secure AI systems are likely to be given greater access and autonomy and therefore require less human supervision, resulting in greater labour displacement. **Higher adoption:** Higher adoption would be required for significant labour displacement. Though labour displacement could still be lower even with higher adoption. | + +### 6. Global cooperation + +#### 6.1. International cooperation: + +International cooperation on AI governance is increasingly fragmented, particularly with the proliferation of open-source models reducing centralised control and subverting governance standards (Science Business, 2025). There have been positive signs of cooperation, including the Bletchley Declaration, the OECD AI Principles, and the Global Partnership on AI. However, competition and national security concerns have led more countries to assert control over AI infrastructure, talent, and governance frameworks (Centre for Geopolitics, JP Morgan Chase, 2025). + +Only US or Chinese labs have the funds, talent, power and infrastructure to achieve AGI in the next 5 years (Federal Reserve Board, 2025). China is pursuing state-led self-reliance and lower cost open-source exports and is expanding its influence via BRICS+ and the Digital Silk Road (International Services Shanghai, 2025; East Asia Forum, 2025). While the US favours private-sector innovation, infrastructure buildout, and defence integration, and is forming tech bilateral alliances through technology prosperity deals (The White House, 2025b; UK Government, 2025a). Currently, the US has a decisive lead over China in terms of compute. However, China’s unmatched energy production capacity will likely provide an advantage over longer timelines (Altman, 2025b). + +The EU has made concerted moves towards tech sovereignty, striving to bolster their own AI capabilities and assert regulatory leadership (Carnegie Endowment for International Peace, 2025). The UK has made similar efforts, but has aligned more closely with the US, for instance, in being the only participating nation to join the US in declining to sign the joint principles of the 2025 Global AI Summit (BBC News, 2025). The Middle East is also building out compute hubs and India is positioning itself as a leader in the Global South. + +#### 6.2. Global conflict: + +Global tensions are at their highest level since the aftermath of World War II, making cooperation more difficult. The Global Peace Index has declined almost every year since 2014, signalling a structural shift toward a more violent and fragmented world (Institute for Economics and Peace, 2025; Peace Research Institute Oslo, 2025). AI itself is also becoming increasingly weaponised as it becomes more integrated into military systems, compressing decision cycles and enhancing battlefield autonomy, such as in Ukraine (Centre for Geopolitics, JP Morgan Chase, 2025). It is possible that an AI-based weapon system disrupts the military balance of power (Ding and Dafoe, 2021). + +#### 6.3. Globalised supply chains: + +AI infrastructure is now a critical strategic resource. While the US, EU, and China have recently invested heavily in domestic chip production, the world is increasingly dependent on Taiwan for its semiconductors. Taiwan produces over 90% of the most advanced semiconductor chips and just 1 of its companies (TSMC) occupies over 70% of global market share (The Economist, 2023). Export controls have become key levers amid this level of concentration. Meanwhile, gulf states are leveraging cheap energy and capital to position themselves as AI hubs, while access to critical minerals, including rare earths and lithium, remains a challenge for resource-poor nations (Data Centre Magazine, 2025). The geographic concentration of the AI supply chain creates significant vulnerability to single points of failure and geopolitical tensions. + +| Relationship to other CUs | | +| --- | --- | +| **Lower global cooperation** | **Higher global cooperation** | +| **Lower security:** Low levels of cooperation will likely mean replication, confusion, or opposition around safety standards, with race dynamics possibly expediting the development of dangerous capabilities and encouraging a disregard for security. **Higher capability:** The development of high capability systems could drive an arms race dynamic and more fragmentation. Though it could also mobilise cooperation. | **Higher security:** Higher levels of cooperation would be required for global safety standards. However, low levels of security could mobilise cooperation. **Higher adoption:** Global cooperation on AI will help to avoid widespread malicious use and build public trust. | + +### 7. Underpinning economic assumptions + +Table 7 below sets out the assumptions underpinning economic elements of the scenarios, bringing together elements of the above paragraphs on critical uncertainties. + +| Scenario | AI deployment capability & adoption | Supply side productivity & investment | Demand side growth & demand | +| --- | --- | --- | --- | +| Slow Burn | AI outperforms humans at a minority of cognitive tasks. AI deployment occurs incrementally, at a similar pace to electrification in the early 20th century. Only a small proportion of firms have successfully integrated AI. | Modest productivity growth; productivity gains are confined to early-adopting sectors. Modest impact on labour market churn. | **Modest growth uplift.** Consumer demand broadly stable, though some deflationary pressure from the real economic spillovers arising from a global market-correction in AI asset prices. | +| Open Frontier | AI outperforms humans at a minority of cognitive tasks. AI deployment occurs moderately fast, at a similar pace to internet adoption in the early 21st century. A sizeable minority of firm have integrated AI into their workflows. | Considerable aggregate productivity gains but high variance across firms and sectors. Considerable impact on labour-market churn. | **Considerable growth uplift.** Consumer demand under pressure from disruption of highly exposed sectors. However, wage and employment gains in some sectors enable economic growth to continue. | +| Augmented Growth | AI matches human-level cognitive performance across most knowledge tasks. AI is adopted at the pace of smartphones in the 2010s. Around half of businesses have adopted the technology. | Substantial economy-wide productivity growth from widespread business adoption of labour augmenting technology. Significant increase in investment in intangible assets and compute infrastructure. | **Economic boom.** Labour-augmenting productivity gains produce real wage increases for the majority, resulting in stable aggregate demand and employment despite labour-market churn. | +| Transformation Economy | AI matches human-level cognitive performance across most knowledge tasks. AI is adopted at an unprecedented pace and is used by the majority of businesses. | Substantial labour-saving efficiency gains for adopting businesses. However, economic gains for UK businesses are partially offset by the cost of accessing models resulting in profit leakage to overseas frontier AI firms. | **Growth is reduced.** Labour displacement from adoption of human-level AI models results in competitive pressure on wages and employment. This reduces consumer spending and destabilises aggregate demand. | +| Take-Off | AI exceeds human experts across virtually all cognitive tasks. AI is adopted at an unprecedentedly fast pace and is used by most businesses. | Transformative productivity gains without recent historical precedent. Fundamental restructuring of labour markets, and large-scale redistribution of workers. Large scale capital investment boom in UK AI infrastructure. | **Rapid growth with K-shaped distribution.** Productivity gains pool heavily at the top; aggregate demand is maintained by elite consumption and investment rather than broad wage growth. Rising prices for in-person services and manual labour. | + +## Annex C: Updates since the April 2025 report + +| New Name | Most similar scenario from Jan. 2024 report | Key Differences Relating to the CUs | +| --- | --- | --- | +| Slow Burn | AI Disappoints | Slow Burn reflects more clearly how AI is unlikely to conclusively disappoint, even if the highest capabilities take longer to develop. For instance, more consideration is given to how new ways could be found to integrate and productise existing AI systems, even without radical improvements in capabilities. Access has been changed from 1 to 2 to reflect how open-weight systems would likely still be reasonably accessible given the focus of developers on productisation and open-weight systems being easier to fine-tune for specific uses and environments. Adoption has also been changed from 2 to 1 to reflect that this scenario shows the lowest plausible level of adoption that could be expected by 2030. | +| Open Frontier | AI Wild West | We have highlighted more of the opportunities highly capable and accessible open-weight systems would present. ‘Capability’ has also been lowered from 3 to 1 to show how this scenario represents a slowdown in AI progress, rather than continuation. | +| Augmented Growth | New scenario replacing Unpredictable Advanced AI | This is a new scenario, in which AI systems are highly capable but have not yet replaced most human workers given they still lack some key cognitive abilities. It also illustrates the possibility of significant progress on making AI systems safer. It has replaced “Unpredictable Advanced AI”, which has been merged into “AI on a Knife Edge” (now called “Take-Off”) given the overlap between these. | +| Transformation Economy | AI Disrupts the Workforce | ‘Adoption’ has been raised from 3 to 5 given how many firms would need to adopt AI for widespread job displacement to occur. | +| Take-Off | AI on a Knife Edge (with themes from Unpredictable Advanced AI) | ‘Safety’ has been lowered from 4 to 1 to more accurately reflect the key themes and to integrate themes from “Unpredictable Advanced AI”, which had 1 on safety. Both previous scenarios covered misalignment risks, but this is now defined more clearly in this scenario. ‘Global Cooperation’ has been lowered from 2 to 1 to reflect the AI race being a key theme of this scenario. | + +## References + +**AI Security Institute** (UK Government) (2025) [Frontier AI trends report](https://www.aisi.gov.uk/frontier-ai-trends-report) + +**Acemoglu, D., Autor, D. and Johnson, S.** (2026) [Building Pro-Worker Artificial Intelligence. MIT Economics (PDF)](https://economics.mit.edu/sites/default/files/2026-03/Building%20Pro-Worker%20Artificial%20Intelligence.pdf) + +**Ada Lovelace Institute** (2025) [How do people feel about AI?](https://attitudestoai.uk/assets/documents/How-do-people-feel-about-AI-2025-Ada-Lovelace-Institute.pdf) + +**Allas, T. and Goodman, A.** (2025) [Not yet productive, already disruptive: AI’s uneven effects on UK jobs and talent. McKinsey & Company](https://www.mckinsey.com/uk/our-insights/the-mckinsey-uk-blog/ai-uneven-effects-on-uk-jobs-and-talent) + +**Altman, S.** (2025a) [Reflections](https://blog.samaltman.com/reflections) + +**Altman, S.** (2025b) [Thoughts on the AI buildout](https://blog.samaltman.com) + +**Amazon** (2025) [Q1 2025 earnings call](https://ir.aboutamazon.com/events/event-details/2025/Q1-2025-Amazoncom-Inc-Earnings-Conference-Call-/default.aspx) + +**Amazon Web Services** (2025) [Amazon Q Developer reaches $260 million milestone](https://aws.amazon.com/blogs) + +**André, C., Bétin, M. and Gal, P.** (2025) [Dynamism in generative AI markets since the release of ChatGPT. CEPR](https://cepr.org/voxeu/columns/dynamism-generative-ai-markets-release-chatgpt) + +**Anthropic** (2025a) [Activating AI Safety Level 3 (ASL 3) protections. May](https://www.anthropic.com/activating-asl3-report) + +**Anthropic** (2025b) [Claude Code](https://code.claude.com/docs/en/overview) + +**Anthropic** (2025c) [Project Vend: can Claude run a small shop?](https://www.anthropic.com/research/project-vend-1) + +**Anthropic** (2025d) [Agentic misalignment](https://www.anthropic.com/research/agentic-misalignment) + +**Anthropic** (2025e) [Disrupting AI-enabled espionage](https://www.anthropic.com/news/disrupting-AI-espionage) + +**Anthropic** (2025f) [Detecting and countering misuse of AI: August 2025](https://www.anthropic.com/news/detecting-countering-misuse-aug-2025) + +**Anthropic** (2025g) [Anthropic Economic Index: AI’s impact on software development. Anthropic Research](https://www.anthropic.com/research/impact-software-development) + +**Anthropic** (2026a) [Pick up where you left off with Claude](https://www.youtube.com) + +**Anthropic** (2026b) [Understanding Claude’s personalization features](https://support.claude.com/en/articles/10185728-understanding-claude-s-personalization-features) + +**Anthropic** (2026c) [Economic Index](https://www.anthropic.com/economic-index#country-usage) + +**Anthropic** (2026d) [Labor market impacts of AI: A new measure and early evidence. Anthropic Research](https://www.anthropic.com/research/labor-market-impacts) + +**Apart Research** (2024) [For-Profit AI Safety](https://apartresearch.com/news/ai-safety-needs-to-scale-and-heres-how-you-can-do-it) + +**Artificial Analysis** (2026a) [AI model performance trends](https://artificialanalysis.ai/trends) + +**Artificial Analysis** (2026b) [Language model inference price](https://artificialanalysis.ai/trends#language-model-inference-price) + +**The Atlantic** (2025) [Elon Musk and Tesla Optimus](https://www.theatlantic.com/technology/2025/11/elon-musk-tesla-optimus/684968/) + +**Bai, Y. et al.** (2022) [Constitutional AI: Harmlessness from AI Feedback](https://arxiv.org/abs/2212.08073) + +**Balesni, M. et al.** (2024) [Towards evaluations-based safety cases for AI scheming](https://arxiv.org/abs/2411.03336) + +**BBC News** (2025) [UK and US refuse to sign international AI declaration](https://www.bbc.co.uk/news/articles/c8edn0n58gwo) + +**BBC News** (2026) [AI firm Anthropic agrees to pay authors $1.5bnto settle piracy lawsuit](https://www.bbc.co.uk/news/articles/c5y4jpg922qo) + +**Bean. A, Payne. R, Parsons. G, et al.** (2026) [Reliability of LLMs as medical assistants for the general public: a randomized preregistered study](https://www.nature.com/articles/s41591-025-04074-y) + +**Bengio, Y. (chair) et al.** (2026) [International AI Safety Report 2026. International AI Safety Report. February](https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026) + +**Benton, G. et al.** (2024) [Sabotage evaluations for frontier models](https://www.anthropic.com/research/sabotage-evaluations) + +**Bondarenko, Y. et al.** (2025) [Demonstrating specification gaming in reasoning models](https://arxiv.org/abs/2502.13295) + +**Brynjolfsson, E. et al.** (2024) [Generative AI at Work](https://arxiv.org/abs/2304.11771) + +**Brynjolfsson, E., Chandar, B. and Chen, R.** (2025) [Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab (Working Paper)](https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/) + +**Burnham, G. and Adamczewski, T.** (2025) [LLMs now accept longer inputs, and the best models can use them more effectively. Epoch AI](https://epoch.ai/data-insights/context-windows) + +**Burton, J.** (2023) [Algorithmic extremism? The securitization of artificial intelligence (AI) and its impact on radicalism, polarization and political violence](https://www.sciencedirect.com/science/article/abs/pii/S0160791X23000672) + +**Carnegie Endowment for International Peace** (2025) [The EU’s AI power play: between deregulation and innovation](https://carnegieendowment.org/) + +**Casper, S. et al.** (2023) Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. Available at: (Accessed: 20 May 2026) + +**Cazzaniga, M., Jaumotte, F., Li, L., Melina, G., Panton, A.J., Pizzinelli, C., Rockall, E.J. and Tavares, M.M.** (2024) [Gen-AI: Artificial Intelligence and the Future of Work. International Monetary Fund (IMF) Staff Discussion Note No. 2024/001](https://www.imf.org/en/publications/staff-discussion-notes/issues/2024/01/14/gen-ai-artificial-intelligence-and-the-future-of-work-542379) + +**Centre for the Future of Intelligence** (n.d.) [ADELE: A dataset for evaluating long-term AI capabilities](https://kinds-of-intelligence-cfi.github.io/ADELE/) + +**Centre for Geopolitics, JP Morgan Chase** (2025) [The geopolitics of AI: decoding the new global operating system. Eurasia Group](https://www.jpmorganchase.com/content/dam/jpmorganchase/documents/center-for-geopolitics/decoding-the-new-global-operating-system.pdf) + +**Cissé, A., Cooper, M.E., Zhu, M., Evangelopoulos, X. and Cooper, A.I.** (2026) [Can we automate scientific reasoning in closed-loop experiments using large language models? Digital Discovery](https://pubs.rsc.org/en/content/articlelanding/2025/dd/d5dd00520e) + +**CNBC** (2025a) [Satya Nadella says up to 30% of Microsoft code is written by AI](http://www.cnbc.com/2025/04/29/satya-nadella-says-as-much-as-30percent-of-microsoft-code-is-written-by-ai.html) + +**CNBC** (2025b) [Scale AI announces multimillion-dollar defence deal, a major step in U.S. military automation](https://www.cnbc.com/2025/03/05/scale-ai-announces-multimillion-dollar-defense-military-deal.html?msockid=172bd1a115326f6b1a27c48011326458) + +**CNBC** (2025c) [Report on OpenAI business user growth](https://www.cnbc.com/2025/06/04/openai-chatgpt-enterprise-ai.html?msockid=0d1007ee6c9a6d03391611976d1f6cdb) + +**Cui, Z. et al.** (2024) [The productivity effects of generative AI: Evidence from a field experiment with GitHub Copilot](https://mit-genai.pubpub.org/pub/v5iixksv/release/2) + +**Data Centre Magazine** (2025) [How the Middle East data centre sector can spur AI dominance](https://www.datacentremagazine.com/) + +**DeepMind** (2024) [AlphaFold](https://deepmind.google/science/alphafold/) + +**Dell’Acqua, F. et al.** (2023) [Navigating the jagged technological frontier](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321) + +**Denison, C. et al.** (2024) [Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models](https://arxiv.org/abs/2406.10162) + +**Department for Science, Innovation and Technology** (2024) [Public attitudes to data and AI: Tracker survey (Wave 4) report](https://www.gov.uk/government/publications/public-attitudes-to-data-and-ai-tracker-survey-wave-4/public-attitudes-to-data-and-ai-tracker-survey-wave-4-report) + +**Department for Science, Innovation & Technology** (2025) [Economic Estimates: Employment in the Digital Sector, January 2024 to December 2024. GOV.UK](https://www.gov.uk/government/statistics/economic-estimates-employment-in-the-digital-sector-january-2024-to-december-2024/economic-estimates-employment-in-the-digital-sector-january-2024-to-december-2024) + +**Department for Science, Innovation and Technology** (2026) [AI adoption research](https://www.gov.uk/government/publications/ai-adoption-research/ai-adoption-research) + +**de Souza, G.** (2025) [Artificial intelligence in the office and the factory: Evidence from administrative software registry data. CEPR / Federal Reserve Bank of Chicago Working Paper WP 2025-11](https://cepr.org/voxeu/columns/artificial-intelligence-office-and-factory-evidence-administrative-software-registry) + +**Ding J. Dafoe, A** (2021) [Engines of Power: Electricity, AI, and General-Purpose Military Transformations](https://arxiv.org/pdf/2106.04338) + +**East Asia Forum** (2025) [China expands AI globally through the Digital Silk Road](https://eastasiaforum.org/2025/04/11/china-expands-ai-globally-through-the-digital-silk-road/) + +**The Economist** (2023) [Taiwan’s dominance of the chip industry makes it more important. The Economist](https://www.economist.com/special-report/2023/03/06/taiwans-dominance-of-the-chip-industry-makes-it-more-important) + +**Engineering and Technology Jobs** (2025) [Surge in demand for AI professionals sends job vacancies soaring](https://engineering-jobs.theiet.org/article/surge-in-demand-for-ai-professionals-sends-job-vacancies-soaring-) + +**Epoch AI** (2023) [How predictable is language model benchmark performance?](https://%20epoch.ai/blog/how-predictable-is-language-model-benchmark-performance) + +**Epoch AI** (2025a) [AI companies’ revenue](https://epoch.ai/data-insights/ai-companies-revenue) + +**Epoch AI** (2025b) [What you need to know about AI data centres](https://epoch.ai/blog/what-you-need-to-know-about-ai-data-centers) + +**Epoch AI** (2025c) [AI in 2030-Extrapolatingcurrent trends](https://www.searchyour.ai/en/ai-in-2030-extrapolating-current-trends) + +**Epoch AI** (2025d) [Open weights vs closed weights models](https://epoch.ai/data-insights/open-weights-vs-closed-weights-models) + +**Epoch AI** (2025e) [How much does it cost to train frontier AI models?](https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models) + +**Epoch AI** (2026a) [MirrorCode: Evidence that AI can already do some weeks long coding tasks. 10 April](https://epoch.ai/blog/mirrorcode-preliminary-results/) + +**Epoch AI** (2026b) [Data on AI companies. Updated 28 April](https://epoch.ai/data/ai-companies) + +**Epoch AI** (2026c) [AI Benchmarking Hub](https://epoch.ai/benchmarks) + +**Epoch AI** (2026d) [GPU clusters database](https://epoch.ai/data/gpu-clusters?view=map&tab=point) + +**Epoch AI** (2026e) [Trends in Artificial Intelligence. Epoch AI](https://epoch.ai/trends) + +**Edelman, Y and Ho, A** (2025) [Compute scaling will slow down due to increasing lead times](https://epoch.ai/gradient-updates/compute-scaling-will-slow-down-due-to-increasing-lead-times) + +**Federal Reserve Board** (2025) [The state of AI competition in advanced economies](https://www.federalreserve.gov/econres/notes/feds-notes/the-state-of-ai-competition-in-advanced-economies-20251006.html) + +**Filiz, I., Judek, J.R., Lorenz, M. and Spiwoks, M.** (2023) [The extent of algorithm aversion in decision-making situations with varying gravity. PLOS ONE, 18(2)](https://pmc.ncbi.nlm.nih.gov/articles/PMC9942970/) + +**Föyen, L.F., Zapel, E., Lekander, M., Hedman-Lagerlöf, E. and Lindsäter, E.** (2025) [Artificial intelligence vs. human expert: Licensed mental health clinicians’ blinded evaluation of AI-generated and expert psychological advice on quality, empathy, and perceived authorship. Internet Interventions, 41, 100841](https://pmc.ncbi.nlm.nih.gov/articles/PMC12169703/) + +**Frank, D.-A., Chrysochou, P., Mitkidis, P., Otterbring, T. and Ariely, D.** (2024) [Navigating uncertainty: Exploring consumer acceptance of artificial intelligence under self-threats and high-stakes decisions. Technology in Society, 79](https://doi.org/10.1016/j.techsoc.2024.102732) + +**Gade, P. et al.** (2024) [BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B](https://arxiv.org/abs/2311.00117) + +**Gemini Team, Google** (2023) [Gemini: A family of highly capable multimodal models](https://arxiv.org/abs/2312.11805) + +**Gmyrek, P., Berg, J., Kamiński, K., Konopczyński, F., Ładna, A., Nafradi, B., Rosłaniec, K. and Troszyński, M.** (2025) [Generative AI and Jobs: A Refined Global Index of Occupational Exposure. International Labour Organization](https://www.ilo.org/publications/generative-ai-)and-jobs-refined-global-index-occupational-exposure) + +**Goldman Sachs** (2024) [AI: to buy, or not to buy, that is the question.](https://www.goldmansachs.com/insights/goldman-sachs-research/global-strategy-paper-ai-to-buy-or-not-to-buy-that-is-the-question) + +**Google Research** (2025) [Introducing nested learning: a new ML paradigm for continual learning.](https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/) + +**Government Office for Science** (2024) [Futures toolkit for policymakers and analysts.](https://www.gov.uk/government/publications/futures-toolkit-for-policy-makers-and-analysts) + +**Government Office for Science** (2025) [AI 2030 scenarios.](https://assets.publishing.service.gov.uk/media/6808fc002a86d6dfb2b52772/AI_2030_Scenarios_Report.pdf) + +**Greenblatt, R. et al.** (2023) [AI Control: Improving Safety Despite Intentional Subversion.](https://arxiv.org/abs/2312.06942) + +**Guan, M. et al.** (2024) [Deliberative Alignment: Reasoning Enables Safer Language Models.](https://arxiv.org/abs/2412.16339). + +**Gundlach, H., Lynch, J., Mertens, M. and Thompson, N.** (2025) [The Price of Progress: Algorithmic Efficiency and the Falling Cost of AI Inference.](https://arxiv.org/abs/2511.23455). + +**Hadfield-Menell, D. et al.** (2016) [The Off-Switch Game.](https://arxiv.org/abs/1611.08219). + +**Hassabis, D., Kumaran, D., Summerfield, C. and Botvinick, M.** (2017) [Neuroscience-inspired artificial intelligence.](https://doi.org/10.1016/j.neuron.2017.06.011) Neuron, 95(2), pp. 245–258. + +**Heitmann, M., Hinrichsen, T., Africa, D. and Sandbrink, J.** (2025) [Understanding AI trajectories: mapping the limitations of current AI systems.](https://www.aisi.gov.uk/research/understanding-ai-trajectories-mapping-the-limitations-of-current-ai-systems) + +**Ho, A and Berg, A** (2025) [Quantifying the algorithmic improvement from reasoning models.](https://epoch.ai/gradient-updates/quantifying-the-algorithmic-improvement-from-reasoning-models) + +**Hobbs, H. et al.** (2026), [“Exploring possible AI trajectories through 2030”](https://doi.org/10.1787/cb41117a-en) OECD Artificial Intelligence Papers, No. 55, OECD Publishing, Paris. + +**Hoffmann, J. et al.** (2022) [Training compute-optimal large language models.](https://arxiv.org/abs/2203.15556) + +**Hu, Y. et al.** (2023) [Toward general-purpose robots via foundation models: a survey and meta-analysis.](https://arxiv.org/abs/2312.08782) + +**Hugging Face** (2025) [Vectara hallucination evaluation model.](https://huggingface.co/vectara/hallucination_evaluation_model) + +**The Independent** (2025) [Google Gemini update adds memory features.](https://www.independent.co.uk/tech/google-gemini-update-ai-memory-b2807453.html) + +**Institute for Economics and Peace** (2025) [Global Peace Index 2025.](https://www.visionofhumanity.org/wp-content/uploads/2025/06/Global-Peace-Index-2025-web.pdf) + +**International Services Shanghai** (2025) [Key alliances launched to boost China–BRICS AI cooperation](https://english.shanghai.gov.cn/en-WAICLatest/20250729/eac167d7c7c6452bb57f05440b6e6345.html) + +**Ipsos** (2024) [Public trust in AI.](https://www.ipsos.com/sites/default/files/ct/news/documents/2024-09/Ipsos%20Public%20Trust%20in%20AI.pdf) + +**Jarviniemi, J. and Hubinger, E.** (2024) [Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant.](https://arxiv.org/abs/2405.01576) + +**Kaplan, J. et al.** (2020) [Scaling laws for neural language models.](https://arxiv.org/abs/2001.08361) + +**Krakovna, V. et al.** (2023) [Power-seeking can be probable and predictive for trained agents.](https://arxiv.org/pdf/2304.06528) + +**Kwa, T., West, B., Becker, J. et al.** (2025) [Measuring AI Ability to Complete Long Software Tasks.](https://arxiv.org/abs/2503.14499) + +**Lambert, N. et al.** (2023) [The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback.](https://arxiv.org/abs/2311.00168) + +**Lewis, P. et al.** (2020) [Retrieval-augmented generation for knowledge-intensive NLP tasks.](https://arxiv.org/abs/2005.11401) In: Advances in Neural Information Processing Systems, 33, pp. 9459–9474. + +**Liu, P., Wang, L. and Vincent, C.** (2020) [Self-driving vehicles against human drivers: Equal safety is far from enough.](https://%20ora.ox.ac.uk/objects/uuid%3Ad39c1bf4-cf54-4c1e-9a1d-d988376f50d0) Journal of Experimental Psychology: Applied, 26(4), pp. 692–704. + +**LMSYS (2026)** [Chatbot Arena leaderboard.](https://lmarena.ai/leaderboard) + +**MacDiarmid, M. et al** (2025) [Natural emergent misalignment from reward hacking.](https://www-cdn.anthropic.com/daad4360a8bdc707f8b22e3e745796ba27e57fb3.pdf) + +**McKinsey & Company** (2025) [The cost of compute power: A $7 trillion race to scale data centers.](https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-cost-of-compute-a-7-trillion-dollar-race-to-scale-data-centers) McKinsey & Company. + +**Meta Platforms** (2025) [Q2 2025 earnings call.](https://investor.atmeta.com/investor-events/event-details/2025/Q2-2025-Earnings-Call/default.aspx) + +**Microsoft** (2024) [Q4 2024 earnings call.](https://www.microsoft.com/en-us/investor/events/fy-2024/earnings-fy-2024-q4?msockid=0d1007ee6c9a6d03391611976d1f6cdb) + +**Mirzadeh, I. et al** (2025) [GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.](https://arxiv.org/abs/2410.05229) + +**Model Evaluation & Threat Research (METR)** (2025) [How does time horizon vary across domains?](https://metr.org/blog/2025-07-14-how-does-time-horizon-vary-across-domains/) + +**Model Evaluation & Threat Research (METR)** (2026) [Task completion time horizons of frontier AI models.](https://metr.org/time-horizons/) + +**Morris, M. et al.** (2024) [Levels of AGI for Operationalizing Progress on the Path to AGI.](https://arxiv.org/abs/2311.02462) + +**Mu, J.** (n.d.) [Statement on AI-driven software engineering (post on X).](https://x.com/) + +**National Bureau of Economic Research** (2025) [How people use ChatGPT (Working Paper w34255).](https://www.nber.org/papers/w34255) + +**National Cyber Security Centre** (2025) [The impact of AI on cyber threat (now to 2027).](https://www.ncsc.gov.uk/report/impact-ai-cyber-threat-now-2027) + +**Newman, S.** (2025) [GPT-5: the case of the missing agent.](https://secondthoughts.ai/p/gpt-5-the-case-of-the-missing-agent) + +**The New York Times** (2025) [OpenAI Uncovers Evidence of A.I.-Powered Chinese Surveillance Tool.](https://www.nytimes.com/2025/02/21/technology/openai-chinese-surveillance.html?unlocked_article_code=1.zE4.VE1v.QxVmvFuBeQtX&smid=url-share) + +**Ngo, R. et al.** (2022) [The Alignment Problem from a Deep Learning Perspective.](https://arxiv.org/abs/2209.00626) (Accessed: 20 May 2026). + +**Noy, S. and Zhang, W.** (2023) [Experimental evidence on the productivity effects of generative AI.](https://www.science.org/doi/10.1126/science.adh2586) + +**Office for National Statistics** (2026) [Business Insights and Conditions Survey (BICS), January 2026.](https://www.ons.gov.uk/economy/economicoutputandproductivity/output/) + +**OpenAI** (2025a) [GPT-4.5 system card.](https://openai.com/index/gpt-4-5-system-card/) + +**OpenAI** (2025b) [o1 preview.](https://openai.com/index/introducing-openai-o1-preview/) + +**OpenAI** (2025c) [ChatGPT Agent.](https://openai.com/index/introducing-chatgpt-agent/) + +**OpenAI** (2025d) [Introducing GPT-OSS.](https://openai.com/index/introducing-gpt-oss/) + +**OpenAI** (2026) [Memory FAQ.](https://help.openai.com/en/articles/8590148-memory-faq) + +**Our World in Data** (2024) [Mobile phone subscriptions data.](https://ourworldindata.org/grapher/number-of-internet-users) + +**Our World in Data** (2026) [External funding for privately held AI companies raising above $1.5 million.](https://ourworldindata.org/grapher/private-investment-in-artificial-intelligence) + +**Peace Research Institute Oslo** (2025) [New data shows conflict at historic high as U.S. signals retreat from world stage.](https://www.prio.org/news/3616) + +**Penadés, J.R., Gottweis, J., He, L. et al.** (2025) [‘AI mirrors experimental science to uncover a mechanism of gene transfer crucial to bacterial evolution’,](https://www.cell.com/cell/fulltext/S0092-8674%2825%2900973-0) Cell, 188(23), pp. 6654–6665.e2. + +**Peng, S., Kalliamvakou, E., Cihon, P. and Demirer, M.** (2023) [The impact of AI on developer productivity: Evidence from GitHub Copilot.](https://%20arxiv.org/abs/2302.06590) + +**Rahman, R. and Owen, D.** (2024) [The training compute of notable AI models has been doubling roughly every six months. Epoch AI.](https://epoch.ai/data-insights/compute-trend-post-2010) + +**RAND Corporation** (2025) [The United States needs to stress test critical infrastructure for different AI adoption scenarios.](https://www.rand.org/pubs/commentary/2025/01/the-united-states-needs-to-stress-test-critical-infrastructure.html) + +**RBC Wealth Management** (2026) [Big tech’s AI expansion: from investment to scalable returns.](https://www.rbcwealthmanagement.com/en-us/insights/big-techs-ai-expansion-from-investment-to-scalable-returns) + +**Ren, J. et al.** (2025) [The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems.](https://arxiv.org/abs/2503.03750) + +**Research and Markets** (2025) [Embodied AI and humanoid robot market research 2024–2025.](https://www.researchandmarkets.com/reports/6078204/embodied-ai-humanoid-robot-market-research) + +**Ritchey, T.** (2011) [‘General Morphological Analysis (GMA)’.](https://link.springer.com/chapter/10.1007/978-3-642-19653-9_2) In: Ritchey, T. Wicked problems – social messes: decision support modelling with morphological analysis. Berlin: Springer, pp. 7–18. + +**Rose Kirk. H et al** (2025) [Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships.](https://arxiv.org/abs/2512.01991) + +**Scheurer, J. et al.** (2023) [Large Language Models can Strategically Deceive their Users when Put Under Pressure.](https://arxiv.org/abs/2311.07590) + +**Science Business** (2025) [Experts warn of fragmentation in global collaboration for innovation.](https://sciencebusiness.net/news/synergies/experts-warn-fragmentation-global-collaboration-innovation#%3A~%3Atext%3DThe%20geopolitical%20context%2C%20changes%20in%2Cus%20more%2C%E2%80%9D%20Flanagan%20said) + +**Shao, Y., Zope, H., Jiang, Y., Pei, J., Nguyen, D., Brynjolfsson, E. and Yang, D.** (2025) [Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce.](https://arxiv.org/abs/2506.06576) + +**Smaldone, A.M. et al.** (2025) [Quantum machine learning in drug discovery: applications in academia and pharmaceutical industries.](https://pubs.acs.org/doi/abs/10.1021/acs.chemrev.4c00678?utm_source=chatgpt.com) + +**Stiefenhofer, P.** (2025) [The Future of Work and Capital: Analyzing AGI in a CES Production Model.](https://arxiv.org/abs/2502.07044) + +**Tech in Asia** (2025) [China’s Agibot targets 5,000 robots this year.](https://www.techinasia.com/news/chinas-agibot-targets-142m-revenue-ceo-says) + +**TechCrunch** (2025a) [Sam Altman says ChatGPT has hit 800m weekly active users.](https://techcrunch.com/2025/10/06/sam-altman-says-chatgpt-has-hit-800m-weekly-active-users/) + +**TechCrunch** (2025b) [Sam Altman says OpenAI will have a legitimate AI researcher by 2028.](https://techcrunch.com/2025/10/28/sam-altman-says-openai-will-have-a-legitimate-ai-researcher-by-2028/) + +**TechSpot** (2025) [Nvidia reaches historic 92% GPU market share.](https://www.techspot.com/news/108225-nvidia-reaches-historic-92-gpu-market-share-leaves.html) + +**Time** (2025) [Parents allege ChatGPT is responsible for their teenage son’s death by suicide.](https://time.com/7312484/chatgpt-openai-suicide-lawsuit/) + +**Toner, H.** (2025). [Taking jaggedness seriously: Why we should expect AI capabilities to remain extremely uneven. Rising Tide (Substack)](https://helentoner.substack.com/p/taking-jaggedness-seriously) + +**Turner, A. et al.** (2023) [Optimal Policies Tend to Seek Power.](https://arxiv.org/abs/1912.01683) + +**UK AI Security Institute** (2025) [Automation of AI R&D (AISI Deep Dive AISI\_2025\_07\_24\_DD\_01).](https://epoch.ai/files/Interviewing_AI_researchers_on_automation_of_AI_R_D.pdf) + +**UK Government** (2025a) [US–UK pact will boost advances in drug discovery, create tens of thousands of jobs and transform lives.](https://www.gov.uk/government/news/us-uk-pact-will-boost-advances-in-drug-discovery-create-tens-of-thousands-of-jobs-and-transform-lives) + +**UK Government** (2025b) [International AI safety report 2025.](https://www.gov.uk/government/publications/international-ai-safety-report-2025) + +**UK Government** (2025c) [International AI safety report 2025](https://internationalaisafetyreport.org/sites/default/files/2025-10/international_ai_safety_report_2025_english.pdf) + +**Vestager, M., Cardell, S., Kanter, J. and Khan, L.** (2024) [Joint statement on competition in generative AI foundation models and AI products. Federal Trade Commission.](https://www.ftc.gov/legal-library/browse/joint-statement-competition-generative-ai-foundation-models-ai-products) + +**Vinyals, O.** (2025) [Social media post on AI systems.](https://x.com/OriolVinyalsML/status/1990854455802343680) + +**The White House** (2023). [Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. Executive Office of the President, 30 October.](https://bidenwhitehouse.archives.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/) + +**The White House** (2025a). [America’s AI Action Plan. Executive Office of the President, July.](https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf) + +**The White House** (2025b) [The United States signs technology prosperity deals with Japan and Korea.](https://www.whitehouse.gov/articles/2025/10/the-united-states-signs-technology-prosperity-deals-with-japan-and-korea/) + +**Whitfill, J. et al.** (2025) [Forecasting AI Time Horizon Under Compute Slowdowns.](https://joel-becker.com/images/publications/forecasting_time_horizon_under_compute_slowdown.pdf) + +**Zheng, X. et al.** (2025) [Lifelong Learning of Large Language Model based Agents: A Roadmap.](https://arxiv.org/abs/2501.07278) + +[Back to top](#contents) + +## Is this page useful? + +* [Maybe](/contact/govuk) +* Yes this page is useful +* No this page is not useful + +Thank you for your feedback + +Report a problem with this page + +## Help us improve GOV.UK + +Do not include personal or financial information like your National Insurance number or credit card details. + +This field is for robots only. Please leave blank + +What were you doing? + +What went wrong? + +Send +Cancel + +## Help us improve GOV.UK + +To help us improve GOV.UK, we’d like to know more about your visit today. +[Please fill in this survey (opens in a new tab and requires JavaScript)](https://www.smartsurvey.co.uk/s/gov-uk-banner/?c=no-js). + +Cancel + +## Services and information + +* [Benefits](/browse/benefits) +* [Births, death, marriages and care](/browse/births-deaths-marriages) +* [Business and self-employed](/browse/business) +* [Childcare and parenting](/browse/childcare-parenting) +* [Citizenship and living in the UK](/browse/citizenship) +* [Crime, justice and the law](/browse/justice) +* [Disabled people](/browse/disabilities) +* [Driving and transport](/browse/driving) +* [Education and learning](/browse/education) +* [Employing people](/browse/employing-people) +* [Environment and countryside](/browse/environment-countryside) +* [Housing and local services](/browse/housing-local-services) +* [Money and tax](/browse/tax) +* [Passports, travel and living abroad](/browse/abroad) +* [Visas and immigration](/browse/visas-immigration) +* [Working, jobs and pensions](/browse/working) + +## Government activity + +* [Departments](/government/organisations) +* [News](/search/news-and-communications) +* [Guidance and regulation](/search/guidance-and-regulation) +* [Research and statistics](/search/research-and-statistics) +* [Policy papers and consultations](/search/policy-papers-and-consultations) +* [Transparency](/search/transparency-and-freedom-of-information-releases) +* [How government works](/government/how-government-works) +* [Get involved](/government/get-involved) + +--- + +## Support links + +* [Help](/help) +* [Privacy](/help/privacy-notice) +* [Cookies](/help/cookies) +* [Accessibility statement](/help/accessibility-statement) +* [Contact](/contact) +* [Terms and conditions](/help/terms-conditions) +* [Rhestr o Wasanaethau Cymraeg](/cymraeg) +* [Government Digital Service](/government/organisations/government-digital-service) + +All content is available under the [Open Government Licence v3.0](https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/), except where otherwise stated + +[© Crown copyright](https://www.nationalarchives.gov.uk/information-management/re-using-public-sector-information/uk-government-licensing-framework/crown-copyright/) diff --git a/sources/resources/leicht-moonshot.md b/sources/resources/leicht-moonshot.md new file mode 100644 index 0000000..3234025 --- /dev/null +++ b/sources/resources/leicht-moonshot.md @@ -0,0 +1,241 @@ +--- +url: https://writing.antonleicht.me/p/the-moonshot +--- + +[![Threading the Needle](https://substackcdn.com/image/fetch/$s_!4SKU!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dff89d-5b66-4169-8ef8-eb2ec9ed94e8_1024x1024.png)](/) + +# [Threading the Needle](/) + +SubscribeSign in + +# The Moonshot + +### Real sovereign AI has never been tried + +[![Anton Leicht's avatar](https://substackcdn.com/image/fetch/$s_!FPyB!,w_36,h_36,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75422da7-aafa-42ab-8fa6-cf4f0df85cf0_3166x3166.jpeg)](https://substack.com/%40antonleicht) + +[Anton Leicht](https://substack.com/%40antonleicht) + +Jun 17, 2026 + +21 + +Share + +This weekend, the U.S. government shut down Fable 5, the leading frontier AI model. The administration was motivated by domestic cybersecurity reasons, but chose export controls as an instrument to force this shutdown: it was the sharpest sword available to it on short notice. The effects on U.S. allies—fear, disorientation, alienation—were not only not the goal, but weren’t even seriously considered. In 2026’s AI policy, the rest of the world is so powerless that it can be **cut off from frontier AI as mere collateral damage of domestic U.S. policy.[1](#footnote-1)** + +Some international attempts to protest or negotiate betray a sad misunderstanding: ‘surely, something as bad for us must be *about us* in some way’. But the lesson from Fable is this: the American AI takeoff, from tech to policy to politics, is simply moving so fast that it is leaving the rest of the world behind. Now, I’ve cautioned not to overreact to the capriciousness of America’s adolescent AI policy, and [I](https://writing.antonleicht.me/p/cut-off) [still](https://writing.antonleicht.me/p/import-imperatives) [think](https://www.thefai.org/posts/an-allied-world-on-the-american-ai-stack-a-strategy-for-export-leadership) there are [ways](https://www.foreignaffairs.com/united-states/ai-divide) to rein this dynamic in. Countries can build leverage, adoption effectiveness, and secure access deals, draw the U.S. ecosystem into mutual codependencies and anchor it in the rest of the world. That remains our best shot at the best possible equilibrium, an allied world turning American AI into real-world abundance and wealth. But we must realise it won’t work in every scenario. + +Because these days, if you can’t see the noose of a politically charged security apparatus wrapping around previously-abundant artificial intelligence, I’m not sure you’re paying attention. Increasingly, **the frontier of AI capability is controlled by a maximally volatile version of the U.S. executive branch**. Evidence is generated ad-hoc or perhaps deep inside the intelligence community, action is taken based on personal loyalties and with little respect for long-run consequences, decisions are biased toward immediate effects and domestic concerns. All this is carried out by an American presidency that considers itself a unitary executive that wields supreme power, and Congress is frozen into inaction by vociferous AI politics and institutional dysfunction. No interdependence and rational economic incentive truly binds this sort of administration—and so **no ally can truly rely on American frontier AI.** + +Subscribe + +Now assume another thing to be true: frontier AI really matters, the best systems are strategically and economically superior to the rest, and the resulting lead gives those at the frontier a decisive economic and military advantage. Frontier systems become so powerful that they threaten the sovereign state’s monopoly on violence, and that their owners become as powerful as any nation. In that world, you either own a frontier system yourself, or you are at the mercy of those who build, own, and control them. From that, **any reasonable country would conclude that it simply needs its own frontier AI, however high the costs.** + +No one contests the logic itself. Critics of the conclusion, myself included, debate the premises: perhaps the frontier AI market will not shake out this way, and then fast-followers are enough; perhaps the U.S. will not be that unreliable and bilateral engagement can succeed; perhaps the technical paradigm will fail and alternative paradigms can catch up. But evidence keeps compounding in support of the worst-case scenario instead: the frontier really seems that important, and America really seems that volatile. In that world, if you want a sovereign strategy—not managed dependency, not a hedge, not a bet—the table stakes are clear: compete on what we know to work. Anything else is sovereignty theatre. + +[![](https://substackcdn.com/image/fetch/$s_!Oyqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png)](https://substackcdn.com/image/fetch/%24s_%21Oyqt%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png) + +## **The Challenge** + +Competing on frontier AI on purely technical terms is difficult, and middle powers are not set up to do it at all—that is the lesson from years of failure to even get close to the frontier. Even firms close to the frontier are struggling, and in ways that make it clear they don’t see a business case in being second-best. xAI is now renting out its compute to Anthropic instead, and Meta’s Superintelligence Lab remains embattled by [internal melodrama](https://www.silicon.co.uk/e-innovation/artificial-intelligence/meta-ai-hires-626529). But insofar as countries treat this as a strategically vital asset, the economic logic *stops mattering* to some extent. They need this asset to avoid unilateral dependency or entrenched vulnerability to anyone who does have frontier AI, economic rival or strategic aggressor. And genuinely willing nation states are powerful actors still: even today, they plausibly have sufficient energy, talent, funds, perhaps even semiconductors, to pull it off if they go all-in. + +But even provided countries’ willingness, this is *not* only a matter of succeeding at the most ambitious infrastructure project in recent history. It’s a matter of succeeding at this project while under **adversarial pressure from the U.S. and strict scrutiny from fickle electorates.** No one voicing the ambition of reaching the frontier is grappling with that further challenge of insulating the project against American coercion and domestic backlash at the same time. + +As a result, **the debate about sovereign frontier AI remains unserious**. The LinkedIn types and Euroboosters, stripped of real resources, are forced to pretend that you can build AGI in a coworking space in Stockholm. EU institutions voice lofty ambitions, just to shape them into massively underfunded subsidies for blue-skies research that would bear its first fruits after American superintelligence has already been deployed. And everyone who is aware of the scope of the challenge shies back from discussing it. Just seriously discussing middle power parity at the frontier has become the marker of a fool. But that won’t cut it anymore. If things continue down this path, sovereign frontier development might yet become the only play. I think we sceptics have dismissed that future too readily. + +There are two ways to read the remainder of this essay; in fact, there are two audiences for it. One will read this as *the plan*: what we should really do if we only had the willpower, a document to send to your superiors and hope that, just maybe, the penny has dropped. The other will read this as a reductio ad absurdum: see, this is what it would take, and that’s why it would never work. Either way, all that’s left to do is to look at how we could pull this off. **What follows is the middle power project.** + +[![](https://substackcdn.com/image/fetch/$s_!Oyqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png)](https://substackcdn.com/image/fetch/%24s_%21Oyqt%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png) + +## **The Coalition** + +*Who* would actually do this? **No single country could.** Because middle powers’ capital markets are too shallow, the financial burden needs to be borne by governments, and no government alone can fund this. And because the project might run counter to American interests, we should expect U.S. pressure, not only through AI-specific interventions, but horizontal escalation to domains like security cooperation or trade. Only a committed bloc of countries could hope to resist such pressure even under the best possible conditions. + +I think **a coalition of liberal democracies** is most likely to strike the right balance between the required scale and cohesion. Among these democracies, calibrating a cohesive coalition structure is difficult. Forces pull in different directions: on one hand, many countries have something meaningful and valuable to contribute. Europe has a still-unrivalled economic, industrial and fiscal base, and, as the biggest established bloc, would be the obvious starting point—yet it lacks the assets that matter most: a frontier lab and the compute to train one. Canada and Britain bring technical talent, and Britain also offers state capacity; Japan and South Korea bring further economic heft and semiconductor supply chain positions; Australia, Canada and Norway could be attractive hosts for compute. + +But **internal strife is a large risk** to a coordinated coalition. Alliance logic would have to structure around all its members’ contributions, and any of them falling away would add further volatility to an already-unstable constellation at a dire political time. Overextending the alliance in the abstract early, considering members that are not fully bought in, only to lose them later is a very dangerous prospect. That’s especially true given U.S. incentives to bilaterally target single participating members to get them to defect. If America offers massive datacenter investments and a broad exemption from export controls to one of its close allies—say, the UK, which seems most likely to clear the bar on security and alignment—can you really count on them refusing? Would they even opt in to begin, or decide to try their luck with bilateral U.S. engagement for now? + +Based on my understanding of the state of the bilateral relationship and the need for security guarantees, it strikes me as unlikely that South Korea, Japan, or the U.K. would reliably join a coalition for middle power AGI. But these things are downstream of domestic politics, and bilateral relationships deteriorate faster than you think—just look to Canada. Who knows who would join after another Fable episode or two? The coalition could expand and admit further participants at a higher price of entry if the worst-case trends in U.S. AI policy continued to unfold. + +So, the likely cast of value-aligned governments with substantial contributions to make that you can get together on short notice is the **EU, Canada and Australia**; with a warm invitation extended to the **U.K., Japan and South Korea** for whenever they’d like to reconsider their American alignment, and cooperation offers made to the usual tag-alongs of Western alignment like Norway, Switzerland, or New Zealand. Defection, of course, remains a risk even in that cluster, as America has strong incentives to buy off and buy out the participants. One way or another, the coalition will need a strong, binding commitment mechanism; perhaps the easiest technical (though politically difficult) path would be to precommit a majority of the resources—funding, talent, infrastructure access—in a way that cannot simply be clawed back by a defector. Further interested middle powers beyond the close alliance could commit funding in exchange for guaranteed access rights one rung below the coalition members, but should not be afforded full constituent voting rights, in an effort to keep the decision-making process somewhat effective. + +[![](https://substackcdn.com/image/fetch/$s_!D9-2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4ff123-dc64-43f7-818d-ee84276b2b0f_2244x701.png)](https://substackcdn.com/image/fetch/%24s_%21D9-2%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/fb4ff123-dc64-43f7-818d-ee84276b2b0f_2244x701.png) + +## **The Vehicle** + +Once that coalition is assembled, what would the actual vehicle be—would it be a private company, a nonprofit, a government agency, a consortium of sorts? The answer is yes, it would need elements of all of these: a private vehicle that draws resources from a consortium of legacy firms, corresponding with government agencies through a layer of special appointees. + +The coalition **does not have the luxury to rely only on pure private sector champions**. Even in America, the trend is moving toward stronger government involvement in the AI labs and their decisionmaking, and America has the advantage of a vital and well-capitalised private ecosystem to begin with. Our coalition does not have the same market structure, and neither does it have the benefit of time. No venture market in the coalition—nor all of them combined—could carry a project at this scale. Most of the burden will fall to the treasuries. + +And a speculative private enterprise alone, listed in one country, suspected to simply be chasing subsidies and profits, cannot rally the kind of governmental support that will be required. And last, if we really do face the securitised future of American AI, by the time that Europe finally has its private sector frontier champions up and running, America will already have furthered the soft nationalisation of its own developers through taking equity stakes and deeply interlinking labs with its intelligence community, and our private sector would be left racing against the U.S. government. + +No—**this needs the full backing of the coalition countries**, including the full support of their national security authorities in protecting the project, the full support of their trade authorities in securing the supply chains, and the full support of their treasuries in financing and backstopping the various expenditures we will require later on. Most of all, it needs the broader geopolitical weight of a coalition. + +#### *Institutional Design* + +The **vehicle to execute the project itself would need to be structured like a private company**: a highly powerful leadership group chosen on the basis of exceptional demonstrated research taste, an ability to stand up and scale up huge organisations, and an ability to communicate with funders—which in this unique case largely means communicating with a broad range of governments. The structure below this team of co-founders would largely resemble today’s most successful frontier AI developers. We do not have the luxury of experimenting with new formats: this is the only structure we know can make frontier AI models, so if we bet big, this is the structure we have to deploy. I expect the question of *where* the vehicle is based to be politically sensitive and the result of difficult negotiations. It’ll be spread across different cities to accommodate talent incentives and reduce lump risk in one jurisdiction. That also has the benefit of spreading out the second-tier effects of the project, like stimulating effects to tech ecosystems, or taxable revenue from salaries and infrastructure. + +[![](https://substackcdn.com/image/fetch/$s_!eUs-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd051d145-9a02-480b-abe9-cc4d085bd89a_2508x627.png)](https://substackcdn.com/image/fetch/%24s_%21eUs-%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/d051d145-9a02-480b-abe9-cc4d085bd89a_2508x627.png) + +Into this vehicle, **existing teams from any interested domestic champions can be appropriated and absorbed**, though with little respect to their current institutional structures. Organisations like Mistral and Cohere/Aleph Alpha, as well as various narrow champions like Black Forest Labs or ElevenLabs are valuable concentrations of talent and initial expertise. But there can be absolutely no over-indexing on their preexisting structure: regrettably, these companies have not been able to reach anything close to frontier parity, and so in the interest of the project’s success, we have to consider them an insufficient model. More to the point, they’re also an insufficient attractor: top-tier research talent will not believe in a continuation of what many of them perceive to be failed players, even if they share their ambitions. For better or for worse, the legacy brands of middle power AI have lost their credibility, and so they cannot serve as the basis for the project. Buying out roughly [half the sector](https://www.vestbee.com/insights/articles/europe-s-most-valuable-genai-startups)—the structures, IP, data, and teams worth having—at near its current valuation perhaps makes up for $25 billion of the total cost. + +Since the vehicle will require massive government support and should be responsive to government interest in its approach to public policy, safety and national security, there will have to be a **complementary government structure**. This structure should be entirely separate from the technical operations of the vehicle itself, lest bureaucratic torpor drag down the project and turn it into a government initiative. What I’d suggest is that **each participating country designate a ‘project czar’.** Czars are senior domestic policy figures at a ministerial or state secretary rank, convinced of the viability of the project and somewhat adept at navigating their respective domestic political institutions. These czars are provided with high-calibre teams of frontier policy operatives, similar in shape and makeup to the U.K. taskforce that became the world’s first AI safety institute. Czars and their teams are liaisons between the project and its constituting governments: they own all communications, translate technical language into government language and back; make sure the project understands when it brushes against non-negotiable policy imperatives and make sure the governments understand when it comes too close to impeding progress. This is an enormously important part of this project: if the government links work well, chances are much higher that political backlash can be contained and defection risk can be minimised. + +#### *Capitalisation* + +Once this vehicle exists, it needs to be equipped with sufficient funding. To give a realistic sense of the challenge, I’ll do my best to give you an informed guess. + +As a rough approximation, the project might cost **$500 billion over four years.** + +[![](https://substackcdn.com/image/fetch/$s_!ohBx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b75834d-779f-4aed-8910-d00c4561c338_1774x887.png)](https://substackcdn.com/image/fetch/%24s_%21ohBx%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/6b75834d-779f-4aed-8910-d00c4561c338_1774x887.png) + +Getting this money is technically simple and politically difficult. Treasuries have deep pockets, middle power debt capacity is still large, and one would have to hope that the product eventually yields revenue. Until then, treasuries, pension funds, wealth funds and so on will simply have to find legislative and executive ways to backstop private sector investments and provide outright funding where necessary. Hopefully, the project would manage to crowd in substantial private sector investment eventually: legacy industries in middle powers have deep cash reserves and a desire to gain exposure to AI, and could perhaps even be offered exclusive access, product fine-tuning and much more in exchange for contributions. But the coalition should not count on abundant private funding—**the expectation should be that the treasuries need to bear much of the initial four-year cost.** + +The upside of this capitalisation approach is an **Airbus model of strategic ownership**: the public carries all the early strategic risk, but in return also *owns* the resulting frontier lab. Frontier labs tend to be worth quite a lot of money, and the sovereign equity this project might generate would be enormously useful for government financing later on. In effect, the project sets up a volatile, overconcentrated sovereign wealth fund instead of just being a pure waste of money. If the project succeeds, the company can be partly privatised to the European market at strategically acceptable levels of private ownership, making for some belated cash backflow to any treasury interested in doing its own mini-IPO. In turn, limitations on clawing back your equity too early serve as a commitment mechanism: either you see the project through with us, or you lose your initial input. + +In that light, the overall price tag is substantial but not insurmountable. It is not dissimilar to what countries like Germany have taken on in [special debt packages](https://www.cleanenergywire.org/factsheets/qa-germanys-eu500-bln-infrastructure-fund-whats-it-climate-and-energy) before, and it would be distributed across a much broader coalition. There’s not much to say about the specifics: this is doable immediately, but only through leaders who decide the trade-off is worth going for. + +[![](https://substackcdn.com/image/fetch/$s_!Oyqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png)](https://substackcdn.com/image/fetch/%24s_%21Oyqt%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png) + +## **The Compute** + +Building frontier models is extremely resource-intensive: the buildout in the required computing power has given rise to the world’s most valuable firm and is shaping up to be the largest infrastructure push in human history. Middle powers have often taken to the pernicious illusion that they can instead sidestep some of this effort—maybe you could find a new paradigm, for instance, or make progress on ‘world models’ (though no one seems to agree on what that’s supposed to be). I think these hopes—as embodied by prominent labs like Ami and LawZero—are not entirely misplaced, but they’re long shots: they might or might not find a new paradigm, and it might or might not be competitive at the frontier of the most important tasks. And they face structural disadvantages: if there is a mold-breaking approach to be found, chances seem high that a compute-rich frontier developer with all its talent would find it. + +Paradigm-shifting bet is a fine part of a broad portfolio, but insufficient for what we are talking about here: a project with a clear path to a strategic payoff that’s legible even to sceptical policymakers and conservative treasuries. **There’s only one paradigm we know can get to the frontier, so that’s the paradigm we pick.** + +#### *The Price Tag* + +If this project is to compete on anything near the current paradigm, this project will need compute. There are two figures to benchmark the ballpark against. First, the good news: the project needs somewhat *less* than an American AI developer over the same time because there isn’t quite as much inference to run for most of it; American developers need to balance training and profitability in a huge market, while the project might be free to research and develop for a while. Comparing hyperscaler spending to the cost of this project is therefore misleading. American big tech needs enough compute to provide software to billions of people; the middle power project, especially in its early stages, does not. + +Then, the bad news: the frequently invoked figures for training costs in Chinese models like Kimi K2 or DeepSeek R1 are also wildly misleading. Even if it is correct that the final training run for a model in 2025 only costs [a few million dollars](https://epoch.ai/gradient-updates/what-went-into-training-deepseek-r1), three factors massively increase the compute draw beyond the actual pre-training run for the frontier model. Specifically, the project needs far more compute for internal deployment of its own proprietary coding agents to accelerate its research; for R&D experiments required to reach and further the frontier; and for pre- and posttraining at scale, which has become much more expensive since 2025. + +I’m not an expert in the semiconductor and datacenter discussions that follow; my numbers might be off by a lot, but I feel reasonably confident they are within the right order of magnitude to grasp the underlying strategic dimensions. My best guess is: the project might need as much as the [currently-scheduled](https://epoch.ai/data/data-centers?view=graph&tab=power) compute buildout for one major American lab. The project needs less inference than that developer, but perhaps more R&D; and I also assume that American labs will add further compute over the next few years beyond what is already planned. If that is true, **the project would at least require about 3 million Blackwell-class chips** over the four-year build, equivalent to roughly [5–6 GW of all-in datacenter power](https://epoch.ai/blog/what-you-need-to-know-about-ai-data-centers), and might plausibly cost around $275–300 billion to buy and put into somewhat secure datacenters. That runs more expensive than what hyperscalers pay for the same chips, but we need to price in our project’s relative inexperience in building the datacenters—which will require paying a premium on datacenter construction expertise—and the pressures of entering into an already-contested compute market. + +[![](https://substackcdn.com/image/fetch/$s_!trf0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7f21c49-b40a-4fa3-a9f6-544e36070442_1983x793.png)](https://substackcdn.com/image/fetch/%24s_%21trf0%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/a7f21c49-b40a-4fa3-a9f6-544e36070442_1983x793.png) + +#### *Building Out* + +The playbook for getting this compute online has been written a few times at this point. My colleagues at the Carnegie Endowment have published a [comprehensive report](https://carnegieendowment.org/research/2026/06/the-compute-coalition-how-to-build-the-future-of-ai-in-the-free-world) on compute buildout in middle powers just last week, and others have done good work on this. The short version is: you pick a combination of middle powers good at building compute for different reasons and push aggressive buildout projects in each of them. You go for a combination of host countries that can go fast so you can commence the actual work as early as you can, and host countries that can go broad so you can actually get the compute depth you need without running into bottlenecks. You probably slightly expand the compute host list based on political considerations—some treasuries will want the investment to be local, some countries will want local instances of secure datacenters for national security exercises, and so on. This is simply an industrial megaproject—difficult, but difficult in well-understood ways that middle powers have sometimes navigated in times of crisis before. As late as 2022’s gas shock, Germany—famously NIMBY-rich and laggard in infrastructure—managed to build LNG terminals [within 10 months](https://www.energyconnects.com/news/gas-lng/2022/november/germany-launches-first-new-floating-lng-terminal-in-record-time/). + +The same is true for the energy required. **Among the middle powers, there still are available energy solutions to get datacenters online quickly**: renewable mixes in Australia, Scandinavia and Iberia; nuclear-powered main grid capacity in France, Finland and perhaps Korea or Japan; deindustrialised high-power zones in Germany and perhaps Britain. Spread the compute between those, and the scarcity in behind-the-meter generation capacity will not be an immediate problem. Still, the project would be a drag on the grid and on overall energy supply, and it will have to make up for it by building out at least some proprietary capacity in the long run; I’d expect us to have to calculate something like $3 billion for energy opex and another $10 billion for grid improvements and power generation capacity. + +In the meantime, there will be a **scramble for rental compute available before the clusters come online.** Two measures in particular would be necessary: consolidating government-owned supercomputers and repurposing them away from the research ecosystem that currently uses them and into the project; and scrounging the spot GPU market for any rentals available. Between expropriating and renting the supercomputers and renting whatever compute we can get in a competitive market to accelerate the start date, I’d expect to spend about another $25 billion—rented frontier compute is very expensive, with Anthropic alone paying [$1.25 billion a month](https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/) to rent xAI’s Colossus cluster. That way, the researchers can commence work on day 1, with a smooth on-ramp into *more* R&D compute than would be available at a major frontier lab. + +#### *Acquisitions* + +That is, believe it or not, the *easy* part of getting the required compute. **The hard part is buying the chips without the Americans intervening**. The Trump administration is capable of remarkable feats of inconsistent policymaking—but surely even they wouldn’t simultaneously decide to export control frontier models, but still allow the unrestrained exports of frontier chips to a middle power project *explicitly aimed at* creating the same frontier capabilities outside the US. And export controlling chips is very easy: there are enough American buyers, the authorities are in place, the supply chain is entirely concentrated through an American firm. + +You’d think the same concerns would apply to exporting compute to China, and yet the administration has been hesitant to restrict that flow. The reason is that, with China, the U.S. concern is fuelling indigenisation by forcing Chinese industry to catch up on semiconductors, making them less reliant in the long run—which means it’s possible that the U.S. might choose to export to China, but not to its allies. The main *unofficial* reason for exporting to China, however, is that Nvidia likes money and a diverse pool of buyers. That reason does translate to middle powers, and it’s the strongest asset on their side: **if Nvidia likes the project, perhaps Nvidia can get the U.S. government to hold off on sabotaging it**. There is an uncomfortable analogy to being cut off from AI models here, but there are also differences: chips are not *directly* harmful in the same way as frontier models, so it’s more difficult to evaluate them for risks wholesale; believing they will be risky requires believing in the success of the project itself, which I assume many Americans will not take seriously; and of course Nvidia is much more influential with the current government than, for instance, Anthropic. + +That alone might not be enough. At some point, Nvidia’s influence might wane. Or Nvidia might genuinely not have enough chips to cater to all their premium customers at the same time. And once the project starts succeeding, the U.S. government might start having second thoughts about the security implications: if intelligence agencies are getting worried about the proliferation of advanced AI capabilities, they might lobby to put a stop to further compute exports very quickly. + +#### *The Coercion Shield* + +This is where the middle powers’ own semiconductor supply chain leverage comes in. Under normal conditions, chokepoints are only of limited use. But **in the context of the project, semiconductor bottlenecks can be used as an anti-coercion instrument:** if America attempts to stop the sale of chips to the project, the project in turn restricts resources—EUV machines, raw materials, memory chips if East Asia joins—that the *Americans* need to develop AGI. I think this tactic is more promising than the usual ASML bluster for two reasons: it is deployed in a context with broad, coalition-wide opt-in from all middle powers, and it is not only aimed at the U.S. government, but at Nvidia, which much more directly depends on access to upstream semiconductor inputs to stay in business. Deployed right, the coercion shield enlists Nvidia as a US-internal advocate: it gives Nvidia everything they’ve ever wanted in terms of avoiding monopsonies and supply chain bottlenecks, clarifies the alternative of a chip market crash, and shows the easy way out of continuing to sell chips. + +The ask is also fairly modest by itself: all the leverage needs to do is convince the U.S. that it’s **the path of least resistance to continue allowing the export of U.S. chips into the project**—no major interventions into U.S. domestic industry required. Neither the U.S. nor the coalition has a sovereign semiconductor supply chain by itself; leveraging that fact accordingly at least increases the odds of retaining uneasy mutual access. Fundamentally, the coercion shield trades on the assumption that the Americans will think themselves ahead, and would rather tolerate a European project they dismiss than a risk to their own AGI supply chain. + +This might work, but it’s a highly risky bet: the coercion shield takes some of the biggest existing sources of leverage and economic participation the coalition has, and it spends them on protecting the project instead of securing access, playing the U.S.-China dichotomy, or anything else. That is another way in which the project represents a full all-in—putting all middle power AI assets into the basket of actually pulling it off. But without a coercion shield, the project would crumble under American pressure before it even had the chips to launch. + +[![](https://substackcdn.com/image/fetch/$s_!Oyqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png)](https://substackcdn.com/image/fetch/%24s_%21Oyqt%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png) + +## **The Talent** + +Once you have the chips, you need someone who knows how to use them to make frontier models. Specifically, two types of talent matter. The leading group of founders needs to be made up of truly exceptional talent that strongly identifies with the project; and the larger group of engineers needs to be up to the standard of the task, and, most importantly, largely consists of former frontier lab employees. + +#### *Talent as Industrial Espionage* + +The latter part is vital and frequently misunderstood. Hiring ex-lab employees is not just a way to get good engineers on board, it’s a well-practiced form of basically somewhat legal industrial espionage. The current dynamic of leading researchers switching teams is part of the American frontier labs’ homeostatic balance: **they take practices with them, deploy them at their new employers, and no secret stays secret for very long**. Within that fact, there’s a source of hope and a deadline for middle powers at once. A source of hope, because you do not have to start from scratch: if you win over a cadre of high-calibre researchers, they can deploy a lot of quasi-proprietary techniques to get the project started. But a deadline, because this channel could close soon: labs are increasingly siloed, advances in internally deployed AI systems make information exchanges less necessary, and so poaching researchers will become less and less effective. + +#### *Recruiting* + +The fundamental logic of pulling off the recruitment is very simple—I’ll spare you the incentive structure analysis and just broadly give you the following: to recruit frontier lab researchers, you need an interesting challenge, a mission they can believe in, and compensation high enough that joining is attractive rather than suicidal. The latter part is in conceptual ways the easiest: you need to pay them as well as they are paid at the American labs—but you need to factor in lab equity, because a European project is a much less promising IPO prospect than an American AI company. All in all, I’d expect the cost of poaching sufficient senior talent for an effort like this to be close to what Meta was [reported](https://techcrunch.com/2025/06/27/meta-is-offering-multimillion-dollar-pay-for-ai-researchers-but-not-100m-signing-bonuses/) to be spending on top researchers, with individual senior pay packages in the tens of millions, plus obvious perks like fast-tracked visas. + +Across the project, **the personnel bill runs to about $65 billion**, in three tiers. Indulge my speculation as justification not for the figure, but the order of magnitude. First, a founding group of perhaps seven, distributed across participating nationalities, of the calibre that currently co-founds a frontier lab and is wealthy on paper through its equity. They would need guaranteed packages in the hundreds of millions each to make walking away rational, call it $2.5 billion in total. Second, a senior research cadre of around 150 on Meta poaching terms, $5–8 million a year apiece, each clearing into the tens of millions over the life of the project: perhaps $10 billion over the four-year build. Third, a technical corps of perhaps 3,000 staff—lower than frontier labs in the mid-thousands, but justifiable because of the product-related functions the project doesn’t yet need. These will be more expensive for the project than for a frontier lab: we have no IPO and no equity lottery to dangle, so we pay in cash what the American labs pay in stock, perhaps about $40 billion for the entirety of staff. We also have to buy out the unvested frontier-lab equity each hire forfeits on the way in, which I expect to cost at least another $10 billion up front. + +[![](https://substackcdn.com/image/fetch/$s_!bWBt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81d5416-12dd-4de2-94b3-c81946db68e5_2172x724.png)](https://substackcdn.com/image/fetch/%24s_%21bWBt%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/c81d5416-12dd-4de2-94b3-c81946db68e5_2172x724.png) + +As for the mission to believe in, I suspect the most effective way to create that is to **empower and recruit respected leaders.** Researchers—correctly—assume that the success of a project depends on the research taste of its leaders, and the extent to which these leaders will be able to act on what that taste suggests. The first step of the project would therefore be to approach these leaders. I suspect many middle power nationals with some patriotic loyalties on the lookout for a new challenge exist, and if the coalition was to approach them with a real commitment to enable their work and secure their independence, they might just be interested to join. From there, the founders are the project’s main ambassadors to the research community: they sell the vision, relay the commitment made to them, and get the teams on board. + +#### *Non-Human Talent* + +**Most workers at a frontier AI developer in 2027 will not be humans, but AI agents.** Already now, these AI agents do much of the actual *work* in a lab: they [write and run the code](https://www.aol.com/articles/anthropic-ceo-says-90-code-030401205.html), communicate the findings, and fill out the forms. The human talent is mostly responsible for coming up with ideas, taking meetings, and switching the model back to Opus when Fable gets cut off. Getting access to these AI agents in the early stages of the project will be the most difficult part of recruiting. Already today, Anthropic has [limited](https://www.anthropic.com/news/claude-fable-5-mythos-5) the use of its most advanced models for frontier LLM development; it doesn’t seem absurd to think that its competitors might follow suit. Because the project would not have its own frontier model from the start, it also would not have its own frontier coding agents from the start, massively slowing down its progress—in effect hamstringing the nonhuman part of its workforce substantially. + +Luckily, this problem resolves itself as the project takes off: once it builds its own coding agents, it can deploy them internally, and hopefully the external dependence diminishes. That is, incidentally, another reason why a sovereign *frontier* developer might be the only stable sovereign project: an eternal fast-follower will never have sufficient internal coding agents to keep up with compounding gains at the frontier. Securing coding agent access is hard but not impossible—perhaps, for the time being, OpenAI wants to brand itself as less inclined to sabotage competitors than Anthropic and therefore sells a big coding agent contract. Or perhaps we can simply offer an attractive deal structure for the first year or so, before the U.S. government takes note that the project might actually succeed. If we do well enough on compute, we can also offer to run some of the foreign AI agents ourselves, further increasing the financial incentives for the U.S. labs to give us access. Expecting to pay a hefty premium, we might plan to spend **$3 billion on foreign AI agents** over the first 18 months. + +[![](https://substackcdn.com/image/fetch/$s_!Oyqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png)](https://substackcdn.com/image/fetch/%24s_%21Oyqt%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png) + +## **Staying Alive** + +Once all these assets are in place, the climb toward a frontier system starts. It will be slow and arduous either way—slightly faster if you have unlimited access to U.S. compute and coding agents, much slower if not. But within months of the first clusters coming online, I’d expect a first rudimentary model a few months behind the frontier to be produced. If things go well, I suspect that **within, say, two years of its launch, the project might be able to actually deploy a frontier model.** + +**Within that time, the project is politically vulnerable**. The opening motivation will dissipate as time goes on. Single members will try to defect. AI, as a technology, will itself grow more and more unpopular, labour impacts and misuse risks will manifest. Not every contributing country will be excited to see its resources flowing into reaching frontier parity if the frontier itself is what is scaring their electorate. It seems impossible to game out a response to this threat in detail, but it’s the biggest uncertainty in this entire process. The naive response of a well-capitalised project is to buy them off ahead of time: ratepayer pledges to constrain individual effects of energy prices, direct AI dividends to anyone asymmetrically affected by the project’s progress, industry electricity price subsidies so that compute buildouts don’t trade off against legacy interests, and host-region payments to guard against local datacenter backlash. I expect the political insulation to require what is effectively **a package of bribes on the order of $80 billion**—but as with many investments in this list, these are incidentally also somewhat defensible instances of government investments that some middle powers have been considering anyway. And put side by side with [hundreds of billions](https://www.bruegel.org/dataset/national-policies-shield-consumers-rising-energy-prices) middle powers have spent shielding their economies from Covid or energy crises, the political shield seems outright modest. + +[![](https://substackcdn.com/image/fetch/$s_!DBQT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cbf3bb5-51e7-413d-9561-5018533a7274_2804x561.png)](https://substackcdn.com/image/fetch/%24s_%21DBQT%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/5cbf3bb5-51e7-413d-9561-5018533a7274_2804x561.png) + +The other part of the response is that the project will have **much better political appeal than American AI developers**: because the project does not face the same kinds of economic pressures to compete for consumer markets and investor stories, many of the most socially corrosive and politically vulnerable application areas (such as rapid labour market uptake, addictive use cases, or child safety issues) aren’t visibly connected to the project, just to its American alternatives. Between these two saving graces, we’ll need a lot of good comms work—but there is some hope that the project is structurally less vulnerable than its current American version. + +#### *Productisation* + +Make it through all this, and you finally and ultimately get a frontier system. That helps for about three months. But the **American AI developers are not just frontier laboratories, they are also sophisticated and capable product firms** with deep inroads into consumer markets and a wide roster of extremely revenue-rich business contracts. Building a frontier model is viable for them because the returns to being in the lead are enormous and can be reinvested into staying at the frontier. No European firm has the same market or sophistication available to them, and it has in fact been a perennial failure mode of non-American firms to scale at a similar level of ambition. + +To begin with, **the project would be the main frontier AI provider to coalition governments and high-security private applications**—perhaps even by necessity due to U.S. models not being available. Beyond that, the coalition might have to take protectionist action to encourage its private sectors to use the project’s models even where American alternatives are available. There is some economic risk to that, but it is also the only way to kickstart a flywheel between the coalition’s critical industrial capacity and the project. Without a genuinely ambitious project, that’s a recipe for [disaster](https://simongrimm.substack.com/p/let-europeans-use-ai) because it compels firms to use a sub-par economic input. But if the project itself succeeds, its privileged access to what would be the largest market in the world could entrench it as a viable economic actor. And then, if the project succeeds in developing a serviceable model, it can begin competing for global markets. There, it has a distinct advantage, as it squares off against capricious Americans and Chinese alternatives widely seen as untrustworthy. + +All that is its own challenge—building a successful frontier AI business is even more difficult than building frontier models. In the beginning, **we should expect to subsidise the project** and its ability to continue providing frontier AI for the foreseeable future, and consider it an expensive investment into a vital defence contractor, with an option to become a breakout economic success. But the upside potential is still high: done well, the project has a greater claim to political legitimacy than the American developers, and a greater claim to international reliability than either of its competitors. If you think breakneck competition with China and domestic political backlash are the greatest risks to the American labs, you might look at the project and think: maybe it could not just survive, but win? + +[![](https://substackcdn.com/image/fetch/$s_!VwWN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd727e20a-c7c2-46fd-a05a-8d928b841983_1254x1254.png)](https://substackcdn.com/image/fetch/%24s_%21VwWN%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/d727e20a-c7c2-46fd-a05a-8d928b841983_1254x1254.png) + +## **Half Measures** + +Vapid inspirational quotes will tell you that if you shoot for the moon and miss, you might at least end up among the stars. Yet in AI, if you shoot for the moon and miss, you end up with a subsidy-chasing fast-follower instead. If a sovereign development project does not succeed at developing a frontier model, it is not worth pursuing as a response to our initially outlined challenges at all. + +The impulsive reaction to all of the proposals above will usually be ‘well, this seems a little bit much’. The sentiment will be: yes, we have an interest in sovereignty and some spending to mobilise, but surely this is the radical starting position, and we can actually find a reasonable compromise at a fraction of the cost here. That attitude is the standard operating procedure for the European Commission in particular, which has time and time again taken good ideas—like the Gigafactories or the frontier AI initiative—and watered them down until they’ve become unrecognisably unambitious. In this case more than in any other, that impulse is fatal: **either you do the fully ambitious version of this, a version you can honestly believe might make it to the frontier, or you might as well not do it**. You might see just how unsatisfying the second-best outcome is by looking at the second rung of AI developers today—some of them do fine on short-term revenue, yet clearly none of them have the standing of genuine strategic assets even to their home countries. + +In this case, the frequently-invoked little brother to the project is to create a ‘fast-follower’ instead. A fast-follower is an AI company that consistently stays some amount of time behind the frontier—it’s what Chinese developers and Europe’s Mistral have tried to differing degrees of success. But **a fast-follower is fundamentally a solution to a very different problem**: it supposes that there will be a visible frontier, it will be cost-effective to chase it at a stable distance of a few months, and that the model that results will be useful. But think back to the top, to our Fable scenario: a world where the United States security apparatus tightly controls and only sparsely distributes critical frontier capabilities, and where those capabilities are key inputs into vital economic and strategic functions. If that scenario is true, a fast-follower doesn’t help: the visible frontier is far out of reach and difficult to emulate, and second place is no consolation. If that scenario is false, I struggle to see why you’d even need a domestic fast-follower at all. If access is *not* locked down, then surely you are just better off importing the models instead, not bearing any of the huge costs outlined above, and investing into more effective adoption and downstream value capture instead. + +Worse, **fast-followers might become less and less feasible the more important frontier AI becomes.** As the closed frontier, available only to the U.S. government and perhaps privileged buyers, becomes less visible to everyone else, it becomes more difficult to emulate: followers cannot distill it, and they cannot even guess at the shape of state-of-the-art models from using them and observing their performance anymore. And the further the limited frontier is from the commercially available frontier, the more AI asymmetrically accelerates the incumbents: if Anthropic’s coding agents are much better than external coding agents, Anthropic will be much better at growing the gap. + +There are, admittedly, pleasant things about having your own fast-follower. Fast-followers like the Chinese open-sourcers hurt American developers’ leverage a bit, provide secure and privileged access to models in domains where the frontier does not matter, allow unlimited specialised posttraining for applications where you don’t want to share your approach because it’s based on privileged data, for instance, and so on. If it’s possible as a matter of economic policy—not a big consolidated project—to encourage middle power fast-followers, I think it should definitely be done. + +But all of that deeply, fundamentally misses the point, which is simply this: usually, middle powers are not well-positioned to develop AI models. They should attempt to do so under extraordinary circumstances of geopolitical need. And it is precisely under dire circumstances that a fast-follower does *not* cut it and a frontier model is needed instead. + +[![](https://substackcdn.com/image/fetch/$s_!Oyqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png)](https://substackcdn.com/image/fetch/%24s_%21Oyqt%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2c0e6d1e-3f54-4c73-819b-6058182a96a5_500x500.png) + +## **Where to Begin?** + +Where does all that leave us? Between two worlds, I think. **One is the world of ‘what should be done under certainty and perfect conditions’.** In that world, the project I describe is the only thing worth doing, and we need to start working on how to pull it off. In practice, that would mean that in many rooms at once, conversations would need to start. Researchers in San Francisco meeting after work to assemble teams, world leaders huddling on the sidelines of the G7 begin forging a political consensus. Teams in capitals are springing into action and standing up taskforces. AI experts hailing from middle powers, but now in American exile, start work on Signal groups and Google documents gathering names, contacts, funders, ways to make this work. Things move fast, but under the surface. The project grows together from both ends: the governments mobilise their motivation in the abstract and start reaching out into the AI world, and the AI world rises to the challenge and offers them a blueprint to work with. + +That’s not our world. **Instead, we live in the world of political realism**; a world that knows just how embattled the resources required for the project are. How rare political ambition of that scale is in the Western world, how scarce energy and public funds already are. And how uncertain that specific future of the highly securitised, weapons-grade frontier still is. In that world, **putting an actual, honest price tag on AI sovereignty makes you realise how unlikely it is to happen.** + +Trying to achieve AI sovereignty is to ask the most structurally conservative governments in the world to hinge their fiscal credibility and political fate on what they still consider a highly speculative technical bet. It’s going into cabinet rooms and asking for 500 *billion* dollars. In the real world, none of this will happen. Perhaps with a sigh of exasperation, you close this tab and return to the political reality of middle power policy. The rest of the world is and will remain behind, but we can manage the dependence. We forge access deals, attract U.S. investment, drive ahead the integration between domestic industries and foreign model providers. It’s a bet, but it’s a decently hedged bet for now: perhaps the market trends favour the effective adopter, perhaps the gains diffuse broadly. Perhaps serious middle power contributions have the U.S. reconsider the question of alliance-wide alignment, and the American project can come to benefit the world’s democracies as well. + +But no matter how much progress we make in that real world, our fate will never quite be in our hands. When we wake up in the European morning, we’ll still check our phones and see what the American evening has brought this time: do we still have access to leading artificial intelligence? Are our systems still secure, are our firms still online—are our citizens still safe? There will always be that uneasy feeling, gnawing at our confidence in piecemeal progress. **If we were truly serious, wouldn’t we launch the project instead?** + +--- + +Subscribe + +[1](#footnote-anchor-1) + +*My thoughts on what middle powers’ direct response to Fable should be can be found [here](https://x.com/anton_d_leicht/status/2065778604110184742).* + +21 + +Share + +TopLatest + +No posts + +### Ready for more? + +Subscribe + +© 2026 Anton Leicht · [Privacy](https://substack.com/privacy) ∙ [Terms](https://substack.com/tos) ∙ [Collection notice](https://substack.com/ccpa#personal-data-collected) + +[Start your Substack](https://substack.com/signup?utm_source=substack&utm_medium=web&utm_content=footer)[Get the app](https://substack.com/app/app-store-redirect?utm_campaign=app-marketing&utm_content=web-footer-button) + +[Substack](https://substack.com) is the home for great culture + +This site requires JavaScript to run correctly. Please [turn on JavaScript](https://enable-javascript.com/) or unblock scripts diff --git a/sources/resources/leicht-roadmap.md b/sources/resources/leicht-roadmap.md new file mode 100644 index 0000000..2b716e9 --- /dev/null +++ b/sources/resources/leicht-roadmap.md @@ -0,0 +1,144 @@ +--- +url: https://writing.antonleicht.me/p/a-roadmap-for-ai-middle-powers +--- + +[![Threading the Needle](https://substackcdn.com/image/fetch/$s_!4SKU!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dff89d-5b66-4169-8ef8-eb2ec9ed94e8_1024x1024.png)](/) + +# [Threading the Needle](/) + +SubscribeSign in + +# A Roadmap For AI Middle Powers + +### First thoughts on leveraging bottlenecks for abundant intelligence + +[![Anton Leicht's avatar](https://substackcdn.com/image/fetch/$s_!FPyB!,w_36,h_36,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75422da7-aafa-42ab-8fa6-cf4f0df85cf0_3166x3166.jpeg)](https://substack.com/%40antonleicht) + +[Anton Leicht](https://substack.com/%40antonleicht) + +Jan 23, 2025 + +1 + +Share + +AI middle powers – that is, most advanced economies that are not the US or China – need to find a strategy for participating in rapid AI progress. Specifically, they need to find a durable connection to AI-driven economic growth and some strategic leverage in negotiations with great AI powers. I argue the most promising path is leveraging bottlenecks: to find a sector that stands between AI progress and real-world effects and make economic and foreign policy around it. + +Subscribe + +## **Here’s Where We Are** + +[Recently](http://antonleicht.me/writing/middlepowers), I argued that what I call ‘**AI middle powers**’ – that is, most advanced economies that are not the US or China – **are in a tough spot**: + +* Their participation in AI-powered growth is contingent on finding a niche +* Their sharing in AI benefits depends on compute & foreign policy +* Their safety from AI-powered threats hinges on access to foreign top-tier defensive capabilities. + +Currently, most middle powers, from Germany and France via India to Japan and South Korea, lack a perspective to avoid being left behind. What’s more, they also lack motivation to make a plan — ‘situational awareness’, as it were. So, any proposal faces a tricky political economy. + +[![](https://substackcdn.com/image/fetch/$s_!OssO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55c8416b-0732-4d2c-b474-3c19dabaca51_1516x998.png)](https://substackcdn.com/image/fetch/%24s_%21OssO%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/55c8416b-0732-4d2c-b474-3c19dabaca51_1516x998.png) + +This post operates on what I take to be a **conservative consensus position** among people working in frontier AI: Over the next decade, AI capabilities will continue to rapidly increase, driven mostly by private companies in the [US and China](https://epoch.ai/data/notable-ai-models#china-compute-trends) supported by the [respective national governments](https://foreignpolicy.com/2023/06/19/us-china-ai-race-regulation-artificial-intelligence/). This will make AI systems meaningful *possible* drivers of growth and scientific progress, as well as central aggressive and defensive tools in issues of national security. I try to make the following arguments compatible with most specific versions of that overall story. + +## **Here’s What Middle Powers Need** + +I think a more AI-centric world economy and the resulting novel strategic paradigm asks two questions of a middle power: **What is their economic contribution to an AI-dominated economy,** and **what is their strategic leverage in an AI-driven world order?** + +They need an economic contribution for two reasons: + +First, to create a **business incentive to proliferate US- or China-built AI services** into their country. If there is no strong, profitable sector that uses AI in your country, it might quickly become unprofitable to frontier AI companies to offer their full suite of products; especially if there are some costs to operating in your country, such as regulatory compliance or localisation. Maybe this incentive can even be your consumer market, if it’s big enough – but that’s much more contingent on (a) a very specific model of AI diffusion with profits driven by ultimately B2C AI products, and (b) on a fairly modest rate of AI-powered growth at home that makes clandestine markets worth the cost of entering them. For one example of how this can go, look at Brussels digital policy, where the size and power of the European economy long [compelled](https://www.politico.eu/article/parliament-tech-eu-european-commission-us-president-donald-trump-meta-mark-zuckerberg/) tech companies to offer their products even despite considerable costs, but has lately [started failing](https://www.ft.com/content/2c1b6bfd-ce73-451d-8123-0df964266ae8). + +[![](https://substackcdn.com/image/fetch/$s_!UDfc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb17f993d-eed0-4c8e-9e13-a20c9cc65576_2400x1634.png)](https://substackcdn.com/image/fetch/%24s_%21UDfc%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/b17f993d-eed0-4c8e-9e13-a20c9cc65576_2400x1634.png) + +Second, more generally, to **participate in the overall economic returns from AI**. Much valuable contributions have recently tried predicting the [overall economic returns](https://inferencemagazine.substack.com/p/how-much-economic-growth-from-ai) of widespread AI adoption — no matter how this might go, it likely would not be a universal trend. As with most major technological revolutions in the past, AI diffusion would likely dramatically increase growth in some parts of the economy – and the world –, and could cause stagnation and even destitution in others. In 1820, you probably don’t want to be a weavers’ town. + +These two factors are often conflated, but somewhat independent: If you reap some early days benefits from AI growth by buying and using low-margin intelligence, that can still fail to structurally incentivize future participation. There was a brief period where the weavers from above felt pretty good about industrial production of their yarn, just before the power looms. And inversely, even if your overall economy stays behind global growth trends, but you at least have a sustainable AI-related sector, that might be all you need to motivate the developers and their national overlords to provide the AI capabilities at large. That’s not very aspirational, but survivable — especially if AI has [deflationary effects](https://www.reuters.com/breakingviews/ais-deflationary-winds-will-blow-away-profits-2023-06-27/). + +Middle powers also need **strategic leverage even beyond these economic entanglements** to make sure they aren’t entirely at the [geopolitical mercy](https://www.hoover.org/research/defense-against-ai-dark-arts-threat-assessment-and-coalition-defense) of their associated great AI power. Economic incentives are fine and sufficient as long as AI remains a predominantly civilian technology, where you can trade goods and leverage comparative advantage in a way that ultimately ensures mutually beneficial proliferation. But that’s only part of the story. Even though obvious security applications have failed to manifest so far, trends point at increasing securitisation. This can lead to a world where countries need access to near-frontier capabilities to stave off AI-powered aggression, but these capabilities are safeguarded by state-run projects in great AI powers. These state-run projects might not be too keen on sharing — [recent](https://www.hoover.org/research/defense-against-ai-dark-arts-threat-assessment-and-coalition-defense) [policy writing](https://www.rand.org/pubs/research_reports/RRA2849-1.html) from the US that second-guesses proliferation even through the Five Eyes intelligence alliance should send shivers down many middle powers’ spines. With all the Manhattan project talk abound, even a lot of close US allies would do well to remember how much wringing it took for nuclear proliferation; and what the US might have learned from it. + +In that setting, a middle power might need to make a **strategic contribution that gives allied great powers a reason to provide access** to leading security-relevant models, participation in the compute supply chain; and that reason should be compelling enough that great powers can’t unilaterally leverage the middle powers’ dependency on their models and compute. + +## **Here’s What I Think Middle Powers Should Do** + +To start, **I don’t think the way to do that is for middle powers to get ambitious about building AI models themselves**. The point has been made elsewhere — in short: In almost all middle powers, the infrastructural gap to the US and China is huge: in terms of energy, compute, and talent, and competition for the latter two is at an all-time high. This is made worse by the shaky economic situation and subsequent investment limitations of many of these powers, and the comparative lack of venture capital that could alleviate that. If [catching up](https://www.bruegel.org/policy-brief/catch-us-or-prosper-below-tech-frontier-eu-artificial-intelligence-strategy#:~:text=The%20EU%20could%20gradually%20build,for%20patents%20and%20skilled%20researchers.) was an option, it would obviously be advisable – but currently, it does not seem realistic. + +[![](https://substackcdn.com/image/fetch/$s_!uz0y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cb939d2-0806-40f5-8f9d-362f31113c62_1690x394.png)](https://substackcdn.com/image/fetch/%24s_%21uz0y%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/1cb939d2-0806-40f5-8f9d-362f31113c62_1690x394.png) + +That [headline](https://www.euractiv.com/section/tech/news/us-dwarfs-eu-ai-infrastucture-push-with-500-billion-investment/) illustrates the depth of misery well. + +With [recent news](https://arxiv.org/abs/2501.12948), you might think that open source can bail you out after all — but that seems risky to count on: Fast-following tends to be easier early in paradigms, securitisation makes top-tier proliferation through open source less likely, and the shift to inference scaling places large infrastructural requirements even on the use of open-source models. + +So what should middle powers do instead? In short: Be valuable on bottlenecks that arise when intelligence becomes cheaper and cheaper. Tyler Cowen [says](https://www.dwarkeshpatel.com/p/tyler-cowen-4): “[Massive increases in intelligence] just means the other constraints in your system become a lot more binding, that the marginal importance of those goes up”. That can be leveraged. + +**On the economic level, that means scanning your economy for bottlenecks for translating AI capability gains to GDP growth. Then, make a substantial contribution to widening them.** + +As to what these bottlenecks will be: That’s a whole area of contentious research I hope to see much more of. For starters, reasonable speculation likely converges on some general takeaways: Don’t over-index on knowledge economy services that can likely be replaced or moved to the great powers at low costs (instead of enhanced). Instead, sit at intersections with natural limitations to AI deployment & automation. Much more has been said [elsewhere](https://inferencemagazine.substack.com/p/how-much-economic-growth-from-ai) on what that might be. Generally, the production of (often physical) **goods either upstream or downstream from AI use seems constraining and relevant**. Specific versions of leveragable bottlenecks to AI-powered growth might be: + +* [Strategic elements](https://www.rand.org/pubs/perspectives/PEA3776-1.htmlhttps%3A//www.rand.org/pubs/perspectives/PEA3776-1.html) of the compute supply chain, from raw materials to lithography — the Netherlands and Taiwan are the obvious examples. +* Novel & privileged [sources of data](https://www.ft.com/content/8187a268-3494-11ea-a6d3-9a26f8c3cba4), esp. around industrial and strategic applications. Israel, Singapore, and some smaller sectors of most industrial states do well on this. +* Robotics & general embodiment; anything that ramps up the ability of AI systems to interface with the real world. South Korea and Japan do well on this. +* Industrial capacity writ large: Many avenues to the purported economic and social [benefits](https://openai.com/global-affairs/openais-economic-blueprint/) of AI run through physically building things. Someone actually has to make the RNA vaccines and superweapons. Germany and Japan are good at this, but e.g. Italy and India also have a decent share. + + + Specifically: **automated production quality** that might reduce friction for implementing AI-driven efficiency gains & shortens iteration cycles for experimenting on AI-generated proposals. + + And: sufficiently **digitized production** that enables comprehensive and real-time data collection for AI-powered optimization, quickfire iteration & troubleshooting. + +A lot of people in AI will tell you that US- and China-based capacities in all these areas will also be boosted by advanced AI, stripping middle powers of niches. The long history of comparative advantage and the starting positions suggest that’s not a given — especially when both great powers are focused on an AI race. + +![](https://substackcdn.com/image/fetch/$s_!JDXu!,w_474,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f26819-b60c-4204-bf45-71a89e17e592_1977x1411.jpeg)![](https://substackcdn.com/image/fetch/$s_!rD58!,w_474,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11705217-4d60-453d-ab08-eda74317d2e8_1200x1200.webp)![](https://substackcdn.com/image/fetch/$s_!LbcV!,w_474,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa61ecd99-75f0-4b76-820b-ccf4334a42e6_1200x1200.jpeg) + +This is not about country specifics, and the economic policy pathways to realize this are a contentious issue themselves. But **for a short illustration**, imagine a country like Germany or Japan refocused policy efforts to get aggressive about further automating its still-substantial industrial capacity, provide the energy to scale it up quickly if needs be, and removed regulatory constraints on leveraging the troves of data that already exist – and *then* implemented highly advanced AI systems on top of that, buying intelligence at cheap market prices. I’m very optimistic that the returns to the quality and quantity of German or Japanese exports would be remarkable – and beyond most conceivable alternative plans with similar political costs. In simple summary: When getting ready for increasing AI capabilities, **you want to be ready to turn cheap intelligence into valuable growth.** + +## **Here’s How That Helps on Strategy & Security** + +Thinking in bottlenecks for intelligence also provides the **necessary strategic leverage.** Take the now-classic great power competition scenario for ~2028, where China and the US both [chase](https://www.theverge.com/2017/8/3/16007736/china-us-ai-artificial-intelligence) increasing AI capabilities. How does actually prevailing in that scenario realistically look like for either of the great powers? Sure, there is some notion of capability progress where the leader gets to an ASI system that does something mind-blowing like deploy nanobots or provide a skeleton key to the world’s infrastructure; but reasonable people suggest that this might be quite infeasible. + +So I take it the more realistic, albeit less exciting, view is that **getting to higher and higher AI capabilities mostly serves as a [force multiplier](https://www.defenseone.com/sponsors/2024/10/solidify-bedrock-cyber-defenses-ai-force-multiplier/400124/)**: It makes your existing forces more and more effective, it helps you get a lot more out of your industrial capacity, it accelerates your R&D a lot, etc. Winning, in that sense, just means that your economy and military gets a lot more bang for its buck due to some distribution of efficiency gains. If that is true, and the great powers chase greater increases of the force multiplier, then there is comparative advantage in middle powers in contributing to the multiplicand: Increasing the basis of capacity that is made more effective and efficient by the increasingly abundant intelligence. + +Practically speaking: **If your AGI makes all industrial output you apply it to many times as valuable, then any allies that add industrial capacity get you a lot further in your great power competition.** And if the US is already heavily invested into winning an AI race, it stands to reason that it won’t also effortlessly outpace its major allies on industrial capacity — it’s not *that* far ahead in terms of growth and GDP just yet. To make this clear by the inverse point: If the US only had its own industrial capacity to leverage, its models would have to be quite a bit better to make up for the gap to Chinese capacity already. And if it ever fell behind, things would look pretty dicey pretty quickly. Sure, maybe AGI helps you rapidly ramp up your capacity. But maybe it doesn’t. Or maybe you won’t win decisively. Maybe the three months of AGI-advantage don’t ramp it up enough. **Would you bet the house on that if you were sitting on the NSC?** + +[![Bild](https://substackcdn.com/image/fetch/$s_!F1H4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80625925-3f75-43b1-a963-8738145c1546_1148x412.jpeg "Bild")](https://substackcdn.com/image/fetch/%24s_%21F1H4%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/80625925-3f75-43b1-a963-8738145c1546_1148x412.jpeg) + +It’s up to the middle powers to shift the US rationale away from [this](https://www.hoover.org/research/defense-against-ai-dark-arts-threat-assessment-and-coalition-defense). + +So, back to middle powers: Imagine a situation in which the US is barely ahead in a tight capability race and broader geopolitical competition, and wants to cement that position. In this situation, Germany or Japan could be able to offer a ‘superchargeable’ industry to the US to capitalize on capability gains through; countries like the Netherlands could offer even more comprehensive access and AI-powered improvements to compute supply, etc. This might be a sufficiently meaningful contribution to geopolitical competition that the US would be motivated to trade for it: By extending its frontier-powered security umbrella and guaranteeing continued access to compute and civilian model use. And if the middle powers caught up mysteriously, or AI turned out a nothingburger after all, they’d still be left with a stronger strategic position tha before. **Great powers will not extent their hand out of the goodness of their heart, so it’s best to think about what one can trade.** + +## **Here’s What’s Next & How To Sell It** + +In summary, here’s a **quick sketch of a roadmap for a middle power** that’s actually interested in avoiding falling behind on autopilot: + +1. **Don’t get into a race you can’t win.** Get enough compute to run crucial onshore inference, that’s it. +2. **Find something in your economy** that: + + 1. Is not directly AI + 2. Bottlenecks returns of rapid AI progress to real-world effects + 3. Can be argued to have some strategic dimension in great power competition +3. **On economic policy**, strongly prioritize making these sectors competitive and ‘AI-compatible’. This can include getting some compute for inference, but should mostly not be about AI per se. Think in months, not decades. Go into debt if you have to. +4. **On foreign policy**, leverage these capabilities in negotiations with an AI great power to get good, but realistic, terms to partake in compute supply, model access and security umbrellas. Do this as early as you can – the negotiation terms will get worse as AI gets bigger. + +I can’t speak to political dynamics everywhere to the same extent, but by and large I think this sort of plan has the **big advantage of a favourable political economy**. In short: All this **does not require a visible pivot to a tech- or AI-driven** economic model. It can be framed as benefitting powerful incumbents, existing political coalitions, and plausibly at-risk demographics whose transformation shocks are being preempted. On the level of general political sentiment — or ‘vibes’ —, the plan can be framed as a **culturally and politically continuous doubling down** on existing strengths, while being entirely responsive to a paradigm shift. + +Unlike more directly AI-related strategies, this **roadmap avoids being characterised as speculative, risky, or captured by industry** — an historically likely fate for much policy that tries to address upcoming transformation. Add to this the plenty of individual incentives to deny the reality of disruptive AI effects, and you see how easily political resistance can be [motivated](https://antonleicht.substack.com/p/inferencepolitics). Any alternative plan for middle powers that puts a more overtly AI-focused push, whether that’s development at the frontier or downstream deployments below it, at its center is much more susceptible to these political hurdles. Given the currently very low uptake of ‘taking AI seriously’ in most of these middle powers, I think this political economy aspect matters. + +**AI middle powers are in big trouble and risk being left behind** on economic benefits and security applications of frontier capabilities. While they might not catch up, they still have options: They can find sectors that benefit from rapid AI progress, equip them for that progress, and leverage them as comparative advantage in negotiations with great powers. I think this is a somewhat robust plan with a somewhat favourable political economy that might be worth considering. + +Subscribe + +1 + +Share + +TopLatest + +No posts + +### Ready for more? + +Subscribe + +© 2026 Anton Leicht · [Privacy](https://substack.com/privacy) ∙ [Terms](https://substack.com/tos) ∙ [Collection notice](https://substack.com/ccpa#personal-data-collected) + +[Start your Substack](https://substack.com/signup?utm_source=substack&utm_medium=web&utm_content=footer)[Get the app](https://substack.com/app/app-store-redirect?utm_campaign=app-marketing&utm_content=web-footer-button) + +[Substack](https://substack.com) is the home for great culture + +This site requires JavaScript to run correctly. Please [turn on JavaScript](https://enable-javascript.com/) or unblock scripts diff --git a/sources/resources/leicht-safety.md b/sources/resources/leicht-safety.md new file mode 100644 index 0000000..63f1583 --- /dev/null +++ b/sources/resources/leicht-safety.md @@ -0,0 +1,215 @@ +--- +url: https://writing.antonleicht.me/p/how-ai-safety-is-getting-middle-powers +--- + +[![Threading the Needle](https://substackcdn.com/image/fetch/$s_!4SKU!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dff89d-5b66-4169-8ef8-eb2ec9ed94e8_1024x1024.png)](/) + +# [Threading the Needle](/) + +SubscribeSign in + +# How AI Safety Is Getting Middle Powers Wrong + +### The case for pivoting from global governance to national interests + +[![Anton Leicht's avatar](https://substackcdn.com/image/fetch/$s_!FPyB!,w_36,h_36,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75422da7-aafa-42ab-8fa6-cf4f0df85cf0_3166x3166.jpeg)](https://substack.com/%40antonleicht) + +[Anton Leicht](https://substack.com/%40antonleicht) + +Jan 22, 2026 + +10 + +Share + +**Geopolitical realities are shifting, and leave most of the world in a tough spot.** That much will become clearer still once the current fault lines are exacerbated by the advent of advanced AI, developed by two great powers alone. Few who take this prospect seriously care about the fate of middle powers; technologists, accelerationists, national security types have moved to US and China policy. Yet one group remains well-positioned to contribute to the middle power question: AI safety advocates.[1](#footnote-1) + +We face a [Gaullist moment](https://www.cbc.ca/news/politics/mark-carney-speech-davos-rules-based-order-9.7053350) of reawakening national ambition in many middle powers. It provides the political preconditions to advocate for ambitious AI strategies the world over. The international branches of the safety movement are well-positioned to seize that moment: from their rare position of expertise on AI, and with a strong national foothold in many middle powers, they could pivot to improving the response of middle powers to the deployment and diffusion of AI technology: building misuse resilience, guarding against economic disempowerment, and making sure the world order in which advanced AI emerges is a stable one. + +Subscribe + +Instead, **much of the safety community is stuck answering the wrong questions.** They think middle powers can and should exert substantial influence on the *development* of frontier AI systems in the US and China – a proposition that was always contentious, but is untenable in the world as it is today. By sticking to it, they’re hurting their own credibility in middle powers that are now more attuned to their own national interest; and undermining the perception of AI safety work in the US. If safety advocates in middle powers abandoned the unrealistic pursuit of leverage and global governance, they could **pivot to making AI deployment go well in middle powers** – where this important work is tractable and neglected. + +[![](https://substackcdn.com/image/fetch/$s_!KJpY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2288f183-9076-47b5-b248-884f54abe762_1000x250.png)](https://substackcdn.com/image/fetch/%24s_%21KJpY%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2288f183-9076-47b5-b248-884f54abe762_1000x250.png) + +# **A Platform In Review** + +**Safety advocates have largely engaged with the middle power conversation from a US- and China-centric point of view:** they assume that critical risks emerge from the development of AI systems, and since that development happens in the US and China, their engagement in middle powers must aim at affecting these two great powers. Some readers may respond that domestic regulation does not fundamentally target development, just market entry conditions. That’s a fair argument about the legal principle, but the compelling force of the market and the infeasibility of developing wholly separate AI products for it makes it practically moot. When meeting market entry conditions for a vital market requires changing how you develop, it is in effect development-facing regulation – and in fact, this is arguably by design as the intended consequence of a ‘[Brussels Effect](https://www.governance.ai/research-paper/brussels-effect-ai)’. + +The variations on this theme are many: AI safety organisations have tried to pass safety-focused regulation in the EU, UK, and elsewhere; they have lobbied middle power governments to influence the US and even China to take AI risks more seriously; they are developing tools of leverage for middle powers to strongarm the US into safety concessions; and they are seeking signatories to a multilateral treaty. The core hypothesis is always the same: maybe one can use middle powers as a lever to affect US development. I believe this thesis was always somewhat flawed, but it is now growing untenable. There are two problems with the approach. + +#### ***It Doesn’t Really Work*** + +First, past attempts to influence development from outside are faltering. Using domestic law to [constrain foreign developers](https://scholarship.law.columbia.edu/faculty_scholarship/271/) has found its landmark example in the [EU AI Act](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai), which is not doing too well: the broader law is a famously unpopular piece of legislation, its enforcement has been [delayed numerous times](https://www.euronews.com/my-europe/2025/11/19/european-commission-delays-full-implementation-of-ai-act-to-2027), and its application to the cutting edge of AI technology is evermore uncertain: the technology moves fast, the commission moves slowly, and the regulation is necessarily a snapshot of 2023’s best and worst assumptions about AI. The weaknesses of the overall act apply least readily to its most safety-coded aspects concerning general-purpose AI, and the safety and security chapter of the [Code of Practice](https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai) is a very effective implementation tool for them. My reservations have to do with the broader economic and geopolitical setting, where substantial future enforcement seems unrealistic, should it become actually burdensome: whenever any stipulation turns out too annoying for a frontier developer, they can quickly enlist the US administration and their own economic leverage to skirt the Code. A safetyist might say the key outcome was the transparency pathways set in place; I’m unsure how great the marginal gain over [SB-53](https://carnegieendowment.org/emissary/2025/10/california-sb-53-frontier-ai-law-what-it-does) is compared to the effort, and arguing this difference warranted all that effort seems extraordinary.[2](#footnote-2) + +--- + +**International policy fares even worse.** I’ve written at length about international policy attempts up to [last year](https://writing.antonleicht.me/p/the-early-death-of-international), which leaves me to discuss the more recent call for [Red Lines](https://red-lines.ai/) at the UN. While safety advocates view this as an achievement, it strikes me that most powerful decisionmakers on AI haven’t even heard of it. For good reason: the UN—whether as a body or a collection of countries—has no real power it can exert on an issue if a major power doesn’t want it to. Safety advocates know this, but some view the Red Lines being in place as valuable nevertheless: once political salience arrives, they argue, the Red Lines provide a framework to hold onto. I don’t have much new to say on this – I see no way the US government lets any other country dictate what it does on AI. It’s too polarised an issue, too easy a domestic political win, and too much of a strategic imperative to exact absolute sovereignty over domestic development. A similar fate befalls the [call](https://superintelligence-statement.org/) for international prohibition of superintelligence: in both cases the burden of proof is on the safetyists to explain how this ever happens. + +[![The Early Death of International AI Governance](https://substackcdn.com/image/fetch/$s_!vymH!,w_140,h_140,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9c65657-9e64-4996-9628-25b58070fe39_1600x1222.jpeg) + +#### The Early Death of International AI Governance + +[Anton Leicht](https://substack.com/profile/113003310-anton-leicht) + +· + +March 31, 2025 + +[Read full story](https://writing.antonleicht.me/p/the-early-death-of-international)](https://writing.antonleicht.me/p/the-early-death-of-international) + +Most have never publicly tried – with perhaps the sole exception of a recent [paper](https://asi-prevention.com/) by prominent AI safety organisations Conjecture and ControlAI, which outlines an admirably honest account of what it would take: middle powers in pursuit of a global mission, outright leveraging the threat of sabotaging critical digital infrastructure against an erstwhile ally. The paper correctly identifies the required scope, but also demonstrates the lack of feasibility: it’s an all-in strategy with no off-ramps and great rewards for defection for any of the involved middle powers, and in a politically heterogeneous setting, it seems like an outright impossible sell. If this is the best actual way to affect development, that’s bad news for the development-focused strategy. + +--- + +**All this will grow even less effective over time.** The specific reason is that legislation leveraging the consumer market for outside effect works best the first time you try it – sneak provisions in on a low, technocratic level, slowly build up compliance frameworks and path dependencies that entrench your law and make it last even against political headwinds in the future. This is perhaps the success story of the GDPR; but with every iteration, the political headwinds happen earlier, and the entrenchment is less and less resilient to opposition from the markets you seek to regulate. The general version of this is that development-focused regulation from abroad works best in a low-salience environment, outside the direct purview of the public and political decision-makers, negotiated and accepted by mid-level technocrats who converge on substantive policy thought. + +[![](https://substackcdn.com/image/fetch/$s_!ZMN7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb406f3-a596-45af-9e0a-18374f90b349_1600x421.png)](https://substackcdn.com/image/fetch/%24s_%21ZMN7%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/adb406f3-a596-45af-9e0a-18374f90b349_1600x421.png) + +This is the [geopolitical backdrop](https://www.nbcnews.com/tech/tech-news/us-rejects-international-ai-oversight-un-general-assembly-rcna233478) we are working with for the foreseeable future. Both OSTP Director Kratsios’ and President Trump’s comments at the UN should be required reading. + +**We’ve now rapidly left this low-salience environment in two unrelated ways.** First, AI is no longer low-salience: governments around the world have identified it as a geopolitically and economically important issue, and they’re hesitating to hand it over to peripheral decisionmakers and institutions. For instance, I suspect that neither the American nor the European electorate will be particularly happy to trust the European Union on this one. Second, international technology regulation is no longer low-salience. The US administration is keeping a watchful eye on who and what compels their technology companies, and European powers are increasingly careful in choosing their battles, because while they might be gearing up for conflict on some fronts, they know they have to fear outright retaliation. I don’t think anything is turning back the clock on this – even a Democratic administration would not be enthusiastic to reaffirm the Europeans to please continue regulating their AI companies after all. The heyday of governing the US from outside its borders is over. + +[![](https://substackcdn.com/image/fetch/$s_!KJpY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2288f183-9076-47b5-b248-884f54abe762_1000x250.png)](https://substackcdn.com/image/fetch/%24s_%21KJpY%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2288f183-9076-47b5-b248-884f54abe762_1000x250.png) + +#### ***In Fact, It Hurts*** + +**Development-focused international work is also increasingly unhelpful to the broader safetyist agenda.** In the US and many techno-optimistic countries, AI safety concerns have grown closely associated with the vehicles and alliances safety advocates have chosen abroad. Insofar as the US administration rejects international governance or the American electorate rejects European policy approaches to technology, there’s increasing opposition to the safety agenda too. Today, safety advocates in America often have to defend against the allegation that they seek to import European and internationalist notions of governance. Being branded as ‘European-style’ is not helpful in the current political environment. + +The reputational harm also affects international governance more generally. AI safety has risen in political salience, even if it hasn’t kept up in political power, and what safetyists say they want sticks in decisionmakers’ minds. International governance was never going to be the main way AI gets regulated. But safetyist organisations publicly announcing that they were using national laws and international treaties to externally constrain the US has made it much harder to get treaties that would also have this incidental effect. Instead of letting a gentle Brussels effect proliferate quietly, the strategy has been broadcast loudly and led to more forceful rejections of any attempt at international governance. + +--- + +**The focus on AI development frequently leads safety advocates to work against the narrow national interest of their respective countries.** It seeks to leverage resources of a middle power not to pursue that middle power’s unique goals, but a broader altruistic mission. That’s a problem, because it weakens the safety movement’s standing in national environments. This is especially true today, when middle power leadership frequently views its own attempts to export values through legislation as a costly mistake of the now-bygone ‘end of history’. A local safety movement whose theory of change runs through leveraging local resources for global solutions with indirect national benefits at best risks political sidelining in today’s environment of national interest. The most obvious issue is national regulation, which asks countries to spend their precious leverage against the US on frontier regulation to benefit the world. But it relates to other narrow interests too – consider how safety advocates have positioned themselves on European compute buildout and frontier model development. They’ve largely been hesitant to support this, not for the sound reason that it’s economically unwise, but for the bad reason that it might create another party to the AI race. With friends like these… + +The Conjecture/ControlAI paper again is admirably honest in this. To be clear, I’m not harping on the paper itself – it’s entirely realistic in its suggestions, in that they are the only way to actually *do* what material like the superintelligence statement seems to *want*. Don’t even imagine it was enacted – just imagine it got traction and became a prominent piece of safety advocacy. It would cause great offense on all sides of the conversation. Middle power governments, strategically and fiscally overextended and desperate to retain their standing, with no spare resources to invest in global missions, would be very skeptical of the advice of a movement that wants them to sacrifice their national interests in service of a greater good.[3](#footnote-3)To the US as the target of this mission, it paints the safety movement as interested in regulating American economic activity by sneaking past the US electorate – this will correctly be read as outright offensive. Luckily, this perception hasn’t yet spread – but if the safety movement doubles down on this as some viable pathway to multilaterally affect development today, a reputation hit is only one or two media campaigns away. + +--- + +#### ***Real Trade-Offs*** + +You might dismiss all this and say that the odds are long and the costs are high, but AI safety is so important that this is still the most effective work you can do in middle powers. My conviction is that this is not true, and that the economic and strategic questions are themselves of utmost importance. But even if you disagree, I think **political and financial capital for AI safety is better spent elsewhere: on advancing the national and collective AI strategy of middle powers writ large.** This directly trades off against the development-focused approach for two reasons. + +The most obvious axis of trade-off is simply limited talent and funding – the safety movement is only so big, only has so many capable operatives, only so much breadth of talent pipelines, and so on, so simply adding the type of work I’ll suggest to the current portfolio seems unlikely to work. + +The second, more important trade-off relates to credibility. Due to the political ramifications of endorsing safetyist domestic regulation, safety advocates lose influence on a range of other issues they’d otherwise be helpful voices on. They could be leading voices on national sovereignty, AI strategies, beneficial deployment, and downstream resilience; but they are much less than they could be, not least because they still operate in the reputational shadow of their attempts to trade off their nations’ interests in favour of the greater good. Some organisations have tried to change this perception, pivoting toward a more sovereignty-focused approach to middle power work. But the reputational effects are great enough that these organisations often operate quietly and dodge affiliation with the rest of the ecosystem. As a result, **the more absurd versions of safety middle power work gain disproportionate attention**, worsening the reputational problems even further. All this is a far cry from an AI safety ecosystem that confidently owns tractable safety-relevant issues in middle powers. They could lead this new conversation if they focused on it – and I believe they should do so. + +[![](https://substackcdn.com/image/fetch/$s_!KJpY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2288f183-9076-47b5-b248-884f54abe762_1000x250.png)](https://substackcdn.com/image/fetch/%24s_%21KJpY%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2288f183-9076-47b5-b248-884f54abe762_1000x250.png) + +# **Why Middle Powers?** + +**Even if middle powers will never influence AI development, I believe working on middle power policy is valuable.** I also believe this for many reasons that don’t have much to do with AI safety in the narrow sense, but instead with capturing and proliferating AI’s benefits throughout the world. But I’m making a case to safety advocates – so I’ll focus on the safety-relevant reasons to work on middle power policy today. + +[![](https://substackcdn.com/image/fetch/$s_!I2Bz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8cebe1-3373-413f-9ef6-e25a07f68584_1280x527.png)](https://substackcdn.com/image/fetch/%24s_%21I2Bz%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/8f8cebe1-3373-413f-9ef6-e25a07f68584_1280x527.png) + +‘And on AI we’re co-operating with like-minded democracies to ensure we won’t ultimately be forced to choose between hegemons and hyperscalers.’, says Canadian PM Carney in a much-lauded recent [speech](https://www.cbc.ca/lite/story/9.7053350). + +--- + +**First, middle powers are where harms from AI misuse might manifest earliest and most dramatically**: their governments are less engaged on AI, so they might fail to adopt defensive measures in homeland and national security in time; but criminals will view rich middle power populations as targets for exploitation, terrorism, and blackmail. This makes work on strengthening misuse resilience in middle powers important and promising: if the threat were made clear to middle power governments, they’d have little choice but to invest in defensive measures with positive spillover effects. If you get the US worried about misuse, they might try to regulate the models; but the EU can’t regulate the models, so if guided well, they’d have to respond by funding defensive measures instead. + +One promising area is *[differential](https://vitalik.eth.limo/general/2025/01/05/dacc2.html#1), [defensively focused](https://www.joinef.com/posts/introducing-def-acc-at-ef/) [acceleration](https://ifp.org/the-launch-sequence/)*: outpacing the development of potentially harmful AI capabilities by hastening the development of technologies that guard against these harms. This is in the interest of many middle powers, both because proliferation of harmful capabilities seems a foregone conclusion and because there is a political window opening. Many middle powers, alarmed by the ongoing rearrangement of geopolitical realities, are building out militaries and relying less on American and Chinese beneficence. The ongoing [rearmament](https://www.consilium.europa.eu/en/policies/defence-numbers/) of Europe (and, to a lesser extent, East Asian countries) provides an obvious context for ambitious projects in this vein. But they need to be framed as linked to national interest. The current pitch suffers from the safety movement’s reputation for prioritising global over national interests: no middle power wants to bankroll tech that lets US AI development speed ahead. But framed as a reaction to the uncontrollable trajectory of frontier AI development, resilience-focused innovation measures could become a highly successful part of middle power policy. + +--- + +**Second, middle powers are where gradual disempowerment and widespread destitution seem most plausible.** In an AI great power, there are policy backstops against disempowerment: taxation and redistribution are possible, and the government can intervene. This plays out differently in a middle power – if AI rips through the workforce, moving revenue and power from domestic workers to US AI labs, there’s little they can do. Discussions about [growth divergence](https://econofact.org/factbrief/fact-check-has-the-economic-gap-between-europe-and-the-united-states-increased-in-the-past-decade) and the ‘Europoors’ are ruefully amusing at 3% growth differences, but become existential at 10+% growth differences. The short version: if a risk of ‘millions end up poor and destitute’ is substantial enough to motivate safetyists, it should motivate them to fix the economic trajectory of middle powers as it relates to AI. + +[![AI, Jobs, and the Rest of the World](https://substackcdn.com/image/fetch/$s_!SEXO!,w_140,h_140,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c9a9985-a994-4284-b4a2-e602eb13efb9_1600x1087.png) + +#### AI, Jobs, and the Rest of the World + +[Anton Leicht](https://substack.com/profile/113003310-anton-leicht) + +· + +September 9, 2025 + +[Read full story](https://writing.antonleicht.me/p/ai-jobs-and-the-rest-of-the-world)](https://writing.antonleicht.me/p/ai-jobs-and-the-rest-of-the-world) + +--- + +**Third, middle powers getting AI wrong can lead to a destabilised world** – making the path to superintelligence more dangerous. The most salient aspect is the threat of conflict. There’s already considerable geopolitical volatility, but it can always get worse. The asymmetrical diffusion of new strategic technologies has often triggered dormant conflicts and exacerbated existing ones: parties attack because they believe themselves at a temporary advantage, or inversely because they think the window to compete is closing. If dormant conflicts suddenly erupt as a result of jagged diffusion of advanced AI, that can quickly destabilise everything from supply chains to a shaky US-China peace. + +But it doesn’t have to go to war. If middle powers feel sufficiently threatened by advanced AI – either by the economic effects or by power shifting toward AI developers – they might attempt to halt AI development. Countries with [substantial positions in the semiconductor supply chain](https://www.state.gov/pax-silica) have real leverage here. Using it would be economically suicidal – but if their next few years go badly enough, they might feel they have no other choice. It will be difficult enough to get the transition to transformative AI right under the best conditions – and the current inadequacy of middle power strategies will destabilise conditions substantially. + +[![The Most Dangerous Time in AI Policy](https://substackcdn.com/image/fetch/$s_!I4jL!,w_140,h_140,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc4f5598-e3c3-43fa-8611-1e906cc6b0ff_1024x640.png) + +#### The Most Dangerous Time in AI Policy + +[Anton Leicht](https://substack.com/profile/113003310-anton-leicht) + +· + +March 5, 2025 + +[Read full story](https://writing.antonleicht.me/p/the-most-dangerous-time-in-ai-policy)](https://writing.antonleicht.me/p/the-most-dangerous-time-in-ai-policy) + +--- + +Fourth, it follows that **contributing to national AI strategies in middle power governments should be a safetyist priority itself** – I believe even if this work reduces these governments’ immediate interest in safety-focused regulation. The above are all safetyist talking points, but they slot into – and require – a broader strategic conversation in middle powers:taking seriously the transformative potential of AI, and grappling with the geopolitical and technical implications. The questions that arise are ones safetyists can answer, and the answers will reduce substantial risks from advanced AI. Not only should answering these questions be a safetyist focus area, but making sure they’re asked as well. It’s an advantage to AI safety when middle powers think clearly about AI strategy, and so it’s worth contributing to that strategic clarity itself. + +All this even opens a potential door back to the development focus – if there is some way to **bring middle powers into a position of strength, they might once again affect AI development,** serve as a real third center of gravity on the question of advanced AI. But the path to this runs through the effective and uncompromising pursuit of their national interest. It will not be reached from a position of economic and strategic weakness they are headed for on their current trajectories. + +[![](https://substackcdn.com/image/fetch/$s_!KJpY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2288f183-9076-47b5-b248-884f54abe762_1000x250.png)](https://substackcdn.com/image/fetch/%24s_%21KJpY%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/2288f183-9076-47b5-b248-884f54abe762_1000x250.png) + +# **What Can You Do?** + +The safety movement can contribute to answering these questions: they require expertise, awareness of AI’s scale, the ability to assemble effective policy operations, and a desire to get this right. Few other players combine these factors: geopolitical and national strategy researchers who take advanced AI seriously are rare, organisations willing to host them rarer still, and these few places lack scale and funding. On these criteria, the safety movement is a capable player, and it could start mobilising resources toward this goal. + +**To the safety movement’s credit, this is already happening in some places.** European policy organisations, for instance, are doing substantial work advancing middle powers’ strategic agendas. But none of this is endorsed from the top, none of it has made its way into the mainstream view of what matters – none of it is reflected in big funding or talent pipelines or presented as a key pathway to safety-minded impact. To grapple with the reputational dimension and deploy resources at scale, a decisive pivot is needed. + +--- + +**Existing policy organisations** would not be the main drivers of this process. Some are positioned to pivot, and pivot they should – including by clearly breaking with their past choices and proposals. The world is changing rapidly, and anyone who says ‘we get it now’ and starts pulling for the sovereign fate of middle powers will be welcomed. That said, the safety movement has gotten into trouble for quick pivots and perceived two-facedness in the past, and it’s probably not realistic for deeply entrenched organisations to reverse course completely. In fact, it’s probably good if some safety organisations with strong positions on this stay where they are: to catch the true believers in global-mission thinking about safety in middle powers, and to promote internal disagreement. It’s probably good if there are organisations to which new middle power orgs can point and say ‘we strongly disagree with them, and we represent a different approach to international safety policy’. + +--- + +**The research ecosystem can do more.** Established scholars already work closely on national sovereignty, but until recently, their work on strategic questions has always been wrapped in layers of interpretation that contextualised it with some development-focused AI safety point. Giving these researchers freedom and encouragement to pursue these questions without a predetermined takeaway, suggesting that whatever strategic pathway they find is of value, could unlock substantial resources. These people have the ideas – make it clear it’s part of the mission and not taboo to get them out there. + +--- + +But the even more promising aspect of the research ecosystem is **talent pipelines.** Until recently, major safety-aligned mentorship programs have mostly focused their project selection in international governance on this development-focused approach. If these programs offered a starting point for middle-power-focused researchers, I’m confident more people would choose this path. The number of bright, ambitious people who understand AI and want to help their home countries, but end up defecting to US-focused work or contorting into development-focused safety work, is staggeringly high. Give them a home. + +--- + +Much of this **comes down to funding.** Major funders could publicly prioritise these cause areas; launch RFPs around them; and incubate organisations that tackle them. In the most ambitious version, they apply the same strategy to the middle power space that they have to the national security conversation, where safety-aligned organisations have made keystone grants to top-tier institutions and cultivated deep expertise between old-school policy hands and capable, entrepreneurial safety advocates who brought cutting-edge knowledge to enrich the discussion. I’d welcome the same for middle power work. Right now, there simply isn’t enough space to do work that grapples with AI and its strategic implications for middle powers. It’s in the safety movement’s interest to step up, and the world would be thankful for it. + +[![](https://substackcdn.com/image/fetch/$s_!844s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f28fc73-ecd5-411a-b329-a706d7c6089c_2142x345.png)](https://substackcdn.com/image/fetch/%24s_%21844s%21%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep/https%3A//substack-post-media.s3.amazonaws.com/public/images/8f28fc73-ecd5-411a-b329-a706d7c6089c_2142x345.png) + +This is the limited extent of non-US AI policy work requested in Coefficient Giving’s headline AI governance [RFP](https://coefficientgiving.org/funds/navigating-transformative-ai/request-for-proposals-ai-governance/), for instance. + +**The safety movement has the people, the institutions, and the resources.** **What it lacks is the right theory of change for middle powers.** The development-focused approach was always a long shot; today it’s actively harmful. The alternative – helping middle powers navigate AI deployment, build resilience, and avoid strategic blunders – is tractable, neglected, and would actually advance safety. The moment for that is now. Seize it with haste. + +Subscribe + +[1](#footnote-anchor-1) + +By which I mean the broad AI safety ecosystem – funders, researchers, policy organisations that primarily focus on substantial and catastrophic risks from very advanced AI systems. It pains me to say that [this polemic](https://www.aipanic.news/p/the-ai-existential-risk-industrial) provides a good overview. + +[2](#footnote-anchor-2) + +Some safety advocates will say this was them making the best of a bad situation – that the AI Act would have happened anyways, and they sought to improve it. I doubt this, both because safety advocates played a substantial role in passing it and because their opposition would surely have been counterfactually influential – but it matters little for the forward-looking evaluation. + +[3](#footnote-anchor-3) + +Yes, I know that some safety advocates believe the reduction of existential risk is also the primary national priority of any smaller government. But even if that’s true, no one’s gotten very far in selling it. + +10 + +Share + +TopLatest + +No posts + +### Ready for more? + +Subscribe + +© 2026 Anton Leicht · [Privacy](https://substack.com/privacy) ∙ [Terms](https://substack.com/tos) ∙ [Collection notice](https://substack.com/ccpa#personal-data-collected) + +[Start your Substack](https://substack.com/signup?utm_source=substack&utm_medium=web&utm_content=footer)[Get the app](https://substack.com/app/app-store-redirect?utm_campaign=app-marketing&utm_content=web-footer-button) + +[Substack](https://substack.com) is the home for great culture + +This site requires JavaScript to run correctly. Please [turn on JavaScript](https://enable-javascript.com/) or unblock scripts diff --git a/sources/resources/national-ai-plan.md b/sources/resources/national-ai-plan.md new file mode 100644 index 0000000..3a9ae32 --- /dev/null +++ b/sources/resources/national-ai-plan.md @@ -0,0 +1,491 @@ +--- +url: https://www.industry.gov.au/publications/national-ai-plan +--- + +You have Javascript disabled and your user experience may not be as intended without enabling it. Please consider enabling Javascript or email us at digital [at] industry.gov.au for further assistance. + +* [Skip to main content](#heading-1) +* [Skip to search](#search-website-link) +* [Skip to main website navigation](#website-navigation) + +[![Australian Government, Department of Industry Science and Resources](/themes/custom/industry_theme/logo.png)](/ "Home") + +## Mega menu + +* [Who we are](/) + + + Entities + - [AI Safety Institute](/science-technology-and-innovation/technology/artificial-intelligence/ai-safety-institute) + - [Anti-Dumping Commission](/anti-dumping-commission) + - [Anti-Dumping Review Panel](/trade/anti-dumping-review-panel) + - [AusIndustry](/science-technology-and-innovation/industry-innovation/ausindustry) + - [Australian Industry Participation Authority](/major-projects-and-procurement/australian-industry-participation) + - [Australian Radioactive Waste Agency](/australian-radioactive-waste-agency) + - [Australian SKA Office](/science-technology-and-innovation/space-and-astronomy/ska-project-australia/australian-ska-office) + - [Australian Space Agency](https://www.space.gov.au/) + - [Australian Space Discovery Centre](https://www.space.gov.au/australian-space-discovery-centre) + - [Critical Minerals Office](/mining-oil-and-gas/minerals/critical-minerals/critical-minerals-office) + - [Industry Innovation and Science Australia](/science-technology-and-innovation/industry-innovation/industry-innovation-and-science-australia) + - [Investor Front Door](/investor-front-door) + - [Major Projects Facilitation Agency](/major-projects-and-procurement/major-projects-facilitation-agency) + - [National AI Centre](/science-technology-and-innovation/technology/artificial-intelligence/national-ai-centre) + - [National Measurement Institute](/national-measurement-institute) + - [National Offshore Petroleum Titles Administrator](https://www.nopta.gov.au/) + - [Office of Supply Chain Resilience](/trade/office-supply-chain-resilience) + - [Office of the Chief Economist](/office-chief-economist) + - [Office of the Chief Scientist](https://www.chiefscientist.gov.au/) + - [Offshore Decommissioning Directorate](/mining-oil-and-gas/oil-and-gas/offshore-decommissioning-directorate) + - [Questacon](https://www.questacon.edu.au/) + + [Ministers](https://www.minister.industry.gov.au/) + + [People](/people) + - [Organisation chart](/people/organisation-chart) + - [Overseas counsellors](/people/overseas-counsellors) + + [Portfolio agencies and boards](https://www.directory.gov.au/portfolios/industry-science-and-resources) + + [Public consultation](https://consult.industry.gov.au/) + + [Accessing information](/accessing-information) + - [Campaigns](/accessing-information/information-and-advertising-campaigns) + - [Freedom of information](/accessing-information/freedom-information-disclosure) + - [Gifts and benefits](/accessing-information/gifts-and-benefits) + - [Grants reporting](/accessing-information/grants-administration) + - [Indexed list of files](/accessing-information/indexed-list-files) + - [Portfolio appointments and vacancies](/accessing-information/portfolio-appointments-and-vacancies) + - [Privacy impact assessments](/accessing-information/privacy-impact-assessments) + + [Careers](/careers) + - [Current vacancies](https://industrycareers.nga.net.au/) + - [Graduates](/work-us/careers-graduates) + + [Corporate governance](/corporate-governance) + - [Annual report](/corporate-governance/annual-reports) + - [Budget statements](/corporate-governance/budget-statements) + - [Corporate plans](/corporate-governance/corporate-plans) + - [Fraud and corruption control](/corporate-governance/fraud-and-corruption-control) + - [Public interest disclosure](/corporate-governance/public-interest-disclosure) + - [Regulatory performance](/corporate-governance/regulatory-performance) + + [Government to government](/government-government) + - [Australian Business Licensing and Information Service](/government-government/australian-business-licensing-and-information-service) + - [business.gov.au](/government-government/businessgovau) + - [Business Grants Hub](/government-government/business-grants-hub) + - [Business Registration Service](/government-government/business-registration-service) + - [Shared Services Hub](/government-government/shared-services-hub) + - [VANguard](/government-government/vanguard) +* [Topics](/) + + + [Major projects and procurement](/major-projects-and-procurement) + - [Australian industry participation](/major-projects-and-procurement/australian-industry-participation) + - [Investor Front Door](/investor-front-door) + - [Major projects facilitation](/major-projects-and-procurement/major-projects-facilitation-agency) + + [Manufacturing](/manufacturing) + - [Batteries](/manufacturing/batteries) + - [Green metals](/manufacturing/green-metals) + + [Measurement](/national-measurement-institute) + - [Australia's measurement system](/national-measurement-institute/australias-measurement-system) + * [Measurement standards](/national-measurement-institute/australias-measurement-system/measurement-standards) + * [Pattern approval](/national-measurement-institute/nmi-services/pattern-approval) + * [Trade measurement](/national-measurement-institute/trade-and-industry) + - [Awards](/national-measurement-institute/about-us/annual-measurement-awards) + - [Chemical and biological measurement services](/national-measurement-institute/chemical-and-biological-measurement-services) + - [Physical measurement services](/national-measurement-institute/physical-measurement-services) + - [Testing and analytical services](/national-measurement-institute/testing-and-analytical-services) + - [Training and assessment](/national-measurement-institute/measurement-training-and-assessment) + + [Mining, oil and gas](/mining-oil-and-gas) + - [Minerals](/mining-oil-and-gas/minerals) + * [Critical minerals](/mining-oil-and-gas/minerals/critical-minerals) + * [Critical Minerals Office](/mining-oil-and-gas/minerals/critical-minerals/critical-minerals-office) + * [Offshore minerals](/mining-oil-and-gas/minerals/regulating-offshore-mineral-exploration-and-mining) + * [Ranger Uranium Mine](/mining-oil-and-gas/minerals/regulating-ranger-uranium-mine) + - [Oil and gas](/mining-oil-and-gas/oil-and-gas) + * [Domestic gas supply](/mining-oil-and-gas/oil-and-gas/securing-australian-domestic-gas-supply) + * [Offshore oil and gas](/mining-oil-and-gas/oil-and-gas/offshore-oil-and-gas) + * [Offshore decommissioning](/mining-oil-and-gas/oil-and-gas/funding-and-investment-opportunities-offshore-decommissioning-australia) + * [Offshore Decommissioning Directorate](/mining-oil-and-gas/oil-and-gas/offshore-decommissioning-directorate) + * [Decommissioning the Northern Endeavour](/mining-oil-and-gas/oil-and-gas/offshore-oil-and-gas/decommissioning-northern-endeavour) + - [Greenhouse gas storage](/mining-oil-and-gas/offshore-greenhouse-gas-storage) + * [Acreage release process](/mining-oil-and-gas/offshore-greenhouse-gas-storage/offshore-greenhouse-gas-storage-acreage-release-process) + - [Taxes, royalties and export controls](/mining-oil-and-gas/taxes-royalties-and-export-controls-minerals-and-petroleum) + * [Combination certificates](/mining-oil-and-gas/taxes-and-royalties-minerals-and-petroleum/applying-combination-certificates-petroleum-projects) + * [Rough diamonds](/mining-oil-and-gas/taxes-and-royalties-minerals-and-petroleum/applying-export-rough-diamonds) + * [Uranium and controlled ores](/mining-oil-and-gas/taxes-and-royalties-minerals-and-petroleum/applying-export-uranium-and-controlled-ores) + + [Radioactive waste management](/australian-radioactive-waste-agency) + + [Science, technology and innovation](/science-technology-and-innovation) + - [Diversity in STEM](/science-technology-and-innovation/diversity-in-stem) + - [Industry innovation](/science-technology-and-innovation/industry-innovation) + * [AusIndustry](/science-technology-and-innovation/industry-innovation/ausindustry) + * [Business Research and Innovation Initiative](/science-technology-and-innovation/industry-innovation/business-research-and-innovation-initiative) + * [Critical Technologies Challenge Program](/science-technology-and-innovation/industry-innovation/critical-technologies-challenge-program) + * [Industry Innovation and Science Australia](/science-technology-and-innovation/industry-innovation-and-science-australia) + * [Industry Growth Program](/science-technology-and-innovation/industry-innovation/industry-growth-program) + * [Research and Development Tax Incentive](/science-technology-and-innovation/industry-innovation/research-and-development-tax-incentive) + - [International science collaboration](/science-technology-and-innovation/international-collaboration-science-and-research) + * [Global](/science-technology-and-innovation/international-collaboration/global-science-and-technology-diplomacy-fund) + * [Europe](/science-technology-and-innovation/international-collaboration/horizon-europe) + * [India](/science-technology-and-innovation/international-collaboration/collaborating-india-science-and-research) + * [China](/science-technology-and-innovation/international-collaboration/collaborating-china-science-and-research) + * [Vietnam](/science-technology-and-innovation/international-science-collaboration/australia-vietnam-science-technology-and-innovation-cooperation-initiative) + - [Science engagement](/science-technology-and-innovation/science-engagement) + * [Australian Space Discovery Centre](https://www.space.gov.au/australian-space-discovery-centre) + * [Measurement awards](/national-measurement-institute/about-us/annual-measurement-awards) + * [Prime Minister's Prizes for Science](/science-technology-and-innovation/science-engagement/prime-ministers-prizes-science) + - [Space and astronomy](/science-technology-and-innovation/space-and-astronomy) + * [Australian Space Agency](https://www.space.gov.au/) + * [SKA](/science-technology-and-innovation/space-and-astronomy/co-hosting-ska-telescope) + * [Optical astronomy](/science-technology-and-innovation/space-and-astronomy/optical-astronomy-australia) + - [Technology](/science-technology-and-innovation/technology) + * [Artificial intelligence](/science-technology-and-innovation/technology/artificial-intelligence) + * [AI Safety Institute](/science-technology-and-innovation/technology/artificial-intelligence/ai-safety-institute) + * [National AI Centre](/science-technology-and-innovation/technology/artificial-intelligence/national-ai-centre) + * [Quantum](/science-technology-and-innovation/technology/quantum) + * [Robotics](/science-technology-and-innovation/technology/robotics) + * [Critical technology standards](/science-technology-and-innovation/technology/critical-technology-standards) + + [Trade](/trade) + - [Anti-dumping](/anti-dumping-commission) + - [Anti-dumping reviews](/trade/anti-dumping-review-panel) + - [Country of origin labelling](/trade/country-origin-labelling) + - [Exporting resources](/mining-oil-and-gas/taxes-royalties-and-export-controls-minerals-and-petroleum#container-7529) + - [Free trade agreements](/trade/growing-australian-industries-through-trade) + - [Reducing technical barriers](/trade/reducing-technical-barriers-trade) + - [Standards and conformance](/trade/australias-standards-and-conformance-infrastructure) + - [Steel tariff rate quotas](/trade/indonesia-steel-tariff-rate-quota-certificates) + - [Supply chain resilience](/trade/office-supply-chain-resilience) + - [Global and domestic supply chain status](/trade/global-and-domestic-supply-chain-status) +* [Publications](/publications) +* [News](/news) +* [Contact us](/contact-us) + +close menu + +* [Who we are](/) + + BACK TO MAIN MENU + + Entities + - [AI Safety Institute](/science-technology-and-innovation/technology/artificial-intelligence/ai-safety-institute) + - [Anti-Dumping Commission](/anti-dumping-commission) + - [Anti-Dumping Review Panel](/trade/anti-dumping-review-panel) + - [AusIndustry](/science-technology-and-innovation/industry-innovation/ausindustry) + - [Australian Industry Participation Authority](/major-projects-and-procurement/australian-industry-participation) + - [Australian Radioactive Waste Agency](/australian-radioactive-waste-agency) + - [Australian SKA Office](/science-technology-and-innovation/space-and-astronomy/ska-project-australia/australian-ska-office) + - [Australian Space Agency](https://www.space.gov.au/) + - [Australian Space Discovery Centre](https://www.space.gov.au/australian-space-discovery-centre) + - [Critical Minerals Office](/mining-oil-and-gas/minerals/critical-minerals/critical-minerals-office) + - [Industry Innovation and Science Australia](/science-technology-and-innovation/industry-innovation/industry-innovation-and-science-australia) + - [Investor Front Door](/investor-front-door) + - [Major Projects Facilitation Agency](/major-projects-and-procurement/major-projects-facilitation-agency) + - [National AI Centre](/science-technology-and-innovation/technology/artificial-intelligence/national-ai-centre) + - [National Measurement Institute](/national-measurement-institute) + - [National Offshore Petroleum Titles Administrator](https://www.nopta.gov.au/) + - [Office of Supply Chain Resilience](/trade/office-supply-chain-resilience) + - [Office of the Chief Economist](/office-chief-economist) + - [Office of the Chief Scientist](https://www.chiefscientist.gov.au/) + - [Offshore Decommissioning Directorate](/mining-oil-and-gas/oil-and-gas/offshore-decommissioning-directorate) + - [Questacon](https://www.questacon.edu.au/) + + [Ministers](https://www.minister.industry.gov.au/) + + [People](/people) + - [Organisation chart](/people/organisation-chart) + - [Overseas counsellors](/people/overseas-counsellors) + + [Portfolio agencies and boards](https://www.directory.gov.au/portfolios/industry-science-and-resources) + + [Public consultation](https://consult.industry.gov.au/) + + [Accessing information](/accessing-information) + - [Campaigns](/accessing-information/information-and-advertising-campaigns) + - [Freedom of information](/accessing-information/freedom-information-disclosure) + - [Gifts and benefits](/accessing-information/gifts-and-benefits) + - [Grants reporting](/accessing-information/grants-administration) + - [Indexed list of files](/accessing-information/indexed-list-files) + - [Portfolio appointments and vacancies](/accessing-information/portfolio-appointments-and-vacancies) + - [Privacy impact assessments](/accessing-information/privacy-impact-assessments) + + [Careers](/careers) + - [Current vacancies](https://industrycareers.nga.net.au/) + - [Graduates](/work-us/careers-graduates) + + [Corporate governance](/corporate-governance) + - [Annual report](/corporate-governance/annual-reports) + - [Budget statements](/corporate-governance/budget-statements) + - [Corporate plans](/corporate-governance/corporate-plans) + - [Fraud and corruption control](/corporate-governance/fraud-and-corruption-control) + - [Public interest disclosure](/corporate-governance/public-interest-disclosure) + - [Regulatory performance](/corporate-governance/regulatory-performance) + + [Government to government](/government-government) + - [Australian Business Licensing and Information Service](/government-government/australian-business-licensing-and-information-service) + - [business.gov.au](/government-government/businessgovau) + - [Business Grants Hub](/government-government/business-grants-hub) + - [Business Registration Service](/government-government/business-registration-service) + - [Shared Services Hub](/government-government/shared-services-hub) + - [VANguard](/government-government/vanguard) +* [Topics](/) + + BACK TO MAIN MENU + + [Major projects and procurement](/major-projects-and-procurement) + - [Australian industry participation](/major-projects-and-procurement/australian-industry-participation) + - [Investor Front Door](/investor-front-door) + - [Major projects facilitation](/major-projects-and-procurement/major-projects-facilitation-agency) + + [Manufacturing](/manufacturing) + - [Batteries](/manufacturing/batteries) + - [Green metals](/manufacturing/green-metals) + + [Measurement](/national-measurement-institute) + - [Australia's measurement system](/national-measurement-institute/australias-measurement-system) + * [Measurement standards](/national-measurement-institute/australias-measurement-system/measurement-standards) + * [Pattern approval](/national-measurement-institute/nmi-services/pattern-approval) + * [Trade measurement](/national-measurement-institute/trade-and-industry) + - [Awards](/national-measurement-institute/about-us/annual-measurement-awards) + - [Chemical and biological measurement services](/national-measurement-institute/chemical-and-biological-measurement-services) + - [Physical measurement services](/national-measurement-institute/physical-measurement-services) + - [Testing and analytical services](/national-measurement-institute/testing-and-analytical-services) + - [Training and assessment](/national-measurement-institute/measurement-training-and-assessment) + + [Mining, oil and gas](/mining-oil-and-gas) + - [Minerals](/mining-oil-and-gas/minerals) + * [Critical minerals](/mining-oil-and-gas/minerals/critical-minerals) + * [Critical Minerals Office](/mining-oil-and-gas/minerals/critical-minerals/critical-minerals-office) + * [Offshore minerals](/mining-oil-and-gas/minerals/regulating-offshore-mineral-exploration-and-mining) + * [Ranger Uranium Mine](/mining-oil-and-gas/minerals/regulating-ranger-uranium-mine) + - [Oil and gas](/mining-oil-and-gas/oil-and-gas) + * [Domestic gas supply](/mining-oil-and-gas/oil-and-gas/securing-australian-domestic-gas-supply) + * [Offshore oil and gas](/mining-oil-and-gas/oil-and-gas/offshore-oil-and-gas) + * [Offshore decommissioning](/mining-oil-and-gas/oil-and-gas/funding-and-investment-opportunities-offshore-decommissioning-australia) + * [Offshore Decommissioning Directorate](/mining-oil-and-gas/oil-and-gas/offshore-decommissioning-directorate) + * [Decommissioning the Northern Endeavour](/mining-oil-and-gas/oil-and-gas/offshore-oil-and-gas/decommissioning-northern-endeavour) + - [Greenhouse gas storage](/mining-oil-and-gas/offshore-greenhouse-gas-storage) + * [Acreage release process](/mining-oil-and-gas/offshore-greenhouse-gas-storage/offshore-greenhouse-gas-storage-acreage-release-process) + - [Taxes, royalties and export controls](/mining-oil-and-gas/taxes-royalties-and-export-controls-minerals-and-petroleum) + * [Combination certificates](/mining-oil-and-gas/taxes-and-royalties-minerals-and-petroleum/applying-combination-certificates-petroleum-projects) + * [Rough diamonds](/mining-oil-and-gas/taxes-and-royalties-minerals-and-petroleum/applying-export-rough-diamonds) + * [Uranium and controlled ores](/mining-oil-and-gas/taxes-and-royalties-minerals-and-petroleum/applying-export-uranium-and-controlled-ores) + + [Radioactive waste management](/australian-radioactive-waste-agency) + + [Science, technology and innovation](/science-technology-and-innovation) + - [Diversity in STEM](/science-technology-and-innovation/diversity-in-stem) + - [Industry innovation](/science-technology-and-innovation/industry-innovation) + * [AusIndustry](/science-technology-and-innovation/industry-innovation/ausindustry) + * [Business Research and Innovation Initiative](/science-technology-and-innovation/industry-innovation/business-research-and-innovation-initiative) + * [Critical Technologies Challenge Program](/science-technology-and-innovation/industry-innovation/critical-technologies-challenge-program) + * [Industry Innovation and Science Australia](/science-technology-and-innovation/industry-innovation-and-science-australia) + * [Industry Growth Program](/science-technology-and-innovation/industry-innovation/industry-growth-program) + * [Research and Development Tax Incentive](/science-technology-and-innovation/industry-innovation/research-and-development-tax-incentive) + - [International science collaboration](/science-technology-and-innovation/international-collaboration-science-and-research) + * [Global](/science-technology-and-innovation/international-collaboration/global-science-and-technology-diplomacy-fund) + * [Europe](/science-technology-and-innovation/international-collaboration/horizon-europe) + * [India](/science-technology-and-innovation/international-collaboration/collaborating-india-science-and-research) + * [China](/science-technology-and-innovation/international-collaboration/collaborating-china-science-and-research) + * [Vietnam](/science-technology-and-innovation/international-science-collaboration/australia-vietnam-science-technology-and-innovation-cooperation-initiative) + - [Science engagement](/science-technology-and-innovation/science-engagement) + * [Australian Space Discovery Centre](https://www.space.gov.au/australian-space-discovery-centre) + * [Measurement awards](/national-measurement-institute/about-us/annual-measurement-awards) + * [Prime Minister's Prizes for Science](/science-technology-and-innovation/science-engagement/prime-ministers-prizes-science) + - [Space and astronomy](/science-technology-and-innovation/space-and-astronomy) + * [Australian Space Agency](https://www.space.gov.au/) + * [SKA](/science-technology-and-innovation/space-and-astronomy/co-hosting-ska-telescope) + * [Optical astronomy](/science-technology-and-innovation/space-and-astronomy/optical-astronomy-australia) + - [Technology](/science-technology-and-innovation/technology) + * [Artificial intelligence](/science-technology-and-innovation/technology/artificial-intelligence) + * [AI Safety Institute](/science-technology-and-innovation/technology/artificial-intelligence/ai-safety-institute) + * [National AI Centre](/science-technology-and-innovation/technology/artificial-intelligence/national-ai-centre) + * [Quantum](/science-technology-and-innovation/technology/quantum) + * [Robotics](/science-technology-and-innovation/technology/robotics) + * [Critical technology standards](/science-technology-and-innovation/technology/critical-technology-standards) + + [Trade](/trade) + - [Anti-dumping](/anti-dumping-commission) + - [Anti-dumping reviews](/trade/anti-dumping-review-panel) + - [Country of origin labelling](/trade/country-origin-labelling) + - [Exporting resources](/mining-oil-and-gas/taxes-royalties-and-export-controls-minerals-and-petroleum#container-7529) + - [Free trade agreements](/trade/growing-australian-industries-through-trade) + - [Reducing technical barriers](/trade/reducing-technical-barriers-trade) + - [Standards and conformance](/trade/australias-standards-and-conformance-infrastructure) + - [Steel tariff rate quotas](/trade/indonesia-steel-tariff-rate-quota-certificates) + - [Supply chain resilience](/trade/office-supply-chain-resilience) + - [Global and domestic supply chain status](/trade/global-and-domestic-supply-chain-status) +* [Publications](/publications) +* [News](/news) +* [Contact us](/contact-us) + +Search this website + +Search + +## Navigate + +* Explore the digital publication + + + [National AI Plan](/publications/national-ai-plan) + + [Acknowledgement of country](/publications/national-ai-plan/acknowledgement-country) + + [Ministers' foreword](/publications/national-ai-plan/ministers-foreword) + + [Introduction](/publications/national-ai-plan/introduction) + + [The National AI Plan on a page](/publications/national-ai-plan/national-ai-plan-page) + + [Capture the opportunities](/publications/national-ai-plan/capture-opportunities) + + [Spread the benefits](/publications/national-ai-plan/spread-benefits) + + [Keep Australians safe](/publications/national-ai-plan/keep-australians-safe) + + [References](/publications/national-ai-plan/references) + +* Explore this publication + + + [National AI Plan](/publications/national-ai-plan) + + [Acknowledgement of country](/publications/national-ai-plan/acknowledgement-country) + + [Ministers' foreword](/publications/national-ai-plan/ministers-foreword) + + [Introduction](/publications/national-ai-plan/introduction) + + [The National AI Plan on a page](/publications/national-ai-plan/national-ai-plan-page) + + [Capture the opportunities](/publications/national-ai-plan/capture-opportunities) + + [Spread the benefits](/publications/national-ai-plan/spread-benefits) + + [Keep Australians safe](/publications/national-ai-plan/keep-australians-safe) + + [References](/publications/national-ai-plan/references) + +[National AI Plan](/publications/national-ai-plan) + +1. [Home](/) + / +2. National AI Plan + +# National AI Plan + +Date published: + +2 December 2025 + +### Download or share + +Downloadable files +Share this page + +Downloads + +[National AI Plan [pdf 1.75 MB]](/sites/default/files/2025-12/national-ai-plan.pdf) + +* [Facebook](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fwww.industry.gov.au%2Fpublications%2Fnational-ai-plan"e=National%20AI%20Plan) +* [X](https://x.com/intent/tweet?text=National%20AI%20Plan&url=https%3A%2F%2Fwww.industry.gov.au%2Fpublications%2Fnational-ai-plan) +* [LinkedIn](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fwww.industry.gov.au%2Fpublications%2Fnational-ai-plan) +* Email + +### Topics + +* [Science, technology and innovation](/publications?pub-topic=2945) +* [Technology](/publications?pub-topic=2963) + +### Publisher + +* [Department of Industry, Science and Resources](/publications?publisher=3026) + +## On this page + +## Introduction + +The National AI Plan is the Australian Government's plan to grow the AI industry in Australia. + +The plan sets out the steps the government will take to support Australia to build an AI-enabled economy that is more competitive, productive and resilient. It aims to make sure that everyone in Australia benefits from the AI opportunity, across all regions, industries and communities. + +The plan has 3 goals: + +* **Capture the opportunity** by building smart infrastructure, backing domestic AI capability and attracting global investment. +* **Spread the benefits** through widespread AI adoption, supporting and training Australian workers, and improved public services. +* **Keep Australians safe** with legislative and regulatory frameworks that mitigate AI harms, while promoting widespread responsible practices and international engagement that upholds Australia’s values. + +## Read the plan + +* ### [Acknowledgement of country](/publications/national-ai-plan/acknowledgement-country) + + Our department recognises the First Peoples of this Nation and their ongoing cultural and spiritual connections to the lands, waters, seas, skies, and communities. +* ### [Ministers' foreword](/publications/national-ai-plan/ministers-foreword) + + Read a foreword by the Minister for Industry and Innovation and Science, and the Assistant Minister for Science, Technology and the Digital Economy. +* ### [Introduction](/publications/national-ai-plan/introduction) + + Read the introduction to the National AI Plan. + +### Our vision + +### The National AI Plan on a page + +[Our goals for AI in Australia](/publications/national-ai-plan/national-ai-plan-page) + +* ### [Capture the opportunities](/publications/national-ai-plan/capture-opportunities) + + Capturing the opportunities of AI will help industry to scale, create high-quality jobs, and let us compete on the global stage. +* ### [Spread the benefits](/publications/national-ai-plan/spread-benefits) + + Every Australian should be able to benefit from AI, regardless of age, location or gender. +* ### [Keep Australians safe](/publications/national-ai-plan/keep-australians-safe) + + Our approach allows us to respond quickly and effectively to emerging risks and harms. + +### Appendix + +* ### [References](/publications/national-ai-plan/references) + + Read a list of sources for the National AI Plan. + +## Download a copy of the plan + +[National AI Plan [pdf 1.75 MB]](/sites/default/files/2025-12/national-ai-plan.pdf) + +## More information + +[The ministers launched the plan on 2 December 2025](https://www.minister.industry.gov.au/ministers/timayres/media-releases/national-ai-plan-empowering-all-australians) + +[Read more about the work we are doing to shape the future of AI in Australia](/science-technology-and-innovation/technology/artificial-intelligence) + +[NAIC is the government’s lead body supporting industry to unlock the economic benefits of AI](/node/93610) + +[Learn about the department's work in progressing AI safety science](/node/94883) + +[Read the Guidance for AI Adoption](/node/95254) + +## Publication previous and next page links + +National AI Plan + +* [Next page: + Acknowledgement of country](/publications/national-ai-plan/acknowledgement-country "Go to next page") + +## Contact us + +Email + +artificial.intelligence@industry.gov.au + +## Explore + +* [Careers](/careers) +* [People](/people) +* [Ministers](https://minister.industry.gov.au) +* [Publications](/publications) +* [News](/news) +* [Accessibility](/site-notices/accessibility) +* [Public consultation](https://consult.industry.gov.au) +* [Corporate governance](/corporate-governance) +* [Accessing information](/accessing-information) +* [Feedback and complaints](/contact-us/feedback-and-complaints) +* [Site notices](/site-notices) +* [Sitemap](/sitemap) + +## Contact us at the department + +Department of Industry, Science and Resources + +[General Enquiries](/contact-us/general-enquiries-form) + ++61 2 6213 6000 + +Industry House, +10 Binara Street, +Canberra + +Department of Industry +ABN: 74 599 608 295 + +## Subscribe for updates at the department + +Sign up for one or more of our email lists to receive the latest news, announcements and opportunities from our department. + +[Subscribe](/contact-us/newsletters) + +## Connect with us at the department + +* [Follow us on LinkedIn](https://www.linkedin.com/company/department-of-industry "Follow us on LinkedIn") +* [Follow us on Facebook](https://www.facebook.com/IndustryGovAu/ "Follow us on Facebook") +* [Follow us on X](https://x.com/IndustryGovAu "Follow us on X") +* [Follow us on Youtube](https://www.youtube.com/c/IndustryGovAu "Follow us on YouTube") +* [Subscribe to our RSS feeds](/contact-us/rss-feeds "Subscribe to our RSS feeds") + +## Acknowledgement of Country + +Our department recognises the First Peoples of this Nation and their ongoing cultural and spiritual connections to the lands, waters, seas, skies, and communities. + +We Acknowledge First Nations Peoples as the Traditional Custodians and Lore Keepers of the oldest living culture and pay respects to their Elders past and present. We extend that respect to all First Nations Peoples. + +*Artwork credit: DISR Journey*, 2024 by Chern’ee Sutton + +[Visit the Acknowledgement of Country page + +![](/themes/custom/industry_theme/images/acknowledgement_image_2024.png)](/node/75691) diff --git a/sources/resources/techpolicy-aiagency.md b/sources/resources/techpolicy-aiagency.md new file mode 100644 index 0000000..2056b23 --- /dev/null +++ b/sources/resources/techpolicy-aiagency.md @@ -0,0 +1,283 @@ +--- +url: https://techpolicy.au/aiagency +--- + +[![](/wp-content/themes/anu/assets/images/main-logo.png)](https://techpolicy.au) + +* [About us](/about-us) +* [Our Work](https://techpolicy.au/our-work) +* [Education](/education) +* [Podcast](/podcast) +* [Blog](/thecommons) +* [News & Events](/news-and-events) +* [Contact Us](/contact-us) + +# Expanding AI Sovereignty to AI Agency + +AI is reshaping global power, prosperity and security but debates about AI sovereignty are often binary, conflated, lacking evidence, and disconnected from the complex trade-offs leaders face. + +**TPDi’s AI Agency Tool, presented here in its final form, offers a practical solution.** + +**The Tool is a structured and repeatable method to assess a nation’s AI maturity, sovereignty and agency across 103 AI capabilities, producing prioritised recommended actions.** + +We applied the tool to produce **Australia’s 2025 AI Agency Assessment**: the first comprehensive, independent, evidence-based, expert-led assessment of Australia’s AI capabilities at the national level. We then mapped the **Australian Government’s 2025 National AI Plan** against the assessment. + +The tool is adaptable and scalable. We invite you to apply the tool so that your country, region, sector, community or organisation to identify your agency and help to proactively shape a technology that is already shaping our world. + +[Report: Expanding AI Sovereignty to AI Agency](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Report_Expanding-AI-sovereignty-to-AI-agency.pdf) +[Analysis: Australia’s National AI Plan: Australia’s AI Agency Assessment](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Analysis_2025-Australia-assessment-and-National-AI-Plan.pdf) +[Assessment: Australia’s 2025 AI Agency Assessment](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Assessment_2025-Australia-AI-agency_C.pdf) + +[![](https://techpolicy.au/wp-content/uploads/2026/06/02626-TPDI-Sovereignty-Report_cvr_FA4-scaled-800x0-c-default.jpg)](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Report_Expanding-AI-sovereignty-to-AI-agency.pdf) + +## + +### + +### + +### Everyone, it seems, wants ‘AI sovereignty’. But what most need is ‘AI agency’. + +The term ‘AI sovereignty’ dominates policy discussions and drives investment decisions but is used to mean everything from strategic self-reliance to cultural preservation and individual autonomy. + +Its use as a binary – where AI is sovereign, or it is not – leaves most countries disempowered. There are also more practical confusions. AI is not one thing, so exactly what AI capabilities are we talking about? How do you measure them? What would sovereignty really mean in each case? Since we published our discussion paper in November 2025, [From AI Sovereignty to AI Agency](https://techpolicy.au/wp-content/uploads/2025/11/TPDi-AI-Agency-Discussion-Paper-November-2025-1.pdf), the debate has gained momentum. Binary notions of AI and sovereignty are increasingly seen as reductive in today’s strategic landscape. + +### AI agency is the power of a country to shape its AI future. + +The global AI supply chain is a complex global web, and power comes from the ability to shape relationships, not retreat from them. TPDi’s AI Agency Tool offers a practical solution. + +**Instead of only asking if a country possesses sovereign control, the tool assesses whether a country has AI agency to steer outcomes, protect and promote national interests, and capture value in a globally connected system.** + +### Sovereignty is an enduring principle. + +The concept of AI agency put forward in this report does not signal a rejection, dilution or abandonment of the concept of sovereignty. On the contrary, the pursuit of AI agency is grounded in the very objectives that animate the pursuit of sovereignty. It offers a pragmatic policy approach that recognises most countries cannot, and need not, lead and control every AI capability. + +## Explore the AI Agency Tool via Podcasts + +Podcast +[Tech Mirror](https://techpolicy.au/podcast/ai-agency-and-mythos-5) +[AI Agency & Mythos 5](https://techpolicy.au/podcast/ai-agency-and-mythos-5) + +How does TPDi’s AI Agency Tool help to bring nuance to the debate about AI dependance? + +Podcast +[ABC Insiders](https://www.abc.net.au/news/2026-06-13/iob-johanna-weaver-ai/106792224) +[Is Australia ready for the AI transformation?](https://www.abc.net.au/news/2026-06-13/iob-johanna-weaver-ai/106792224) + +An overview of Australia’s National AI Agency Assessment + +# The AI Agency Tool + +## 6 Layers and 103 Capabilities + +![](https://techpolicy.au/wp-content/uploads/2026/06/Screenshot-2026-06-08-at-12.37.03-pm.png) + +![](https://techpolicy.au/wp-content/uploads/2026/06/Layers-round-scaled.jpg) + +**The aim is not to lead in all 103 capabilities – but, rather, to understand strengths, reduce critical dependencies, & build leverage where national advantages exist.** **The AI Agency Tool provides an evidence-based framework to do this, grounded in data rather than instinct or spin.** + +## How to use the tool + +**What is it?** The AI Agency Tool is a structured and repeatable method to assess a nation’s AI maturity, sovereignty and agency across 103 AI capabilities, producing prioritised recommended actions. + +**Why does it matter?** The tool equips decision-makers to pursue AI agency: the capacity to steer outcomes, protect and promote national interests, and capture value in a globally connected system. Used well, the tool enables nuanced strategies, better-targeted investment, and deeper understandings of trade-offs. The tool equips leaders to identify their leverage in high agency capabilities and harness it to offset their vulnerabilities. + +**Who is it for?** + +* **Policymakers** of countries of all sizes and stages of AI maturity, particularly where strategic dependence is high, and choices are constrained (Australia’s application from page 14). +* **Business leaders** navigating geopolitical risk, supply chains and long-term investment decisions (uses for enterprise assessments on page 21). +* **Researchers** conducting national assessments, tracking progress over time and holding governments accountable (tool guide from page 25 and methodology from page 33). + +**How does it work?** A step-by-step process to gather evidence and make assessments to both inform and analyse strategy. + +⬇️ [**Report: Expanding AI Sovereignty to AI Agency**](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Report_Expanding-AI-sovereignty-to-AI-agency.pdf) + +![](https://techpolicy.au/wp-content/uploads/2026/06/Screenshot-2026-06-08-at-12.36.02-pm.png) + +## + +# Overview of the AI Agency Tool + +The AI Agency Tool starts by producing a **maturity rating** for 103 AI capabilities defined in the **typology**. It then situates the traditional binary objective of *sovereign control* over those capabilities within an expanded spectrum that is fit for purpose in today’s strategic landscape. The traditional *control* framing is expanded to also consider the management of international partnerships (*access*), the importance of resilience (*choice*) and pursuit of competitive advantage (*leverage*). We call this the **sovereignty spectrum***.* An assessment of each element of the spectrum produces a **sovereignty rating**. + +The maturity and sovereignty ratings, and an assessment of global scarcity, are then combined into a single view to produce an **AI agency score** – a measure of national competitive advantage for each AI capability. + +**AI agency thus offers a pragmatic pathway for advancing sovereignty:** pursuing enduring principles of control and self-determination, while equipping leaders to navigate a world of international interdependence (including shared infrastructure and global supply chains) and geopolitical competition. + +The visualisation below shows how sovereignty is embedded as a core element of the AI Agency Tool. + +![](https://techpolicy.au/wp-content/uploads/2026/06/Overview-of-Natiional-AI-Agency-Assessment-Process-scaled-e1780888476135.jpg) + +**The tool is adaptable and scalable. The right mix of strategic capabilities will differ country to country. We invite you to apply the tool so that your country, region, sector, community or organisation.** + +**A digital version of the tool is coming soon: watch this space.** + +# The Tool in Practice: Australia’s 2025 AI Agency Assessment + +![](https://techpolicy.au/wp-content/uploads/2026/06/Screenshot-2026-06-08-at-12.36.35-pm.png) + +## + +## Strategic silences in the National AI Plan align with areas of low national agency. + +**Not all omissions from the plan should be viewed as gaps.** For many capabilities the assessment’s recommended action is ‘maintain and monitor’. Many of the areas receiving limited attention in the plan fall into this category. They are capabilities where Australia currently has low maturity and low agency, or where government intervention is less necessary. This includes: + +* accelerator manufacturing +* frontier model development +* most forms of private-sector AI capability. + +The government has reasonably adopted a targeted approach in the plan, focusing public investment on what the government has assessed to be foundational capabilities while allowing the private sector to lead where appropriate. + +## There is untapped potential in Australia’s highest agency capabilities. + +**While the plan establishes strong foundations, it does not fully capitalise on all of Australia’s areas of highest agency – the capabilities in which Australia has competitive advantage.** The assessment identifies opportunities to better leverage Australia’s strengths in critical minerals, strategic data assets and model development in computer vision. These are areas where Australia already possesses very high agency and which could be used more deliberately to strengthen national capability, address weaknesses and close critical gaps elsewhere in the AI ecosystem, while increasing Australia’s international leverage. + +*Australia’s 2025 AI Agency Assessment* is the first application of the AI Agency Tool. It assesses Australia’s AI maturity, sovereignty, agency (as at December 2025), and produces recommended actions across 103 capabilities. We then mapped the Australian Government’s 2025 National AI Plan against the assessment. + +## Australia’s highest areas of agency + +The assessment found Australia has 8 capabilities that fall within the highest band: very high agency.  This includes: + +* our rich endowment in strategic and critical minerals +* 5 domain specific datasets (medical, geospatial, environment and resources, demographic and infrastructure) +* our expertise in developing computer vision models +* our proven impact in international engagement (influence and norm shaping). + +## Every significant commitment in the National AI Plan aligns with the assessment’s recommendations. + +Governments have finite resources and must make tough decisions about what to prioritise. **It is noteworthy that the plan is strongly aligned with the assessment’s findings on where Australia should leverage, build or maintain agency.** The plan’s major commitments focus on areas where Australia already possesses meaningful maturity and agency. This includes: + +* data centres and supporting infrastructure +* public cloud +* general AI applications +* government and small to medium enterprise (SME) adoption +* international engagement. + +Leaning into these capabilities harnesses existing strengths while also offering enabling benefits across the whole ecosystem. + +The assessment found only 2 capabilities in which Australia has the lowest level of agency, these being manufacturing and packaging of accelerators (AI chips). + +## + +## There are critical gaps to close + +**The assessment also identifies several globally scarce capabilities that are important to the public interest but remain underdeveloped in Australia – these capabilities should be the focus of the next wave of government prioritisation.** This includes: + +* public sector and public interest compute infrastructure +* key data lifecycle management capabilities (such as copyright, sourcing, validation and annotation) +* culturally and nationally inclusive models +* discerning, inclusive, and trusted AI adoption +* general AI literacy +* several specialised AI skills +* regulatory and oversight capability. + +Addressing these gaps would strengthen Australia’s ability to support innovation, research, public services and national resilience, while ensuring AI systems reflect Australian values, cultures and identities and, perhaps most importantly, ensure that the benefits of AI are widely distributed. + +⬇️ [Assessment: Australia’s 2025 AI Agency Assessment](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Assessment_2025-Australia-AI-agency_C.pdf) +⬇️ [Analysis: Australia’s National AI Plan cf: Australia’s AI Agency Assessment](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Analysis_2025-Australia-assessment-and-National-AI-Plan.pdf) + +# Read the Analysis of Australia’s 2025 Assessment and National AI Plan + +# Acknowledgments + +![](https://techpolicy.au/wp-content/uploads/2026/01/1759125649931-1000x1000-c-default.webp) + +![](https://techpolicy.au/wp-content/uploads/2026/01/1759125648635-1000x1000-c-default.webp) + +![](https://techpolicy.au/wp-content/uploads/2026/01/1763954271110-1000x1000-c-default.webp) + +## **Sovereignty was never ceded** + +TPDi acknowledges the Ngunnawal and Ngambri people who are the Traditional Owners of the land upon which this report was prepared in Canberra, Australia. We pay our respects to Elders past and present. + +The authors affirm that sovereignty has never been ceded by First Nations peoples living on the continent now known as Australia. We recognise Indigenous Sovereignty as enduring and inherent, as well as fundamentally different to the new concept of ‘AI sovereignty’ (to which this report is responding). + +Any national conversation about AI should reinforce, rather than distract from, the distinct and profoundly important conversations about Indigenous Sovereignty and governance. The AI Agency Tool emphasises the importance of pursuing both Indigenous Sovereignty and AI agency in parallel. Indeed, they can be understood as reinforcing policy goals. + +The tool explicitly highlights the importance of Indigenous Sovereignty. First Nations interests and sensitivities are profound and intersecting across all 103 AI capabilities identified in the typology. A review of the typology through lenses of cultural priority, sovereignty beyond a nation-state framing, lived experience, and the risk of structural harm, revealed that at least 38 AI capabilities have particular implications for First Nations peoples. The tool frames inclusive, empowering and rights-respecting approaches to these capabilities as indicators of AI maturity. + +The concept of AI agency will look different in practice for different stakeholders, including First Nations peoples compared to governments and organisations. The tool is intended to be adopted and applied by different groups, shaped by their distinct perspectives and priorities in implementation. In this way, it acts as both an assessment and a transparency tool. + +This approach is intended to support the empowerment of Indigenous voices, leadership and agency in future-proofing communities and shaping Australia’s future. + +See pages 4 and 46 in the Report for further discussion on these issues. + +## A collective endeavour + +### Authors + +Zoe Jay Hawkins, Co-Founder and Deputy Executive Director, TPDi; Johanna Weaver, Co-Founder and Executive Director, TPDi; Rebecca Razavi, Visiting Policy Fellow, Oxford Internet Institute; Executive Director, UNSW Public Policy Institute; Meredith Hodgman, Head of Strategy and Engagement, TPDi; Vili Lehdonvirta, Professor of Technology Policy, Aalto University; Senior Fellow, Oxford Internet Institute; Mercedes Page, Visiting Scholar, Massachusetts Institute of Technology Sloan School of Management and Murray Lab. With invaluable contributions from Dorina Wittmann, Heidi Brockway and Helen Portillo-Castro. + +### Peer Reviewers + +We are grateful for the invaluable feedback we received throughout the project from our peer reviewers (affiliations at time of review): Andrew Brodie, Eva Hopewell and Todd Phillips (Deadly Coders), Belinda Dennett (AirTrunk), Brendan Hopper (Commonwealth Bank of Australia), Clare Beaton-Wells (United Nations Youth Australia), Christina Wiremu-Brook (Kuria), Dave Lemphers (Maincode), David Masters (Atlassian), Dot West, Lyndon Ormand-Parker and Daniel Featherstone (First Nations Digital Inclusion Advisory Group), Prof Elanor Huntington (CSIRO), Ian Opperman (ServiceGen), Jamie Morse (Macquarie Technology Group), Josh Griggs and Pauline Fetaui (Australian Computer Society), Kate Conroy (Queensland University of Technology), Kyle Turner (Paul Ramsay Foundation), Lee Hickin (National AI Centre), Liam Carroll, Bill Simpson Young, Alistair Reid and Tiberio Caetano (Gradient Institute), Mark Stickells (Pawsey Supercomputing Research Centre), Nadia Court, Tanya Saad and Anna Gurevich (Semiconductor Sector Service Bureau), Rosie Hicks (Australian Research Data Commons), Sally Ann Williams (Australian Academy of Technological Sciences & Engineering), Scott Winch (Sax Institute), Simon Kriss (Sovereign AI Australia), Simon Spencer (Trideca), Dr Sue Keay (Robotics Australia Group), Taylor Dee Hawkins (Foundations for Tomorrow), Dr Tobias Feakin (Protostar Strategy), Tim Carton (CDC Data Centres), and Tim Moriarty (Muzo.ai). + +### Consultations + +TPDi gratefully acknowledges the contributions of experts and organisations who participated in workshops, surveys, interviews and peer reviews that informed this project. Their participation reflects wide engagement across Australia’s technology, research, policy and civil society communities. + +In September 2025, TPDi brought together over 250 individuals and 187 organisations across five cities, and online, to survey and map Australia’s AI strengths and define our path forward. + +A special acknowledgment to the industry and research bodies that helped amplify the consultation process, with executive participation and opportunities for their members to contribute, including the Australian Academy of Science (AAS), Australian Computer Society (ACS), Australian Council of Learned Academies (ACOLA), Electronic Frontiers Australia Inc, IoT Alliance Australia (IoTAA), Digital Rights Watch (DRW), the Kingston AI Group, ARC Centre of Excellence for Automated Decision-Making and Society (ADM+S), Gradient Institute, Science and Technology Australia (STA), Tech Council of Australia (TCA), and the UNSW AI Institute. + +**A full list of the individuals and organisation who participated in the consultations in an Annex of the Report – with thanks from team TPDi to each and every person.** + +### Building on previous work + +The Final Report, Assessment and Analysis presented on this webpage build on the [Discussion Paper,](https://techpolicy.au/wp-content/uploads/2025/11/TPDi-AI-Agency-Discussion-Paper-November-2025-1.pdf) and the [Draft Australia AI Capability Assessment](https://techpolicy.au/wp-content/uploads/2026/04/TPDi-AI-Agency-Tool-Australian-Stocktake-Discussion-Draft-Web-Doublespread-CC.pdf), published by TPDi in November 2025. + +This work builds on longstanding research by TPDi Co-Founder [Zoe Jay Hawkins](https://www.linkedin.com/in/zoe-hawkins/) and colleagues [Vili Lehdonvirta](https://www.linkedin.com/in/vililehdonvirta/) and [Bóxī Wú](https://www.linkedin.com/in/boxiwu/) ([Oxford Internet Institute, University of Oxford](https://www.linkedin.com/company/oxford-internet-institute/) & [Aalto University](https://www.linkedin.com/company/aalto-university/)) on compute sovereignty, which has been featured by The New York Times, the Australian Financial Review, TIME Magazine, BBC, The Economist, The Financial Times, Forbes, and POLITICO (among others). + +⬇️ [Read the article in the New York Times](https://www.nytimes.com/interactive/2025/06/23/technology/ai-computing-global-divide.html) + +⬇️ [Read the research paper: Hawkins, ZJ., Lehdonvirta, V., and Wu, B. (2025). AI Compute Sovereignty: Infrastructure Control Across Territories, Cloud Providers, and Accelerators.](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5312977) + +⬇️ [Read the research paper:Lehdonvirta, V., Wú, B., & Hawkins, ZJ. (2024). Compute North vs. Compute South: The Uneven Possibilities of Compute-based AI Governance Around the Globe](https://ojs.aaai.org/index.php/AIES/article/view/31683) + +![](https://techpolicy.au/wp-content/uploads/2025/08/2-1-scaled-1000x1000-c-default.webp) + +![](https://techpolicy.au/wp-content/uploads/2025/08/3-1-scaled-1000x1000-c-default.webp) + +![](https://techpolicy.au/wp-content/uploads/2025/08/1-1-scaled-1000x1000-c-default.webp) + +## A thanks to our supporters + +This project was made possible by the generous support of the **Australian Computer Society** and the **Department of Industry, Science and Resources**. + +**TPDi’s independence is our most valuable asset.** As a registered not-for-profit, our work is supported by external funding. We only accept funding from entities that agree to be disclosed publicly and commit to respect and promote TPDi’s independence. **TPDi does not represent the views of any of our funders; all outputs represent solely the views of TPDi and/or the authors. [Learn more](https://techpolicy.au/funders) about TPDi’s blind trust funding model and how this protects and promotes our independence.** + +![](https://techpolicy.au/wp-content/uploads/2025/08/ACS-1-scaled-1000x1000-c-default.webp) + +![](https://techpolicy.au/wp-content/uploads/2025/08/DISR-1-scaled-1000x1000-c-default.webp) + +⬇️ [Report: Expanding AI Sovereignty to AI Agency](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Report_Expanding-AI-sovereignty-to-AI-agency.pdf) | ⬇️ [Assessment: Australia’s 2025 AI Agency Assessment](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Assessment_2025-Australia-AI-Agency.pdf) | ⬇️ [Analysis: Australia’s National AI Plan: Australia’s AI Agency Assessment](https://techpolicy.au/wp-content/uploads/2026/06/TPDi-2026-Analysis_2025-Australia-assessment-and-National-AI-Plan.pdf)| ⬇️ [Read the Media Release](https://techpolicy.au/news/launch) + +[![](/wp-content/themes/anu/assets/images/main-logo.png)](https://techpolicy.au) + +* Learn more +* [Projects](https://techpolicy.au/projects) +* [News and Events](https://techpolicy.au/new-news-and-events) +* [Podcast](https://techpolicy.au/tech-mirror-podcast) +* [Contact Us](https://techpolicy.au/contact-us) + +* Follow us on +* [Follow us on LinkedIn](https://www.linkedin.com/company/techpolicydesigncentre/) + +### Subscribing to TPDIs mailing list + +Δ + +Name(Required) + +First + +Last + +Email + +Submit + +Copyright Tech Policy Design Institute 2026 + +* [Privacy Policy](https://techpolicy.au/privacy-policy) +* [Disclaimer](https://techpolicy.au/disclaimer) diff --git a/tree.svg b/tree.svg new file mode 100644 index 0000000..0a51cb2 --- /dev/null +++ b/tree.svg @@ -0,0 +1,143 @@ + + +Choose a path + + + + + + + + + + + + + + + + + + + +stalls + +race / deal + +control lost + +control kept + +seized abroad + +seized at home + +power stays spread + +no public stake + +set our price + +host it + +stay supplier + +concert + +US bloc + +public stake + +no stake +1. Prepare early — 2026-28: royalties, biosecurity, allies + +1. Prepare early +2026-28: royalties, biosecurity, allies + +2. Build the people — 2026-29: safety testers, inspectors + +2. Build the people +2026-29: safety testers, inspectors + +2029: the US chooses — go-slow deal with China, or race? + +2029: the US chooses +go-slow deal with China, or race? + +VII. Boring decade — ~10% + +VII. Boring decade +~10% + +RISK: AI escapes control — ~75% racing, ~25% in the deal + +RISK: AI escapes control +~75% racing, ~25% in the deal + +I. Doom — ~50% + +I. Doom +~50% + +RISK: power seized — one actor holds the AI + +RISK: power seized +one actor holds the AI + +II. Someone else's empire — ~25% — seized abroad + +II. Someone else's empire +~25% — seized abroad + +RISK: power grabbed at home — seized in our own country + +RISK: power grabbed at home +seized in our own country + +3. Sign, or set our price? — the deal reaches us, 2029 + +3. Sign, or set our price? +the deal reaches us, 2029 + +IX. Homegrown autocracy — citizens stop being needed + +IX. Homegrown autocracy +citizens stop being needed + +6. If the deal collapses — US bloc, concert, or neutral? + +6. If the deal collapses +US bloc, concert, or neutral? + +4. Host compute here? — 2030-32 + +4. Host compute here? +2030-32 + +VIII. Middle-power concert + +VIII. Middle-power concert + +V. Allied compute host + +V. Allied compute host + +VI. Garrison ally + +VI. Garrison ally + +5. Tax it, own a share? — 2031-34 + +5. Tax it, own a share? +2031-34 + +III. Dividend commonwealth + +III. Dividend commonwealth + +IV. Quarry economy — drift default + +IV. Quarry economy +drift default + + \ No newline at end of file