AI-caused catastrophe
Prices whether an actual AI-caused catastrophe occurs given that models already have the capability to cause one, with control mitigations as the remaining line of defence.
In Risks
Risk dossier · event
P(this | the condition below) = 0.1–0.45
Conditional on
- Capability thresholds crossed — 0.7–0.95 by 2030
Composed, all-in: 0.07–0.43 — the intervals multiplied along the full `requires` chain, assuming the conditions are independent, which they are not entirely.
What would move it
- A frontier lab publishes a safety case that rests on control mitigations rather than on models lacking dangerous capability.
How the interval was set
Conditioned on capability thresholds already being crossed — i.e. models that do possess the capability to cause a catastrophe — how likely is an actual AI-caused catastrophe?
The conditioning removes the single strongest protective argument on record: the Anthropic safety-cases quote says "the cleanest argument that current-day AI models will not cause a catastrophe is probably that they lack the capability to do so." Once that inability argument is void, the burden shifts entirely to the second line of defence named in the rogue-eval quote — "mitigations that would prevent LLMs from causing a catastrophe even if they tried" — which is exactly the `ai-control-mitigations` influence, and it pushes the interval down substantially from where a bare capability premise alone would put it. Bengio's FAQ pushes up: the attack surface is cheap ("it suffices that an AI has access to the internet and strong cybersecurity skills"), so capability plus modest misalignment or misuse is close to sufficient for large-scale harm. The interpretability quote (arXiv 2306.06924) is weaker, generic evidence about opacity perpetuating harm; it widens rather than shifts.
Two irreducible uncertainties keep the band wide. First, "catastrophe" is undefined here — a coordinated critical-infrastructure cyber event and an existential outcome differ by orders of magnitude in probability, and the sources use the word for both. Second, control mitigations are untested at the relevant capability level; whether they hold against a capable adversarial model is an open empirical question, not a settled one. Capability also does not imply propensity: misuse gating, deployment restrictions and deterrence intervene. Published expert numbers in this literature cluster around 10–20% unconditionally; conditioning on capability raises that, mitigations pull it back, and the definitional slack forbids anything narrower than roughly one-in-ten to near-even. No year is stated in the node or dominant in the evidence, so no horizon is asserted; implicitly this is the decade or two following threshold crossing.
Grounded in
https://alignment.anthropic.com/2024/rogue-eval — 1 quoted claim
- “we may also try to build mitigations that would prevent LLMs from causing a catastrophe even if they tried to cause one”
https://alignment.anthropic.com/2024/safety-cases — 1 quoted claim
- “The cleanest argument that current-day AI models will not cause a catastrophe is probably that they lack the capability to do so.”
https://arxiv.org/abs/2306.06924 — 1 quoted claim
- “trying to explain black box models, rather than creating models that are interpretable in the first place, is likely to perpetuate bad practices and can potentially cause catastrophic harm to society”
https://yoshuabengio.org/2023/06/24/faq-on-catastrophic-ai-risks/ — 1 quoted claim
- “It suffices that an AI has access to the internet and strong cybersecurity skills to already do a lot of damage, especially if these attacks are coordinated”
Also phrased across sources as: “LLM-caused catastrophe” · “Catastrophic harm to society” · “Catastrophic AI-caused damage”
Assessed probabilities are the model’s knowledge, not verification — 2026-09-05 · assess-risk@3 · claude-opus-5.
This node asks a single narrow question: if AI systems already have the ability to cause a catastrophe, does one actually happen — and the bar it sets is an actual event in the world, not a close call or a demonstrated capability.
Most of the hard questions were answered before this step. The usual reassurance about today’s systems is that they simply cannot do the damage; Anthropic’s safety-cases write-up puts it directly: “The cleanest argument that current-day AI models will not cause a catastrophe is probably that they lack the capability to do so.” This node starts after that argument has expired. What is left is the second line of defence from the same body of work — “mitigations that would prevent LLMs from causing a catastrophe even if they tried to cause one.” That is the AI control mitigations influence in the subtree below, and it is the main thing holding the interval above down from where a bare capability premise would leave it.
What pushes the estimate up is how little else is needed once the ability exists. Bengio’s FAQ on catastrophic risks argues the attack surface is cheap: “It suffices that an AI has access to the internet and strong cybersecurity skills to already do a lot of damage, especially if these attacks are coordinated.” On that account, capability plus a modest amount of misuse or misalignment is close to enough. The interpretability literature adds a weaker, more general worry — that explaining opaque models after the fact, rather than building understandable ones, “is likely to perpetuate bad practices and can potentially cause catastrophic harm to society.” That argument does not move the estimate in one direction so much as widen it: it says we may not see the failure coming.
Two things keep the band wide and forbid anything tighter. First, “catastrophe” is not pinned down here. A coordinated attack on critical infrastructure and an outcome that ends human control differ enormously in likelihood, and the sources use the same word for both. Second, control mitigations have not been tested against a model at the capability level this node assumes. Whether they hold against a capable adversary that is actively working around them is an open empirical question. And having the ability is not the same as having the motive: deployment limits, access gating and ordinary deterrence all sit in between. No year is fixed in the evidence, so no horizon is asserted; read this as the period following the threshold crossing rather than a dated forecast.
The subtree
The diagram is an illustration; the model is the record.