Prices whether AI control mitigations — monitoring, honeypots, deployment-time defences — will be strong enough by 2030 to stop a misaligned model that is actively trying to cause catastrophe.
Tagged
risk-model
12 pieces.
Prices whether it becomes possible and affordable by 2070 to build AI systems that are highly capable, plan on their own, and model their own situation — the step every other state here rests on.
Prices whether an actual AI-caused catastrophe occurs given that models already have the capability to cause one, with control mitigations as the remaining line of defence.
Prices whether frontier AI models cross capability thresholds sufficient to make a catastrophe physically possible by 2030, on the strict reading its downstream edge requires.
Prices whether permanent human disempowerment by power-seeking misaligned AI would itself destroy humanity's long-term potential, given every upstream condition holds by 2070.
Prices whether misaligned AI power-seeking, granted it is already happening at high impact, scales to the permanent disempowerment of essentially all humanity by 2070.
Prices whether strong incentives to build powerful agentic AI will exist by 2070, given that such systems are technically and economically feasible.
Prices the chance that humanity loses meaningful control over AI systems and the trajectory they set, unconditionally, by 2035.
Prices whether deployed misaligned AI systems actually seek power over people in high-impact ways, given that such systems are feasible, misaligned by default, and built anyway.
Prices whether, given feasible powerful agentic AI and strong incentives to build it, aligned systems turn out much harder to build than misaligned ones that still look worth deploying.
Prices whether a frontier model, by 2030, possesses the capability to sabotage oversight — disabling monitors, sandbagging evaluations, subverting supervision — with no parent conditions attached.
Prices whether AI models will deliberately mislead their overseers by 2035, and why the estimate turns on how strictly that bar is read.