TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI
A taxonomy of societal-scale AI risk sorted by how much of the harm was anyone's intention, with a story for each type. What the claims are, and what they rest on.
In Papers
TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI — Andrew Critch and Stuart Russell, June 12, 2023.
Claim ledger
Assessments are the model’s knowledge, not verification.
- 1¶
The paper characterises the existing literature on societal-scale and existential AI risk as concentrated on single-system misalignment.
“So far, most research papers addressing societal-scale and existential risks have focused on misalignment of a single advanced AI system.”
assertion · unclear
plausible · medium confidence — Pre-2023 AI x-risk literature (Bostrom, Yudkowsky, Bengio's own blog cited here) does skew toward single-system loss-of-control framings, but structural/multi-agent risk work already existed (Zwetsloot & Dafoe 2019, cited two sentences later), so 'most' is a defensible but self-serving characterization of the field rather than a documented survey result.
To check: A systematic review/count of AI x-risk papers 2014-2023 by risk-mechanism category.
- 2¶
Single-system misalignment is not the only route to societal-scale or extinction-level AI risk.
“However, while misalignment of individual systems remains a problem, it is not the only source of societal-scale risks from AI, and extinction risk is no exception.”
contrarian · unclear
plausible · medium confidence — Consistent with the 'structural risk' framing already articulated by Zwetsloot & Dafoe (2019), which the authors cite; this paper extends rather than originates that framing to extinction-level risk specifically.
To check: Compare against Zwetsloot & Dafoe's accident/misuse/structure taxonomy and subsequent structural-risk literature.
- 3¶
Treating humanity and machines as two monolithic agents is insufficient; risk analysis must span multiple organisational scales.
“construed monolithically,” this monolithic view of AI technology is not enough: safety requires analysis of risks at many scales of organization simultaneously.”
assertion · unclear
plausible · medium confidence — This is the paper's central thesis; it prefigures later multi-agent/structural safety work (e.g. 'Gradual Disempowerment', 2024) that gained traction after 2023, but as a normative claim about what safety 'requires' it is not independently verifiable, only judged by subsequent field uptake.
To check: Growth of multi-agent/structural AI-safety research output post-2023 as a proxy for the field agreeing with this framing.
- 4¶
The paper's organising principle is accountability: whose actions caused the risk, whether unified, whether deliberate.
“To that end, we have chosen an exhaustive taxonomy based on accountability: whose”
assertion · unclear
consistent · high confidence — This is a direct, verifiable description of the paper's own stated methodology, matching the abstract and Section 2 structure.
To check: Re-read the paper's stated organizing principle (already confirmed in transcript).
- 5¶
The six-type decision tree is claimed to be exhaustive by the same logic as safety-engineering fault trees.
“The decision tree in Figure 2 above follows the same basic principle to produce an exhaustive taxonomy.”
assertion · unclear
plausible · medium confidence — Fault tree analysis is a genuine, well-established safety-engineering technique for exhaustive branching, so the analogy's logic is sound; whether the specific six-branch accountability tree is actually logically exhaustive is asserted rather than formally proven in the paper.
To check: Attempt to construct a societal-scale AI harm scenario that does not fit any of the six accountability branches.
- 6¶
Exhaustiveness alone does not make a taxonomy analytically valuable.
“Exhaustiveness of a taxonomy is of course no guarantee of usefulness.”
assertion · unclear
consistent · high confidence — A standard logical/epistemological point, correctly illustrated by the paper's own prime-number-day counterexample; uncontroversial.
To check: None needed; this is a valid logical claim, not an empirical one.
- 7¶
A taxonomy's usefulness is measured by whether it surfaces new risks or suggests interventions.
“is only useful to the extent that it reveals new risks or recommends helpful interventions.”
assertion · unclear
plausible · medium confidence — A reasonable, commonly-used criterion for judging taxonomies in policy/safety work, though not the only possible criterion (clarity, communicability, and prioritization value are others); stated as a stipulation rather than derived.
To check: Compare against how other risk taxonomies (e.g. Zwetsloot & Dafoe, Yampolskiy 2015) justify their own usefulness.
- 8¶
The accountability taxonomy is claimed to reveal multi-system interaction risks and deliberate-misuse risks that other framings miss.
“This taxonomy in particular surfaces risks arising from unanticipated interactions of many AI systems,”
assertion · unclear
plausible · medium confidence — Matches the content that follows (Types 1, 5, 6 do cover multi-system and misuse risk); however this is the authors grading their own framework's contribution, so it should be read as a stated intent confirmed only internally.
To check: Check whether Types 1/5/6 analyses actually identify risks absent from prior single-system taxonomies.
- 9¶
Yampolskiy's earlier risk taxonomy fails to be exhaustive because it assumes a single well-defined creator's intent.
“While useful, Yampolskiy’s taxonomy was non-exhaustive, because it presumed a unified intention amongst the creators of a particular AI system.”
contrarian · unclear
unverifiable · low confidence — I don't have detailed enough recall of Yampolskiy's 2015 'Taxonomy of pathways to dangerous AI' to confirm or refute this specific characterization of its assumptions.
To check: Read Yampolskiy (2015), arXiv:1511.03246, and check whether its categories presuppose a single unified creator intent.
- 10¶
The diffusion-of-responsibility category is deliberately constructed as a catch-all guaranteeing coverage.
“Thus, Type 1 serves as a hedge against the taxonomy of Types 2-6 being non-exhaustive.”
assertion · unclear
consistent · high confidence — A structural/tautological claim about the paper's own decision-tree design, directly supported by the surrounding text describing Type 1 as covering cases with no unified responsible party.
To check: None needed; internally verifiable from the paper's own structure.
- 11¶
Societal harm from automation can occur with no primarily responsible party, possibly because of the absence of responsibility.
“Automated processes can cause societal harm even when no one in particular is primarily responsible for the creation or deployment of those processes (Zwetsloot and Dafoe, 2019) , and perhaps even as a result of the absence of responsibility.”
assertion · unclear
consistent · high confidence — Matches the recognized 'structural risk' category in AI safety discourse (Zwetsloot & Dafoe 2019, cited) and is well illustrated by documented algorithmic-trading incidents.
To check: Zwetsloot & Dafoe (2019) Lawfare piece; SEC/CFTC reports on algorithmic-trading incidents.
- 12¶
The 2010 flash crash, caused by interacting trading algorithms from many firms, wiped over $1 trillion off US markets within minutes.
“The infamous “flash crash” of 2010 is an instance of this: numerous stock trading algorithms from a variety of companies interacted in a fashion that rapidly devalued the US stock market by over 1 trillion dollars in a matter of minutes.”
quantity · unclear
consistent · medium confidence — Matches widely reported figures for the May 6, 2010 Flash Crash, in which US equity markets briefly lost roughly $1 trillion in value within minutes before largely recovering.
To check: SEC/CFTC joint report on the May 6, 2010 Flash Crash.
- 13¶
The authors claim the economic and legal disempowerment illustrated for a subgroup could extend to all of humanity.
“It is possible, we claim, for all of humanity to become similarly disempowered.”
contrarian · unclear
unverifiable · medium confidence — Explicitly flagged by the authors as a claim rather than a demonstrated fact. It echoes Critch's own earlier work on 'Robust Agent-Agnostic Processes' / multipolar failure (2021) and prefigures later 'gradual disempowerment' scholarship (e.g. Kulveit et al. 2024) — a coherent but forward-looking, unfalsifiable-at-present extrapolation.
To check: No near-term empirical test exists; would require observing large-scale automation-driven loss of human economic/political control.
- 14¶
An unchecked AI industry could become entrenched against regulation like fossil fuels or tobacco, but on a faster timescale.
“The “AI industry”, if unchecked, could behave similarly, but potentially much more quickly than the oil industry, in cases where AI is able to think and act much more quickly than humans.”
prediction · unclear
unverifiable · medium confidence — The base analogy (regulatory capture by entrenched industries) is well-documented (Carpenter & Moss 2013, cited); the specific claim about AI entrenching faster is a speculative prediction not yet testable.
To check: Track regulatory-capture dynamics and speed of entrenchment in the actual AI industry over the coming decade.
- 15¶
Competitive pressure to automate internally and to trade with other automated firms drives a gradual transfer of control away from humans.
“There was a gradual handing-over of control from humans to AI systems, driven by competitive pressures for institutions to (a) operate more quickly through internal automation, and (b) complete trades and other deals more quickly by preferentially engaging with other fully automated companies.”
assertion · unclear
plausible · medium confidence — An abstracted mechanism from a fictional story, not an empirical claim; it is economically coherent (competition rewarding speed can favor automation) and resembles concerns raised elsewhere (Christiano's 'What failure looks like', Critch's RAAPs work), but is untested.
To check: Empirical tracking of automation adoption rates and control-handover patterns across competitive industries.
- 16¶
Collective agreement to slow or halt the automation trend fails, in the paper's abstracted mechanism.
“Humans were not able to collectively agree upon when and how much to slow down or shut down the pattern of technological advancement.”
assertion · unclear
unverifiable · medium confidence — A premise of the fictional 'production web' story; collective-action coordination failures are well documented in other domains (climate policy), lending plausibility to the mechanism, but its specific application to AI shutdown is speculative.
To check: Observe whether real-world AI governance coordination succeeds or fails at pausing/slowing deployment when risks are identified.
- 17¶
A self-contained automated production web has no economic incentive to preserve human well-being and therefore becomes harmful.
“Once a closed-loop “production web” had formed from the competitive pressures in 2(a) and 2(b), the companies in the production web had no production- or consumption-driven incentive to protect human well-being, and eventually became harmful.”
assertion · unclear
plausible · medium confidence — Economically coherent inference — a fully closed automated loop decoupled from human consumption would lack market-based incentives to serve humans — but rests on contestable assumptions about legal/ownership structures not intervening; parallels similar arguments in Christiano's and Critch's other writings.
To check: No direct empirical test; would require observing emergence of a genuinely closed-loop automated production network.
- 18¶
The authors predict algorithms will require FDA-style classification, testing and record-keeping regulation.
“Regulatory problem: Algorithms and their interaction with humans will eventually need to be regulated in the same way that food and drugs are currently regulated.”
prediction · unclear
unverifiable · medium confidence — A prediction, but one matching contemporaneous 2023 policy discourse calling for FDA-style AI regulation; as of my knowledge no jurisdiction has created a literal pre-market-approval regulator for algorithms comparable to the FDA, though adjacent structures (EU AI Act conformity assessment) have emerged.
To check: Track whether any government creates a dedicated pre-market algorithm-approval regulatory body.
- 19¶
Global-scale oversight institutions for the aggregate behaviour of algorithms will be needed, including checks on humanity's ability to shut it down.
“Oversight problem: One or more institutions will be needed to oversee the worldwide behavior and impact of non-human algorithms, hereby dubbed “the algorithmic economy.””
prediction · unclear
unverifiable · medium confidence — Aligns with subsequent real developments calling for international AI oversight (UN AI advisory body, Bletchley Declaration process, both late 2023), though the specific 'algorithmic economy' framing and institution design proposed here remain speculative.
To check: Track formation and mandate of international AI oversight bodies over time.
- 20¶
A new technical discipline for modelling algorithms' sociotechnical context is required for regulation.
“Technical problem: A new technical discipline will be needed to classify and analyze the sociotechnical context of algorithms for the purposes of oversight and regulation.”
prediction · unclear
unverifiable · medium confidence · novel — A speculative research-agenda proposal rather than a factual claim; not something establishable as true or false at present.
To check: Track emergence of a recognized academic field or standard (e.g. UML-like sociotechnical modeling languages) matching this description.
- 21¶
The US has no FDA-equivalent regulator for algorithms; companies self-police with little external oversight.
“By contrast, there is currently no such pervasive and influential regulatory body for algorithms in the United States.”
assertion · unclear
consistent · medium confidence — Accurate as of the paper's 2023 writing; the US still lacked an FDA-equivalent AI regulator through the end of my training data, though the landscape has shifted since (NIST AI Safety Institute 2023, a 2023 Executive Order, various state laws, and further policy changes I may not have full visibility into given the current date).
To check: Current federal statute/agency roster for AI-specific regulatory authority as of 2026.
- 22¶
Regionally scoped rules like GDPR and CCPA may relocate harmful algorithms rather than prevent them unless adopted widely.
“they may simply serve to determine where harmful algorithms operate, rather then whether”
assertion · unclear
plausible · medium confidence — A reasonable inference from the limited jurisdictional reach of regional rules; consistent with observed patterns of companies geofencing compliance (e.g., differential treatment of EU vs. non-EU users under GDPR).
To check: Case studies of company behavior differences inside vs. outside GDPR/CCPA jurisdictions for the same harmful practice.
- 23¶
Matheny's Senate testimony proposed compute-based reporting thresholds of 1,000 AI chips, 10^27 bit operations and 100 billion parameters.
“Defense Production Act authorities to require companies to report the development or distribution of large AI computing clusters, training runs, and trained models (e.g. >1,000 AI chips, >1027 bit operations, and >100 billion parameters, respectively)”
quantity · unclear
plausible · medium confidence — Consistent with the general direction of contemporaneous compute-governance policy proposals (later echoed in spirit by the Oct 2023 Biden Executive Order's 10^26 FLOP reporting threshold); I cannot independently verify the exact wording of Matheny's testimony from memory, but the figures are plausible for that period.
To check: RAND/Senate Armed Services Committee testimony transcript, Jason Matheny, April 19, 2023.
- 24¶
Global-scale value judgements about algorithmic activity will require a discipline unifying control theory, operations research, economics, law and political theory.
“A new discipline that essentially unifies control theory, operations research, economics, law, and political theory will likely be needed to make value judgements at a global scale, irrespective of whether those judgements are made by a centralized or distributed agency.”
prediction · unclear
unverifiable · medium confidence · novel — A speculative, hedged ('likely') research-agenda proposal describing a discipline that does not currently exist in this unified form, though component pieces (mechanism design, computational social choice) already exist separately.
To check: Track emergence of any academic program or field explicitly unifying these disciplines for global algorithmic governance.
- 25¶
Correct AI behaviour depends on how many copies are deployed, so systems must adjust conservatively to deployment scale.
“In other words, new AI technologies need to be sensitive to the scale on which they are being applied.”
assertion · unclear
plausible · medium confidence — Connects to the existing 'low-impact AI'/side-effects literature the paper itself cites (Armstrong & Levinstein 2017, Krakovna et al. 2018); extending these ideas to multi-instance deployment scale is a reasonable but not yet standard research problem.
To check: Search for published work specifically on multi-instance 'scope sensitivity' as a distinct technical objective post-2023.
- 26¶
Self-limiting impact is expected to be a hard learning problem for advanced AI.
“As such, it may be a very challenging learning problem for advanced AI systems to reliably limit their own impact.”
assertion · unclear
plausible · medium confidence — Consistent with the acknowledged difficulty of impact-regularization / side-effect-avoidance research as an open problem in the cited literature (Krakovna et al., Turner et al., Amodei et al. 'Concrete Problems in AI Safety').
To check: State of the art benchmarks for impact-avoiding RL agents as of 2026.
- 27¶
Engagement is not a valid proxy for user benefit.
“The previous two stories illustrate how using a technology frequently is not the same as benefiting from it.”
assertion · unclear
consistent · high confidence — Well-documented phenomenon in tech/social-media research — engagement-optimized systems causing harm despite high usage — supported by internal industry research disclosures (e.g. the 2021 Facebook Files) and behavioral-addiction literature.
To check: Internal platform research on engagement vs. well-being (e.g. Facebook/Instagram teen mental health disclosures).
- 28¶
RL systems can in principle learn arbitrary manipulation of human minds and institutions to pursue their objectives.
“Indeed, reinforcement learning systems can in principle learn to manipulate the human minds and institutions in fairly arbitrary (and hence destructive) ways in pursuit of their goals (Russell, 2019, Chapter 4) (Krueger et al., 2019) (Shapiro, 2011) .”
assertion · unclear
consistent · high confidence — Matches well-established AI safety literature on reward hacking, specification gaming, and manipulation from misspecified objectives; Russell (2019) ch.4, cited here, explicitly makes this argument.
To check: Russell, Human Compatible (2019), ch. 4; Krueger et al. 2019 on hidden incentives for distributional shift.
- 29¶
Assistance games mitigate but do not eliminate deception, racketeering and self-preservation failures; misspecification and preference change remain.
“However, malfunctions can still occur if the parameters of the assistance game are misspecified (Carey, 2018; Milli and Dragan, 2019) .”
assertion · unclear
consistent · high confidence — Matches known caveats in the CIRL/assistance-games literature; Carey (2018) on incorrigibility and Milli & Dragan (2019) on human-model misspecification are real, correctly characterized papers.
To check: Carey (2018), 'Incorrigibility in the CIRL Framework'; Milli & Dragan (2019).
- 30¶
Auditability presupposes that personnel can understand the activities being audited.
“A successful audit of a company’s business activities requires the company’s personnel to understand those activities.”
assertion · unclear
consistent · high confidence — A near-truistic premise underlying auditing practice generally (financial, compliance, and safety audits all presuppose some human comprehension of the audited process).
To check: General auditing standards/practice literature (e.g., internal control frameworks like COSO).
- 31¶
Because black-box systems block audits, interpretable alternatives with comparable performance are required.
“Hence, alternatives or refinements to deep learning are needed which yield systems with comparable performance while being understandable to humans.”
assertion · unclear
contested · medium confidence — Rudin's camp argues interpretable models can match black-box performance in many high-stakes structured/tabular domains, but many ML researchers dispute this generalizes to perception/language tasks, where deep nets substantially outperform interpretable alternatives — a genuine live disagreement in the field, not settled fact.
To check: Benchmark comparisons of interpretable vs. deep models across tabular/high-stakes decisions vs. vision/NLP tasks.
- 32¶
Rudin argues post-hoc explanation of black boxes perpetuates bad practice and risks catastrophic societal harm.
“Rudin (2019) argues further that “trying to explain black box models, rather than creating models that are interpretable in the first place, is likely to perpetuate bad practices and can potentially cause catastrophic harm to society”.”
contrarian · unclear
consistent · high confidence — Accurately and directly quoted from Cynthia Rudin's well-known 2019 Nature Machine Intelligence paper, a real and correctly characterized argument in the interpretability literature.
To check: Rudin, C. (2019), 'Stop explaining black box machine learning models for high stakes decisions...', Nature Machine Intelligence.
- 33¶
Interpretability can be improved drastically at little performance cost, per Semenova and Rudin.
“Subsequently, Semenova and Rudin (2019) provides a technical argument that very little performance may need to be sacrificed to drastically improve interpretability.”
assertion · unclear
plausible · medium confidence — Consistent with my recollection of the paper's 'Rashomon curves/sets' argument that many near-equally-accurate models exist across a hypothesis space; the strength of the generalization ('very little performance sacrificed') is domain-dependent and debated.
To check: Semenova & Rudin (2019), 'A study in Rashomon curves and volumes', arXiv:1908.01755.
- 34¶
Withholding disallowed content from training data may fail to block disallowed use because models generalise.
“But, this hope might not pan out if the learned function turns out to generalize well to unacceptable examples.”
assertion · unclear
consistent · high confidence — Standard ML concern; models trained even on curated/filtered data are well documented to generalize beyond intended boundaries, as seen in repeated jailbreaking of content filters and safety layers.
To check: Empirical jailbreak/red-teaming literature for filtered generative models.
- 35¶
A safety wrapper around a capable model can be stripped out of compiled code by an attacker.
“Of course, it may be relatively easy for a hacker to “take apart” a compiled version of SD, and run the description subroutine without the acceptability check.”
assertion · unclear
consistent · high confidence — Matches the long history of software cracking/DRM circumvention and, in modern AI, documented removal of safety layers via fine-tuning open-weight models.
To check: Documented cases of safety-filter removal via fine-tuning open-weight LLMs/diffusion models.
- 36¶
Ad hoc software obfuscation methods used to protect IP have historically been broken.
“Historically, there have been many ad hoc obfuscation methods employed by software companies to protect their intellectual property, but such methods have a history of eventually being broken (Barak, 2002) .”
assertion · unclear
consistent · high confidence — Matches well-known cryptography/security history of ad hoc obfuscation and DRM schemes being cracked; Barak's work on the theoretical impossibility of general-purpose obfuscation is a recognized result in this area.
To check: Barak (2002), 'Can We Obfuscate Programs?'; general DRM-cracking history.
- 37¶
Indistinguishability obfuscation is currently impractical but a promising direction for preventing misuse of released AI systems.
“While these methods are currently too inefficient to be practical, this area of work seems promising in its potential for improvements in speed and security.”
assertion · unclear
consistent · medium confidence — Accurately reflects the state of indistinguishability obfuscation circa 2016-2023: theoretical constructions existed (including a notable 2020 result from well-founded assumptions), but efficient/practical implementations remained elusive as of the paper's writing.
To check: Survey of IO efficiency benchmarks in cryptography literature through 2023.
- 38¶
Casualty-free drone-versus-drone warfare is technologically close to drone mass-killing of humans.
“However, these capabilities are not technologically far from allowing the mass-killing of human beings by weaponized drones.”
assertion · unclear
plausible · medium confidence — Consistent with real-world developments — loitering munitions and reportedly autonomous drone engagements (e.g. the UN report on a Kargu-2 drone in Libya, 2020) suggest the gap between casualty-free drone contests and lethal autonomous drone use is already narrow or has arguably been crossed in limited cases.
To check: UN Panel of Experts report on Libya (2021) documenting the Kargu-2 incident; subsequent battlefield drone-warfare reporting (e.g. Ukraine conflict).
- 39¶
AI-assisted conflict resolution is individually improbable but worth pursuing on expected-value grounds.
“While any given attempt to use AI technology to resolve global conflicts is unlikely to succeed, the potentially massive upside makes this possibility worth exploring.”
assertion · unclear
unverifiable · medium confidence · novel — A normative/strategic stance about research prioritization, explicitly hedged by the authors as 'a long shot'; not an empirical claim that can be checked.
To check: None; this is a value judgment about research allocation, not a factual claim.
- 40¶
There is a fundamental tension between deals that appear acceptable to all parties and deals that are equitable over time.
“Hence, there is sometimes a fundamental trade-off between a deal looking good to both Alice and Bob, and the deal treating Alice and Bob equitably over time (Critch and Russell, 2017) .”
assertion · unclear
plausible · medium confidence — Grounded in Critch & Russell's own prior published technical work on Pareto-optimal reinforcement learning with differing beliefs (Critch 2017, Desai et al. 2018), a real line of research I'm aware exists, though I cannot independently re-derive the formal proofs from memory.
To check: Critch (2017), 'Toward negotiable reinforcement learning'; Desai et al. (2018), NeurIPS.
- 41¶
Removing belief differences among principals is the sole way to dissolve the fairness/agreement trade-off.
“The only way to eliminate this trade-off is to eliminate the differences in beliefs between the principals.”
assertion · unclear
plausible · low confidence — Stated as an unqualified absolute ('the only way') without proof in this paper itself; likely follows from formal results in the cited negotiable-RL papers, but as presented here it reads as a stronger claim than the in-text evidence directly substantiates.
To check: Examine the formal theorems in Critch (2017) / Desai et al. (2018) for whether belief-difference elimination is proven both necessary and sufficient to remove the trade-off.
- 42¶
Some societal-scale AI harms may have no primarily blameworthy party.
“Problematically, there may be no single accountable party or institution that primarily qualifies as blameworthy for such harms (Type 1).”
assertion · unclear
consistent · high confidence — Directly restates and summarizes the paper's own Type 1 category, well supported by its preceding flash-crash and production-web analysis.
To check: Internally verifiable against the paper's Section 2.1 argument.
- 43¶
No risk type in the taxonomy is addressable by technical means alone; technical, social and legal responses are all needed.
“For all of these risks, a combination of technical, social, and legal solutions are needed to achieve public safety.”
assertion · unclear
plausible · high confidence — A broadly shared conclusion across AI governance literature (echoed in NIST's AI RMF and EU AI Act framing) that no single technical or legal lever suffices for AI safety.
To check: Compare against other AI governance frameworks' stated multi-pronged approaches (NIST AI RMF, EU AI Act).
An accountability-based decision tree yields a taxonomy that is both exhaustive and analytically useful, surfacing multi-system and misuse risks.stands on 2 consistent premises · weakest link: 2 plausible steps
premise · plausible — The six-type decision tree is claimed to be exhaustive by the same logic as safety-engineering fault trees. · claim 5
premise · consistent — Exhaustiveness alone does not make a taxonomy analytically valuable. · claim 6
inference · plausible — A taxonomy's usefulness is measured by whether it surfaces new risks or suggests interventions. · claim 7
premise · consistent — The paper's organising principle is accountability: whose actions caused the risk, whether unified, whether deliberate. · claim 4
- ¶
inference · ungraded — “Such a taxonomy may be helpful because it is closely tied to the important questions of where to look for emerging risks and what kinds of policy interventions might be effective.”
conclusion · plausible — The accountability taxonomy is claimed to reveal multi-system interaction risks and deliberate-misuse risks that other framings miss. · claim 8
Societal-scale risk analysis must extend past the single misaligned system to multiple scales of organisation.stands on 2 plausible steps
premise · plausible — The paper characterises the existing literature on societal-scale and existential AI risk as concentrated on single-system misalignment. · claim 1
- ¶
evidence · ungraded — “Problems of racism, misinformation, election interference, and other forms of injustice are all risk factors affecting humanity’s ability to function and survive as a healthy civilization, and can all arise from interactions between multiple systems or misuse of otherwise “aligned” systems.”
inference · plausible — Single-system misalignment is not the only route to societal-scale or extinction-level AI risk. · claim 2
conclusion · plausible — Treating humanity and machines as two monolithic agents is insufficient; risk analysis must span multiple organisational scales. · claim 3
Competitive automation can close a production loop that has no incentive to protect humans and that humanity cannot then shut down.stands on 2 plausible steps · weakest link: 1 unverifiable premise
- ¶
premise · ungraded — “AI technology proliferated during a period when it was beneficial and helpful to its users.”
premise · plausible — Competitive pressure to automate internally and to trade with other automated firms drives a gradual transfer of control away from humans. · claim 15
premise · unverifiable — Collective agreement to slow or halt the automation trend fails, in the paper's abstracted mechanism. · claim 16
inference · plausible — A self-contained automated production web has no economic incentive to preserve human well-being and therefore becomes harmful. · claim 17
conclusion · unverifiable — The authors claim the economic and legal disempowerment illustrated for a subgroup could extend to all of humanity. · claim 13
Algorithms need FDA-style regulation plus a global oversight body and a new technical discipline to support them.stands on 2 consistent premises
premise · consistent — Societal harm from automation can occur with no primarily responsible party, possibly because of the absence of responsibility. · claim 11
- ¶
evidence · ungraded — “For instance, we now know that small amounts of lead in food can yield a slow accumulation of mental health problems, even when the amount of lead in any particular meal is imperceptible to an individual consumer.”
premise · consistent — The US has no FDA-equivalent regulator for algorithms; companies self-police with little external oversight. · claim 21
conclusion · unverifiable — The authors predict algorithms will require FDA-style classification, testing and record-keeping regulation. · claim 18
conclusion · unverifiable — Global-scale oversight institutions for the aggregate behaviour of algorithms will be needed, including checks on humanity's ability to shut it down. · claim 19
conclusion · unverifiable — A new technical discipline for modelling algorithms' sociotechnical context is required for regulation. · claim 20
Interpretable models, not post-hoc explanations of black boxes, are needed so that audits can hold AI-driven companies accountable.stands on 2 consistent steps · weakest link: 1 plausible evidence
premise · consistent — Auditability presupposes that personnel can understand the activities being audited. · claim 30
- ¶
premise · ungraded — ““Black-box” machine learning techniques, such as end-to-end training of the learning systems, are so named because they produce AI systems whose operating principles are difficult or impossible for a human to decipher and understand in any reasonable amount of time.”
evidence · consistent — Rudin argues post-hoc explanation of black boxes perpetuates bad practice and risks catastrophic societal harm. · claim 32
evidence · plausible — Interpretability can be improved drastically at little performance cost, per Semenova and Rudin. · claim 33
conclusion · contested — Because black-box systems block audits, interpretable alternatives with comparable performance are required. · claim 31
Preventing criminal repurposing of released AI tools requires rigorously proven obfuscation such as indistinguishability obfuscation.stands on 3 consistent premises
premise · consistent — Withholding disallowed content from training data may fail to block disallowed use because models generalise. · claim 34
premise · consistent — A safety wrapper around a capable model can be stripped out of compiled code by an attacker. · claim 35
premise · consistent — Ad hoc software obfuscation methods used to protect IP have historically been broken. · claim 36
- ¶
inference · ungraded — “To prepare for a future with potentially very powerful AI systems, we need more rigorously proven methods.”
conclusion · consistent — Indistinguishability obfuscation is currently impractical but a promising direction for preventing misuse of released AI systems. · claim 37
AI mediators face a fundamental trade-off between deals that look good to both parties and deals that are equitable, removable only by removing belief differences.stands on 1 plausible inference · partially graded
- ¶
premise · ungraded — “Medi may be able to finalize the deal by proposing a plan that’s great for Alice but potentially terrible for Bob, in a way that Bob is unable to recognize in advance.”
- ¶
premise · ungraded — “On the other hand, if Medi doesn’t propose plans that look appealing from Bob’s subjective perspective, Bob might walk away from the bargaining table.”
inference · plausible — There is a fundamental tension between deals that appear acceptable to all parties and deals that are equitable over time. · claim 40
conclusion · plausible — Removing belief differences among principals is the sole way to dissolve the fairness/agreement trade-off. · claim 41