A formal argument that alignment methods assuming static preferences are unsound once an AI can change what people want, with a framework for reasoning about it. What the claims are, and what they rest on.
Space
Papers
Close readings of the literature the risk model is built from — what a paper claims, and what those claims rest on.
The literature this publication reasons from: preprints and reports on AI risk, alignment, capability and governance, read one at a time. Most of the risk model’s conditionals were mined from papers like these, and until now the conclusions were public while the sources were not.
A piece here carries the same parts as a recap, in the same order: the claims the paper makes, each one quoted, graded and assessed, and the argument chains they form — written so that someone who never opens the PDF can argue with it, claim by claim.
What the reading must add for this source class is the evidence class. A number from a measured experiment, a step in a formal argument and a restatement of somebody else’s finding are three different kinds of thing, and a claim is only as strong as which one it is. Not peer review, and not a ranking by venue — most of this is arXiv, and a paper is weighed by what it shows.
Conventions
Followed, not enforced. Nothing here checks them.
- Every piece reads exactly one paper, declared as its `paper:` source.
- The title is the paper's own title, so a piece is findable by the work it reads rather than by what someone decided to call the reading.
- A paper is read for what it argues and what that rests on — a measurement, a formal argument, a survey of others' work — never for where it was published.
6 pieces, newest first
A taxonomy of societal-scale AI risk sorted by how much of the harm was anyone's intention, with a story for each type. What the claims are, and what they rest on.
A six-premise argument, each premise given a probability, for how likely it is that power-seeking AI causes an existential catastrophe by 2070. What the claims are, and what they rest on.
Four axes for saying whether an AI system is manipulating someone, and the argument that the absence of a definition is itself the risk. What the claims are, and what they rest on.
A pre-registered audit of Twitter's engagement-based timeline against a reverse-chronological baseline — measured on real users, not modelled. What the claims are, and what they rest on.
A position paper arguing that the way we train large models pushes them toward situational awareness, deceptive alignment and power-seeking. What the claims are, and what they rest on.