Person
Tom Everitt
Cited in two posts, AI Alignment with Changing and Influenceable… and Characterizing Manipulation from AI Systems, since September 2026
In their words
“Using the notation from Everitt et al. (2021a) , a CID is a graphical model that distinguishes decision nodes where an AI system makes a decision, structure nodes which capture important variables in an environment and their effects on each other, and utility nodes which the AI system is trained to optimize.”
· Characterizing Manipulation from AI Systems · machine-resolved
“MC would like to thank (in no particular order) Tom Everitt, Cam Allen, Cassidy Laidlaw, Nora Amman, Rohin Shah, Alan Chan, Richard Pettigrew, Marcus Pivato, Orr Paradise, Ann He, Henri Wadsworth, Alex Pan, Erik Jones, Riqui Zhong, Tom Gilbert, Atoosa Kasirzadeh, Vincent Cognitzer, Daniel Kilov, Iason Gabriel, and the members of the Center for Human Compatible AI (CHAI) and InterAct lab.”
· AI Alignment with Changing and Influenceable Reward Functions · machine-resolved
In the claim ledger
4 promoted claims about them. Assessments are the model’s knowledge, not verification.
Having an incentive is not the same as pursuing it.
consistent · Characterizing Manipulation from AI Systems
The paper's influence-incentive definition deliberately covers accidental side effects, unlike instrumental control incentives.
consistent · AI Alignment with Changing and Influenceable Reward Functions
The proposed DR-MDP objectives are, with exceptions, tractable with existing RL methods.
plausible · AI Alignment with Changing and Influenceable Reward Functions
The paper claims Everitt et al.'s TI-unaware reward modelling algorithm removes direct influence incentives but not influence incentives in the paper's broader sense.
plausible · AI Alignment with Changing and Influenceable Reward Functions