Person
Richard Pettigrew
Cited in one post, AI Alignment with Changing and Influenceable…, since September 2026
In their words
“MC would like to thank (in no particular order) Tom Everitt, Cam Allen, Cassidy Laidlaw, Nora Amman, Rohin Shah, Alan Chan, Richard Pettigrew, Marcus Pivato, Orr Paradise, Ann He, Henri Wadsworth, Alex Pan, Erik Jones, Riqui Zhong, Tom Gilbert, Atoosa Kasirzadeh, Vincent Cognitzer, Daniel Kilov, Iason Gabriel, and the members of the Center for Human Compatible AI (CHAI) and InterAct lab.”
· AI Alignment with Changing and Influenceable Reward Functions · machine-resolved
In the claim ledger
5 promoted claims about them. Assessments are the model’s knowledge, not verification.
There may be no uniquely correct notion of trajectory-level optimality for a person with changing preferences.
plausible · AI Alignment with Changing and Influenceable Reward Functions
The paper endorses Pettigrew's weighted-average-of-selves theory as progress, against Paul's dissent.
unverifiable · AI Alignment with Changing and Influenceable Reward Functions
The paper claims Pettigrew's weighting scheme risks producing inconsistent multi-step plans if weights are re-evaluated at each node.
plausible · AI Alignment with Changing and Influenceable Reward Functions
The paper reports Pettigrew's stronger before-and-after agreement criterion, offered because the post-hoc test can be gamed by preference-altering manipulation.
unverifiable · AI Alignment with Changing and Influenceable Reward Functions
The paper positions its Unambiguous Desirability property as the multi-timestep, AI-policy generalization of Pettigrew's nudge heuristic.
plausible · AI Alignment with Changing and Influenceable Reward Functions