Person
Derek Parfit
Cited in one post, AI Alignment with Changing and Influenceable…, since September 2026
In their words
“Figure 2: Writer’s curse (adapted from Parfit (1984) , p.”
· AI Alignment with Changing and Influenceable Reward Functions · machine-resolved
In the claim ledger
4 promoted claims about them. Assessments are the model’s knowledge, not verification.
Headline conclusion, hedged: there may be no principled definition of optimal AI behaviour once preferences change.
plausible · AI Alignment with Changing and Influenceable Reward Functions
In the writer's-curse example, maximising the initial reward function requires pushing the user into a state whose new reward function rejects it — influence 'away from' the optimized preferences.
consistent · AI Alignment with Changing and Influenceable Reward Functions
There may be no uniquely correct notion of trajectory-level optimality for a person with changing preferences.
plausible · AI Alignment with Changing and Influenceable Reward Functions
The paper reports Parfit's rejection of timeless evaluation of a life, and takes it as motivation for the difficulty of specifying AI objectives over changing selves.
plausible · AI Alignment with Changing and Influenceable Reward Functions