Person
Zachary Kenton
Cited in one post, Characterizing Manipulation from AI Systems, since September 2026
In their words
“In the context of language models, Kenton et al. (2021) ’s definition of manipulation also requires that the response of the human benefits the AI system in some way, which can be thought of as a notion of incentives.”
· Characterizing Manipulation from AI Systems · machine-resolved
In the claim ledger
3 promoted claims about them. Assessments are the model’s knowledge, not verification.
Existing definitions of AI manipulation fail on either implementability or generality.
plausible · Characterizing Manipulation from AI Systems
Incentives cannot be read off the objective alone; the whole training setup matters.
plausible · Characterizing Manipulation from AI Systems
RLHF creates an incentive to win labeller approval through manipulation absent behavioural constraints.
plausible · Characterizing Manipulation from AI Systems