Person
Stephen Casper
Cited in one post, AI Alignment with Changing and Influenceable…, since September 2026
In their words
“More generally, it seems challenging to say the least to explicitly hardcode full models of “human bias” mathematically, as reflected by the limitations of current reward learning methods (McKinney et al., 2023; Tien et al., 2023; Casper et al., 2023) .”
· AI Alignment with Changing and Influenceable Reward Functions · machine-resolved