Person
Yuntao Bai
Cited in one post, AI Alignment with Changing and Influenceable…, since September 2026
In their words
“While reward functions may not be the best way to encode certain targets of alignment, such as norms or contractualist values (Hadfield-Menell & Hadfield, 2018; Zhi-Xuan et al., 2024; Bai et al., 2022) , they are still sufficiently expressive to encode any desired behavior (while potentially requiring to drop the Markovian assumption, as discussed in Section A.7).”
· AI Alignment with Changing and Influenceable Reward Functions · machine-resolved