Person
Andrew Barto
Cited in one post, AI Alignment with Changing and Influenceable…, since September 2026
In the claim ledger
1 promoted claim about them. Assessments are the model’s knowledge, not verification.
Iteratively retrained myopic optimization of long-term metrics converges to the non-myopic RL optimum.
consistent · AI Alignment with Changing and Influenceable Reward Functions