Person
Richard Sutton
Cited in two posts, AI Alignment with Changing and Influenceable… and DHH: Future of Programming, AI, Agentic…, since September 2026
Brought up by David Heinemeier Hansson and Lex Fridman
In their words
“This was one of the-- I think it was an interview with Sutton or maybe one of the other original guys talking about this sense that intelligence is perhaps not as constrained just to the specific one spot.”
[4:07:37] · DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux | Lex Fridman Podcast #501 · machine-resolved
In the claim ledger
1 promoted claim about them. Assessments are the model’s knowledge, not verification.
Iteratively retrained myopic optimization of long-term metrics converges to the non-myopic RL optimum.
consistent · AI Alignment with Changing and Influenceable Reward Functions