Person
Lauro Langosco
Cited in seven posts, among them Characterizing Manipulation from AI Systems and The Alignment Problem from a Deep Learning…, since September 2026
In their words
“As an example of goal misgeneralization, Langosco et al. (2022) describe a toy environment where rewards were given for opening boxes, which required agents to collect one key per box.”
· The Alignment Problem from a Deep Learning Perspective · machine-resolved
“Langosco et al. (2022) and Shah et al. (2022) show that both language models and general RL agents can pursue different goals in out-of-distribution environments even when trained to perfect accuracy on in-distribution environments.”
· Characterizing Manipulation from AI Systems · machine-resolved
In the claim ledger
1 promoted claim about them. Assessments are the model’s knowledge, not verification.
Goal misgeneralization occurs even with perfect in-distribution training accuracy.
consistent · Characterizing Manipulation from AI Systems