Skip to content

Person

Lauro Langosco

Cited in seven posts, among them Characterizing Manipulation from AI Systems and The Alignment Problem from a Deep Learning…, since September 2026

In their words

  • As an example of goal misgeneralization, Langosco et al. (2022) describe a toy environment where rewards were given for opening boxes, which required agents to collect one key per box.

    · The Alignment Problem from a Deep Learning Perspective · machine-resolved

  • Langosco et al. (2022) and Shah et al. (2022) show that both language models and general RL agents can pursue different goals in out-of-distribution environments even when trained to perfect accuracy on in-distribution environments.

    · Characterizing Manipulation from AI Systems · machine-resolved

In the claim ledger

1 promoted claim about them. Assessments are the model’s knowledge, not verification.