Person
Max Jaderberg
Cited in one post, Characterizing Manipulation from AI Systems, since September 2026
In their words
“In terms of determining the primitives upon which nodes in a CID can be constructed, interpretability tools may help: for example, Jaderberg et al. (2019) finds that RL agents trained to play capture-the-flag have neural activation patterns that correspond to important concepts in game, such as the status of the flag.”
· Characterizing Manipulation from AI Systems · machine-resolved
In the claim ledger
1 promoted claim about them. Assessments are the model’s knowledge, not verification.
RL agents develop internal representations corresponding to human-legible game concepts.
consistent · Characterizing Manipulation from AI Systems