Person
Kenneth Li
Cited in one post, Characterizing Manipulation from AI Systems, since September 2026
In their words
“Li et al. (2023) provide evidence from interpretability tools that language models trained only on transcripts of board game play can learn to model the underlying board state of the game.”
· Characterizing Manipulation from AI Systems · machine-resolved
In the claim ledger
1 promoted claim about them. Assessments are the model’s knowledge, not verification.
Interpretability work shows sequence models trained on game transcripts represent latent board state.
consistent · Characterizing Manipulation from AI Systems