Person
Rohin Shah
cited by 12 posts · in three sources · first cited September 2026
In their words
“Rohin Shah”
· · machine-resolved
“Shah et al. (2022) provide a speculative larger-scale example, conjecturing that InstructGPT’s competent responses to questions its developers didn’t intend it to answer (such as questions about how to commit crimes) resulted from goal misgeneralization (rather than reward misspecification).”
· · machine-resolved
“Thanks to Rohin Shah for discussion of the humans example.”
· Existential Risk from Power-Seeking AI · machine-resolved