Person
Alex Pan
Cited in seven posts, among them AI Alignment with Changing and Influenceable… and The Alignment Problem from a Deep Learning…, since September 2026
In their words
“Across diverse text-based social environments, Pan et al. (2023) find that language models fine-tuned to maximize the game-reward take the most power-seeking actions.”
· The Alignment Problem from a Deep Learning Perspective · machine-resolved
“MC would like to thank (in no particular order) Tom Everitt, Cam Allen, Cassidy Laidlaw, Nora Amman, Rohin Shah, Alan Chan, Richard Pettigrew, Marcus Pivato, Orr Paradise, Ann He, Henri Wadsworth, Alex Pan, Erik Jones, Riqui Zhong, Tom Gilbert, Atoosa Kasirzadeh, Vincent Cognitzer, Daniel Kilov, Iason Gabriel, and the members of the Center for Human Compatible AI (CHAI) and InterAct lab.”
· AI Alignment with Changing and Influenceable Reward Functions · machine-resolved