AI Alignment & Safety Reading List
Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026
As capabilities grow, ensuring systems do what we intend becomes the hard problem. This list gathers concrete problems, scalable oversight ideas, and empirical approaches to building agents that remain corrigible and helpful.
AI Alignment & Safety: 10 key papers
- Concrete Problems in AI Safety
Amodei et al. arXiv 2016.
- Scalable agent alignment via reward modeling: a research direction
Leike et al. arXiv 2018.
- Risks from Learned Optimization in Advanced Machine Learning Systems
Hubinger et al. arXiv 2019.
- Artificial Intelligence, Values and Alignment
Gabriel. arXiv 2020.
- AI Safety Gridworlds
Leike et al. arXiv 2017.
- The Off-Switch Game
Hadfield-Menell et al. arXiv 2016.
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Burns et al. arXiv 2023.
- Goal Misgeneralization in Deep Reinforcement Learning
Langosco et al. arXiv 2021.
- Unsolved Problems in ML Safety
Hendrycks et al. arXiv 2021.
- Alignment of Language Agents
Kenton et al. arXiv 2021.
← Back to main page