AI Alignment & Safety Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

As capabilities grow, ensuring systems do what we intend becomes the hard problem. This list gathers concrete problems, scalable oversight ideas, and empirical approaches to building agents that remain corrigible and helpful.

AI Alignment & Safety: 10 key papers

  1. Concrete Problems in AI Safety
    Amodei et al. arXiv 2016.
  2. Scalable agent alignment via reward modeling: a research direction
    Leike et al. arXiv 2018.
  3. Risks from Learned Optimization in Advanced Machine Learning Systems
    Hubinger et al. arXiv 2019.
  4. Artificial Intelligence, Values and Alignment
    Gabriel. arXiv 2020.
  5. AI Safety Gridworlds
    Leike et al. arXiv 2017.
  6. The Off-Switch Game
    Hadfield-Menell et al. arXiv 2016.
  7. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
    Burns et al. arXiv 2023.
  8. Goal Misgeneralization in Deep Reinforcement Learning
    Langosco et al. arXiv 2021.
  9. Unsolved Problems in ML Safety
    Hendrycks et al. arXiv 2021.
  10. Alignment of Language Agents
    Kenton et al. arXiv 2021.
← Back to main page