Safe Reinforcement Learning & Constrained MDPs Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

Safety in reinforcement learning is formalized through constrained MDPs and algorithmic guarantees. This list covers Lagrangian methods, trust-region projections, safety critics, and modern benchmarks that keep exploration within feasible regions.

Safe Reinforcement Learning & Constrained MDPs: 10 key papers

  1. Constrained Policy Optimization
    Achiam et al. arXiv 2017.
  2. Safe Exploration in Continuous Action Spaces
    Dalal et al. arXiv 2018.
  3. Lyapunov-based Safe Policy Optimization for Continuous Control
    Chow et al. arXiv 2019.
  4. A Lyapunov-based Approach to Safe Reinforcement Learning
    Chow et al. arXiv 2018.
  5. Safe Model-based Reinforcement Learning with Stability Guarantees
    Berkenkamp et al. arXiv 2017.
  6. Reward Constrained Policy Optimization
    Tessler et al. arXiv 2018.
  7. Projection-Based Constrained Policy Optimization
    Yang et al. arXiv 2020.
  8. Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
    Stooke et al. arXiv 2020.
  9. Safe Reinforcement Learning via Shielding
    Alshiekh et al. arXiv 2017.
  10. A Review of Safe Reinforcement Learning: Methods, Theory and Applications
    Gu et al. arXiv 2022.
← Back to main page