Safe Reinforcement Learning & Constrained MDPs Reading List
Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026
Safety in reinforcement learning is formalized through constrained MDPs and algorithmic guarantees. This list covers Lagrangian methods, trust-region projections, safety critics, and modern benchmarks that keep exploration within feasible regions.
Safe Reinforcement Learning & Constrained MDPs: 10 key papers
- Constrained Policy Optimization
Achiam et al. arXiv 2017.
- Safe Exploration in Continuous Action Spaces
Dalal et al. arXiv 2018.
- Lyapunov-based Safe Policy Optimization for Continuous Control
Chow et al. arXiv 2019.
- A Lyapunov-based Approach to Safe Reinforcement Learning
Chow et al. arXiv 2018.
- Safe Model-based Reinforcement Learning with Stability Guarantees
Berkenkamp et al. arXiv 2017.
- Reward Constrained Policy Optimization
Tessler et al. arXiv 2018.
- Projection-Based Constrained Policy Optimization
Yang et al. arXiv 2020.
- Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
Stooke et al. arXiv 2020.
- Safe Reinforcement Learning via Shielding
Alshiekh et al. arXiv 2017.
- A Review of Safe Reinforcement Learning: Methods, Theory and Applications
Gu et al. arXiv 2022.
← Back to main page