Adversarial Reinforcement Learning & Robustness to Adversaries Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

Core papers on adversarial reinforcement learning and robustness — from early demonstrations of policy vulnerability to principled robust MDP formulations and certified defenses.

Adversarial Reinforcement Learning & Robustness to Adversaries: 10 key papers

  1. Robust Adversarial Reinforcement Learning
    Pinto et al. arXiv 2017.
  2. Adversarial Attacks on Neural Network Policies
    Huang et al. arXiv 2017.
  3. Delving into adversarial attacks on deep policies
    Kos and Song. arXiv 2017.
  4. Robust Deep Reinforcement Learning with Adversarial Attacks
    Pattanaik et al. arXiv 2017.
  5. Robust Reinforcement Learning using Adversarial Populations
    Vinitsky et al. arXiv 2020.
  6. Certifiable Robustness to Adversarial State Uncertainty in Deep Reinforcement Learning
    Everett et al. arXiv 2020.
  7. Characterizing Attacks on Deep Reinforcement Learning
    Pan et al. arXiv 2019.
  8. Adversarial Policies: Attacking Deep Reinforcement Learning
    Gleave et al. arXiv 2019.
  9. Robust Reinforcement Learning on State Observations with Learned Optimal Adversary
    Zhang et al. arXiv 2021.
  10. Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RL
    Sun et al. arXiv 2021.
← Back to main page