Self-Play & Population-Based Training Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

Competing against past versions of itself, an agent can bootstrap from scratch to superhuman play. This list covers the theory and practice of self-play, league training, and population-based optimization that powered landmark systems.

Self-Play & Population-Based Training: 10 key papers

  1. Mastering the game of Go without human knowledge
    Silver et al. Nature 2017.
  2. Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
    Silver et al. arXiv 2017.
  3. Emergent Complexity via Multi-Agent Competition
    Bansal et al. arXiv 2017.
  4. Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
    Heinrich and Silver. arXiv 2016.
  5. Asymmetric self-play for automatic goal discovery in robotic manipulation
    Plappert et al. arXiv 2021.
  6. Fictitious Self-Play in Extensive-Form Games
    Heinrich et al. International Conference on Machine Learning 2015.
  7. Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
    Brown et al. arXiv 2020.
  8. Student of Games: A unified learning algorithm for both perfect and imperfect information games
    Schmid et al. arXiv 2021.
  9. Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning
    Perolat et al. arXiv 2022.
  10. Learning to Play No-Press Diplomacy with Best Response Policy Iteration
    Anthony et al. arXiv 2020.
← Back to main page