Multi-Armed Bandits & Best-Arm Identification Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

From the finite-time analysis of UCB to modern best-arm identification, this list traces the core ideas behind sequential decision-making under uncertainty. It emphasizes regret bounds, information-theoretic limits, and practical algorithms that balance exploration and exploitation.

Multi-Armed Bandits & Best-Arm Identification: 10 key papers

  1. Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
    Bubeck and Cesa-Bianchi. arXiv 2012.
  2. A Tutorial on Thompson Sampling
    Russo et al. arXiv 2017.
  3. Linearly Parameterized Bandits
    Rusmevichientong and Tsitsiklis. arXiv 2008.
  4. Introduction to Multi-Armed Bandits
    Slivkins. arXiv 2019.
  5. Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling
    Riquelme et al. arXiv 2018.
  6. Minimal Exploration in Structured Stochastic Bandits
    Combes et al. arXiv 2017.
  7. Analysis of Thompson Sampling for the multi-armed bandit problem
    Agrawal and Goyal. arXiv 2011.
  8. Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits
    Agarwal et al. arXiv 2014.
  9. Thompson Sampling: An Asymptotically Optimal Finite Time Analysis
    Kaufmann et al. arXiv 2012.
  10. lil' UCB : An Optimal Exploration Algorithm for Multi-Armed Bandits
    Jamieson et al. arXiv 2013.
← Back to main page