Multi-Armed Bandits & Best-Arm Identification Reading List
Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026
From the finite-time analysis of UCB to modern best-arm identification, this list traces the core ideas behind sequential decision-making under uncertainty. It emphasizes regret bounds, information-theoretic limits, and practical algorithms that balance exploration and exploitation.
Multi-Armed Bandits & Best-Arm Identification: 10 key papers
- Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Bubeck and Cesa-Bianchi. arXiv 2012.
- A Tutorial on Thompson Sampling
Russo et al. arXiv 2017.
- Linearly Parameterized Bandits
Rusmevichientong and Tsitsiklis. arXiv 2008.
- Introduction to Multi-Armed Bandits
Slivkins. arXiv 2019.
- Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling
Riquelme et al. arXiv 2018.
- Minimal Exploration in Structured Stochastic Bandits
Combes et al. arXiv 2017.
- Analysis of Thompson Sampling for the multi-armed bandit problem
Agrawal and Goyal. arXiv 2011.
- Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits
Agarwal et al. arXiv 2014.
- Thompson Sampling: An Asymptotically Optimal Finite Time Analysis
Kaufmann et al. arXiv 2012.
- lil' UCB : An Optimal Exploration Algorithm for Multi-Armed Bandits
Jamieson et al. arXiv 2013.
← Back to main page