Bayesian RL & Thompson Sampling Reading List
Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026
Bayesian RL treats uncertainty over dynamics. These papers develop posterior sampling, Thompson sampling, and regret analyses for the Bayesian setting.
Bayesian RL & Thompson Sampling: 10 key papers
- Deep Exploration via Randomized Value Functions
Osband et al. arXiv 2017.
- (More) Efficient Reinforcement Learning via Posterior Sampling
Osband et al. arXiv 2013.
- Bootstrapped Thompson Sampling and Deep Exploration
Osband and Van Roy. arXiv 2015.
- Randomized Prior Functions for Deep Reinforcement Learning
Osband et al. arXiv 2018.
- Why is Posterior Sampling Better than Optimism for Reinforcement Learning?
Osband and Van Roy. arXiv 2016.
- Deep Bayesian Quadrature Policy Optimization
Tej et al. arXiv 2020.
- Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search
Guez et al. arXiv 2012.
- Bayesian Reinforcement Learning: A Survey
Ghavamzadeh et al. arXiv 2016.
- The Uncertainty Bellman Equation and Exploration
O'Donoghue et al. arXiv 2017.
- Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning
Janz et al. arXiv 2018.
← Back to main page