Bayesian RL & Thompson Sampling Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

Bayesian RL treats uncertainty over dynamics. These papers develop posterior sampling, Thompson sampling, and regret analyses for the Bayesian setting.

Bayesian RL & Thompson Sampling: 10 key papers

  1. Deep Exploration via Randomized Value Functions
    Osband et al. arXiv 2017.
  2. (More) Efficient Reinforcement Learning via Posterior Sampling
    Osband et al. arXiv 2013.
  3. Bootstrapped Thompson Sampling and Deep Exploration
    Osband and Van Roy. arXiv 2015.
  4. Randomized Prior Functions for Deep Reinforcement Learning
    Osband et al. arXiv 2018.
  5. Why is Posterior Sampling Better than Optimism for Reinforcement Learning?
    Osband and Van Roy. arXiv 2016.
  6. Deep Bayesian Quadrature Policy Optimization
    Tej et al. arXiv 2020.
  7. Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search
    Guez et al. arXiv 2012.
  8. Bayesian Reinforcement Learning: A Survey
    Ghavamzadeh et al. arXiv 2016.
  9. The Uncertainty Bellman Equation and Exploration
    O'Donoghue et al. arXiv 2017.
  10. Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning
    Janz et al. arXiv 2018.
← Back to main page