Optimization for Deep Learning Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

Training deep networks is an optimization problem with peculiar geometry. These papers explain why adaptive methods, normalization, and learning-rate schedules work, and where their limits lie.

Optimization for Deep Learning: 10 key papers

  1. Adam: A Method for Stochastic Optimization
    Kingma and Ba. arXiv 2014.
  2. On the Convergence of Adam and Beyond
    Reddi et al. arXiv 2019.
  3. Decoupled Weight Decay Regularization
    Loshchilov and Hutter. arXiv 2017.
  4. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
    Ioffe and Szegedy. arXiv 2015.
  5. Layer Normalization
    Ba et al. arXiv 2016.
  6. Visualizing the Loss Landscape of Neural Nets
    Li et al. arXiv 2017.
  7. Large Batch Training of Convolutional Networks
    You et al. arXiv 2017.
  8. SGDR: Stochastic Gradient Descent with Warm Restarts
    Loshchilov and Hutter. arXiv 2016.
  9. Why Momentum Really Works.
    Goh. Distill 2017.
  10. Sharpness-Aware Minimization for Efficiently Improving Generalization
    Foret et al. arXiv 2020.
← Back to main page