Optimization for Deep Learning Reading List
Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026
Training deep networks is an optimization problem with peculiar geometry. These papers explain why adaptive methods, normalization, and learning-rate schedules work, and where their limits lie.
Optimization for Deep Learning: 10 key papers
- Adam: A Method for Stochastic Optimization
Kingma and Ba. arXiv 2014.
- On the Convergence of Adam and Beyond
Reddi et al. arXiv 2019.
- Decoupled Weight Decay Regularization
Loshchilov and Hutter. arXiv 2017.
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe and Szegedy. arXiv 2015.
- Layer Normalization
Ba et al. arXiv 2016.
- Visualizing the Loss Landscape of Neural Nets
Li et al. arXiv 2017.
- Large Batch Training of Convolutional Networks
You et al. arXiv 2017.
- SGDR: Stochastic Gradient Descent with Warm Restarts
Loshchilov and Hutter. arXiv 2016.
- Why Momentum Really Works.
Goh. Distill 2017.
- Sharpness-Aware Minimization for Efficiently Improving Generalization
Foret et al. arXiv 2020.
← Back to main page