Scaling Laws and Emergent Capabilities Reading List
Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026
How model performance scales with data, compute, and parameters, and what emerges at scale.
Scaling Laws and Emergent Capabilities: 10 key papers
- Scaling Laws for Neural Language Models
Kaplan et al. arXiv 2020.
- Training Compute-Optimal Large Language Models
Hoffmann et al. arXiv 2022.
- Deep Learning Scaling is Predictable, Empirically
Hestness et al. arXiv 2017.
- Explaining Neural Scaling Laws
Bahri et al. arXiv 2021.
- Scaling Laws for Autoregressive Generative Modeling
Henighan et al. arXiv 2020.
- Scaling Laws for Transfer
Hernandez et al. arXiv 2021.
- Beyond neural scaling laws: beating power law scaling via data pruning
Sorscher et al. arXiv 2022.
- Emergent Abilities of Large Language Models
Wei et al. arXiv 2022.
- Are Emergent Abilities of Large Language Models a Mirage?
Schaeffer et al. arXiv 2023.
- A Solvable Model of Neural Scaling Laws
Maloney et al. arXiv 2022.
← Back to main page