Mechanistic Interpretability Reading List
Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026
What computations do neural networks actually implement? Mechanistic interpretability opens the black box to find circuits, features, and algorithms hidden in weights, turning mystery into mechanism.
Mechanistic Interpretability: 10 key papers
- Toy Models of Superposition
Elhage et al. arXiv 2022.
- In-context Learning and Induction Heads
Olsson et al. arXiv 2022.
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Wang et al. arXiv 2022.
- Progress measures for grokking via mechanistic interpretability
Nanda et al. arXiv 2023.
- Towards Automated Circuit Discovery for Mechanistic Interpretability
Conmy et al. arXiv 2023.
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
Cunningham et al. arXiv 2023.
- Locating and Editing Factual Associations in GPT
Meng et al. arXiv 2022.
- Transformer Feed-Forward Layers Are Key-Value Memories
Geva et al. arXiv 2020.
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models
Nanda et al. arXiv 2023.
- The Quantization Model of Neural Scaling
Michaud et al. arXiv 2023.
← Back to main page