Mechanistic Interpretability Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

What computations do neural networks actually implement? Mechanistic interpretability opens the black box to find circuits, features, and algorithms hidden in weights, turning mystery into mechanism.

Mechanistic Interpretability: 10 key papers

  1. Toy Models of Superposition
    Elhage et al. arXiv 2022.
  2. In-context Learning and Induction Heads
    Olsson et al. arXiv 2022.
  3. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
    Wang et al. arXiv 2022.
  4. Progress measures for grokking via mechanistic interpretability
    Nanda et al. arXiv 2023.
  5. Towards Automated Circuit Discovery for Mechanistic Interpretability
    Conmy et al. arXiv 2023.
  6. Sparse Autoencoders Find Highly Interpretable Features in Language Models
    Cunningham et al. arXiv 2023.
  7. Locating and Editing Factual Associations in GPT
    Meng et al. arXiv 2022.
  8. Transformer Feed-Forward Layers Are Key-Value Memories
    Geva et al. arXiv 2020.
  9. Emergent Linear Representations in World Models of Self-Supervised Sequence Models
    Nanda et al. arXiv 2023.
  10. The Quantization Model of Neural Scaling
    Michaud et al. arXiv 2023.
← Back to main page