Explainable AI & Interpretability Reading List

Curated by Mouhssine Rifaki | Stanford Electrical Engineering | Last updated August 2026

Core papers on explainable AI and interpretability, from model-agnostic explanations to mechanistic understanding.

Explainable AI & Interpretability: 10 key papers

  1. "Why Should I Trust You?": Explaining the Predictions of Any Classifier
    Ribeiro et al. arXiv 2016.
  2. A Unified Approach to Interpreting Model Predictions
    Lundberg and Lee. arXiv 2017.
  3. Axiomatic Attribution for Deep Networks
    Sundararajan et al. arXiv 2017.
  4. Towards A Rigorous Science of Interpretable Machine Learning
    Doshi-Velez and Kim. arXiv 2017.
  5. The Mythos of Model Interpretability
    Lipton. arXiv 2016.
  6. Sanity Checks for Saliency Maps
    Adebayo et al. arXiv 2018.
  7. Learning Deep Features for Discriminative Localization
    Zhou et al. arXiv 2015.
  8. Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
    Selvaraju et al. arXiv 2016.
  9. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead
    Rudin. arXiv 2018.
  10. Explanation in Artificial Intelligence: Insights from the Social Sciences
    Miller. arXiv 2017.
← Back to main page