Autumn 2026 ✦ Stanford University
MS&E 319: Efficient Generative Models
Given a modeling goal and a limited computational budget, how should one choose the training objective, model architecture, and inference algorithm?
- Meets Tuesdays and Thursdays, 3:00–4:20 PM; room TBD
- Prerequisites CS 224N, CS 324, CS 336, or equivalent preparation. We assume some familiarity with deep learning and gradient-based optimization, as well as the analysis of algorithms and basic proofs. Mathematical and algorithmic maturity at the level of CS 161/261 is recommended.
- Instructors Amin Karbasi, Anay Mehrotra, Amin Saberi, and Grigoris Velegkas
- Note May be repeated for credit.
Overview
Modern large language models are powerful partly because of their scale, but equally because of the many careful algorithmic choices behind them. This course studies the foundations of those choices. Given a modeling goal and a limited computational budget, how should one choose the training objective, model architecture, and inference algorithm?
We begin by reviewing language models and transformers, then build on these foundations to study recent research on efficient training and inference. The course moves through five themes: efficiency in pre-training; efficiency through architecture; efficiency at inference time; efficiency in post-training; and risks and vulnerabilities.
The goal is to understand not only what these methods do, but why they work, what tradeoffs they introduce, and what research questions they open. Students will complete two homework assignments and one group project.
Who this is for
- Students who want to read recent papers carefully and build a principled view of efficient generative modeling.
- Students comfortable with mathematical and algorithmic reasoning who want to move beyond using LLMs as black boxes, and understand the choices that make them trainable, deployable, and reliable.
Tentative Lecture Outline
Topics and their order may change as the course develops.
Course overview
Overview of the course and LLMs
Efficiency in Pre-training
Introduction to the pre-training pipeline
Lecture contents
- LLM architectures: then and now
- Training objective and causal masking
- Optimizers
- Scaling laws
Advances in pre-training
Lecture contents
- Training objectives: infilling (fill-in-the-middle) and multi-token prediction
- Document packing, multi-document attention, and cross-attention
- Advanced Optimization: Muon
Efficiency via Architecture
Mixtures of experts
Lecture contents
- Sparse expert routing: classic MoE, GShard, and Switch Transformers
- Expert specialization: shared and fine-grained experts
- Load balancing: auxiliary losses and auxiliary-loss-free methods
Innovations in Attention I: Alternate Architectures
Lecture contents
- Attention heads and KV sharing: MHA, MQA, and GQA
- Multi-head latent attention (MLA)
- Learned sparse attention: DSA, lightning indexer, and IndexCache
Innovations in Attention II: Sparser, Approximate, Faster
Lecture contents
- Sparse attention: structured patterns and token selection
- Linear and approximate attention
- Efficient attention computation: FlashAttention
Efficiency at Inference Time
Efficiency via Quantization
Lecture contents
- Quantizing pretrained models
- Continued pre-training of quantized models
- KV-cache quantization
Techniques for Efficient Inference
Lecture contents
- Speculative decoding and learned draft models
- Model orchestration: routing, retrieval, and agentic workflows
- Reusing computation: prefix and response caching
Efficiency in Post-training
Post-training I: Objectives and Supervision
Lecture contents
- From pre-training to supervised fine-tuning (SFT)
- RL objectives: expected reward and reference-policy regularization
- Reward signals: learned preference models and verifiable rewards
- Direct preference optimization (DPO)
Post-training II: Policy Gradients and Optimization
Lecture contents
- Policy gradients and REINFORCE
- Advantage estimation: baselines, actor–critic methods, RLOO, and GRPO
- Sampling and exploration: on-policy and off-policy data
- Controlling policy updates: TRPO and PPO
Advanced Topics in Post-training: Distillation and Learning to Reason Longer
Lecture contents
- Distillation and self-distillation
- Learning to reason longer and test-time scaling
- Iterative and recursive self-improvement: possibilities and limits
Risks and Vulnerabilities
Risks and Vulnerabilities of Generative AI
Looking Ahead
Where do we go from here?
Evaluation
Evaluation will consist of two homework assignments and one group project. Further details, including grading weights, deadlines, and group sizes, will be announced.
Homework assignments may include analytical problems, implementation exercises, or both.
Research projects are encouraged. Projects may be entirely theoretical, empirical, or a combination of both, and should address a topic related to the course. Students with relevant ongoing research are encouraged to connect their course project to it. Groups may also complete a focused survey of a course-related topic.
Compute resources may be available to support course projects (details will be announced).
Instructors
Amin Karbasi
VP and Chief AI Scientist at Cisco; Adjunct Professor, Stanford University
Office hours: TBD.