|
Stanford MS&E 226 – Fundamentals of Data Science
Class description
This course provides an introduction to applied data analysis, with an emphasis on providing a conceptual framework for thinking about uncertainty from machine learning, statistical, and causal perspectives. The class involves developing an interconnected understanding across these ways of thinking. Class lectures will be supplemented by data-driven problem sets and a project. Prerequisites: multivariable calculus; probability.
Outline of topics
Prediction. Train-test-validate; cross validation; binary classification; using optimization to build predictive models (maximum likelihood; linear and logistic regression; regularization, lasso, and ridge; other methods); model complexity and the bias-variance decomposition.
Inference. Frequentism and sampling distributions; p-values, confidence intervals, and hypothesis testing; application to linear and logistic regression; bootstrap; multiple hypothesis testing; post-selection inference.
Causality. The Rubin causal model, potential outcomes, and counterfactuals; randomized experiments; causal inference from observational data.
Bayesian statistics and decision-making. Basics of Bayesian statistics; priors and posteriors; Bayesian vs. frequentist statistics; a Bayesian approach to decision-making.
Course info
All logistical information about the course is available in the syllabus linked from the navigation (Stanford login required).
Enrolled students should use Ed Discussion via Canvas for course announcements.
Professor
Ramesh Johari
|