Teaching Machine learning

DS 3000, Artificial Intelligence Systems Engineering

Introduction to Machine Learning

How a model is fitted, how its performance is estimated honestly, and how to tell a result that will hold from one that will not.

Course content

Step 0. Loss: Adam 0.00, Momentum 0.00, SGD 0.00
Same surface, same start point, three optimizers.

About the course

The course builds machine learning from the ground up: fitting a first model by hand, estimating its performance honestly, and moving through classification, trees and ensembles to neural networks, deep learning and the methods that work without labels. It closes with two questions every model must answer: how uncertain its predictions are, and for whom it fails.

Each chapter is organized around a single question. The lectures run live code in the browser, so each idea can be tried as soon as it is introduced.

Offered as
DS 3000, undergraduate
Program
Artificial Intelligence Systems Engineering
Instructor
Fadi AlMahamid, Ph.D.
Content
12 chapters, 70 topics
Lecture slides
Enrolled students, with a username and password

Course content

The 12 chapters progress from a straight line fitted to five points to deep networks, unlabelled data and the limits of what a model can claim.

  1. Chapter 1

    Introduction to Machine Learning

    What is a model, and what is it actually searching for?

    In this chapter

    1. The landscape
    2. The three learning paradigms
    3. From data to a result you trust
    4. Linear regression, your first model
    5. Finding the best line
    Slides for chapter 1
  2. Chapter 2

    Multiple Linear Regression

    What changes when a prediction depends on more than one thing?

    In this chapter

    1. From one feature to many
    2. Solving for the weights
    3. Choosing a loss
    4. Judging generalization
    5. Exploring a real dataset
    Slides for chapter 2
  3. Chapter 3

    Probability and Maximum Likelihood

    Which model makes the data you actually saw most plausible?

    In this chapter

    1. Why probability
    2. Random variables
    3. Discrete variables
    4. Continuous variables
    5. Joint distributions
    6. Likelihood and MLE
    7. Back to regression
    Slides for chapter 3
  4. Chapter 4

    Supervised Learning: Classification

    How do you predict a label rather than a number?

    In this chapter

    1. The problem
    2. The linear score
    3. Decision boundaries
    4. Logistic regression
    5. Maximum likelihood
    6. SVM and KNN
    7. Loss and evaluation
    Slides for chapter 4
  5. Chapter 5

    Model Selection, Feature Engineering and Regularization

    How do you choose between models without fooling yourself?

    In this chapter

    1. Bias and variance
    2. Choosing a model
    3. Cross-validation
    4. Feature construction
    5. Feature selection
    6. Regularization
    Slides for chapter 5
  6. Chapter 6

    Decision Trees

    What can you learn from asking one question at a time?

    In this chapter

    1. Inside a tree
    2. Growing a tree
    3. Depth, and its price
    Slides for chapter 6
  7. Chapter 7

    Ensemble Learning

    Why do many weak models beat one strong one?

    In this chapter

    1. Why combine
    2. Voting
    3. Bagging
    4. Random forests
    5. Boosting
    6. Stacking, and choosing
    Slides for chapter 7
  8. Chapter 8

    Neural Networks

    What does a network compute that a linear model cannot?

    In this chapter

    1. The computational unit
    2. Perceptron to network
    3. Backpropagation
    4. Designing a network
    5. Networks on pixels
    6. Learned representations
    7. Training challenges
    Slides for chapter 8
  9. Chapter 9

    Deep Learning

    What changes when the network gets deep?

    In this chapter

    1. Why deep learning
    2. How convolution works
    3. Building a deep CNN
    4. Modeling sequences
    5. Large language models
    6. Autoencoders
    Slides for chapter 9
  10. Chapter 10

    Unsupervised Learning

    What can you find when nothing is labeled?

    In this chapter

    1. Clustering
    2. K-means
    3. Judging a clustering
    4. Gaussian mixtures
    5. Hierarchical clustering
    6. Anomaly detection
    Slides for chapter 10
  11. Chapter 11

    Dimensionality Reduction

    How do you keep the signal when you throw columns away?

    In this chapter

    1. Why fewer dimensions
    2. Spread and covariance
    3. Principal component analysis
    4. Choosing the dimension
    5. Nonlinear reduction
    Slides for chapter 11
  12. Chapter 12

    Uncertainty, Bias and Causality

    How wrong might this be, and who does it fail?

    In this chapter

    1. Why uncertainty matters
    2. Confidence intervals
    3. The bootstrap
    4. Uncertainty in deep models
    5. Where bias comes from
    6. Measuring and fixing fairness
    7. Causality
    Slides for chapter 12

Lecture slides

The lectures are interactive decks with animated figures that you can play and step through, and Python cells that run in the browser. They are available to enrolled students, who sign in with the username and password provided by the instructor.

Reading

Primary material

The lecture slides

There is no required textbook. The lecture slides are the primary material, and each chapter names its own sources.

Further reading

  • The Elements of Statistical Learning

    Hastie, Tibshirani and Friedman

    The standard reference. Heavier than this course, and the place to go when a chapter is not enough.

  • Pattern Recognition and Machine Learning

    Bishop

    A second voice on the same material, from the probabilistic side.

  • Machine Learning: A Probabilistic Perspective

    Murphy

    Broad and encyclopedic. Useful for looking one thing up.