Course: Introduction to Machine Learning, DS 3000
Three optimizers descend the same loss surface.
Training a model amounts to descending a loss surface. Here SGD, Momentum and Adam start from the same point; how and why their paths differ is a central lesson of his machine learning course.
See the course