Week 3 — Gradient Methods
[0:05] Gradient descent converges linearly for strongly convex functions.
[0:31] The learning rate must be smaller than two over the Lipschitz constant.
[1:15] We now turn to the stochastic variants of the method.
[2:47] Variance reduction recovers the linear rate under mild assumptions.