Learning path

Full curriculum

Full curriculum

Arrows go from each prerequisite to the units that depend on it. Hover or focus a unit to highlight its path.

Unit content

Stochastic gradient descent and mini-batch optimization

When an objective is an average over many data points,

$$J(\theta)=\frac1N\sum_{i=1}^N L_i(\theta),$$

computing the full gradient at every update can be expensive.

Stochastic gradient descent estimates the gradient from one randomly selected example. Mini-batch methods use a small subset:

$$g_B(\theta)=\frac1{|B|}\sum_{i\in B}\nabla L_i(\theta).$$

The update

$$\theta_{k+1}=\theta_k-\eta g_B(\theta_k)$$

is cheaper but noisy: two batches generally produce different gradient estimates.

The noise can make progress irregular, so learning-rate schedules and averaging matter. Larger batches reduce gradient noise but cost more per update.

Stochastic optimization is especially useful when the full objective is a large sum, as in machine learning, but the idea is broader than any particular model family.