Unit content
Stochastic gradient descent and mini-batch optimization
When an objective is an average over many data points,
$$J(\theta)=\frac1N\sum_{i=1}^N L_i(\theta),$$
computing the full gradient at every update can be expensive.
Stochastic gradient descent estimates the gradient from one randomly selected example. Mini-batch methods use a small subset:
$$g_B(\theta)=\frac1{|B|}\sum_{i\in B}\nabla L_i(\theta).$$
The update
$$\theta_{k+1}=\theta_k-\eta g_B(\theta_k)$$
is cheaper but noisy: two batches generally produce different gradient estimates.
The noise can make progress irregular, so learning-rate schedules and averaging matter. Larger batches reduce gradient noise but cost more per update.
Stochastic optimization is especially useful when the full objective is a large sum, as in machine learning, but the idea is broader than any particular model family.