Unit content
Backpropagation and computational graphs
Training a neural network requires derivatives of the loss with respect to parameters in every layer. Backpropagation computes those derivatives efficiently by applying the chain rule backward through the composed computation.
Computational graph
A forward pass can be viewed as a graph of intermediate values:
input → layer → activation → layer → prediction → loss
Each node depends on earlier values.
Local derivatives
During the backward pass, each operation receives a derivative of the final loss with respect to its output and combines it with its own local derivative.
If
$$y=f(u),\qquad u=g(x),$$
then
$$\frac{dy}{dx}=\frac{dy}{du}\frac{du}{dx}.$$
The same principle applies repeatedly through a large network.
Reusing intermediate results
Backpropagation is efficient because shared intermediate derivatives are reused rather than recomputed independently for every parameter.
Gradients and updates
Backpropagation computes gradients; it does not decide how parameters change. An optimizer such as gradient descent uses those gradients to perform the update.
Deep-learning training is therefore a cycle: forward computation, loss evaluation, backward differentiation and parameter update.