Learning path

Full curriculum

Full curriculum

Unit content

Backpropagation and computational graphs

Training a neural network requires derivatives of the loss with respect to parameters in every layer. Backpropagation computes those derivatives efficiently by applying the chain rule backward through the composed computation.

Computational graph

A forward pass can be viewed as a graph of intermediate values:

input → layer → activation → layer → prediction → loss

Each node depends on earlier values.

Local derivatives

During the backward pass, each operation receives a derivative of the final loss with respect to its output and combines it with its own local derivative.

If

$$y=f(u),\qquad u=g(x),$$

then

$$\frac{dy}{dx}=\frac{dy}{du}\frac{du}{dx}.$$

The same principle applies repeatedly through a large network.

Reusing intermediate results

Backpropagation is efficient because shared intermediate derivatives are reused rather than recomputed independently for every parameter.

Gradients and updates

Backpropagation computes gradients; it does not decide how parameters change. An optimizer such as gradient descent uses those gradients to perform the update.

Deep-learning training is therefore a cycle: forward computation, loss evaluation, backward differentiation and parameter update.