Backpropagation
Backpropagation computes how a neural network's loss changes with each learned parameter. It applies the chain rule backward through the computation; an optimizer then uses those gradients to update the parameters.
Training first runs a forward calculation, producing a prediction and a loss. Backpropagation works backward through those operations to compute gradients: how sensitive that loss is to each weight and bias. The optimizer chooses how to use those gradients.
For a small example, let the prediction be y = w × x and loss be L = (y − target)². With input 2, weight 3 and target 4, the prediction is 6 and loss is 4. The chain rule gives dL/dw = 2(y − target) × x = 8. A gradient-descent step with learning rate 0.1 changes the weight to 2.2. This example is illustrative, not a general choice of learning rate.
The distinction matters when debugging training. Computing the gradient does not guarantee a useful update, a global optimum or good performance on new examples. Ordinary serving inference uses learned parameters without training them; fine-tuning adds a training loop that can use backpropagation.
Sources
- PyTorch: Automatic differentiation with torch.autograd — Shows forward computation, backward gradient propagation, the chain rule and disabling gradient tracking for inference.
Go deeper
- 3Blue1Brown: Backpropagation, intuitively video
See how an error signal changes weights through several layers.