Gradient Boosting
Gradient boosting builds an ensemble in stages, adding a learner that fits the negative loss gradients of the current predictions. For squared-error regression, those gradients correspond to residual errors; other losses produce different fitting targets.
Also known as: GBM
For a house with a target price of 300,000 and a current prediction of 280,000, the residual is 20,000. In a squared-error setup, a new tree learns residual-like targets across the training examples. Its prediction is multiplied by a learning rate and added to the ensemble’s previous prediction. The numbers illustrate the procedure, not a measured housing model.
The general method uses negative gradients of the chosen loss, rather than always using simple target-minus-prediction residuals. Standard boosting keeps earlier learners fixed while adding new ones. Trees are common learners, but the boosting principle is not limited to them.
This differs from independently fitting and averaging trees in a random forest. More stages can improve fit but can also overfit. Use validation to choose the learning rate, tree size and stopping point. The distinction matters when diagnosing an ensemble: each stage corrects the combined prediction under a particular objective, rather than casting an independent vote.
Sources
- scikit-learn: Ensemble methods, gradient boosting — Derives additive prediction updates, negative-gradient fitting and learning-rate shrinkage.
Go deeper
- scikit-learn: Gradient boosting regression docs
Fit a regression example and inspect how training and test error change across stages.