AI Glossary

Gradient Boosting

Gradient boosting builds an ensemble in stages, adding a learner that fits the negative loss gradients of the current predictions. For squared-error regression, those gradients correspond to residual errors; other losses produce different fitting targets.

Also known as: GBM

· Updated · Chain of Thought

For a house with a target price of 300,000 and a current prediction of 280,000, the residual is 20,000. In a squared-error setup, a new tree learns residual-like targets across the training examples. Its prediction is multiplied by a learning rate and added to the ensemble’s previous prediction. The numbers illustrate the procedure, not a measured housing model.

The general method uses negative gradients of the chosen loss, rather than always using simple target-minus-prediction residuals. Standard boosting keeps earlier learners fixed while adding new ones. Trees are common learners, but the boosting principle is not limited to them.

This differs from independently fitting and averaging trees in a random forest. More stages can improve fit but can also overfit. Use validation to choose the learning rate, tree size and stopping point. The distinction matters when diagnosing an ensemble: each stage corrects the combined prediction under a particular objective, rather than casting an independent vote.

Sources

Go deeper