Boosting
Boosting builds a prediction model by adding learners in stages, each responding to the combined model's remaining errors or loss. The learners' predictions are combined into an ensemble rather than used independently.
A house-price model can start with a rough prediction, then add small decision trees that adjust it for patterns the model has missed. Each tree contributes a correction; the final prediction combines those contributions. Gradient boosting fits new learners to the direction that reduces a chosen loss, not always simply to raw prediction errors.
Adaptive boosting (AdaBoost) uses a different mechanism: it increases the influence of examples the previous stage misclassified. Both are boosting. Bagging instead trains learners on resampled data and aggregates them; their fitting need not depend on the previous learner’s mistakes.
This matters because the number of stages, the strength of each contribution and the learner’s complexity affect both fit and overfitting. Adding trees can lower training loss while making new predictions worse. Use validation to choose these settings, and compare with simpler models on the same data rather than assuming an ensemble must win.
Sources
- scikit-learn: Ensemble methods — Explains staged gradient boosting, AdaBoost example weighting and the contrast with bagging.
Go deeper
- scikit-learn: Gradient Boosting regression docs
Fit a regression example and inspect training and test behavior as stages are added.