Unit content
Boosting and additive ensembles
Boosting builds an ensemble sequentially, with each new weak learner chosen to improve the mistakes of the current ensemble.
An additive model has the form
$$F_M(x)=\sum_{m=1}^{M}\alpha_m h_m(x),$$
where each $h_m$ is a relatively simple learner, often a shallow decision tree.
In AdaBoost-style classification, misclassified training examples receive greater weight so that later learners focus more strongly on them. In gradient boosting, each new learner is fitted to a direction that reduces the current loss, analogous to taking a functional gradient step.
The important contrast with bagging is structural:
- bagging fits component models largely independently and averages them;
- boosting fits components sequentially so that later models correct the current ensemble.
A sequence of individually weak predictors can therefore form a strong nonlinear model. But because boosting deliberately keeps adapting to residual errors, excessive depth, too many rounds or overly aggressive step sizes can overfit. Number of learners, tree complexity and learning rate are therefore hyperparameters selected with held-out data.