Learning path

Full curriculum

Full curriculum

Unit content

Bagging and bootstrap aggregation

Bagging, short for bootstrap aggregation, reduces the variance of an unstable learning algorithm by averaging many predictors fitted to different bootstrap resamples of the training data.

The procedure is:

  1. draw many bootstrap samples;
  2. fit one predictor to each sample;
  3. average their numerical predictions or vote among their class predictions.

If individual predictors make partly independent errors, aggregation cancels some of their variation. This is especially effective for deep decision trees, whose fitted structures can change substantially when the sample changes slightly.

Suppose ten tree regressors predict values around 18, 22, 19, 24 and so on. Any one tree may be sensitive to its bootstrap sample, while their average can vary much less from one training dataset to another.

Bagging primarily attacks variance, not systematic bias shared by all component models. If every learner is too simple to capture the true relationship, averaging many copies does not repair that limitation.

The broader ensemble principle is that diversity among component errors can make an aggregate predictor more stable than its individual members.