Unit content
Random forests
A random forest is an ensemble of decision trees designed to make bagged trees less correlated with one another.
Each tree is trained on a bootstrap sample. In addition, when a tree considers a split, it is allowed to examine only a random subset of the available features.
Why add this extra randomness? Suppose one very strong predictor tends to dominate the first split of almost every ordinary tree. Bagging may then produce many similar trees whose errors remain correlated. Random feature selection forces different trees to explore alternative structures.
Predictions are aggregated by averaging for regression or voting for classification. The forest can therefore keep the flexibility of deep trees while substantially reducing their variance.
Two distinct sources of randomness should be remembered:
- bootstrap sampling varies the training examples;
- feature subsampling varies the candidate splits.
Random forests usually require little feature scaling and handle nonlinear interactions naturally. Their cost is reduced interpretability compared with a single tree and greater computation and memory.