Unit content
Tree pruning and complexity control
A decision tree can keep splitting until leaves contain very few training examples, often driving training error down while harming generalization.
Pruning controls this complexity by removing splits whose predictive benefit does not justify the additional structure.
One approach is pre-pruning: stop growth when a node is too small, a depth limit is reached, or the best split improves impurity by too little. Another is post-pruning: first grow a larger tree, then collapse subtrees using validation performance or a complexity-penalized objective.
A common form of cost-complexity pruning balances fit against leaf count:
$$R_\alpha(T)=R(T)+\alpha|T|,$$
where $R(T)$ measures training error or impurity and $|T|$ measures tree size. Larger $\alpha$ favors smaller trees.
Pruning makes the same bias-variance tradeoff seen elsewhere in machine learning. A very small tree may miss real interactions; an unrestricted tree may memorize idiosyncratic training cases. The appropriate complexity must be chosen using held-out data rather than by looking only at training fit.