Unit content
Decision trees for prediction
A decision tree predicts by recursively splitting the feature space into regions and assigning a prediction to each resulting leaf.
For example, a tree predicting whether a customer will buy might begin with
age < 30?
and then split one branch using income < 40,000?. Each internal node asks a feature question; each leaf produces a class or numerical prediction.
Training chooses splits that make child nodes more homogeneous in their targets. For classification, common impurity measures include entropy
$$H=-\sum_c p_c\log p_c$$
and Gini impurity
$$G=1-\sum_c p_c^2.$$
A candidate split is useful when it reduces weighted impurity. If a node with mixed classes is divided into two nearly pure children, the reduction is large.
Trees naturally model nonlinear relationships and interactions and usually do not require numerical feature scaling. Their weakness is instability: a small change in data can produce a different sequence of splits. Deep trees can also overfit, which motivates pruning and ensemble methods.