Unit content
Bias-variance decomposition for squared prediction error
The bias-variance tradeoff can be made precise for squared-error prediction.
Assume
$$Y=f(x)+\varepsilon,$$
where $\mathbb E[\varepsilon]=0$ and $\operatorname{Var}(\varepsilon)=\sigma^2$. Different training samples produce different fitted predictors $\hat f(x)$.
At a fixed input $x$, the expected squared prediction error decomposes as
$$\mathbb E[(Y-\hat f(x))^2] =\sigma^2+\big(\mathbb E[\hat f(x)]-f(x)\big)^2+\operatorname{Var}(\hat f(x)).$$
The terms are:
- irreducible noise $\sigma^2$;
- squared bias, measuring systematic error of the average fitted predictor;
- variance, measuring sensitivity to the particular training sample.
Suppose a very rigid model always predicts almost the same straight trend across repeated datasets. Its variance may be low but its bias high if the true relationship is curved. A highly flexible model may track each dataset closely, lowering bias while making predictions vary substantially between samples.
The decomposition does not imply that every real learning problem has one simple scalar complexity knob, but it explains why reducing training error alone can increase expected test error.