Learning path

Full curriculum

Full curriculum

Unit content

Hypothesis spaces and empirical risk minimization

A supervised learning algorithm does not search over every imaginable input-output rule. It chooses from a hypothesis space $\mathcal H$: the family of predictors its representation allows.

Given training examples $(x_i,y_i)$ and loss $L$, empirical risk minimization (ERM) chooses a hypothesis that minimizes average training loss:

$$\hat h\in\arg\min_{h\in\mathcal H}\frac1n\sum_{i=1}^n L(h(x_i),y_i).$$

For a model restricted to affine functions of chosen features, $\mathcal H$ contains only those affine predictors. A different representation or structural constraint defines a different hypothesis space.

Training therefore contains two distinct choices:

  1. choose the hypothesis space and representation;
  2. optimize within that space using observed data.

A low empirical risk does not by itself guarantee good generalization. A hypothesis space flexible enough to memorize the sample can achieve near-zero training error while behaving poorly elsewhere.

Penalties, architecture, feature design and explicit complexity limits can all restrict or prefer parts of the hypothesis space. This viewpoint unifies many seemingly different machine-learning algorithms.