Unit content
Supervised learning and loss functions
In supervised learning, a model learns a mapping from inputs to outputs using examples for which the desired output is known.
A dataset can be written as
$$(x_1,y_1),\ldots,(x_n,y_n).$$
A parameterized model $f(x;\theta)$ produces a prediction from each input.
Loss functions
A loss function measures how poorly a prediction matches its target. For regression, squared error is common:
$$L=(y-\hat y)^2.$$
For probabilistic classification, cross-entropy is a natural loss because it penalizes assigning low probability to the observed class.
Training and generalization
Training adjusts the parameters to reduce average loss on observed data. The real goal, however, is not to memorize the training examples but to perform well on new data drawn from the same relevant problem distribution.
A model can therefore have low training loss and still generalize poorly.
Model, objective and optimizer
Three ideas should be kept distinct:
- the model determines which predictions are possible;
- the loss defines what counts as a better prediction;
- the optimizer determines how the parameters are changed.
Deep learning uses neural networks as the model family, but the same supervised-learning structure applies to much simpler models as well.