Unit content
Hyperparameters and model selection
A learning algorithm usually has settings that are not fitted in the same way as ordinary model parameters. These hyperparameters control the learning procedure or the family of predictors being considered.
Examples include a neighborhood size, a maximum structural depth, a penalty strength $\lambda$, or the number of latent groups a model is allowed to represent.
A sound model-selection procedure is:
- define candidate models or hyperparameter settings;
- fit each candidate using training data;
- compare candidates on validation data or by cross-validation;
- choose the candidate using the predetermined evaluation criterion;
- assess the final chosen procedure on untouched test data.
Searching many configurations can itself overfit the validation process. The best validation score among hundreds of noisy trials is partly selected because of favorable noise. Nested cross-validation or a final test set helps separate selection from final assessment.
A parameter such as a fitted coefficient is learned from training examples inside one model. A hyperparameter helps define which fitting problem is solved. Keeping that distinction clear prevents evaluation data from quietly entering model construction.