Unit content
Training, validation, and test sets
A model is useful only if it performs well on data that did not determine its fitted parameters or modeling choices. For that reason, a dataset is commonly divided into three roles.
- The training set is used to fit model parameters.
- The validation set is used to choose among models or hyperparameter settings.
- The test set is held back until the end to estimate performance of the chosen procedure.
Suppose 10,000 labeled examples are split into 7,000 training, 1,500 validation and 1,500 test examples. Training may compare many parameter values using the training set, and validation may select the best model. The test set should then answer a different question: how well does that already-chosen procedure perform on unseen examples?
If the test set is repeatedly consulted while choosing features, thresholds or hyperparameters, it effectively becomes part of model selection. Its reported performance then becomes optimistically biased.
The important separation is therefore conceptual rather than merely three files: information used to fit, information used to choose, and information used only to evaluate the final choice must be kept distinct.