Unit content
Cross-validation
A single validation split can be noisy, especially when the dataset is not large. $k$-fold cross-validation reuses the available training data to estimate how a modeling choice behaves across several held-out subsets.
Split the available development data into $k$ folds. For each fold in turn:
- fit the complete training procedure on the other $k-1$ folds;
- evaluate it on the held-out fold.
The $k$ validation scores are then averaged. With five folds, every example is used for validation once and for training four times.
Crucially, the complete procedure must be refitted inside each fold. If standardization, imputation or feature selection is fitted once using all folds before cross-validation, information from each validation fold leaks into its own training process.
Cross-validation is usually used for model selection or performance estimation inside the development data. A final untouched test set can still be kept for one last evaluation after choices have been made.
The folds must respect the data-generating structure. Time series generally require chronological splits, and repeated measurements from one person should usually remain in the same fold.