Learning path

Full curriculum

Full curriculum

Unit content

Baseline predictors and meaningful model comparisons

A machine-learning score is informative only relative to a sensible reference. A baseline is a deliberately simple predictor used to establish what performance is achievable without the proposed model's complexity.

For regression under squared error, predicting the training-set mean is a natural baseline. Under absolute error, the median is natural. For classification, one baseline may always predict the most frequent class; another may use a simple domain rule.

Suppose a model achieves 91% accuracy on a dataset whose largest class already accounts for 90% of examples. The one-point improvement is very different from achieving 91% when a simple baseline reaches only 55%.

Comparisons must use the same held-out examples and metric. Changing the test set or preprocessing between models confounds model quality with evaluation conditions.

A strong experiment often uses several baselines: a trivial constant predictor, a simple interpretable model and the current production system if one exists. If a sophisticated model cannot reliably beat simpler alternatives, its extra complexity may not be justified.

Baselines turn an isolated score into evidence about whether learning useful structure has actually occurred.