Unit content
Dataset bias and subgroup evaluation in machine learning
A model learns from the examples, labels and measurement process it is given. Systematic problems in that process can therefore become systematic model behavior.
Selection bias can occur when the training sample does not represent the population on which the model will be used. Measurement bias can occur when a recorded feature or label measures different things across groups. Historical labels can also encode past decisions rather than an objective ground truth.
Aggregate performance can hide these effects. Suppose a classifier is 92% accurate overall but has 98% recall for one subgroup and 65% for another. Reporting only the aggregate obscures a deployment-relevant failure mode.
A useful evaluation therefore examines appropriate metrics across relevant subgroups and asks whether sample sizes are large enough for those comparisons to be meaningful.
No single numerical fairness metric resolves every normative question. Equalizing false-positive rates, false-negative rates, calibration or selection rates can represent different goals and can conflict when base rates differ.
Technical evaluation must therefore make the data-generating process and decision consequences explicit. A model cannot repair an ill-defined target or an unrepresentative dataset merely by optimizing its loss more successfully.