Learning path

Full curriculum

Full curriculum

Unit content

Class imbalance, thresholds, and asymmetric costs

When one class is much rarer than another, ordinary accuracy can hide poor performance on the rare class.

Suppose only 1% of transactions are fraudulent. A classifier that always predicts “legitimate” is 99% accurate but detects no fraud.

Two ideas must be separated:

  • class imbalance describes the data distribution;
  • decision cost describes the consequences of different mistakes.

A probabilistic classifier can use a decision threshold chosen for those costs rather than automatically using $0.5$. If missing fraud is much more expensive than investigating a false alarm, lowering the threshold may be rational.

Training can also use class weights, resampling or losses that emphasize rare examples, but these choices can alter probability calibration. Evaluation should therefore report metrics appropriate to the goal, such as precision, recall and precision-recall curves, rather than one aggregate accuracy.

The correct operating point depends on deployment prevalence and costs. A threshold tuned on a balanced research dataset may be inappropriate when the real positive rate is very different.