Learning path

Full curriculum

Full curriculum

Unit content

Probability calibration and proper scoring rules

A classifier that outputs probabilities should be judged not only by which class it ranks highest, but by whether those probabilities mean what they claim.

A model is calibrated if, among cases assigned probability about $0.7$, roughly 70% actually belong to the positive class. Calibration is different from discrimination: a model can rank positives above negatives well while producing probabilities that are systematically too confident.

One useful scoring rule is log loss for a binary target $y\in{0,1}$ and predicted probability $p$:

$$-\big[y\log p+(1-y)\log(1-p)\big].$$

Another is the Brier score,

$$(p-y)^2.$$

These are proper scoring rules: in expectation, they reward reporting one's true probability rather than strategically distorting it.

Calibration matters whenever probabilities feed downstream decisions. A hospital may act very differently on a reliable 5% risk than on a reliable 60% risk even if both models rank patients in the same order. Reliability diagrams and held-out recalibration methods can diagnose and correct systematic probability distortion.