Unit content
Confusion matrices, precision, recall, and F1
For binary classification, accuracy alone can hide which kinds of mistakes a model makes. A confusion matrix separates predictions into four counts:
- true positives $TP$;
- false positives $FP$;
- true negatives $TN$;
- false negatives $FN$.
From these counts,
$$\text{precision}=\frac{TP}{TP+FP}$$
measures how often a positive prediction is correct, while
$$\text{recall}=\frac{TP}{TP+FN}$$
measures how many actual positives are found.
The F1 score is their harmonic mean:
$$F_1=2\frac{\text{precision}\cdot\text{recall}}{\text{precision}+\text{recall}}.$$
Suppose a screening model finds 80 of 100 diseased patients and incorrectly flags 20 healthy patients. Then recall is $80/100=0.8$ and precision is $80/(80+20)=0.8$.
Which quantity matters depends on the decision. Missing a dangerous disease can make recall critical; sending every harmless email to a spam folder can make precision more important. Metrics therefore encode priorities, not just mathematical bookkeeping.