Learning path

Full curriculum

Full curriculum

Unit content

Naive Bayes classification

A Naive Bayes classifier predicts a class by combining Bayes' theorem with a simplifying conditional-independence assumption.

For class $C$ and features $x_1,\ldots,x_d$,

$$P(C\mid x_1,\ldots,x_d)\propto P(C)\prod_{j=1}^d P(x_j\mid C).$$

The assumption is that the features are conditionally independent given the class. This is often unrealistic, but the resulting classifier can still work well.

Suppose an email classifier uses whether the words invoice and lottery occur. For each class, estimate the prior probability of spam and the conditional probabilities of those words. For a new email, multiply the class prior by the corresponding likelihood terms and choose the class with the larger posterior score.

In practice, products of many small probabilities can underflow numerically, so implementations usually sum log-probabilities instead:

$$\log P(C)+\sum_j\log P(x_j\mid C).$$

Different variants choose probability models appropriate to the features, such as Bernoulli features for presence/absence or Gaussian features for continuous measurements. The defining idea is the same: a generative class-conditional model plus Bayes' rule.