Learning path

Full curriculum

Full curriculum

Arrows go from each prerequisite to the units that depend on it. Hover or focus a unit to highlight its path.

Unit content

Kullback-Leibler divergence

The Kullback-Leibler divergence compares two probability distributions by measuring the extra expected log-loss incurred when a model distribution $q$ is used in place of a reference distribution $p$.

For discrete distributions,

$$D_{KL}(p\lVert q)=\sum_i p_i\log\frac{p_i}{q_i}.$$

Relationship to cross-entropy

Using the same logarithm base,

$$D_{KL}(p\lVert q)=H(p,q)-H(p).$$

The entropy $H(p)$ is the irreducible expected information of outcomes under $p$, while the additional term measures the penalty for using the mismatched distribution $q$.

Nonnegative but not a distance

KL divergence satisfies

$$D_{KL}(p\lVert q)\ge0,$$

with equality when the distributions agree almost everywhere under the relevant conditions.

However,

$$D_{KL}(p\lVert q)\ne D_{KL}(q\lVert p)$$

in general, so it is not a symmetric geometric distance.

Support matters

If $p$ assigns positive probability to an outcome that $q$ declares impossible, the divergence becomes infinite. A model that rules out an event that can actually occur can therefore incur unbounded log-loss.

KL divergence appears throughout statistics, information theory and machine learning as a way to compare probabilistic descriptions.