Learning path

Full curriculum

Full curriculum

Unit content

Cross-entropy

Suppose data are generated according to a true distribution $p$, but a model assigns probabilities according to another distribution $q$. Cross-entropy measures the expected information cost of using $q$ to describe outcomes generated by $p$.

For discrete distributions,

$$H(p,q)=-\sum_i p_i\log_2 q_i.$$

The model is judged on actual outcomes

The weighting uses $p_i$ because that determines how often outcome $i$ occurs. The logarithm uses $q_i$ because that is the probability assigned by the model.

A model pays a large penalty when an outcome occurs frequently under $p$ but the model assigns it very small probability.

Relationship to entropy

If the model distribution is exactly correct,

$$q=p,$$

then cross-entropy reduces to Shannon entropy:

$$H(p,p)=H(p).$$

For any other $q$, the cross-entropy is at least as large under the usual assumptions.

Machine-learning interpretation

In probabilistic classification, minimizing empirical cross-entropy encourages the model to assign high probability to observed classes. The loss is not arbitrary: it corresponds to negative log-likelihood for common categorical models.