Unit content
Cross-entropy
Suppose data are generated according to a true distribution $p$, but a model assigns probabilities according to another distribution $q$. Cross-entropy measures the expected information cost of using $q$ to describe outcomes generated by $p$.
For discrete distributions,
$$H(p,q)=-\sum_i p_i\log_2 q_i.$$
The model is judged on actual outcomes
The weighting uses $p_i$ because that determines how often outcome $i$ occurs. The logarithm uses $q_i$ because that is the probability assigned by the model.
A model pays a large penalty when an outcome occurs frequently under $p$ but the model assigns it very small probability.
Relationship to entropy
If the model distribution is exactly correct,
$$q=p,$$
then cross-entropy reduces to Shannon entropy:
$$H(p,p)=H(p).$$
For any other $q$, the cross-entropy is at least as large under the usual assumptions.
Machine-learning interpretation
In probabilistic classification, minimizing empirical cross-entropy encourages the model to assign high probability to observed classes. The loss is not arbitrary: it corresponds to negative log-likelihood for common categorical models.