Unit content
Shannon entropy
A random variable can produce outcomes with different probabilities. Shannon entropy is the expected self-information of its outcome.
For a discrete distribution with probabilities $p_1,\ldots,p_n$,
$$H(X)=-\sum_i p_i\log_2 p_i.$$
Entropy is small when one outcome is overwhelmingly likely and larger when probability is spread more evenly among alternatives.
For a fair binary variable,
$$H(X)=1\text{ bit},$$
while a variable whose outcome is certain has zero entropy.
Entropy belongs to the probability distribution, not to one particular observed outcome. Self-information measures the surprise of a specific outcome; entropy averages that surprise over all possible outcomes.
Because predictable sources have lower entropy than equally likely alternatives, entropy becomes the fundamental uncertainty quantity used in lossless source coding and mutual information.
Information entropy and thermodynamic entropy have mathematical connections in statistical physics, but they are not interchangeable definitions. Here $H$ describes uncertainty in a probability distribution.