Learning path

Full curriculum

Full curriculum

Unit content

Conditional entropy

When two random variables are related, observing one can reduce uncertainty about the other.

For discrete variables $X$ and $Y$, the conditional entropy of $X$ given $Y$ is

$$H(X\mid Y)=\sum_y p(y)H(X\mid Y=y).$$

Equivalently,

$$H(X\mid Y)=-\sum_{x,y}p(x,y)\log_2 p(x\mid y).$$

It measures the average uncertainty that remains about $X$ after the value of $Y$ is known.

If $Y$ determines $X$ exactly, then

$$H(X\mid Y)=0.$$

If $X$ and $Y$ are independent, learning $Y$ does not reduce uncertainty and

$$H(X\mid Y)=H(X).$$

Conditional entropy also satisfies the chain rule

$$H(X,Y)=H(Y)+H(X\mid Y),$$

which separates joint uncertainty into uncertainty about one variable plus the remaining uncertainty about the other.