Unit content
Softmax and categorical probabilities
A model that assigns one score to each of several mutually exclusive classes often produces unconstrained real numbers called logits. The softmax function converts those logits into a probability distribution.
For logits $z_1,\ldots,z_K$,
$$p_i=\frac{e^{z_i}}{\sum_{j=1}^K e^{z_j}}.$$
Each $p_i$ is positive and
$$\sum_i p_i=1.$$
Relative scores matter
Adding the same constant to every logit does not change the probabilities. Softmax depends on score differences rather than an arbitrary absolute zero.
Sharpening and flattening
Large differences between logits produce a distribution concentrated on the largest scores. Smaller differences produce a flatter distribution.
A temperature parameter $T>0$ can make this explicit:
$$p_i=\frac{e^{z_i/T}}{\sum_j e^{z_j/T}}.$$
Higher temperature produces a flatter distribution; lower temperature makes it more concentrated.
Classification and attention
In classification, softmax turns class scores into predicted categorical probabilities. In attention, it converts compatibility scores into weights that sum to one.
Softmax is therefore a normalization of relative exponential scores, not a rule that makes arbitrary model outputs automatically well calibrated.