Unit content
Attention and self-attention
An attention mechanism lets a representation combine information from other representations according to learned relevance scores rather than through one fixed local connection pattern.
Given query, key and value vectors, scaled dot-product attention computes
$$\operatorname{Attention}(Q,K,V)=\operatorname{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V.$$
Queries, keys and values
A query describes what one position is looking for. Keys describe what each candidate position offers. Their dot products produce compatibility scores.
Softmax converts those scores into nonnegative weights that sum to one, and the output is a weighted combination of the value vectors.
Self-attention
In self-attention, queries, keys and values are derived from the same sequence. Each position can therefore gather information from other positions in that sequence.
Multiple heads
Multi-head attention performs several attention operations with different learned projections. Different heads can specialize in different relationships before their outputs are combined.
Order information
Self-attention by itself does not know the original order of sequence positions. Transformer models therefore add positional information or otherwise encode position explicitly.
Attention is a content-dependent routing mechanism: the data determine which other representations contribute most strongly to each output.