Unit content
Transformer architecture
A transformer builds sequence representations primarily from self-attention and position-wise feed-forward transformations rather than recurrent state updates.
Transformer block
A typical block contains multi-head self-attention followed by a feed-forward network. Residual connections and normalization surround these transformations so information and gradients can propagate through many stacked blocks.
Conceptually:
representations
↓
self-attention
↓
residual + normalization
↓
feed-forward network
↓
residual + normalization
Parallel sequence processing
Self-attention can compare many sequence positions in parallel. Unlike a simple recurrent model, information does not have to pass through every intermediate time step to connect distant positions.
Causal masking
For autoregressive prediction, a position must not attend to future tokens. A causal mask removes those future connections so each prediction only depends on allowed previous context.
Encoder and decoder patterns
The original transformer used an encoder-decoder architecture. Many later models use encoder-only or decoder-only stacks depending on whether the task emphasizes representation, generation or both.
Cost of attention
Ordinary self-attention forms pairwise interactions between sequence positions, so its memory and compute cost grow roughly quadratically with sequence length.