Learning path

Full curriculum

Full curriculum

Unit content

Transformer architecture

A transformer builds sequence representations primarily from self-attention and position-wise feed-forward transformations rather than recurrent state updates.

Transformer block

A typical block contains multi-head self-attention followed by a feed-forward network. Residual connections and normalization surround these transformations so information and gradients can propagate through many stacked blocks.

Conceptually:

representations
   ↓
self-attention
   ↓
residual + normalization
   ↓
feed-forward network
   ↓
residual + normalization

Parallel sequence processing

Self-attention can compare many sequence positions in parallel. Unlike a simple recurrent model, information does not have to pass through every intermediate time step to connect distant positions.

Causal masking

For autoregressive prediction, a position must not attend to future tokens. A causal mask removes those future connections so each prediction only depends on allowed previous context.

Encoder and decoder patterns

The original transformer used an encoder-decoder architecture. Many later models use encoder-only or decoder-only stacks depending on whether the task emphasizes representation, generation or both.

Cost of attention

Ordinary self-attention forms pairwise interactions between sequence positions, so its memory and compute cost grow roughly quadratically with sequence length.