Learning path

Full curriculum

Full curriculum

Unit content

Sequence alignment as a hypothesis of positional homology

A sequence alignment arranges biological sequences so that positions placed in the same column are hypothesized to descend from corresponding positions in an ancestral sequence.

Consider three DNA sequences:

A: A C G T A C
B: A C G T G C
C: A C G C G C

Because they have equal length, they can be compared directly column by column. But insertions and deletions complicate correspondence. Suppose instead one sequence is

A: A C G T A C
D: A C G   A C

An alignment can represent the missing position with a gap:

A: A C G T A C
D: A C G - A C

The dash is not a nucleotide. It represents the hypothesis that an insertion occurred in one lineage or a deletion occurred in the other.

This distinction is important: an alignment is not merely visual spacing. Each aligned column makes a claim of positional homology—that the residues in that column are evolutionarily comparable.

Different alignments can sometimes explain the same sequences, especially in repetitive regions. Choosing among them requires evidence or an explicit scoring/modeling method.

Alignments may involve two sequences (pairwise alignment) or many sequences (multiple sequence alignment). Once homologous positions are aligned, substitutions and indels can be compared across lineages.

Thus the conceptual sequence is

$$\boxed{\text{raw sequences} \rightarrow \text{alignment hypothesis} \rightarrow \text{comparable homologous positions}.}$$

Algorithms for finding optimal alignments are a separate problem; this concept concerns what an alignment means biologically.