Unit content
Molecular sequence evidence for phylogenetic inference
DNA, RNA and protein sequences provide large numbers of heritable characters for reconstructing evolutionary relationships.
Once homologous positions have been aligned, each column can be treated as evidence about sequence change along candidate lineages. Closely related lineages often retain more sequence similarity because less time has passed for differences to accumulate, although evolutionary rates vary.
Simple example
Consider an aligned DNA region:
A: A C G T A C
B: A C G T G C
C: A C G C G C
A and B differ at one aligned position, B and C differ at one, and A and C differ at two. These differences provide evidence about possible relationships, but raw difference counts alone do not prove a tree: substitutions can occur repeatedly at the same site, and different sites can evolve at different rates.
Shared sequence changes can support clades when they are best explained as changes inherited from a common ancestor. Across many sites, candidate trees can be compared by how well they explain the observed pattern of states.
Modern molecular phylogenetics often uses explicit statistical models of sequence evolution and large datasets containing many genes or whole genomes. These models account for complications such as unequal substitution rates and repeated changes at the same site.
Molecular evidence does not replace morphology or fossils. Independent evidence can be combined, and disagreement can reveal convergence, rapid evolution, gene duplication or other complications.
The important principle is that molecular sequences are historical records with noise: they contain information about common ancestry, but that information must be interpreted using explicit hypotheses about sequence change.