Learning path

Full curriculum

Full curriculum

Unit content

The genetic code, codons and reading frames

In protein-coding messenger RNA (mRNA), nucleotide sequence is interpreted in groups of three bases called codons. Each codon specifies an amino acid or a translation stop signal.

Because RNA has four common bases, there are

$$4^3=64$$

possible codons.

The genetic code is redundant or degenerate: several codons can specify the same amino acid. But within the standard code, a particular codon does not normally specify several different amino acids.

Start and stop codons

The codon AUG commonly serves as a translation start codon and specifies methionine. The standard stop codons are

UAA   UAG   UGA

Stop codons do not encode an amino acid. They signal termination of translation.

Reading frame

A nucleotide sequence can be divided into triplets in different ways depending on where counting begins. A reading frame is one particular grouping into consecutive codons.

For example,

5'-AUG GCU UAC UGA-3'

read from the first base gives

AUG | GCU | UAC | UGA
Met | Ala | Tyr | Stop

where Met, Ala and Tyr abbreviate methionine, alanine and tyrosine.

Starting one nucleotide later would produce a different set of codons and therefore a different interpretation.

During translation, the ribosome establishes a reading frame at initiation and then advances along the mRNA codon by codon in the $5'\rightarrow3'$ direction.

The code connects two molecular alphabets

Codons are nucleotide sequences; amino acids are chemically different molecules. The genetic code is therefore a mapping between the nucleotide alphabet of RNA and the amino-acid alphabet of proteins.

The codon table does not itself explain how an amino acid is physically matched to a codon. That molecular decoding step is performed through transfer RNAs and aminoacyl-tRNA synthetases.