Genetic Code – Definition, Codon Table, Properties, and Examples

Summarise with AI:

Genetic code is the set of rules by which the nucleotide sequence present in messenger RNA (mRNA) determines the amino acid sequence during protein synthesis. It is read in groups of three nucleotides, known as codons. Each codon specifies an amino acid or gives a signal for termination of translation.

There are 64 codons in the standard genetic code. Among these, 61 codons code for amino acids and three codons, UAA, UAG and UGA, are stop codons. AUG commonly acts as the start codon and also codes for methionine. The genetic code is nearly universal, but some organisms and mitochondria show variations in codon meaning.

What is the Genetic Code?

Genetic code is the system by which the information present in the nucleotide sequence of messenger RNA (mRNA) is used for protein synthesis. The sequence is read in a group of three nucleotides called a codon. Each codon codes for an amino acid or gives a signal for termination of translation.

There are 64 codons in the standard genetic code. Of these, 61 codons code for amino acids while UAA, UAG and UGA are stop codons. AUG codes for methionine and generally acts as the start codon during protein synthesis. The genetic code is nearly universal, but some variations are also found.

Genetic information refers to the information present in the sequence of genetic material. It determines different structure and functions of the cell or organism. In cells, this information is present in DNA and is transferred from DNA to RNA during gene expression.

The genetic code is different. It is used during translation, where the nucleotide sequence of mRNA determines the amino acid sequence of protein. Thus, genetic information is the information present in the genetic material, while genetic code is used for reading this information during protein synthesis.

A gene is a segment of DNA and is a basic functional unit of heredity. It contains information for the formation of a protein or a functional RNA molecule. Different genes are present within the genetic material of an organism.

A genome refers to the complete genetic material of an organism. It contains genes as well as other DNA sequences. The genetic code, on the other hand, is not a gene or genome. It is the system by which the codons of mRNA are related to amino acids during protein synthesis.

What is a Codon?

A codon is a sequence of three nucleotides present in the messenger RNA (mRNA), that codes for an amino acid during protein synthesis. The process takes place during translation. In this process, the ribosomes read the codons present in mRNA and amino acids are added according to the codon sequence. Some codons do not code for an amino acid and are used for stopping the process of translation.

There are 64 possible codons in the standard genetic code. Of these, 61 codons are used for the 20 amino acids. The remaining three, UAA, UAG and UGA, are the stop codons. More than one codon can code for the same amino acid. AUG codes for methionine and generally acts as the initiation codon. The codons in mRNA are read in 5′ to 3′ direction during translation.

Structure of a Codon

  • A codon is made up of three nucleotides present in the mRNA. These three nucleotides are present one after another and form a single group.
  • There are four different nucleotides used in mRNA; adenine (A), uracil (U), guanine (G), and cytosine (C). Different arrangement of these four nucleotides forms different codons. There are 64 possible codons.
  • The three nucleotides of a codon have three positions, the first position, second position and third position. A change in the nucleotide at these positions can form another codon.
  • The first two positions are important in the recognition of an amino acid. In many codons the third position can be different, but the same amino acid is produced. This is common in the genetic code.
  • Codons present in mRNA are read in the 5′ to 3′ direction, three nucleotides at a time.
  • During translation, a codon of mRNA pairs with the complementary anticodon present in transfer RNA (tRNA). The tRNA carries an amino acid. In this process, amino acids are added one after another during protein synthesis.
Structure of a Codon
Structure of a Codon

Why is the Genetic Code a Triplet Code?

The genetic code is a triplet code because a codon is formed of three nucleotide bases. There are four different bases in RNA. With one base only 4 combinations are formed, while two bases can form 16 combinations. This number is not sufficient for 20 amino acids.

Three bases can form 64 different combinations. Thus, three nucleotides are used in a codon. These three nucleotide groups are referred to as triplet codons. The codons are read one after another during protein synthesis.

The triplet nature of genetic code was also shown by changes in the number of nucleotides. Addition or removal of one or two nucleotides changes the normal reading. When changes occur in a group of three, the reading can be restored. This supported that the genetic information is read in groups of three nucleotides.

Properties or Characteristics of the Genetic Code

The genetic code has some characteristic features by which the information of mRNA is used for synthesis of proteins. The following are the important properties of genetic code-

Properties or Characteristics of the Genetic Code
Properties or Characteristics of the Genetic Code
  • Genetic code is triplet– Each codon is made up of three nucleotide bases. With four different bases, a total of 64 codons are possible (4³ = 64). The triplet nature of the code was established from frameshift experiments.
  • It is degenerate– More than one codon can code for the same amino acid. Thus, 61 sense codons are present for 20 standard amino acids. Leucine, serine and arginine have six codons each, while methionine and tryptophan have only one codon. The codons specifying the same amino acid are referred to as synonymous codons.
  • The code is unambiguous– A particular sense codon normally specifies only one amino acid in a given genetic code. It does not code for two different amino acids at the same time. The different synonymous codons, however, can specify one amino acid.
  • Genetic code is non-overlapping– During normal translation, one nucleotide belongs to only one codon of a reading frame. The next codon starts after completion of the previous three bases. Thus, the codons are read one after another. Programmed frameshifting is an exception found in some genes.
  • The code is commaless– No comma or special nucleotide is present between two successive codons. Once the reading frame is fixed, codons are read continuously in groups of three up to the termination signal.
  • Genetic code is nearly universal– The same codons have the same amino acid meaning in most organisms. It is not completely universal. Mitochondria and some organisms show changes in codon assignment. For example, UGA codes for tryptophan and AUA for methionine in mammalian mitochondria instead of their assignments in the standard code.
  • The code has polarity– Codons of mRNA are read in the 5′ → 3′ direction during translation. The reading takes place in a fixed direction and does not proceed backward.
  • Start and stop signals are presentAUG is the usual initiation codon and starts the coding region in most cases. Three codons, UAA, UAG and UGA, act as termination codons in the standard genetic code. These stop codons are recognized by release factors rather than normal aminoacyl-tRNAs.
  • The code shows colinearity– The sequence of codons in a coding region corresponds to the sequence of amino acids in the polypeptide. Change in the position of codons therefore changes the corresponding position in the protein, depending on the type of mutation.
  • Wobble is present at codon recognition– Pairing is more flexible at the third base of many codons. Due to this, a single tRNA can recognize more than one synonymous codon. This is referred to as the wobble hypothesis, proposed by Francis Crick.

Genetic Code Table

The genetic code table shows the amino acid specified by each codon of mRNA. There are 64 codons. Among these, 61 codons code for amino acids and three codons are stop codons.

First baseSecond base USecond base CSecond base ASecond base G
UUUU, UUC – Phenylalanine
UUA, UUG – Leucine
UCU, UCC, UCA, UCG – SerineUAU, UAC – Tyrosine
UAA, UAG – Stop
UGU, UGC – Cysteine
UGA – Stop
UGG – Tryptophan
CCUU, CUC, CUA, CUG – LeucineCCU, CCC, CCA, CCG – ProlineCAU, CAC – Histidine
CAA, CAG – Glutamine
CGU, CGC, CGA, CGG – Arginine
AAUU, AUC, AUA – Isoleucine
AUG – Methionine (Start)
ACU, ACC, ACA, ACG – ThreonineAAU, AAC – Asparagine
AAA, AAG – Lysine
AGU, AGC – Serine
AGA, AGG – Arginine
GGUU, GUC, GUA, GUG – ValineGCU, GCC, GCA, GCG – AlanineGAU, GAC – Aspartic acid
GAA, GAG – Glutamic acid
GGU, GGC, GGA, GGG – Glycine

The codons are read in 5′ to 3′ direction. AUG generally acts as the start codon. UAA, UAG and UGA are the three stop codons.

Genetic Code Table diagram
Genetic Code Table diagram

How to Read the Genetic Code Table

The genetic code table is used to find the amino acid for a particular codon. For reading the table, the codon is taken from mRNA and read in 5′ to 3′ direction. Three bases are followed one after another.

First, the first base of codon is selected from the left side of table. The second base is then located from the upper part. At this point, the two bases give a group of possible codons. The third base is used finally which gives the particular codon and its amino acid.

For example, AUG can be taken. A is present at first position, U at second and G at third position. Following these three bases in the table gives methionine. AUG is also commonly used as the initiation codon.

Another example is UUU. Here all three positions contain uracil. It codes for phenylalanine.

Some codons do not give an amino acid. UAA, UAG and UGA are such codons and these are referred to as stop codons. They are used during termination of translation.

Worked Example of Reading mRNA Codons

Consider the following mRNA sequence

5′-AUG GCU UUU GAA UGA-3′

The mRNA is read in 5′ to 3′ direction. First, the sequence is divided into groups of three nucleotides. These groups are the codons.

AUG | GCU | UUU | GAA | UGA

The first codon is AUG. From the genetic code table, AUG codes for methionine. GCU gives alanine. The next codon UUU codes for phenylalanine, while GAA codes for glutamic acid. These codons are read one after another.

The last codon is UGA. It is a stop codon. No amino acid is specified by this codon and the translation is terminated at this position.

Therefore, the amino acid sequence formed before termination is-

Methionine – Alanine – Phenylalanine – Glutamic acid

DNA Codon Table vs RNA Codon Table

FeaturesDNA Codon TableRNA Codon Table
MeaningShows the nucleotide triplets present in DNA.Shows the codons present in mRNA.
Bases presentContains A, T, G and C.Contains A, U, G and C.
Thymine / UracilThymine (T) is present.Uracil (U) is present instead of thymine.
Triplet exampleATGAUG
Start sequenceATG in the coding DNA strand corresponds to the start codon.AUG commonly acts as the start codon.
Stop sequencesTAA, TAG and TGA in coding DNA correspond to stop codons.UAA, UAG and UGA are stop codons.
UseIt is used to show the relation of DNA triplets with the corresponding RNA codons.It is mainly used during translation to find the amino acid for a codon.
ReadingThe coding DNA sequence is written in 5′ to 3′ direction for direct comparison with mRNA.mRNA codons are read in 5′ to 3′ direction.
ExampleDNA: 5′-ATG GCT TTT-3′mRNA: 5′-AUG GCU UUU-3′
Amino acidsATG → Methionine, GCT → Alanine, TTT → PhenylalanineAUG → Methionine, GCU → Alanine, UUU → Phenylalanine

The main difference is the presence of T in DNA and U in RNA. In most genetic code tables, the codons are written using mRNA sequences.

Types of Codons in the Genetic Code

There are total 64 codons in the genetic code. Among these, 61 codons code for amino acids and three codons are used for termination. AUG has a special role in initiation of translation. The following are the types of codons-

Types of Codons in the Genetic Code
Types of Codons in the Genetic Code
  1. Sense codons– These are the codons that code for amino acids. There are 61 sense codons. As 61 codons are present for 20 standard amino acids, most amino acids are coded by more than one codon. This is referred to as degeneracy of genetic code. AUG codes for methionine whereas UGG codes for tryptophan, both have only one codon.
  2. Start or initiation codonAUG is the usual start codon. Translation generally begins from this codon and it also codes for methionine. Thus, AUG is also a sense codon. In bacteria, GUG and UUG can also work as start codons in some genes. The first amino acid added in bacterial protein synthesis is generally N-formylmethionine (fMet).
  3. Stop or nonsense codonsUAA, UAG and UGA are the three stop codons. These codons normally do not code for any amino acid. When these codons are reached during translation, formation of the polypeptide chain is stopped. They are also referred to as termination codons. UAA is called ochre, UAG is amber and UGA is opal or umber codon.

In some cases, the stop codons can have a different function. UGA can code for selenocysteine and UAG for pyrrolysine in particular organisms.

Degeneracy of the Genetic Code

Degeneracy of the Genetic Code
Degeneracy of the Genetic Code

What Does Degenerate Genetic Code Mean?

The genetic code is said to be degenerate because an amino acid can be coded by more than one codon. There are 61 sense codons for 20 standard amino acids.

These different codons for a single amino acid are called synonymous codons.

Methionine and tryptophan are exceptions. Methionine has AUG and tryptophan has UGG, only one codon for each. Other amino acids have two, three, four or six codons.

The difference between synonymous codons is commonly found at the third nucleotide. The third position is associated with wobble during codon-anticodon pairing.

Degeneracy is a property of the genetic code. It does not indicate that one codon normally gives several amino acids.

Examples of Degenerate Codons

Some of the examples of degenerate codons are-

  • Leucine– Six codons are present, UUA, UUG, CUU, CUC, CUA and CUG.
  • Serine– It is also represented by six codons, UCU, UCC, UCA, UCG, AGU and AGC.
  • Arginine has CGU, CGC, CGA, CGG, AGA and AGG.
  • Isoleucine– Three codons AUU, AUC and AUA code for isoleucine.
  • Four codons GGU, GGC, GGA and GGG are for glycine.
  • Alanine is coded by GCU, GCC, GCA and GCG.
  • Two codons are present for lysine, AAA and AAG. Phenylalanine also has two, UUU and UUC.
  • Sixfold, fourfold, threefold and twofold codon groups are therefore present in the standard genetic code.

Complete and Partial Degeneracy

Degeneracy may be complete or partial depending on the changes at third position of the codon.

  1. Complete degeneracy– In this type, change of the third nucleotide to any of the four bases does not change the amino acid. GGU, GGC, GGA and GGG all code for glycine. The same condition is present in GCU, GCC, GCA and GCG for alanine. These are also called fourfold degenerate codons.
  2. Partial degeneracy– Only some of the bases at third position can be changed without changing the amino acid. AAA and AAG both code for lysine. AAU and AAC, however, code for asparagine. Twofold degenerate codon groups are common in the genetic code.

Isoleucine has three codons. Leucine, serine and arginine contain six codons and their codons occur in more than one codon group.

Biological Importance of Degeneracy

  • Degeneracy can reduce the effect of some mutations. Change of a nucleotide, particularly at some third codon positions, may still give the same amino acid.
  • Such nucleotide substitutions without a change in amino acid are referred to as synonymous substitutions.
  • Wobble pairing occurs between the third base of codon and first base of tRNA anticodon. Because of this, some tRNAs can recognize more than one codon.
  • All synonymous codons are not used at the same frequency. This unequal use of synonymous codons is called codon usage bias.
  • Synonymous changes are not always without biological effect. Changes in codon use can affect translation and other processes related to gene expression.

Degenerate vs Unambiguous Genetic Code

  • Degenerate genetic code– Several codons may code for one amino acid. Glycine, for example, has GGU, GGC, GGA and GGG.
  • Unambiguous genetic code– A codon normally specifies one particular amino acid. GGU codes for glycine, not for different amino acids.
  • The genetic code has both these properties. More than one codon may be present for an amino acid, whereas one codon normally has one amino acid meaning.

Wobble Hypothesis and Genetic Code

Wobble Hypothesis and Genetic Code
Wobble Hypothesis and Genetic Code

What is the Wobble Hypothesis?

Wobble hypothesis was proposed by Francis Crick in 1966. It explains the less strict base pairing between some codons of mRNA and the anticodon of tRNA.

During translation, the codon pairs with complementary anticodon. The first two bases of a codon generally show more strict pairing. At the third codon base, some non-standard pairing can occur. This is called wobble.

Due to wobble, one tRNA can recognize more than one codon in some cases. A different tRNA is not needed for every sense codon.

G-U pairing is one common wobble pairing. Inosine (I) present in some tRNA anticodons can pair with U, C or A according to the classical wobble rules.

Modified bases are common at this position of tRNA. These modifications can change the codons which are recognized by a tRNA.

Wobble Position in Codon-Anticodon Pairing

  • The wobble position is the third nucleotide of the mRNA codon. In the tRNA anticodon, it pairs with the first nucleotide or position 34.
  • Codon and anticodon are arranged in opposite directions. The third base at the 3′ side of codon therefore comes against the 5′ base of anticodon.
  • At the first two codon positions, normal Watson-Crick pairing is mainly followed. More flexibility is present at the third position.
  • When G occurs at the wobble position of anticodon, it may recognize codons ending with C or U. In classical bacterial wobble, U at this position may pair with A or G.
  • Inosine is a modified purine found in some anticodons. It can pair with A, C or U at the third position of codon. One tRNA can in this way read three related codons.

Relationship Between Wobble and Codon Degeneracy

  • The genetic code is degenerate, because most amino acids have more than one codon. These codons are referred to as synonymous codons.
  • Wobble is associated with this degeneracy. Different synonymous codons, mainly differing at the third base, can be read by the same tRNA in many cases.
  • For example, a wobble G-U pair can permit one tRNA to recognize codons having C or U at their third position. Four-codon families can also be read with fewer tRNAs.
  • Wobble and degeneracy are not the same. Degeneracy is the occurrence of more than one codon for an amino acid, while wobble is the flexible codon-anticodon pairing during translation.
  • Degeneracy can also involve different tRNAs for different synonymous codons. Wobble explains why all the 61 sense codons do not require 61 separate tRNAs.

How Does the Genetic Code Work During Protein Synthesis?

During protein synthesis, the codons present in mRNA are used for arranging amino acids in a protein. The mRNA is read in 5′ to 3′ direction. Three bases make one codon.

How Does the Genetic Code Work During Protein Synthesis

The role of genetic code during this process can be given as follows-

  • Start codon– The reading generally starts from AUG. It is the initiation codon and codes for methionine. In bacteria, the initiating amino acid is N-formylmethionine (fMet).
  • Reading of codons– After initiation, codons are read one after another by the ribosome. Each group of three nucleotides is taken as one codon. The reading frame remains fixed during normal protein synthesis.
  • tRNA and codon recognitiontRNA contains an anticodon and carries amino acid. The anticodon pairs with its corresponding codon of mRNA. For example, a codon for glycine will be recognized by a tRNA carrying glycine. The amino acid is then used in the growing polypeptide chain.
  • Degenerate codons– Several amino acids have more than one codon. GGU, GGC, GGA and GGG all represent glycine. Some related codons can also be recognized by the same tRNA because of wobble pairing at the third codon position.
  • Order of amino acids– The codons occur in a particular sequence in mRNA. Amino acids are also added according to this order. A change in codon may therefore change the amino acid, although some changes give the same amino acid due to degeneracy.
  • Stop codonsUAA, UAG and UGA are termination codons. No amino acid is normally added for these codons. When any one of these is reached, release of the polypeptide takes place and reading of the coding sequence ends.

Reading Frame and the Genetic Code

Reading Frame and the Genetic Code
Reading Frame and the Genetic Code
  • A reading frame is the grouping of nucleotide sequence into successive groups of three bases. Each group forms one codon. The position from where this grouping starts decides the reading frame of mRNA.
  • The same mRNA sequence can have three possible reading frames. In the first frame, reading starts from nucleotide 1. The second starts from nucleotide 2, whereas third frame starts from nucleotide 3. Different codons are obtained in each case and therefore a different amino acid sequence can be formed.
  • For example, a sequence AUGGCUACG… can be grouped as AUG-GCU-ACG in one frame. Starting from the next nucleotide gives UGG-CUA…, while another frame will give GGC-UAC…. Thus, shifting the starting position by even one nucleotide changes the codons.
  • During normal protein synthesis, only the proper reading frame of a coding region is used. Initiation of translation places the ribosome at the start site and the codons are then read successively without changing this frame.
  • An open reading frame (ORF) is a series of codons starting with an initiation codon and ending at a termination codon. No in-frame stop codon is present between them. ORFs are therefore used to identify possible protein-coding regions in nucleotide sequences.
  • A long ORF can indicate a protein-coding sequence, but the presence of an ORF alone does not always prove that it forms a protein. Short and alternative ORFs are also present in many mRNAs.
  • Frameshift mutation occurs when nucleotide insertion or deletion changes the normal grouping of codons. Addition or removal of bases which is not a multiple of three causes this type of mutation.
  • After a frameshift, codons present after the mutation are generally changed. Many different amino acids may now be formed and a premature stop codon is also commonly produced. The protein can therefore become shortened or non-functional.
  • Addition or deletion of three nucleotides, or another multiple of three, does not shift the reading frame. One or more amino acids are added or removed, but the codons present after that region remain in the original frame.

Discovery and Deciphering of the Genetic Code

The genetic code was mainly deciphered during the early 1960s. Different genetic experiments, cell-free protein synthesis and synthetic RNA molecules were used for this work. Some of the important experiments are as follows-

  • In 1961, Francis Crick, Leslie Barnett, Sydney Brenner and R. J. Watts-Tobin studied frameshift mutations in T4 bacteriophage. Addition or deletion of one or two nucleotides disturbed the reading frame. Three such changes in suitable combination could restore it. From these experiments, the code was shown to be read in groups of three nucleotides.
  • Marshall Nirenberg and J. Heinrich Matthaei used a cell-free system of Escherichia coli in 1961. Synthetic poly-U RNA, containing only uracil, was added to the system. It produced a polypeptide containing phenylalanine. UUU was therefore the first codon to be identified, coding for phenylalanine.
  • After the poly-U experiment, other synthetic RNAs were used for finding more codons. Different nucleotide compositions gave different amino acid products. The work gradually gave assignments for many codons.
  • In 1964, Philip Leder and Marshall Nirenberg developed the triplet binding method. Short synthetic trinucleotides were used instead of long RNA molecules. A particular trinucleotide caused binding of its corresponding aminoacyl-tRNA to the ribosome. This method was used for direct identification of a large number of codons.
  • Khorana and his coworkers prepared synthetic polyribonucleotides having known repeating nucleotide sequences. These RNAs were used in protein-synthesizing systems. Several remaining codons and degenerate codon groups were determined from these experiments.
  • Holley worked on transfer RNA and determined the nucleotide sequence of alanine tRNA. The complete sequence was established during the mid-1960s. This work gave important information on the adaptor molecule involved between codon and amino acid during protein synthesis.
  • By the middle of the 1960s, almost all the codons and their amino acid meanings had been established. The genetic code table was completed mainly from the works of Nirenberg and Khorana together with other researchers working on protein synthesis.
  • In 1968, Marshall W. Nirenberg, Har Gobind Khorana and Robert W. Holley were jointly awarded the Nobel Prize in Physiology or Medicine for their work on interpretation of the genetic code and its function in protein synthesis.

Exceptions to the Standard Genetic Code

The genetic code is nearly universal, but it is not completely same in all organisms and cellular systems. Some codons have different meaning in mitochondria, certain bacteria, yeasts and protozoans. These are referred to as non-standard or variant genetic codes.

The following are some of the important exceptions-

  1. Mitochondria– Mitochondria show several changes from the standard genetic code. In animal mitochondria, UGA codes for tryptophan instead of stop. AUA is used for methionine in most metazoan mitochondria, whereas in the standard code it codes for isoleucine. Other changes are also present and these are not same in all mitochondrial groups.
  2. Mycoplasma– In Mycoplasma capricolum, UGA is read as tryptophan. UGA is normally a termination codon in the standard genetic code. UGG and UGA both can therefore code for tryptophan in this bacterium.
  3. Ciliates– Some ciliated protozoans have changes in stop codon meaning. In Tetrahymena and Paramecium, UAA and UAG code for glutamine. UGA mainly acts as stop codon in these organisms.
  4. Euplotes– In Euplotes, UGA codes for cysteine. UAA and UAG are used for termination. Thus, UGA does not have the same meaning in all organisms.
  5. CTG clade yeasts– In several yeasts of CTG clade, the CUG codon is reassigned from leucine towards serine. In Candida albicans, CUG is mainly translated as serine. A small amount of leucine incorporation can also occur.
  6. SelenocysteineUGA can also be used for incorporation of selenocysteine (Sec) in particular mRNAs. Specific tRNA, protein factors and a SECIS element are required for this process. UGA therefore acts as a special recoded codon in this condition.
  7. PyrrolysineUAG can code for pyrrolysine (Pyl) in some methanogenic archaea and some bacteria. UAG is normally an amber stop codon. A specific tRNA and pyrrolysyl-tRNA synthetase are involved in this process.

The standard genetic code is highly conserved, but these exceptions are present in some organisms and organelles. Most changes involve stop codons, while changes in some sense codons are also known.

Genetic Code and Mutations

A mutation is a change in the nucleotide sequence of DNA. In the coding region, this change may alter a codon and also the amino acid sequence of protein. But every nucleotide change does not change amino acid because several amino acids have more than one codon.

Genetic Code and Mutations
Genetic Code and Mutations

Some of the mutations related with genetic code are-

  • Silent mutation– It is a mutation where nucleotide change produces another codon for the same amino acid. For example, GAA and GAG both code for glutamate. Hence, amino acid remains glutamate. These are also referred to as synonymous mutations. Some synonymous mutations can affect mRNA stability and translation.
  • Missense mutation– In this mutation, changed codon codes for another amino acid. A different amino acid is therefore present in the protein. The effect can be small or it may affect the activity of protein, depending upon the amino acid and its position.
  • Nonsense mutation– Here, a sense codon is changed into a stop codon. UAA, UAG or UGA is formed. Protein synthesis now stops before its normal position and a shorter polypeptide can be produced. mRNA containing premature termination codon may also be removed by nonsense-mediated decay (NMD).
  • Frameshift mutation– Addition or deletion of nucleotide changes the reading frame when number of bases is not in multiple of three. Codons after this position become different. Many amino acids may be changed and an early stop codon can also occur.
  • In-frame mutation– Addition or deletion of three bases, or its multiples, does not shift the reading frame. One or more amino acids are added or removed from protein. The remaining sequence is read in its original frame.
  • Start codon mutation– Mutation of the normal start codon can affect initiation of protein synthesis. Translation may fail to start from this position. In some cases another downstream initiation codon is used, producing a shorter protein.
  • Stop-loss mutation– The normal stop codon is changed into a codon for amino acid. Translation continues beyond the normal termination position until another stop codon occurs. An extended protein is formed in this condition.
  • Degeneracy and mutation– The genetic code is degenerate. Due to this property, some base substitutions have no change in the amino acid. This is frequently found at the third base of codon, but all changes at third position do not give the same amino acid.

Importance of the Genetic Code

The genetic code is important for reading the genetic information present in mRNA and formation of proteins. It is also used in mutation study, molecular genetics and different biotechnology works.

Some of the important roles of genetic code are-

  • Protein synthesis– Genetic code is used during translation for arrangement of amino acids in a polypeptide. The codons of mRNA are read and the corresponding amino acids are added.
  • Genetic information– The nucleotide information of a protein-coding gene is finally expressed as an amino acid sequence. Here, each codon has its amino acid meaning or termination function. The genetic code is required for this conversion of nucleotide sequence into protein sequence.
  • Mutation study– Change in nucleotide can also change the codon. It may give the same amino acid, another amino acid, or sometimes a termination codon. Genetic code is therefore important for finding the effect of mutations present within protein-coding regions.
  • Evolutionary study– The genetic code is nearly same among different groups of living organisms, although some exceptions are present. This highly conserved nature is useful during comparison of genes and proteins of different organisms.
  • Recombinant proteins– Genes from one organism can be expressed in another organism using suitable expression systems. The common codon assignments are important here. Codon preference of host, however, can affect the amount of protein formed and codon optimization is commonly used in recombinant protein work.
  • Codon usage– Synonymous codons do not always behave exactly same during gene expression. Changes in their use can affect translation and in some proteins also the folding of newly formed polypeptide. This has importance during design of synthetic and recombinant genes.
  • Genetic engineering– The genetic code can also be experimentally modified. Stop codons or selected sense codons are reassigned for insertion of non-canonical amino acids into proteins. It is used in protein engineering, protein labelling and study of protein structure and function.
  • Disease genetics– Many genetic changes occurring in coding regions are studied according to the codon formed after mutation. Missense and nonsense changes alter the protein sequence directly, whereas some synonymous changes also can have biological effects.

Genetic Code vs Codon vs Anticodon

FeatureGenetic CodeCodonAnticodon
DefinitionGenetic code is the set of rules by which nucleotide sequence of mRNA determines the amino acid sequence of protein.A codon is a sequence of three nucleotides present in mRNA.An anticodon is a sequence of three nucleotides present in tRNA.
LocationIt is represented through the codon arrangement in genetic information.Codons are present in mRNA.Anticodons are present in tRNA.
Number of basesThe genetic code is made up of different triplet codons.Each codon contains three nucleotide bases.Each anticodon also contains three nucleotide bases.
Main functionIt determines which amino acid is represented by each codon.A codon specifies an amino acid or acts as a start or stop signal.Anticodon recognizes and pairs with its corresponding codon during translation.
Role in protein synthesisIt provides the basis for conversion of nucleotide information into amino acid sequence.Codons are read one after another by the ribosome.Anticodons help tRNA to bring the proper amino acid according to the codon.
DirectionThe coding information in mRNA is read in 5′ to 3′ direction.Codon is written in 5′ to 3′ direction.Anticodon pairs with codon in an antiparallel manner.
ExampleUUU codes for phenylalanine, AUG for methionine and UGG for tryptophan.AUG is a codon for methionine and generally acts as start codon.The anticodon complementary to AUG can be written as 3′-UAC-5′.
NumberThe standard genetic code contains 64 codons.There are 64 possible codons, of which 61 are sense codons and 3 are stop codons.The number of anticodons is lower than 61 in many organisms because one tRNA can recognize more than one codon in some cases.
Special featureThe genetic code is degenerate, nearly universal and unambiguous.More than one codon may code for the same amino acid.Wobble pairing can occur at the first base of anticodon with the third base of codon.

FAQ

What is genetic code?

The genetic code is the set of rules by which the nucleotide sequence of mRNA is read during protein synthesis. The information is present in the form of three nucleotide codons. Each codon specifies an amino acid or acts as a start or stop signal.

What are the properties of genetic code?

  1. It is a triplet code.
  2. The genetic code is degenerate.
  3. It is unambiguous.
  4. Codons are generally non-overlapping.
  5. The code is nearly universal.
  6. It is read in 5′ to 3′ direction.
  7. Start and stop codons are present.

Why are there 64 codons?

Four nucleotide bases are present in RNA, A, U, G and C. Each codon contains three nucleotide positions.

4 × 4 × 4 = 4³ = 64 codons

Thus, 64 different triplet combinations are possible.

What are the three stop codons?

The three stop codons are UAA, UAG and UGA. These codons normally do not code for an amino acid and are used for termination of protein synthesis.

What is the start codon?

AUG is the usual start codon in the standard genetic code and also specifies methionine. It generally marks the position from where translation starts.

References

  1. Agris, P. F., Eruysal, E. R., Narendran, A., Väre, V. Y. P., Vangaveti, S., & Ranganathan, S. V. (2018). Celebrating wobble decoding: Half a century and still much is new. RNA Biology, 15(4–5), 537–553. https://doi.org/10.1080/15476286.2017.1356562
  2. Agris, P. F., Narendran, A., Sarachan, K., Väre, V. Y. P., & Eruysal, E. (2017). The importance of being modified: The role of RNA modifications in translational fidelity. The Enzymes, 41, 1–50. https://doi.org/10.1016/bs.enz.2017.03.005
  3. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). From RNA to protein. In Molecular biology of the cell (4th ed.). Garland Science. https://www.ncbi.nlm.nih.gov/books/NBK26829/
  4. Bezerra, A. R., Guimarães, A. R., & Santos, M. A. S. (2015). Non-standard genetic codes define new concepts for protein engineering. Life, 5(4), 1610–1628. https://doi.org/10.3390/life5041610
  5. Bezerra, A. R., Simões, J., Lee, W., Rung, J., Weil, T., Gut, I. G., Gut, M., Bayés, M., Rizzetto, L., Cavalieri, D., Giovannini, G., Bozza, S., Romani, L., Kapushesky, M., Moura, G. R., & Santos, M. A. S. (2021). The role of non-standard translation in Candida albicans adaptation to environmental stress. FEMS Yeast Research, 21(4), foab032. https://doi.org/10.1093/femsyr/foab032
  6. Blanchet, S., & Ranjan, N. (2022). Translation phases in eukaryotes. In K.-D. Entian (Ed.), Ribosome biogenesis: Methods and protocols (pp. 255–280). Humana. https://doi.org/10.1007/978-1-0716-2501-9_13
  7. Brown, T. A. (2002). Genomes (2nd ed.). Wiley-Liss. https://www.ncbi.nlm.nih.gov/books/NBK21121/
  8. Brown, T. A. (2002). Synthesis and processing of the proteome. In Genomes (2nd ed.). Wiley-Liss. https://www.ncbi.nlm.nih.gov/books/NBK21111/
  9. Caskey, C. T., & Leder, P. (2014). The RNA code: Nature’s Rosetta Stone. Proceedings of the National Academy of Sciences of the United States of America, 111(16), 5758–5759. https://doi.org/10.1073/pnas.1404819111
  10. Cooper, G. M. (2000). Expression of genetic information. In The cell: A molecular approach (2nd ed.). Sinauer Associates. https://www.ncbi.nlm.nih.gov/books/NBK9842/
  11. Cooper, G. M. (2000). Translation of mRNA. In The cell: A molecular approach (2nd ed.). Sinauer Associates. https://www.ncbi.nlm.nih.gov/books/NBK9849/
  12. Crick, F. H. C. (1966). Codon-anticodon pairing: The wobble hypothesis. Journal of Molecular Biology, 19(2), 548–555. https://pubmed.ncbi.nlm.nih.gov/5969078/
  13. Crick, F. H. C., Barnett, L., Brenner, S., & Watts-Tobin, R. J. (1961). General nature of the genetic code for proteins. Nature, 192, 1227–1232. https://doi.org/10.1038/1921227a0
  14. Dunkle, J. A., & Dunham, C. M. (2015). Mechanisms of mRNA frame maintenance and its subversion during translation of the genetic code. Biochimie, 114, 90–96. https://doi.org/10.1016/j.biochi.2015.02.007
  15. Freistroffer, D. V., Kwiatkowski, M., Buckingham, R. H., & Ehrenberg, M. (2000). The accuracy of codon recognition by polypeptide release factors. Proceedings of the National Academy of Sciences of the United States of America, 97(5), 2046–2051. https://doi.org/10.1073/pnas.030541097
  16. Fu, H., Liang, Y., Zhong, X., Pan, Z., Huang, L., Zhang, H., Xu, Y., Zhou, W., & Liu, Z. (2020). Codon optimization with deep learning to enhance protein expression. Scientific Reports, 10, 17617. https://doi.org/10.1038/s41598-020-74091-z
  17. Gan, Q., Lehman, B. P., Bobik, T. A., & Fan, C. (2016). Expanding the genetic code of Salmonella with non-canonical amino acids. Scientific Reports, 6, 39920. https://doi.org/10.1038/srep39920
  18. Ganesh, R. B., & Maerkl, S. J. (2022). Biochemistry of aminoacyl tRNA synthetase and tRNAs and their engineering for cell-free and synthetic cell applications. Frontiers in Bioengineering and Biotechnology, 10, 918659. https://doi.org/10.3389/fbioe.2022.918659
  19. Geyer, R., Madany Mamlouk, A., & Kötter, T. (2018). On the efficiency of the genetic code after frameshift mutations. PeerJ, 6, e4825. https://doi.org/10.7717/peerj.4825
  20. Harfe, B. D., & Jinks-Robertson, S. (1999). Removal of frameshift intermediates by mismatch repair proteins in Saccharomyces cerevisiae. Molecular and Cellular Biology, 19(7), 4766–4773. https://pubmed.ncbi.nlm.nih.gov/10373526/
  21. Jestin, J.-L. (2006). Degeneracy in the genetic code and its symmetries by base substitutions. Comptes Rendus Biologies, 329(3), 168–171. https://doi.org/10.1016/j.crvi.2006.01.003
  22. Köhrer, C., & RajBhandary, U. L. (2013). Proteins carrying one or more unnatural amino acids. In Madame Curie Bioscience Database. Landes Bioscience. https://www.ncbi.nlm.nih.gov/books/NBK6136/
  23. Komar, A. A. (2016). The yin and yang of codon usage. Human Molecular Genetics, 25(R2), R77–R85. https://pubmed.ncbi.nlm.nih.gov/27354349/
  24. Leong, A. Z. X., Lee, P. Y., Mohtar, M. A., & Syafruddin, S. E. (2022). Short open reading frames (sORFs) and microproteins: An update on their identification and validation measures. Journal of Biomedical Science, 29, 19. https://pubmed.ncbi.nlm.nih.gov/35300685/
  25. Liu, Y. (2020). A code within the genetic code: Codon usage regulates co-translational protein folding. Cell Communication and Signaling, 18, 145. https://doi.org/10.1186/s12964-020-00642-6
  26. Lukashenko, N. P. (2010). Amino acids 21 and 22—Selenocysteine and pyrrolysine. Biochemistry (Moscow), 75, 921–931. https://pubmed.ncbi.nlm.nih.gov/20873198/
  27. McGowan, J., et al. (2023). Identification of a non-canonical ciliate nuclear genetic code where UAA and UAG encode different amino acids. PLOS Genetics, 19(10), e1010913. https://doi.org/10.1371/journal.pgen.1010913
  28. Morgan, A. R., Wells, R. D., & Khorana, H. G. (1966). Studies on polynucleotides, LIX. Further codon assignments from amino acid incorporation directed by ribopolynucleotides containing repeating trinucleotide sequences. Proceedings of the National Academy of Sciences of the United States of America, 56(6), 1899–1906. https://doi.org/10.1073/pnas.56.6.1899
  29. Nirenberg, M. W., & Leder, P. (1964). RNA codewords and protein synthesis: The effect of trinucleotides upon the binding of sRNA to ribosomes. Science, 145(3639), 1399–1407. https://pubmed.ncbi.nlm.nih.gov/14172630/
  30. Nirenberg, M. W., Leder, P., Bernfield, M., Brimacombe, R., Trupin, J., Rottman, F., & O’Neal, C. (1965). RNA codewords and protein synthesis, VII. On the general nature of the RNA code. Proceedings of the National Academy of Sciences of the United States of America, 53(5), 1161–1168. https://doi.org/10.1073/pnas.53.5.1161
  31. Nirenberg, M. W., & Matthaei, J. H. (1961). The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides. Proceedings of the National Academy of Sciences of the United States of America, 47(10), 1588–1602. https://doi.org/10.1073/pnas.47.10.1588
  32. Oelschlaeger, P. (2024). Molecular mechanisms and the significance of synonymous mutations. International Journal of Molecular Sciences, 25. https://pubmed.ncbi.nlm.nih.gov/38275761/
  33. Ohama, T., Inagaki, Y., Bessho, Y., & Osawa, S. (2008). Evolving genetic code. Proceedings of the Japan Academy, Series B, Physical and Biological Sciences, 84(2), 58–74. https://pubmed.ncbi.nlm.nih.gov/18941287/
  34. Parvathy, S. T., Udayasuriyan, V., & Bhadana, V. (2022). Codon usage bias. Molecular Biology Reports, 49(1), 539–565. https://doi.org/10.1007/s11033-021-06749-4
  35. Popovic, A., & Orrick, J. A. (2024). Biochemistry, mutation. In StatPearls. StatPearls Publishing. https://www.ncbi.nlm.nih.gov/books/NBK576397/
  36. Scolnick, E. M., & Caskey, C. T. (1969). Peptide chain termination. V. The role of release factors in mRNA terminator codon recognition. Proceedings of the National Academy of Sciences of the United States of America, 64(4), 1235–1241. https://doi.org/10.1073/pnas.64.4.1235
  37. Subramaniam, A. R., Pan, T., & Cluzel, P. (2013). Environmental perturbations lift the degeneracy of the genetic code to regulate protein levels in bacteria. Proceedings of the National Academy of Sciences of the United States of America, 110(6), 2419–2424. https://doi.org/10.1073/pnas.1211077110
  38. Suzuki, K. (1999). Nature of mutations in genetic disorders. In G. J. Siegel, B. W. Agranoff, R. W. Albers, S. K. Fisher, & M. D. Uhler (Eds.), Basic neurochemistry: Molecular, cellular and medical aspects (6th ed.). Lippincott-Raven. https://www.ncbi.nlm.nih.gov/books/NBK27942/
  39. Tourancheau, A. B., Tsao, N., Klobutcher, L. A., Pearlman, R. E., & Adoutte, A. (1995). Genetic code deviations in the ciliates: Evidence for multiple and independent events. The EMBO Journal, 14(13), 3262–3267. https://doi.org/10.1002/j.1460-2075.1995.tb07329.x
  40. Walsh, I. M., Bowman, M. A., Soto Santarriaga, I. F., Rodriguez, A., & Clark, P. L. (2020). Synonymous codon substitutions perturb cotranslational protein folding in vivo and impair cell fitness. Proceedings of the National Academy of Sciences of the United States of America, 117(7), 3528–3534. https://doi.org/10.1073/pnas.1907126117
  41. Watanabe, K. (2011). tRNA modification and genetic code variations in animal mitochondria. Journal of Nucleic Acids, 2011, 623095. https://doi.org/10.4061/2011/623095
  42. Watford, S., & Warrington, S. J. (2023). Bacterial DNA mutations. In StatPearls. StatPearls Publishing. https://www.ncbi.nlm.nih.gov/books/NBK459274/
  43. Wong, J. T.-F. (1976). The evolution of a universal genetic code. Proceedings of the National Academy of Sciences of the United States of America, 73(7), 2336–2340. https://doi.org/10.1073/pnas.73.7.2336
  44. Yamao, F., Muto, A., Kawauchi, Y., Iwami, M., Iwagami, S., Azumi, Y., & Osawa, S. (1985). UGA is read as tryptophan in Mycoplasma capricolum. Proceedings of the National Academy of Sciences of the United States of America, 82(8), 2306–2309. https://doi.org/10.1073/pnas.82.8.2306
  45. Yu, J. (2007). A content-centric organization of the genetic code. Genomics, Proteomics & Bioinformatics, 5(1), 1–6. https://doi.org/10.1016/S1672-0229(07)60008-4
  46. Yuan, J., O’Donoghue, P., Ambrogelly, A., Gundllapalli, S., Sherrer, R. L., Palioura, S., Simonović, M., & Söll, D. (2010). Distinct genetic code expansion strategies for selenocysteine and pyrrolysine are reflected in different aminoacyl-tRNA formation systems. FEBS Letters, 584(2), 342–349. https://doi.org/10.1016/j.febslet.2009.11.005

Start Asking Questions