SARS-CoV-2 is an enveloped, positive-sense single-stranded RNA (+ssRNA) virus having a roughly spherical to ellipsoidal virion and a genome of about 30 kb.
The viral particles are commonly around 80-100 nm in diameter as shown by cryo-electron tomography. Variation in size and shape of individual virions is also present. Inside the envelope, the RNA genome is associated with nucleocapsid (N) proteins and packed into ribonucleoprotein (RNP) complexes.
The structure of SARS-CoV-2 is mainly formed by four structural proteins, i.e., spike (S), membrane (M), envelope (E), and nucleocapsid (N) proteins. The S, M and E proteins are associated with the lipid envelope. N protein is present inside the virion and remains bound with the viral RNA.
Spike (S) proteins occur as glycosylated trimers projecting out from the surface of the virus. These projections give the characteristic “corona-like” appearance. The spike also contains the receptor-binding region, which is used for attachment to the host-cell ACE2 receptor.

Structure of SARS-CoV-2
- SARS-CoV-2 is an enveloped, roughly spherical to ellipsoidal virus particle. The shape is not completely uniform. Cryo-electron tomography of intact virions showed particles around 80–100 nm in size, with considerable difference between the individual virus particles.
- Outside the virus particle is a lipid bilayer envelope. In this membrane are three structural proteins, the spike (S), membrane (M), and envelope (E) proteins. The nucleocapsid (N) protein and viral RNA occur within the envelope.
- Large spike (S) glycoproteins project out from the viral surface. They occur as glycosylated trimers and are not placed as a completely regular layer over the membrane. Intact SARS-CoV-2 particles examined by cryo-electron tomography contained an average of about 24 spike trimers, although their number differs among virions. The spikes can also tilt considerably because of flexibility in their stalk region. These projecting structures give coronavirus particles their characteristic “crown-like” appearance.
- Each S protein is formed of the S1 and S2 regions. S1 contains the receptor-binding domain (RBD), while S2 forms the membrane-associated fusion part and contains the transmembrane region by which the spike remains fixed into the viral envelope. The three S molecules together make one spike.
- Membrane (M) protein is present abundantly in the viral envelope. It contains three transmembrane helices with a larger C-terminal region facing towards the inside of the virion. M proteins form dimers and take part in organizing the viral membrane and incorporation of other structural components during formation of the particle.
- The envelope (E) protein is much smaller than S and M. It is an integral membrane protein with a hydrophobic transmembrane region and is incorporated into the viral envelope in relatively small amount. E is associated with the membrane and participates with M during particle assembly and budding.
- Within the envelope, the viral genome is associated with nucleocapsid (N) proteins. N binds with the RNA and produces viral ribonucleoprotein (vRNP) complexes. These are not present as one simple continuous nucleocapsid structure. Multiple RNP assemblies occur throughout the viral lumen, with many also lying close to the inner surface of the membrane.
- The genome packed inside the particle is a positive-sense single-stranded RNA (+ssRNA) of about 30 kb. Packing of such a long RNA takes place around multiple N-containing RNP complexes. Cryo-electron tomography has shown these RNPs as compact structures of about 15 nm in diameter inside intact SARS-CoV-2 particles.

Virion Structure of SARS-CoV-2
Overall Virion Architecture and Lipid Envelope
- SARS-CoV-2 virion is an enveloped particle, generally spherical to ellipsoidal in shape. The shape and size are not completely uniform among the virus particles. In intact virions studied by cryo-electron tomography, the envelope diameter is commonly around 80–100 nm.
- Around the particle is a lipid bilayer envelope. Spike (S), membrane (M), and envelope (E) proteins are associated with this membrane. The nucleocapsid (N) protein is present inside the envelope, associated with the genomic RNA.
- Inside the virion, the positive-sense single-stranded RNA is not present freely. It remains packed with N proteins in the form of viral ribonucleoprotein (vRNP) complexes. These RNP structures occupy much of the internal space and are also found close to the inner surface of the viral membrane.
- Surface spikes- Large spike projections come out from the envelope and are distributed over the virus particle. Under electron microscopy, these projections give the particle a crown-like appearance. The name “coronavirus” is derived from this appearance.
Spike (S) Glycoprotein
- Spike (S) glycoprotein is present on the outer surface of SARS-CoV-2 as a glycosylated homotrimer. Three S protein molecules come together to form one spike. It has a large ectodomain projecting outside and a membrane-spanning part by which the protein is anchored into the lipid envelope.
- On intact virus particles, individual spikes can tilt and take different orientations on the membrane. The S protein is divided mainly into the S1 and S2 regions. S1 contains the receptor-binding structures, whereas the membrane-associated S2 contains the fusion machinery.
S1 Receptor-Binding Region
- S1 region- The receptor-binding domain (RBD) is located in S1. This domain can occur in relatively closed or exposed conformations, and in the exposed state it can directly bind with the angiotensin-converting enzyme 2 (ACE2) receptor.
S2 Fusion Region
- The S2 region forms the central membrane-associated portion of the spike. It contains the fusion peptide, heptad-repeat regions and the transmembrane segment. Following activation of the spike, S2 undergoes structural rearrangement involved in fusion of the viral and cellular membranes.
Membrane (M) Protein
- Membrane (M) protein- It is the most abundant structural protein present in the SARS-CoV-2 envelope. M has three transmembrane helices within the membrane and a comparatively larger C-terminal domain facing towards the inside of the virion.
- M proteins occur mainly as dimers. The dimer has a compact “mushroom-shaped” structure, having transmembrane bundles and an intravirion domain. These dimers can also associate into larger assemblies.
- During formation of the virus particle, M interacts with other structural components. Its inward-facing part can interact with N protein and RNA, bringing the nucleocapsid-RNA material into association with the surrounding viral membrane.
Envelope (E) Protein
- Envelope (E) protein is the smallest among the four major structural proteins. Only a small amount of this protein becomes incorporated into the mature virion.
- It is a membrane-associated protein having a single transmembrane region. During virus production, most E protein remains associated with intracellular membranes. Its transmembrane region can form oligomeric structures in lipid bilayers.
- Along with M and other structural proteins, E takes part in particle assembly and budding. Experiments using SARS-CoV-2 structural proteins have also shown its contribution in efficient formation of virus-like particles.
Nucleocapsid (N) Protein and RNA Packaging
- Nucleocapsid (N) protein- Unlike S, M and E, the N protein is located inside the viral envelope. It binds with the SARS-CoV-2 genomic RNA and forms viral ribonucleoprotein (vRNP) complexes.
- The nearly 30 kb RNA genome is packed by association of N proteins with RNA and with one another. Compact RNP assemblies are formed in this process, fitting the long viral genome within the virion.
- Multiple vRNP units are present inside the viral particle. Cryo-electron tomography has shown these as separate cylindrical assemblies, including many positioned close to the inner side of the envelope.
- Between the packaged genome and viral membrane, a structural association is also present through the M protein. The intravirion region of M can interact with the N-containing RNP material.

SARS-CoV-2 Genome Type and Size
Positive-Sense Single-Stranded RNA Genome
- SARS-CoV-2 contains RNA as its genetic material, not DNA. The genome is a positive-sense single-stranded RNA (+ssRNA). It is present as a single continuous, non-segmented RNA molecule.
- SARS-CoV-2 is the name given to the virus, while COVID-19 is the disease caused by SARS-CoV-2. DNA or RNA is considered for SARS-CoV-2 virus and not for the disease COVID-19.
- Positive-sense means that the genomic RNA of the virus has the same polarity as messenger RNA (mRNA). After entering the host cell, the RNA can be directly used by cellular ribosomes for synthesis of viral proteins. Conversion of the RNA into DNA is not required before translation.
- The SARS-CoV-2 genome is non-segmented. Its genetic material is not divided into different RNA segments. A single long RNA molecule carries the viral genetic information.
Genome Length: About 29.9 kb
- The SARS-CoV-2 genome is approximately 29.9 kilobases (kb) in length. It is one of the large RNA genomes found among RNA viruses.
- The commonly used Wuhan-Hu-1 reference genome contains 29,903 nucleotides (nt). This number represents the length of this particular reference sequence. Every SARS-CoV-2 genome record does not necessarily contain exactly 29,903 nucleotides.
- Slight variation in the reported genome length may be present among individual sequence records. During sequencing, terminal nucleotides may not always be completely recovered. Some terminal primer-derived sequences can be removed during processing, while the poly(A) tail is also not always included in the reported sequence length.
5′ Cap, Untranslated Regions, and 3′ Poly(A) Tail
- At the 5′ end, the SARS-CoV-2 genomic RNA contains a methylated cap followed by the 5′ untranslated region (5′ UTR). The cap provides an mRNA-like feature to the viral RNA. It supports RNA stability and its translation by the host-cell machinery.
- Untranslated regions occur at both ends of the genome. These are referred to as the 5′ UTR and 3′ UTR. They do not code for viral proteins. RNA structures and regulatory sequences present in these regions are involved in viral RNA synthesis, translation and maintenance of the RNA genome.
- The 3′ end contains a poly(A) tail after the 3′ UTR. It consists of a stretch of adenine nucleotides. The length of this tail can vary between different viral RNA molecules.
- SARS-CoV-2 genomic RNA has mRNA-like features, including the 5′ cap and 3′ poly(A) tail. These structures are involved in protection of the RNA from degradation and allow efficient use of the genome during viral protein synthesis.




Genomic Organization of SARS-CoV-2
- The SARS-CoV-2 genome is organized from 5′ to 3′ as 5′ UTR-ORF1a/ORF1b-S-ORF3a-E-M-ORF6-ORF7a-ORF7b-ORF8-N-3′ UTR. Some additional overlapping open reading frames are also present. A major portion of the genome is occupied by ORF1a and ORF1b, while structural and accessory genes are mostly located towards the 3′ end.
- ORF1a and ORF1b together form the replicase region and cover nearly two-thirds of the viral genome. ORF1a is translated first, producing the polyprotein pp1a. At the ORF1a-ORF1b junction, some translating ribosomes undergo a programmed “−1 ribosomal frameshift”. These ribosomes continue translation into ORF1b and produce the longer pp1ab polyprotein.
- pp1a and pp1ab are large precursor proteins. They are cleaved by viral proteases into different non-structural proteins (nsps). pp1a produces proteins up to nsp11, whereas pp1ab extends the proteins up to nsp16. These nsps include viral proteases, RNA-dependent RNA polymerase (RdRp), helicase and several other proteins required during viral RNA synthesis.
- After the replicase region, genes encoding the major structural proteins are present. They occur mainly in the order spike (S), envelope (E), membrane (M), and nucleocapsid (N). Accessory ORFs are located between these structural genes or overlap with their regions.
- ORF3a is present after the S gene. ORF6, ORF7a, ORF7b and ORF8 occur farther along the genome around the E, M and N gene regions. These accessory genes occupy comparatively smaller portions of the genome than ORF1a/ORF1b and the major structural genes.
- Not all coding sequences occupy completely separate regions. Some genes overlap with another gene by using a different reading frame. ORF9b, for example, is encoded within the N-gene region using an alternative reading frame.
- Annotation of some very small ORFs has varied between different studies. ORF10 was included in the original reference annotation. Comparative genomic studies, however, do not support ORF10 as a conserved protein-coding gene. Other small overlapping ORFs have also been described.
- Structural and accessory proteins present towards the 3′ region are mainly produced from a group of subgenomic RNAs (sgRNAs). They are not produced by translation of the complete genome from one end to another. These sgRNAs contain a common 5′ leader sequence joined with different downstream regions of the genome.
- The genomic RNA itself is directly used for translation of ORF1a and ORF1b. The downstream subgenomic RNAs provide separate RNA templates, from which many structural and accessory proteins are translated.

SARS-CoV-2 Genome Expression Strategy
The SARS-CoV-2 genome is expressed through direct translation, polyprotein processing, genome replication, and formation of subgenomic RNAs (sgRNAs). Being a positive-sense RNA genome, the genomic RNA itself can first work as messenger RNA. The following are the major steps of SARS-CoV-2 genome expression.
Step 1: Viral genomic RNA is directly used for translation
Once the positive-sense genomic RNA becomes available in the cytoplasm, its 5′ end is recognized by the cellular translation machinery. Ribosomes begin translation from ORF1a, located in the 5′ region of the genome. No DNA intermediate is formed in this process.
Translation of ORF1a normally produces the large polyprotein pp1a.
Step 2: Some ribosomes shift into ORF1b and produce pp1ab
When the translating ribosome reaches the junction between ORF1a and ORF1b, two ways are possible. Most translation terminates after ORF1a. In some translating ribosomes, a programmed −1 ribosomal frameshift (−1 PRF) occurs at a specific slippery sequence and the ribosome moves into another reading frame.
The ribosome then continues through ORF1b. A much longer polyprotein, pp1ab, is produced instead of pp1a. A stimulatory RNA pseudoknot present near the frameshift site helps in this shifting of the ribosome.
Step 3: pp1a and pp1ab are cut into non-structural proteins
The two polyproteins do not remain as a single protein. They are cleaved at different positions by the viral proteases present within the polyproteins themselves.
Two proteases mainly carry out this processing, the papain-like protease (PLpro) of nsp3 and main protease (Mpro or 3CLpro) of nsp5. From pp1a and pp1ab, the set of non-structural proteins (nsps) needed for viral RNA synthesis is produced. pp1a provides nsps up to nsp11, while expression through pp1ab provides the ORF1b-encoded proteins nsp12 to nsp16, including the RNA-dependent RNA polymerase and helicase.
Step 4: Non-structural proteins form the replication-transcription complex
After proteolytic processing, several nsps come together with viral RNA and modified intracellular membranes. The replication-transcription complex (RTC) is formed. Coronavirus replication is closely associated with membrane structures, particularly double-membrane vesicles derived from cellular membranes.
Among the proteins of this complex, nsp12 is the RNA-dependent RNA polymerase (RdRp). nsp7 and nsp8 associate with the polymerase, while other nsps perform additional functions required during RNA synthesis and processing.
Step 5: A full-length negative-sense RNA copy is synthesized
The positive-sense genomic RNA now also acts as a template for RNA synthesis. Using it, the RTC produces a full-length negative-sense RNA (−gRNA).
This negative RNA is complementary to the viral genome. It is not the principal RNA translated into viral proteins. Instead, it serves as template from which new full-length positive-sense genomic RNAs can be synthesized.
Step 6: New positive-sense genomic RNAs are produced
The full-length negative-sense RNA is copied by the viral RNA-synthesizing machinery and new positive-sense genomic RNA (+gRNA) molecules are formed.
These newly formed genomic RNAs can take part again in viral RNA expression and replication. Some of them later become the genomes incorporated into newly formed virus particles.
Step 7: Subgenomic negative-sense RNAs are formed by discontinuous transcription
Expression of the genes present towards the 3′ region follows a different process. During synthesis of negative-strand RNA, the RTC encounters body transcription-regulatory sequences (TRS-B) located before several downstream viral genes.
At these regions, RNA synthesis may stop and the newly forming negative RNA is transferred towards the 5′ end of the genomic template. It pairs with the leader transcription-regulatory sequence (TRS-L). This template switching joins the leader-related sequence with different downstream regions and forms a set of negative-sense subgenomic RNAs. This process is referred to as discontinuous transcription.
Step 8: Negative-sense sgRNAs are copied into positive-sense subgenomic mRNAs
The negative-sense subgenomic RNAs then work as templates. From them, a nested group of positive-sense subgenomic mRNAs (sg-mRNAs) is synthesized.
These sgRNAs have a common 5′ leader sequence derived from the 5′ end of the genome, but contain different portions of the downstream 3′ genomic region. Although several ORFs can be present on a single subgenomic RNA, translation mainly occurs from the 5′-most accessible ORF of the particular sgRNA.
Step 9: Structural and accessory proteins are translated from the sgRNAs
The positive-sense sgRNAs are used by host ribosomes for synthesis of the downstream viral proteins. The spike (S), envelope (E), membrane (M), and nucleocapsid (N) proteins are expressed mainly through these subgenomic RNAs. Several accessory proteins are also produced from the sgRNA set.
In this way ORF1a/ORF1b is translated directly from the genomic RNA, whereas much of the 3′ structural and accessory region is expressed using the separately produced subgenomic mRNAs

References
- Arya, R., Kumari, S., Pandey, B., Mistry, H., Bihani, S. C., Das, A., Prashar, V., Gupta, G. D., Panicker, L., & Kumar, M. (2021). Structural insights into SARS-CoV-2 proteins. Journal of Molecular Biology, 433(2), 166725. https://doi.org/10.1016/j.jmb.2020.11.024
- Banerjee, A., Nasir, J. A., Budylowski, P., Yip, L., Aftanas, P., Christie, N., Ghalami, A., Baid, K., Raphenya, A. R., Hirota, J. A., Miller, M. S., McGeer, A. J., Ostrowski, M., Kozak, R. A., McArthur, A. G., Mossman, K., & Mubareka, S. (2020). Isolation, sequence, infectivity, and replication kinetics of severe acute respiratory syndrome coronavirus 2. Emerging Infectious Diseases, 26(9), 2054–2063. https://doi.org/10.3201/eid2609.201495
- Benton, D. J., Wrobel, A. G., Xu, P., Roustan, C., Martin, S. R., Rosenthal, P. B., Skehel, J. J., & Gamblin, S. J. (2020). Receptor binding and priming of the spike protein of SARS-CoV-2 for membrane fusion. Nature, 588(7837), 327–330. https://doi.org/10.1038/s41586-020-2772-0
- Brant, A. C., Tian, W., Majerciak, V., Yang, W., & Zheng, Z.-M. (2021). SARS-CoV-2: From its discovery to genome structure, transcription, and replication. Cell & Bioscience, 11, 136. https://doi.org/10.1186/s13578-021-00643-z
- Dai, L., & Gao, G. F. (2021). Viral targets for vaccines against COVID-19. Nature Reviews Immunology, 21, 73–82. https://doi.org/10.1038/s41577-020-00480-0
- Duart, G., García-Murria, M. J., & Mingarro, I. (2021). The SARS-CoV-2 envelope (E) protein has evolved towards membrane topology robustness. Biochimica et Biophysica Acta (BBA) – Biomembranes, 1863(7), 183608. https://doi.org/10.1016/j.bbamem.2021.183608
- Ismail, A. M., & Elfiky, A. A. (2020). SARS-CoV-2 spike behavior in situ: A Cryo-EM images for a better understanding of the COVID-19 pandemic. Signal Transduction and Targeted Therapy, 5, 252. https://doi.org/10.1038/s41392-020-00365-7
- Jackson, C. B., Farzan, M., Chen, B., & Choe, H. (2022). Mechanisms of SARS-CoV-2 entry into cells. Nature Reviews Molecular Cell Biology, 23(1), 3–20. https://doi.org/10.1038/s41580-021-00418-x
- Ke, Z., Oton, J., Qu, K., Cortese, M., Zila, V., McKeane, L., Nakane, T., Zivanov, J., Neufeldt, C. J., Cerikan, B., Lu, J. M., Peukes, J., Xiong, X., Kräusslich, H.-G., Scheres, S. H. W., Bartenschlager, R., & Briggs, J. A. G. (2020). Structures and distributions of SARS-CoV-2 spike proteins on intact virions. Nature, 588(7838), 498–502. https://doi.org/10.1038/s41586-020-2665-2
- Kim, D., Lee, J.-Y., Yang, J.-S., Kim, J. W., Kim, V. N., & Chang, H. (2020). The architecture of SARS-CoV-2 transcriptome. Cell, 181(4), 914–921.e10. https://doi.org/10.1016/j.cell.2020.04.011
- Klein, S., Cortese, M., Winter, S. L., Wachsmuth-Melm, M., Neufeldt, C. J., Cerikan, B., Stanifer, M. L., Boulant, S., Bartenschlager, R., & Chlanda, P. (2020). SARS-CoV-2 structure and replication characterized by in situ cryo-electron tomography. Nature Communications, 11, 5885. https://doi.org/10.1038/s41467-020-19619-7
- Kung, Y.-A., Lee, K.-M., Chiang, H.-J., Huang, S.-Y., Wu, C.-J., & Shih, S.-R. (2022). Molecular virology of SARS-CoV-2 and related coronaviruses. Microbiology and Molecular Biology Reviews, 86(2), e00026-21. https://doi.org/10.1128/mmbr.00026-21
- Lan, J., Ge, J., Yu, J., Shan, S., Zhou, H., Fan, S., Zhang, Q., Shi, X., Wang, Q., Zhang, L., & Wang, X. (2020). Structure of the SARS-CoV-2 spike receptor-binding domain bound to the ACE2 receptor. Nature, 581(7807), 215–220. https://doi.org/10.1038/s41586-020-2180-5
- Long, S. (2021). SARS-CoV-2 subgenomic RNAs: Characterization, utility, and perspectives. Viruses, 13(10), 1923. https://doi.org/10.3390/v13101923
- Malone, B., Urakova, N., Snijder, E. J., & Campbell, E. A. (2022). Structures and functions of coronavirus replication-transcription complexes and their relevance for SARS-CoV-2 drug design. Nature Reviews Molecular Cell Biology, 23(1), 21–39. https://doi.org/10.1038/s41580-021-00432-z
- Mandala, V. S., McKay, M. J., Shcherbakov, A. A., Dregni, A. J., Kolocouris, A., & Hong, M. (2020). Structure and drug binding of the SARS-CoV-2 envelope protein transmembrane domain in lipid bilayers. Nature Structural & Molecular Biology, 27(12), 1202–1208. https://doi.org/10.1038/s41594-020-00536-8
- Manfredonia, I., & Incarnato, D. (2021). Structure and regulation of coronavirus genomes: State-of-the-art and novel insights from SARS-CoV-2 studies. Biochemical Society Transactions, 49(1), 341–352. https://doi.org/10.1042/BST20200670
- Mariano, G., Farthing, R. J., Lale-Farjat, S. L. M., & Bergeron, J. R. C. (2020). Structural characterization of SARS-CoV-2: Where we are, and where we need to be. Frontiers in Molecular Biosciences, 7, 605236. https://doi.org/10.3389/fmolb.2020.605236
- Mathew, B. J., Gupta, S., Nema, R. K., Vyas, A. K., Khare, P., Biswas, D., & Singh, A. K. (2022). Genomic, proteomic and metabolomic profiling of severe acute respiratory syndrome-Coronavirus-2. In A. Parihar, R. Khan, A. Kumar, A. K. Kaushik, & H. A. Gohel (Eds.), Computational approaches for novel therapeutic and diagnostic designing to mitigate SARS-CoV2 infection: Revolutionary strategies to combat pandemics (pp. 49–76). Academic Press. https://doi.org/10.1016/B978-0-323-91172-6.00019-4
- Nakagawa, S., & Miyazawa, T. (2020). Genome evolution of SARS-CoV-2 and its virological characteristics. Inflammation and Regeneration, 40, 17. https://doi.org/10.1186/s41232-020-00126-7
- Nikonova, A. A., Faizuloev, E. B., Gracheva, A. V., Isakov, I. Y., & Zverev, V. V. (2021). Genetic diversity and evolution of the biological features of the pandemic SARS-CoV-2. Acta Naturae, 13(3), 77–88. https://doi.org/10.32607/actanaturae.11337
- V’kovski, P., Kratzel, A., Steiner, S., Stalder, H., & Thiel, V. (2021). Coronavirus biology and replication: Implications for SARS-CoV-2. Nature Reviews Microbiology, 19, 155–170. https://doi.org/10.1038/s41579-020-00468-6
- Wu, F., Zhao, S., Yu, B., Chen, Y.-M., Wang, W., Song, Z.-G., Hu, Y., Tao, Z.-W., Tian, J.-H., Pei, Y.-Y., Yuan, M.-L., Zhang, Y.-L., Dai, F.-H., Liu, Y., Wang, Q.-M., Zheng, J.-J., Xu, L., Holmes, E. C., & Zhang, Y.-Z. (2020). A new coronavirus associated with human respiratory disease in China. Nature, 579(7798), 265–269. https://doi.org/10.1038/s41586-020-2008-3
- Yao, H., Song, Y., Chen, Y., Wu, N., Xu, J., Sun, C., Zhang, J., Weng, T., Zhang, Z., Wu, Z., Cheng, L., Shi, D., Lu, X., Lei, J., Crispin, M., Shi, Y., Li, L., & Li, S. (2020). Molecular architecture of the SARS-CoV-2 virus. Cell, 183(3), 730–738.e13. https://doi.org/10.1016/j.cell.2020.09.018
- Zhang, Z., Nomura, N., Muramoto, Y., Ekimoto, T., Uemura, T., Liu, K., Yui, M., Kono, N., Aoki, J., Ikeguchi, M., Noda, T., Iwata, S., Ohto, U., & Shimizu, T. (2022). Structure of SARS-CoV-2 membrane protein essential for virus assembly. Nature Communications, 13, 4399. https://doi.org/10.1038/s41467-022-32019-3