Comprehensive Glycosylation Profiling of 693,699 SARS-CoV-2 Spike Proteins
Description
This dataset presents large-scale glycosylation mapping across 693,699 SARS-CoV-2 spike sequences, revealing near-universal presence of both N- and O-linked glycans. A total of over 10 million N-glycosylation and 17 million O-glycosylation sites were identified, spanning 205 and 472 unique positions, respectively. The total sequences subjected to analysis were 9.3 Million approximately. All unique sequences were further assessed for QC. The findings highlight strong evolutionary conservation of the glycan shield despite RBD structural drift, emphasizing its key role in spike stability, immune evasion, and host adaptation. Data processed locally via jupyter notebook & curator: Tahir HB
Files
Steps to reproduce
1- Data collection from NCBI / GISAID/ Nextstrain 2- Alignment / Trimming / QC 3- Manually written python scripts over Jupiter notebook for identification of glycan sites and summary generation.