Prediction of potential NS3-based T cell epitopes of Hepatitis C virus by in-silico and in-vitro approach
Description
The data contains the docking and refinement analysis part of the study. This study aimed to identify conserved NS3-based T-cell epitopes of HCV genotype 3 using combined in-silico and in-vitro approaches. NS3 sequences from 36 HCV 3a-infected patients were analysed to generate a consensus sequence, which was subjected to the Expasy-translate tool and IEDB server, respectively, for epitope prediction. Epitope prediction via the IEDB server (percentile rank < 0.5) yielded 119 MHC-I and 436 MHC-II candidates, which were filtered for antigenicity (VaxiJen > 0.4), non-allergenicity, non-toxicity, low human homology (E-value> 2), and cytokine-inducing potential (IFN-γ, IL-4, IL-10). Conservancy analysis across 485 global NS3 sequences (HCV-3a) shortlisted 11 MHC-I and 6 MHC-II epitopes. These epitopes were subjected to galaxypepdock server (https://galaxy.seoklab.org/cgi-bin/submit.cgi?type=PEPDOCK) for docking and the galaxyrefine server ( https://galaxy.seoklab.org/cgi-bin/submit.cgi?type=REFINE). This Dataset contains the docking and refinement analysis yielded models and results. The data contains different HLA allele+epitope combinations, denoted by folder names. Inside the folders, possible models can be found for each combination along with an Excel file which contains the overview and technical details (results) of that specific combination.
Files
Steps to reproduce
Each MHC-I epitope was subjected to docking with 23 different types of MHC-I alleles (prevalent in the Asian population), and each MHC-II epitope was subjected to docking with 12 different sub-chains of MHC-II alleles (prevalent in the Asian population). The allelic prevalence information was collected from the Allele Frequency Net Database [ http://www.allelefrequencies.net/]. The structures of 23 different MHC-I and 12 different sub-chains of MHC-II alleles were downloaded from RCSB PDB. These structures were then prepared for docking by removing ligands, water molecules, other molecules and duplicated residues or alleles, with the help of Discovery studio software [ https://discover.3ds.com/discovery-studio-visualizer-download ] and missing loops were modelled with Modellar 10.3 software [ https://salilab.org/modeller/ ]. Cleaned (.pdb) structures and epitope sequences in (.fasta) format were then analysed for protein-peptide docking based on interaction similarity. For MHC-II molecules, each sub-chain was subjected to docking with MHC-II epitopes separately. The analyses were done by GalaxyPepDock Server [ https://galaxy.seoklab.org/cgi-bin/help.cgi?key=METHOD&type=PEPDOCK ]. Results were visualised in Discovery Studio Software, and the models which have the highest template modelling score (TM-score) were refined with GalaxyRefineComplex server [ https://galaxy.seoklab.org/cgi-bin/submit.cgi?type=COMPLEX ] for further refinement and RMSD calculation. The chosen 23 different MHC-I alleles and their respective PDB IDS were HLA_A_01_01 (6AT9), HLA_A_02_01(4U6Y), HLA_A_02_06(3OXR), HLA_A_02_07(3OXS), HLA_A_03_01(3RL1), HLA_A_11_01(7S8S), HLA_A_24_02(7JYV), HLA_A_68_01(6PBH), HLA_B_07_02(6AT5), HLA_B_14_02(3BXL), HLA_B_15_01(5TXS), HLA_B_15_02(6VB2), HLA_B_18_01(4XXC), HLA_B_35_01(1A1N), HLA_B_35_08(3BWA), HLA_B_40_01(6IEX), HLA_B_40_02(5IEH), HLA_B_44_03(3DX7), HLA_B_44_05(6MTL), HLA_B_51_01(1E27), HLA_B_52_02(3W39), HLA_B_57_01(6BXP), and HLA_B_58_01(5VWH). The chosen 12 different sub chains of MHC_II alleles and their respective PDB IDS were HLA_DQ1(1S9V), HLA_DQB(1S9V), HLA_DMA(2BC4), HLA_DMB(2BC4), HLA_DQA1(5KSV), HLA_DQB1(5KSV), HLA_DRA1(3C5J), HLA_DRB3(3C5J), HLA_DPB1(3WEX), HLA_DPA1(3WEX), HLA_DRB5(1H15), and HLA_DRB1(SV4M).
Institutions
- National Institute of Cholera and Enteric Diseases