Fasciola hepatica genome re-annotation and validation of selected gene models
Description
Fasciolosis is caused by liver flukes: F. hepatica, and a sister species – F. gigantica. A growing concern with controlling the disease is resistance to triclabendazole (TCBZ), the only drug shown to kill both adult and immature liver flukes. Currently, F. hepatica mechanism of resistance to TCBZ is not clearly understood and there is no effective commercially available vaccine. Previous work proposed three mechanisms associated with TCBZ mode of action and resistance: tubulin binding activity, drug uptake mechanisms, and drug metabolism mechanism. Exploring evolutionary forces acting on F. hepatica genes associated with TCBZ mode of action and resistance could explain how the parasite develops resistance to the drug, enable identification of potential drug targets, and facilitate development of new drugs. A re-annotation of the current F. hepatica genome was done using an updated version of the published F. hepatica draft genome (assembly GCA_000947175.1, BioProject PRJEB6687). Subsequently, the current annotation (Fasciola_10x_pilon, GCA_900302435.1 WormBase Parasite Version 15) was compared and critically assessed with the newly reannotated version. Using coding sequences (CDS) of three well-described annotated gene families, manual validation of the annotation was done. A total of 15,879 F. hepatica genes were identified in this project compared to the 9,401 genes in the current annotation, while differences noticed in both annotations include gene fragmentation, missing exons, and missing genes.
Files
Steps to reproduce
F1 – Braker ab initio annotation output files, including the 74,307 predicted gene models from the RNAseq dataset 1. F2 – Braker ab initio annotation output files including the 53,729 predicted gene models from the RNAseq dataset 2 F3 – This folder includes output files of the Maker annotation (after the 4th SNAP training of gene models). A total of 15,879 genes were predicted. F4 – This folder includes output files of Transdecoder annotation to predict genes. A total of 9,401 genes were predicted. F5 – This folder includes a “draft” annotation of the F. gigantica genome from the NorthEastern Hill University, India (ASM286751v3, BioProject: PRJNA339660, BioSample: SAMN05601579). Using Augustus, a total of 11,947 F. gigantica genes were predicted. F6 – Files include the output of the RepeatMasking process of the analysis. Files describe repetitive elements identified and their genome coordinates. F7 – Orthomcl output file of the orthologous grouping of the F. hepatica annotation described earlier (See F3), F. gigantica “draft annotation (See F4), Paragonimus westermani (ASM850834v1), Echinostoma caproni (E_caproni_Egypt_0011_upd), Schistosoma mansoni (Smansoni_v7), Opisthorchis viverrini (OpiViv1.0), and Clonorchis sinensis (C_sinensis-2.0). F8 – Files describe the functional annotation of the predicted gene models. These output files were generated using PANNZER and include each predicted gene’s functional description (DE) and Gene Ontology (GO). F9 – GhostKOALA annotation output of the predicted gene models (See F3) against nonredundant set of KEGG GENES to assign K numbers to the F. hepatica query gene models. F10 – Screenshots from IGV showing the various genes aligned with the F. hepatica genome.
Institutions
- University of Liverpool