Data supporting the molecular reconstruction of the Anisakis simplex Ani s 14-containing locus and protein
Description
This dataset contains the sequence, genomic-coordinate, RNA-seq, proteomic and comparative-genomic resources supporting the reconstruction of the Anisakis simplex Ani s 14-containing locus. The reconstructed model comprises 76 exons and 75 introns and encodes a candidate 3,841-aa ShKT-containing cysteine-rich protein whose C-terminal 217-aa region corresponds to canonical Ani s 14. The dataset includes reconstructed transcript, CDS and protein sequences, exon–intron coordinate maps, splice-junction support, proteomic peptide evidence, comparative exon–intron conservation, and evidence supporting the experimentally determined 3′ end. A historical 3′-RACE product defines a 648-nt post-stop region terminating in a non-templated poly(A) tail, with independent support from LC027371.1 and RNA-seq data. The candidate SL1-containing transcript extends to 12,212 nt up to the experimentally supported cleavage/polyadenylation site, excluding the poly(A) tail. The proposed SL1-containing 5′ end remains qualified because no dedicated full-length 5′-end assay was performed.
Files
Steps to reproduce
The dataset contains the sequence, structural, transcriptomic, proteomic and comparative-genomic evidence used to reconstruct and characterize the Ani s 14-containing protein of Anisakis simplex. The reconstructed transcript and protein models were generated by integrating genomic sequence information with RNA-seq evidence. Exon representation and splice-junction support were assessed by mapping available RNA-seq libraries to the reconstructed 76-exon model. Split-read alignments were used to evaluate the 75 predicted introns and to examine local support for the E43/J43 reconstruction. Reads compatible with an SL1 trans-splicing event were additionally analysed to characterize the candidate 5′ end of the mature transcript. Protein-level analyses included identification of the 11 internal repeat modules (M1–M11), localization of the canonical Ani s 14 region within the C-terminal portion of M11, and comparison with previously annotated A. simplex proteins. Proteomic evidence was incorporated by mapping experimentally identified peptides to the reconstructed protein sequence. Comparative analyses with homologous genomic/protein sequences from related ascaridoid nematodes were used to assess conservation of protein sequence and exon–intron architecture. The deposited files include the reconstructed sequences, exon–intron coordinates, splice-junction evidence, proteomic peptide mappings, comparative analyses and the intermediate/tabulated datasets used to generate the corresponding manuscript figures and supplementary tables. File names and worksheet headings identify the analysis represented in each file. Software versions, parameters and analysis-specific procedures are provided in the associated manuscript and/or accompanying README file.
Institutions
- Universidade de Santiago de CompostelaGalicia, Santiago de Compostela