A comprehensive database of amniote tropoelastin sequences
Description
The database contains genomic sequence and exon identification information for tropoelastins from more than 80 species of amniotes, with approximately equal numbers of synapsid and sauropsid species and broad representation from all sub-groups of amniote species. The database was created to provide reliable, curated, exon-by-exon protein sequences for tropoelastins but may also be useful for a variety of other analyses, including codon usage, regulation of alternate splicing, and intron- and UTR-based sequences affecting expression. This database has been used to identify characteristics of tropoelastins conserved through >300 million years of evolution, including preservation of both regions of positional sequence and collective or compositional characteristics derived from but not strictly dependent on positional sequence.
Files
Steps to reproduce
Genomic sequences, RefSeqs, ESTs, transcriptome sequences, etc. were harvested mostly from the US National Center for Biotechnology Information database (https://www.ncbi.nlm.nih.gov), with a smaller number sourced from the EnsemblGenomes database (http://ensemblgenomes.org). Accession information for all sequences is provided in the data tables. Exon boundaries were predicted using the NNSplice v. 9.0 tool (http://www.fruitfly.org/seq_tools/splice.html). Signal peptide cleavage sites were predicted using SignalP-6.0 (https://biolib.com/DTU/SignalP-6/), using the protein sequence of the first 4-5 domains of tropoelastins, and choosing the Eukarya option. Curations notes are provided with the sequences.
Institutions
- Hospital for Sick Children
- University of Toronto