Data for Discovery of a Geographically Distinct blaTEM Lineage in the Urban Public Transit Network of Quito, Ecuador

Published: 27 July 2026| Version 2 | DOI: 10.17632/6vvh58g85b.2
Contributor:

Description

This dataset contains the supplementary analytical files, reference sequences, and custom scripts supporting the genomic characterization and phylogeographic analysis of blaTEM extended-spectrum beta-lactamase (ESBL) lineages identified within the "Trolebús" mass transit network in Quito, Ecuador. The study utilized Oxford Nanopore Technologies (ONT) long-read environmental sequencing combined with Bayesian phylodynamic inference to track plasmid-mediated antimicrobial resistance genes (ARGs) across high-touch urban surfaces. This repository houses the downstream alignment, profiling, and geographical mapping files used to generate the evolutionary coalescent dynamics and consensus networks discussed in the manuscript. Note: The raw environmental consensus sequences generated via ONT are not included in this repository and have been deposited separately in the NCBI BioProject database.

Files

Steps to reproduce

1) Antimicrobial Resistance (AMR) Profiling: Analyze the environmental blaTEM consensus sequences using the Resistance Gene Identifier (RGI) integrated with the Comprehensive Antibiotic Resistance Database (CARD). Apply the "loose" algorithm parameter to maximize the detection of novel or variant alleles, which generates the profiling output rgi_output.csv 2) Reference Compilation & Alignment: Retrieve the curated TEM beta-lactamase sequence sub-variants ([ARO_3000014]-...-nucleotide.fas) from the CARD ontology at https://card.mcmaster.ca/ontology/36023. Align and concatenate these reference models with the newly generated Ecuadorian transit network sequences to create the master dataset (CARD_trole.fas). 3) Bayesian Phylodynamic Inference: Import the concatenated master dataset into BEAST (Bayesian Evolutionary Analysis Sampling Trees). Configure the mathematical framework using the K2+G substitution model, setting the chain length to 20 million generations alongside a 200,000-generation burn-in period. 4) Geographic Metadata Extraction: Execute the custom Python script (geo_location.py) to query the NCBI database for country-level attributes assigned to the global reference sequences. For sequences with missing or ambiguous metadata, manually verify the NCBI accession attributes—specifically targeting the "journal," "comment," and "title" fields—to compile the final geographic mapping document (geolocated_samples.txt). 5) Phylogeographic Network Construction: Export the posterior distribution of trees generated by the BEAST analysis into the SplitsTree platform. Construct a Bayesian consensus network to visualize evolutionary lineages, applying the geographic metadata to confirm the spatial clustering of the Ecuadorian isolates against the global sequences.

Institutions

Categories

Microbiology, Environmental Science, Antimicrobial Resistance, Bacterial Immunology

Licence