Twelve-Year Epidemiological Trends, Toxin Characterization, and Bacterial Vaginosis Associations of Enteric Pathogens Detected by Multiplex PCR
Description
This collection contains five de-identified Excel workbooks (Datasets 1–5) plus two CSV key‐files, documenting molecular diagnostic results, demographic metadata, and derived analytics for gastrointestinal (GIT) pathogen testing at Medical Diagnostic Laboratories LLC. The date range spans January 2014 to March 2025. All direct patient and sample identifiers (MDLNo and Patient-ID) have been replaced with anonymized codes. Researchers may reproduce published analyses (χ² cross-tabulations, Poisson harmonic regression, STL decomposition, co-occurrence networks, and machine‐learning models) from Osei Sekyere et al. (2025, BioRxiv). This enables studies of prevalence, seasonality, and predictive modeling across an 121+ year period. 2. File Contents Dataset 1. GIT-No-toxins_deidentified.xlsx Sheet: Dataset-No-toxins Description: All multiplex GIT panel PCR tests excluding toxin genes. Each row is one test on a unique or repeat sample. The raw date‐collected is omitted (monthly aggregation available in Dataset 5). Dataset 2. CD-Toxins_deidentified.xlsx Sheet: C.difficile+toxins Description: Toxin subtyping data for C. difficile. One molecular test per row; includes both species presence and toxin gene presence. Dataset 3. Ecoli-Shigella-toxins_deidentified.xlsx Sheet: E.coli+Shigella+toxins Description: Subtyping assays for E. coli (O157/Shiga‐toxin) and Shigella spp. Each row is one assay result. Dataset 4. Statistics & ML_deidentified.xlsx Sheet(s): (preserve original sheet names, e.g., ML-Results, Feature-Importances, etc.) Description: Consolidated analytic outputs for machine learning models (Logistic Regression, Random Forest, XGBoost). Contains per-sample predictions (anonymized), model performance metrics, and feature‐importance data. Use to reproduce Figures D1–D5 and associated Supplementary Figures. Dataset 5. Seasonality-Temporal dynamics_deidentified.xlsx Sheet: Monthly_Positive_Counts Description: Derived monthly aggregation and summary statistics for Poisson harmonic regression and STL decomposition. Reproduce all seasonal plots and numeric tables. 3. Data Provenance & Methods Laboratory: Medical Diagnostic Laboratories LLC, Hamilton Township, NJ, USA. Time Period: January 2014–March 2025. Assays: Multiplex GIT panel (bacterial, viral, protozoal pathogens; no toxin genes). C. difficile toxin PCR for toxin A/B genes. E. coli/Shigella subtyping (rfbA, stx1, stx2 genes Laboratory: Medical Diagnostic Laboratories LLC, Hamilton Township, NJ, USA. Time Period: January 2014–March 2025. License: Data are CC0 (public domain), permitting unrestricted reuse for research and publication. Contact: Dr. John Osei, jod14139@yahoo.com, for questions about data. Keywords: Gastrointestinal pathogens; multiplex PCR; Clostridioides difficile; Escherichia coli; Shigella; seasonality; Poisson regression; STL decomposition; machine learning; de-identified clinical data
Files
Steps to reproduce
Study Design and Data Acquisition This retrospective cross-sectional study analyzed gastrointestinal (GIT) multiplex PCR results from Medical Diagnostic Laboratories (MDL) between January 2014 and March 2025. Three primary datasets were extracted from the laboratory information system: (1) multiplex GIT panel results for 14 bacterial, viral, and protozoal targets, (2) Clostridioides difficile toxin assays (tcdA/tcdB), and (3) Escherichia coli/Shigella subtyping (rfbA, stx1, stx2). Each row in the raw data corresponds to one PCR test on a unique specimen (MDLNo). Patient demographics (age, gender, ethnicity), specimen type, and specimen source were also captured. Unique and repeat tests were identified by MDLNo to distinguish within-sample repeats from distinct specimens. Data Cleaning and Transformation All Excel files were read into Python using pandas v2.2.2. Blank or NA entries (e.g., “\N”, empty strings) were standardized. Column names were harmonized across datasets. Duplicate or invalid records were removed. Patient age was converted to years (months for infants normalized to fractions of a year) and binned into 10-year categories (0–9, 10–19, …, ≥100). Ethnicity entries marked as “Unknown” were retained for completeness but interpreted cautiously due to high missingness. Specimen and source categories were collapsed into consistent, non-overlapping labels (e.g., stool, swab, blood, biopsy). A binary presence/absence matrix (0/1) was built to indicate detection of each pathogen per sample (MDLNo). Descriptive Statistics and Cross-Tabulations Basic summary statistics (counts, percentages) were computed for patient demographics and specimen metadata. Associations between categorical variables (e.g., pathogen vs. specimen type, age group, gender, ethnicity) were tested using Pearson’s χ² or Fisher’s exact test when expected counts < 5. Bonferroni correction controlled family-wise error across multiple comparisons (α = 0.05/number of tests). Age distributions across groups were compared via one-way ANOVA. Temporal and Seasonal Trend Analysis Monthly positive detection counts were calculated by aggregating all “P” results for each pathogen between January 2014 and March 2025. For each pathogen, a Poisson generalized linear model (GLM) was fitted to the monthly counts with a linear time trend and 12-month harmonic terms (sin(2πt/12), cos(2πt/12)). A likelihood-ratio test (2 d.f.) compared the full model against a trend-only model (harmonic coefficients = 0), yielding a seasonality p-value; Bonferroni adjustment for 14 pathogens set α = 3.6 × 10⁻³. Monthly series were also decomposed using seasonal-trend decomposition via Loess (STL) with period = 12 to extract and quantify the seasonal component (peak-to-trough amplitude). All GLM and STL analyses used statsmodels 0.14. All data processing, statistical modeling, and plotting scripts are available on GitHub (see link below):
Institutions
- University of Pretoria
- Medical Diagnostic Laboratories LLC