E-cigarette flavoring additives oral-respiratory hazard prioritization data and code

Published: 14 September 2026| Version 4 | DOI: 10.17632/wn774d8bck.4
Contributor:
Zhongfeng Shi

Description

This dataset provides analysis code, processed data, and reproducibility materials for a computational toxicology study of e-cigarette flavoring additives and carrier solvents. Version 4 adds the corrected replicated molecular-dynamics release used in the substantially revised manuscript. It includes results from three independently seeded 100 ns trajectories for each of three ligand-bound complexes, trajectory-level structural metrics, nine per-trajectory MM-PBSA frame tables, across-trajectory summaries, corrected Table S1 and Figure S2, GROMACS parameter files, and analysis scripts. The revised workflow retains all production frames, uses periodic-boundary-aware minimum protein-ligand heavy-atom distance, and reconstructs the four-chain TNF assembly into temporally continuous periodic images before structural analysis. MM-PBSA was calculated separately for each trajectory using 80 frames from 80.00 to 99.75 ns. Large third-party datasets and approximately 50 GB of raw trajectory files are not redistributed. The included derived data reproduce the published replicate summaries and Figure S2.

Files

Steps to reproduce

All analyses were performed on a Windows 10/11 workstation. The main software and versions used: Python 3.10+ (with RDKit, pandas, requests, scipy) R 4.2+ (with clusterProfiler, Seurat, TwoSampleMR, GSVA, ggplot2) GROMACS 2023+ (for molecular dynamics simulations) AlphaFold 3 Server (https://alphafoldserver.com) Cytoscape 3.10 (for network visualization) AutoDock Vina 1.2 (for molecular docking validation) Reproduction workflow: Run target prediction scripts (SwissTargetPrediction, SEA, PharmMapper, SuperPred) to obtain compound targets. Run disease gene query scripts (GeneCards, DisGeNET, CTD) to collect disease-associated genes. Execute merge and intersection scripts to identify overlapping targets. Perform enrichment analysis using the R scripts (GO, KEGG, GSEA). Conduct AlphaFold 3 docking via the web server; score and filter results using the provided Python scripts. Run GROMACS MD simulations following standard protocols described in the manuscript; extract metrics with the provided analysis scripts. Perform Mendelian randomization using the R scripts with eQTLGen eQTL and FinnGen R12 GWAS summary statistics. Run single-cell RNA-seq re-analysis using the Seurat-based R scripts. Detailed parameters and step-by-step instructions are provided in the README.txt file included in the dataset.

Institutions

Categories

Toxicology, Pharmacology, Bioinformatics, Computational Biology

Licence