Role of environmental pollutants in breast cancer: Emphasis on Methyl-4-hydroxybenzoate and Triple-negative breast cancer

Published: 9 March 2026| Version 1 | DOI: 10.17632/5m6npj27j9.1
Contributor:
孙 永明

Description

Research Hypothesis We hypothesized that Methyl-4-hydroxybenzoate (MEP), a common preservative and emerging environmental contaminant, contributes to triple-negative breast cancer (TNBC) pathogenesis through specific molecular interactions. A network toxicology framework was employed to systematically investigate MEP exposure and TNBC risk. Data Contents This dataset contains supplementary tables supporting the main findings: Table S1: 561 potential MEP targets from ChEMBL, STITCH-5.0, and SwissTargetPrediction, standardized to gene symbols via UniProt. Table S2: 2,447 TNBC-associated targets from GeneCards, OMIM, and TTD. Tables S3–S5: Differentially expressed genes (DEGs) from GSE76250 (GEO) and TCGA datasets, including 319 and 1,442 DEGs respectively, and 192 intersecting TNBC-related DEGs. Tables S6–S8: WGCNA results identifying TNBC-associated modules (blue from GEO, yellow from TCGA) and 60 key TNBC-related DEGs. Key Findings Integrative analysis identified three hub genes (EZH2, NEK2, TYMS) significantly enriched in cancer-related pathways with high diagnostic and prognostic value. These data form the foundation for conclusions that MEP may influence TNBC through these targets. Data Generation & Usage MEP targets were predicted using ChEMBL, STITCH, and SwissTargetPrediction. TNBC targets were retrieved from GeneCards, OMIM, and TTD. Transcriptomic data from GEO and TCGA underwent differential expression analysis (limma/DESeq2) with thresholds of adjusted P < 0.05 and |log₂FC| > 1 (GEO) or > 2 (TCGA). WGCNA was performed to identify co-expression modules. Researchers can use these files for replication, extended analyses, pathway enrichment, or meta-analyses. Data are provided in tab-delimited format compatible with standard bioinformatics tools.

Files

Steps to reproduce

Step 1: MEP Toxicity Prediction The SMILES structure of Methyl-4-hydroxybenzoate (MEP) was obtained from PubChem. Toxicity profiling was performed using two online platforms: ProTox-3.0 (for acute toxicity and toxicological endpoints) and ADMETlab-3.0 (for absorption, distribution, metabolism, excretion, and toxicity properties). These tools classified MEP as a Class IV toxicity substance with potential effects on the blood‑brain barrier, liver, and kidney. Step 2: MEP Target Identification Potential MEP targets were retrieved from three databases: ChEMBL (bioactive compound targets) STITCH-5.0 (chemical‑protein interactions) SwissTargetPrediction (ligand‑based target prediction, probability threshold > 0.1) Results were merged, deduplicated, and standardized to official gene symbols using the UniProt database. Step 3: TNBC Disease Target Acquisition TNBC‑associated targets were obtained by searching "Triple‑negative breast cancer" in: GeneCards (comprehensive gene annotation) OMIM (genetic disease variants) TTD (therapeutic target database) All results were merged, deduplicated, and mapped to official gene symbols via UniProt. Step 4: Transcriptomic Data Processing GEO dataset (GSE76250): Raw CEL files were downloaded and normalized using the affy R package (RMA method). Differential expression analysis was performed with limma (adjusted P < 0.05, |log₂FC| > 1). TCGA data: Raw count matrices and clinical metadata were downloaded. Male samples were excluded, and TNBC samples were extracted based on clinical information. Low‑quality FFPE samples were removed. Differential expression analysis was conducted with DESeq2 (adjusted P < 0.05, |log₂FC| > 2). Intersection of GEO and TCGA DEGs yielded 192 TNBC‑related DEGs. Step 5: Weighted Gene Co‑expression Network Analysis (WGCNA) Co‑expression networks were constructed using the WGCNA R package on TPM‑normalized expression matrices from both datasets. Modules significantly correlated with TNBC phenotype were identified (blue module from GEO, yellow module from TCGA). Genes were further filtered by |gene significance| > 0.2 and |module membership| > 0.8. Intersection with TNBC‑related DEGs produced 60 key TNBC‑related DEGs. Software & Tools: All analyses were performed in R with key packages including limma, DESeq2, WGCNA, affy, clusterProfiler, and ggplot2. Molecular docking (reported in the main article) was performed using AutoDock Tools and visualized with PyMOL.

Categories

Environmental Toxicology, Breast Cancer, Gene Expression

Funders

Licence