The use of a metabologenomics approach for the discovery of antimicrobial natural products - Thesis Supplementary Data

Published: 23 July 2026| Version 1 | DOI: 10.17632/gtm274bsfc.1
Contributors:
Shuaib Hendricks,
,
,
,

Description

This dataset supports a multi-omics investigation into the biosynthetic potential and metabolite production of selected microbial strains under varying culture conditions. The primary hypothesis of this study is that strain repositioning combined with the OSMAC (One Strain Many Compounds) approach and integrated metabolomic and genomic analyses can reveal hidden metabolic diversity in actinobacteria. The dataset contains processed LC-MS/MS metabolomics data, whole genome sequencing (WGS) assemblies, genome annotations, and integrated metabologenomics outputs. Metabolomic data were generated using LC-MS/MS and processed into feature-based molecular networks using GNPS2. These data include MGF files, feature tables, metadata, and annotation outputs from tools such as SIRIUS and CANOPUS. Multivariate and univariate statistical analyses (e.g., PCA, PCoA, random forest classification, ANOVA, PERMANOVA) were performed to identify metabolite patterns associated with specific strains and culture conditions. Genomic data include polished genome assemblies, quality control metrics, and functional annotations, alongside biosynthetic gene cluster (BGC) predictions and comparative genomics analyses. Phylogenomic and taxonomic analyses (ANI, AAI, dDDH) were used to contextualize strain identity and novelty. BiG-SCAPE outputs were used to group BGCs into gene cluster families for comparative analysis. Metabologenomics integration was performed to link metabolomic features with predicted BGCs, enabling the identification of candidate gene-metabolite relationships. These results are provided as putative links which could be prioritised for further drug discovery. The data demonstrate that cultivation conditions significantly influence metabolite production profiles, and that integrating metabolomics with genomics enhances the interpretation of microbial chemical diversity. This dataset can be used for reanalysis of molecular networks, validation of genomic predictions, or further development of metabolite annotation and gene cluster linking approaches. Full methodological details are provided in the associated thesis.

Files

Steps to reproduce

All processed metabolomics and genomics data are provided as supplementary material. Full experimental methods are documented in the Materials and Methods section of the associated thesis. This repository contains the data files that support those methods. The folder structure maps directly to the four thesis chapters: - /Strain_Repositioning/ → Chapter 3 (strain source documentation) - /Microbial_Metabolomics/ → Chapter 4 (LC-MS/MS, molecular networking, statistical analyses) - /Microbial_Genomics/ → Chapter 5 (genome assemblies, annotations, phylogenomics, BGC prediction) - /Metabologenomics/ → Chapter 6 (genome-metabolome integration) Each main folder contains a README.txt file with folder-specific details and file descriptions. Raw LC-MS/MS (.mzML) and sequencing (.fastq) data are available from the corresponding author upon request and will be deposited in a public repository upon publication.

Institutions

Categories

Chemistry, Biochemistry, Molecular Biology, Analytical Chemistry, Drug Discovery, Bioinformatics, Mass Spectrometry, Microbial Genomics, Natural Product, Actinomycete, Untargeted Metabolomics, Multiomics

Funders

Licence