Data and code for “A provenance-aware workflow for species-resolved microbiome candidate prioritisation across heterogeneous public evidence: an application to alcohol-associated hepatitis”
Description
This dataset contains version-matched code and distributable data supporting the manuscript “A provenance-aware workflow for species-resolved microbiome candidate prioritisation across heterogeneous public evidence: an application to alcohol-associated hepatitis”. The workflow uses a clinical genus-level anchor, resolves retained genera to species, applies the genus-adaptive ED-MVE filter and a gutMEGA criterion, separates candidate selection from downstream public-knowledge-base annotation, and summarises Convergent Evidence Profiles and host-expression support in three liver transcriptome datasets. Version 2 adds a deterministic five-scenario leave-one-cohort-out analysis of clinical-anchor cohort dependence, its machine-readable CSV and JSON outputs, and regression tests. The packages include R and Python source code, environment and dependency records, run-order documentation, distributable inputs, intermediate and final result tables, provenance records, and scripts for reconstructing the seven current manuscript figures. Molecular docking is retained as supplementary plausibility analysis and does not determine candidate selection. GeneCards raw or score-bearing material and CTD raw or row-level exports are not redistributed because redistribution rights were not established. Original data-package content is released under CC BY 4.0, original code is released under the MIT License within the code archive, and third-party content remains subject to its source terms. Version 2 replaces the Version 1 code and data archives with the version-matched v1.1.0 release prepared for the provenance-aware workflow. It adds the deterministic clinical-anchor leave-one-cohort-out cohort-dependence analysis, registered CSV and JSON outputs, focused regression tests, updated run documentation, and the current seven-figure reconstruction scripts. The candidate-selection values, frozen species-level scores, existing downstream results and scientific evidence boundaries are unchanged; the new analysis re-applies the registered genus rules without querying external databases or recalculating the frozen species-level scores. Release-facing wording was also normalised to remove manuscript-development terminology.
Files
Steps to reproduce
Download and extract `AH_microbiome_pipeline.zip` and `AH_microbiome_data.zip` side by side as `AH_microbiome_pipeline/` and `AH_microbiome_data/`. Create the environment from `environment.yml`, activate it, and run `Rscript scripts/install_cran_cheminformatics.R`. Set `PIPELINE_DIR`, `DATA_DIR` and `RETICULATE_PYTHON`, then run `AH_DATA_DIR="$DATA_DIR" Rscript "$PIPELINE_DIR/tests/run_tests.R"` for the focused R regression suite and `AH_DATA_DIR="$DATA_DIR" "$RETICULATE_PYTHON" "$PIPELINE_DIR/tests/test_clinical_anchor_loo.py"` for the clinical-anchor leave-one-cohort-out regression. Recompute the leave-one-cohort-out outputs with `python/clinical_anchor_loo.py`, recompute the rule-matched membership ablations with `python/internal_ablation_analysis.py`, and rebuild the seven current manuscript figures with `scripts/main_figures/build_all_presentation_figures.py --data-dir "$DATA_DIR" --output-dir <output_directory>`. Follow `README.md`, `QUICKSTART.md` and `MANUAL_STEPS.md` for complete commands and the documented external-platform boundaries; database and web-platform stages require current source exports, and rights-sensitive GeneCards and CTD inputs must be obtained separately under their source terms.
Institutions
- Nanjing University of Chinese MedicineJiangsu, Nanjing