Data and code for mapping human nutrient-overload programs onto cell states across MASLD progression

Published: 6 September 2026| Version 1 | DOI: 10.17632/5f78c3yz6n.1
Contributor:
Biao Gao

Description

This dataset contains the study-generated derived data, frozen analytical resources, publication source data, quality-control outputs, audit materials, and versioned R code supporting an integrative public-data study of human nutrient-overload programs across metabolic dysfunction-associated steatotic liver disease (MASLD). The core analytical framework links four public omics resources with distinct evidentiary roles. GSE200418 was used as a direct human precision-cut liver slice perturbation anchor for deriving nutrient-response programs. GSE202379 provided donor-aware single-nucleus RNA-sequencing data for mapping these frozen programs across human MASLD cell types and cross-sectional disease states, with donors rather than nuclei or tissue specimens treated as the independent biological units. PXD051911 provided hierarchical same-study multicompartment human proteomic support. GSE312698 provided a bounded, panel-restricted, FOV/reconstructed-specimen-level descriptive spatial projection in human liver CosMx data. The release includes ranked programs and directional gene sets, identifier harmonization resources, donor-aware cross-layer results, proteomic evidence summaries, bounded spatial-projection outputs, publication figures and tables with source data, reproducibility and quality-control records, and the retained analysis code chain. Primary repository data are not redistributed. They should be obtained directly from NCBI GEO and ProteomeXchange/PRIDE under accession numbers GSE200418, GSE202379, GSE312698, and PXD051911. Controlled-access materials are not included. The data support cross-dataset correspondence and orthogonal consistency analyses. They should not be interpreted as evidence of longitudinal disease progression, dietary causality, therapeutic efficacy, or an optimal intervention window.

Files

Steps to reproduce

1. Obtain the primary public datasets from their original repositories: GSE200418, GSE202379 and GSE312698 from NCBI GEO, and PXD051911 from ProteomeXchange/PRIDE. Primary repository data are not redistributed in this package. 2. Download and extract the eight release archives. Begin with 00_README_manifest_environment_v1.0.zip and consult README.md, dataset_accession_manifest.csv, file_manifest_sha256.csv, data_dictionary_file_catalog.csv, and script_execution_order.csv. 3. Follow script_execution_order.csv for the retained R workflow. The code is organized by analytical layer: GSE200418 perturbation-program derivation, GSE202379 donor-aware pseudobulk and cross-layer mapping, PXD051911 proteomic support, GSE312698 bounded spatial projection, and publication-asset generation. 4. Use the frozen ranked programs, directional gene sets, identifier crosswalks and analysis-ready derived objects supplied in the corresponding module archives. Do not reselect programs using downstream disease or spatial results. 5. Treat biological donors, rather than nuclei, cells, sequencing runs, tissue lobes or spatial FOVs, according to the inferential-unit definitions documented in the supplied manifests and audit files. 6. Verify file integrity against the supplied SHA-256 manifests before analysis. 7. Reproduce manuscript figures and tables using the retained publication scripts and compare outputs against the source-data and publication-asset crosswalk supplied in the release. 8. Interpret GSE312698 only as a panel-restricted descriptive spatial projection and PXD051911 as hierarchical same-study multicompartment support; neither constitutes an independent patient-level validation cohort.

Categories

Gastroenterology, Genomics, Proteomics, Metabolism, Bioinformatician

Licence