Data and code for “Uncertainty-aware integration of multi-target prediction and dual-state docking prioritizes generated 5-HT2A candidates

Published: 26 August 2026| Version 1 | DOI: 10.17632/zcmjctx4py.1
Contributors:
,

Description

This dataset provides the data, code, trained models, computational outputs, and provenance documentation supporting the manuscript “Uncertainty-aware integration of multi-target prediction and dual-state docking prioritizes generated 5-HT2A candidates”. The release includes raw and processed ChEMBL-derived activity data, the strict binding-Ki dataset comprising 14,587 target-structure records across 5-HT2A, 5-HT2B, 5-HT2C, and 5-HT1A, scaffold-disjoint train/validation/test partitions, LightGBM and Random Forest model outputs, per-repeat predictions and uncertainty summaries, applicability-domain analyses, archived molecular prior and reinforcement-learning checkpoints, fixed-seed generative-model resampling results, candidate-level strict rescoring and chemical-quality assessments, dual-state docking and native-ligand redocking outputs, and source data for the manuscript figures and tables. Historical reinforcement-learning checkpoints and the archived candidate pool are preserved unchanged as provenance artifacts. The original stochastic RL training seeds and historical candidate-pool sampling seed were not recorded and are not retrospectively inferred. A separate deterministic reproducibility layer provides seed-controlled rerun implementations for future RL training and candidate sampling, with isolated output paths so that historical manuscript artifacts cannot be overwritten. The archive also contains environment specifications, a data dictionary, model and provenance documentation, a complete file manifest, and SHA-256 checksums for integrity verification. The four molecules cand_77, cand_164, cand_76, and cand_134 are computationally prioritized candidates and have not been experimentally validated.

Files

Steps to reproduce

Download JMGM_reproducibility_package_v2.1.0.zip and extract the archive. Read README.md, REPRODUCIBILITY.md, DATA_DICTIONARY.md, and PROVENANCE_LIMITATIONS.md before running the workflow. Recreate the documented computational environment using the environment specification included in the archive. From the package root, run: python src/verify_complete_release.py --verify and python src/verify_traceability_patch.py Both verification procedures should complete successfully for the released package. The strict binding-Ki modeling, scaffold-split evaluation, applicability-domain analysis, candidate rescoring, docking summaries, and publication source-data workflows can be reproduced using the scripts and commands documented in README.md and REPRODUCIBILITY.md. The archived RL checkpoints and historical candidate pool are authoritative provenance artifacts for the manuscript results. Their original stochastic RL training seeds and historical candidate-sampling seed were not recorded and cannot be reconstructed retrospectively. For deterministic future reruns of the historical generative workflow, use the scripts under src/reproducible_rerun/. These use an explicit rerun seed and write results only to isolated reproducibility directories; they do not overwrite or redefine the archived manuscript artifacts. File identities and integrity can be checked against MANIFEST.tsv, CHECKSUMS.sha256, and the SHA-256 file supplied for the complete ZIP archive.

Institutions

Categories

Cheminformatics, Computational Chemistry

Licence