Dataset supporting MOLSMARTS: Benchmarking of a virtual Pharmacophore-based SMARTS Reaction Chemical Library
Description
1. Research introduction A pharmacophore-based SMARTS reaction-driven virtual chemical library (MOLSMARTS and atom base library generation (MOLHYBRID), which both can generate chemically valid, synthetically meaningful, and drug-like molecules. The MOLSMARTS generated about 51 Million unique molecules and MOLYHBRID generated about 56,000 molecules. The generated molecules were benchmarked against the available online databases. Our finding suggest that some novel drug-like molecules might be found within these datasets. The proof-of-concept application of MOLHYBRID for drug discovery were demonstrated through virtual screening and using various computational tools targeting cancer and Tuberculosis proteins. 2. Overview of the dataset used These dataset contains subset (not full database) of molecules and the computed properties. These chemical libraries are 1. CHEMBL - Bioactive compounds with experimentally validated activity 2. ZINC - Virtual and purchasable compounds 3. COCONUT - Natural product compounds 4. SANCDB - South African natural product compounds 5. MOLSMARTS - Pharmacophore-based SMARTS reaction virtual compounds 6. MOLHYBRID - Pharmacophore virtual compounds randomly generated by randomization 7. GDB13- The enumerated molecules 8. ZINC FDA APPROVED - Drug-like and clinically approved compounds 9. DRUGBANK - Approved and investigated drugs 10. PUBCHEM- Large database that is integrated with public chemical database 3. What the Dataset Contains These dataset contains the input SMILES of each database reported in the paper "MOLSMARTS: Benchmarking of a virtual Pharmacophore-based SMARTS Reaction Chemical Library " (not yet published at the date of publication of this dataset). The CSV files contains the following computed properties Bemis-Murcko scaffolds, structural alerts, physicochemical properties, drug-likeness rules, Synthentic accessibility scores (SAS), Quantitative estimation of drug-likeness (QED), Principal moment of inertia (PMI), plane of best fit (PBF) and fraction of sp3 hybridized carbons. 4. How the data was generated The MOLSMARTS and MOLHYBRID datasets were generated by the fully packaged programs available freely for use on GitHub repository, available respectively on these links https://github.com/CMCDD/molsmarts and https://github.com/CMCDD/molhybrid. These programs work on the Linux environment. These programs further helps in elimination of static chemical libraries which require huge memory by allowing users to generate virtual chemical libraries based on demand. The researchers without access to higher performance computing (HPC) can also generate datasets from their smaller computers. 5. The key scientific findings and contribution The dataset provides a benchmarking framework that can allow direct comparison between the datasets. it further validates that the generated molecules obey valency rules and have potential for synthesizability. The available programs allow limitless creativity.
Files
Steps to reproduce
Explained in README files in GitHub and the associated paper.
Institutions
- Rhodes UniversityEastern Cape, Grahamstown
Categories
Funders
- ADA & BERTIE LEVENSTEIN SCHOLARSHIP, RHODES UNIVERSITY