Data for Large-scale assessment of regression modeling practices in phytomedicine extraction process optimization
Description
This dataset supports the study “Large-scale assessment of regression modeling practices in phytomedicine extraction process optimization.” The purpose of the dataset is to enable transparent reanalysis of model selection, predictive validation, diagnostic evaluation, and optimization-stability assessment in phytomedicine extraction-process studies. The deposited data were curated from published phytomedicine extraction-process optimization studies identified through systematic literature screening. Each extraction-process dataset contains experimental runs, process factors, and a continuous response variable extracted from tables, text, or other recoverable numerical information in the source publications. The data were standardized into tabular formats suitable for regression modeling and machine-learning-based prediction, with process variables used as predictors and the extraction outcome used as the response. The main analytical dataset contains 1,148 structurally eligible extraction datasets that met the requirements for unified quadratic regression modeling. These datasets were used to compare seven candidate models: full quadratic regression, two subset-selected parsimonious regression models, quadratic Ridge regression, support vector regression, partial least squares regression, and Gaussian process regression. Model performance was assessed using fitting metrics, leave-one-out cross-validation metrics, diagnostic indicators, and model-predicted optimum stability. An additional set of 251 datasets that did not meet the structural criterion for full quadratic regression is also provided for machine-learning-only sensitivity analysis. The dataset-level summary files report information used to reproduce the main analyses, including dataset structure, experimental design type, sample size, number of factors, model-fitting results, cross-validation performance, diagnostic metrics, optimization results, sensitivity analyses, and stratified robustness analyses. These files can be used to verify the reported findings, compare alternative modeling strategies, examine the behavior of regression and machine learning models under small-sample extraction settings, or conduct secondary methodological research on model-dependent optimization uncertainty. The data should be interpreted as literature-extracted, retrospective benchmark data rather than newly generated experimental measurements. Model-predicted optima derived from these datasets should be regarded as candidate operating regions requiring experimental confirmation, not as definitive production conditions.
Files
Institutions
- Zhejiang Chinese Medical UniversityZhejiang, Hangzhou