Meta-Feature Driven Medical Algorithm Performance Dataset (Medical | Healthcare)
Description
This dataset is a specialized meta-learning dataset designed for algorithm selection in medical informatics. It contains 420 records derived from 42 distinct clinical datasets obtained from the OpenML repository. Each record represents the evaluation of one machine learning algorithm applied to a specific medical dataset, resulting in a structured collection of performance outcomes and statistical descriptors. The dataset integrates meta-features that describe the intrinsic properties of each clinical dataset, alongside performance metrics of multiple algorithm families. These meta-features include statistical signatures such as entropy, class imbalance, skewness, kurtosis, correlation structure, and feature composition ratios. Together, they characterize the geometry and complexity of the underlying medical data. In addition to statistical descriptors, the dataset tracks algorithm-level performance using metrics such as F1-score, accuracy, and execution time. Each record identifies the evaluated algorithm, its performance, and whether it achieved the best result for the corresponding dataset. A comparative reference layer is also included, storing the performance of all algorithm families for each dataset, enabling direct benchmarking. The dataset further incorporates ETL audit information, including data cleaning impact, number of removed constant features, and final feature counts. Structural properties such as number of rows and columns are also included, providing insight into dataset scale and dimensionality. Organized in CSV format, the dataset is compatible with standard analytical tools such as Python, R, and AutoML systems. All data is fully anonymized and derived from publicly available benchmarks, ensuring no exposure of patient-identifiable information. This dataset provides a unified foundation for studying the relationship between dataset characteristics and algorithm performance, supporting the development of automated model selection systems in healthcare analytics.
Files
Steps to reproduce
The Gold Layer is produced through a structured Medallion architecture pipeline consisting of data ingestion, transformation, integration, and final consolidation. The process begins in the Bronze Layer, where 42 clinical datasets are ingested from the OpenML repository in their raw format. These datasets are stored without modification to preserve original structure and ensure traceability. Each dataset is linked with metadata including identifiers, source URLs, and storage paths. In the Silver Layer, preprocessing and feature engineering are performed. Raw datasets are cleaned by removing missing values, redundant attributes, and constant features. Data types are standardized, and statistical meta-features are computed, including entropy, imbalance ratio, skewness, kurtosis, and correlation measures. Structural attributes such as number of rows and columns are also derived. Feature ratios (numerical vs categorical) and null density are calculated to describe dataset composition. Next, algorithm evaluation is conducted. Multiple machine learning algorithm families are applied to each dataset. Performance metrics such as F1-score, accuracy, and execution time are recorded. The best-performing algorithm is identified and labeled. Additionally, a full reference layer is generated, storing the performance of all algorithms for each dataset. In the integration stage, all processed records are unified into a consistent schema. Each record combines dataset metadata, statistical signatures, ETL audit features, and algorithm performance metrics. A denormalized structure is created to ensure efficient analytical access. In the Gold Layer, the final dataset is constructed as a clean, analysis-ready table. All records are validated to ensure completeness and consistency. Missing values are eliminated, and final feature selection is applied. The dataset is then exported in csv format and made accessible for downstream analytics, supporting meta-learning and automated model selection workflows.
Institutions
- New Mansoura UniversityDakahlia, Al Mansurah
- Mansoura UniversityDakahlia, Al Mansurah