Nomige Dermatogenomics Dataset: Genetic, Lifestyle, and Skin Phenotype Variables for Personalized Skin Health Modeling
Description
This dataset contains de-identified dermatogenomics and lifestyle records used to study personalized skin-health phenotypes through machine learning. The dataset integrates genetic markers, lifestyle/environmental variables, and ordinal skin concern severity labels related to pigmentation, scarring, dryness, sensitivity, acne, and redness. It was used in multiple studies on multimodal skin-health modeling, phenotype clustering, multi-output prediction, explainable AI, and privacy-preserving dermatological assessment. The dataset supports research in dermatogenomics, personalized skincare, precision dermatology, explainable machine learning, multi-output prediction, gene-environment interaction modeling, and AI-assisted dermatological assessment.
Files
Steps to reproduce
The dataset contains de-identified records integrating genetic markers, demographic/lifestyle/exposure variables, and ordinal skin phenotype labels. To reproduce the machine-learning analyses reported in the related publications, researchers should: 1. Load the dataset file and consult the accompanying data dictionary. 2. Treat the genetic variables MMP-1, MMP-3, SOD2, GPX1, AQP3, and FLG as encoded genetic marker variables. 3. Use the lifestyle, exposure, allergy, and demographic variables as input predictors. 4. Use Pigmentation, Dryness, Sensitivity, Scarring, Acne, and Redness as ordinal target variables. 5. Apply appropriate preprocessing, including train/test splitting or cross-validation, while avoiding data leakage. 6. Train supervised machine-learning models for single-output or multi-output prediction, depending on the research objective. 7. Evaluate model performance using suitable classification or ordinal prediction metrics, such as accuracy, F1-score, MAE, QWK, or related metrics. 8. Use explainability methods such as feature importance or SHAP analysis to interpret genetic, lifestyle, and environmental contributions. Because the dataset contains individual-level genetic, lifestyle, age, and phenotype variables, access should follow the restricted-access terms defined by the data owner.
Institutions
- Higher Colleges of TechnologyAbu Dhabi, Abu Dhabi