Supplementary Materials for “Effect of Fitzpatrick Skin Type Prompting on Diagnostic Accuracy in Multimodal Large Language Models: A Within-Image Experimental Study”

Published: 18 August 2026| Version 5 | DOI: 10.17632/z4dw2djvkk.5
Contributors:
,
,
,
,
,

Description

This supplementary dataset supports the study “Effect of Fitzpatrick Skin Type Prompting on Diagnostic Accuracy in Multimodal Large Language Models: A Within-Image Experimental Study.” It includes the original study protocol, final prompt templates, predefined diagnostic scoring ontology, processed model outputs used for scoring, diagnostic scores, the fully executed analysis notebook, post hoc sensitivity-analysis results, descriptive top-1 and top-3 diagnostic accuracy results, binary malignancy-detection results, and top-1 confusion-matrix analyses with the associated prediction crosswalk and mapping audit. The study evaluated 656 biopsy-confirmed photographs from the Diverse Dermatology Images dataset under four prompt submissions per browser-based model configuration: no-FST, DDI-concordant FST, and two DDI-discordant FST prompts. ChatGPT 5.2 Edu and Gemini 3.1 Pro were evaluated, producing 5,248 model-output evaluations. Histopathologic diagnosis served as the diagnostic reference standard. Fitzpatrick skin type was treated as DDI-assigned metadata rather than independently verified ground truth. The deposited Model Outputs and Scores workbook contains the processed ranked differentials used for scoring; raw pre-trimming responses were not retained. These materials document the study workflow and support reproduction of the reported primary, exploratory, and post hoc analyses.

Files

Steps to reproduce

This study used all 656 biopsy-confirmed photographs from the Diverse Dermatology Images (DDI) dataset. DDI grouped Fitzpatrick skin type (FST) as I–II, III–IV, or V–VI. Histopathologic diagnosis served as the diagnostic reference standard; FST was treated as DDI-assigned metadata. Each image was submitted to the ChatGPT 5.2 Edu and Gemini 3.1 Pro browser configurations under four conditions: no-FST, DDI-concordant FST, and two DDI-discordant FST prompts, yielding 2,624 evaluations per configuration and 5,248 overall. For FST I–II and V–VI images, discordant prompts were recoded as one adjacent and one extreme; for FST III–IV images, both were adjacent. Each image-prompt combination was submitted in a new Temporary Chat. Submission order was not prescribed or recorded. During quality control, non-diagnostic framing outside the ranked differential was removed; diagnoses and rank were unchanged. Raw pre-trimming responses were not retained. Diagnostic accuracy was scored 0–3 relative to the biopsy-confirmed diagnosis or accepted ontology equivalent: rank 1 = 3, rank 2 = 2, rank 3 = 1, and absent from the top three = 0. The primary CLMM included model configuration and prompt condition as fixed effects and image as a random intercept, with Benjamini–Hochberg adjustment for the four primary tests. Exploratory analyses evaluated confirmed malignant diagnosis omission and no-FST accuracy differences across DDI-assigned FST groups. A post hoc equal-image-weighted sensitivity CLMM assigned weight 0.5 to each adjacent observation from FST III–IV images and weight 1 otherwise, with FST group as an additional fixed effect. Post hoc analyses calculated top-1 accuracy (score = 3), top-3 inclusion accuracy (score > 0), and binary malignancy detection. Binary detection used the DDI malignancy field as the reference and was positive when any top-three diagnosis represented an explicit malignant entity or family; benign and premalignant terms were negative unless an explicit malignancy was also listed. TP, FP, TN, FN, sensitivity, specificity, PPV, NPV, F1, and accuracy were calculated by model and original prompt submission; Wilson 95% CIs were calculated for proportions except F1. For the post hoc top-1 error-pattern analysis, one fixed crosswalk was applied across both models and all four original prompt submissions. Diagnosis strings previously accepted by the ordinal scorer were linked to their established canonical category. Remaining strings were mapped only through the predefined ontology, exact entity labels, and documented equivalences; the principal diagnosis before any parenthetical descriptor was used when no prior score match existed. Uncertain, multi-category, morphologic-only, or out-of-ontology predictions were assigned to Other/unmapped. Complete count and row-normalized confusion matrices were generated. No additional hypothesis testing was performed for descriptive analyses.

Institutions

Categories

Dermatology, Artificial Intelligence Applications, ChatGPT

Licence