Comprehensive Evaluation of Six Large Language Models on Four Core Medical School Courses

Published: 1 July 2026| Version 1 | DOI: 10.17632/cfwkcsxjc5.1
Contributor:
wei zheng

Description

Large language models (LLMs) have demonstrated remarkable potential in medical education, yet their performance on discipline-specific medical school course examinations remains incompletely characterized. This study evaluated six contemporary LLMs on final examinations for four core medical school courses: Surgery, Musculoskeletal System Diseases, Digestive System Diseases, and Respiratory System Diseases. A total of 400 multiple-choice questions (100 per course) were administered to each model. Responses were scored against official answer keys and compared to the performance of 312 medical students. Each question was tested three times per model to assess response consistency and reproducibility.

Files

Categories

Medical Education

Licence