Comprehensive Evaluation of Six Large Language Models on Four Core Medical School Courses
Published: 1 July 2026| Version 1 | DOI: 10.17632/cfwkcsxjc5.1
Contributor:
wei zhengDescription
Large language models (LLMs) have demonstrated remarkable potential in medical education, yet their performance on discipline-specific medical school course examinations remains incompletely characterized. This study evaluated six contemporary LLMs on final examinations for four core medical school courses: Surgery, Musculoskeletal System Diseases, Digestive System Diseases, and Respiratory System Diseases. A total of 400 multiple-choice questions (100 per course) were administered to each model. Responses were scored against official answer keys and compared to the performance of 312 medical students. Each question was tested three times per model to assess response consistency and reproducibility.
Files
Categories
Medical Education