Assessment of Large Language Models in Addressing Questions Related to Oral Health Science: Which Model Can Provide Reliable Answers to Users’ Inquiries About Oral Health?

Published: 19 August 2026| Version 1 | DOI: 10.17632/tkykpx8jrm.1
Contributor:
Ruyi Zhao

Description

To evaluate the reliability of large language models (LLMs) in the field of oral health popular science, we collected and screened 100 common oral diseases and issues frequently encountered by patients during the course of illness, requiring assessors to score responses across four dimensions: comprehensiveness of content, professional accuracy, appropriateness of risk communication, and suitability for doctor-patient communication. Before inputting the questions to the LLMs, no instructions were provided, and each question was posed in a separate chat window to ensure the completeness and independence of each response. For the same question, we used bilingual prompting in both Chinese and English to compare response performance under Chinese- versus English-language inputs. After obtaining all answers, we compiled the responses from Doubao, DeepSeek, Kimi, Qwen, Spark Desk, and ERNIE Bot, and ChatGPT-5.2, and conducted anonymous cross-review among the LLMs, with an additional blind review by three oral health experts.

Files

Categories

Dentistry, Artificial Intelligence

Licence