Evaluating readability, reliability, and scientific quality of large language models in polyendocrine metabolic ovarian syndrome: Cross-Sectional Study

Published: 3 August 2026| Version 1 | DOI: 10.17632/tktstpj99s.1
Contributors:
,
,
, Ye Mo

Description

This cross-sectional comparative study aimed to systematically assess and compare the comprehensive performance of four mainstream LLMs, including GPT-5.6sol, Gemini3.1pro, Kimi K3, and DeepSeek-v4pro, in generating professional medical information regarding PMOS. A standardized evaluation framework was established referring to authoritative clinical guidelines and commonly concerned clinical questions about PMOS, which were summarized and screened from high-attention medical platforms and clinical consultation scenarios. The evaluation scope covered core clinical and cognitive dimensions of PMOS, and a total of 20 representative questions were screened and classified into five standardized domains: Basic Cognition, Symptoms and Pathogenic Mechanisms, Examinations and Diagnostic Criteria, Treatment and Daily Care, as well as Complications and Long-Term Risks.

Files

Institutions

Categories

Artificial Intelligence, Obstetrics

Licence