About this trial
This multicenter retrospective study aims to evaluate the diagnostic and therapeutic performance of three large language models-ChatGPT, Gemini and Deepseek-using 800 archived inpatient medical records from urology departments across four tertiary hospitals. The study will focus on the accuracy and applicability of these models in disease recognition, preliminary diagnosis and treatment recommendation generation, in order to explore their potential value and limitations in supporting clinical decision-making in real-world settings.
Eligibility criteria
Qualifiers
The case data is sourced from the four hospitals involved in the study, with complete and authentic diagnosis and treatment records.
Patients must be 18 years or older, with no gender restrictions.
Complete medical records, including the following core information: patient' s basic information, present illness history, past medical history, physical examination, and auxiliary examinations (including laboratory and imaging tests).
A clear discharge diagnosis and treatment plan (including therapeutic measures and follow-up arrangements).
Disqualifiers
Medical records with significant missing information, such as key clinical details (present illness history, diagnostic or treatment records, etc.).
Cases where the diagnosis or treatment plan is unclear, or where treatment has not been fully completed for an initial diagnosis.
Cases where the primary diagnosis is not urological.
Cases with major errors or inconsistencies in the records that could affect further assessment.
Trial design
Treatments tested in this trial
- Large Language Model Assessment (ChatGPT, Gemini, DeepSeek)