Fluctuating Performance and Discordance of LLMs in Radiology Over Time AIPHI
Chris Kaufmann · MD, MS; The University of Texas at Austin, Dell Medical School, Department of Diagnostic Medicine; Oden Institute for Computational Engineering and Sciences
Objective. Since the introduction of large language models (LLMs), near expert level performance in medical specialties such as radiology has been demonstrated. However, there is limited to no comparative information of model performance, accuracy, and reliability over time in these medical specialty domains.
Methods. LLMs (GPT-4, GPT-3.5, Claude, and Google Bard) were queried monthly from November 2023 to January 2024, utilizing ACR Diagnostic in Training Exam (DXIT) practice questions. Model overall accuracy and by subspecialty category was assessed over time. Internal consistency was evaluated through answer mismatch or intra-model discordance between trials.
Results. GPT-4 had the highest accuracy (78 ± 4.1 %), followed by Google Bard (73 ± 2.9 %), Claude (71 ± 1.5 %), and GPT-3.5 (63 ± 6.9 %). Models demonstrated temporal performance fluctuations, with intra-model discordance rates decreasing for all systems and variable subspecialty performance.
Conclusion. LLMs, except GPT-3.5, performed above 70%, demonstrating substantial subject-specific knowledge. However, performance fluctuated over time, underscoring the need for continuous, radiology-specific standardized benchmarking metrics to gauge LLM reliability before clinical use.