Evaluating the Potential of Large Language Models for Vestibular Rehabilitation Education: A Comparison of ChatGPT, Google Gemini, and Clinicians.
Yael Arbel, Yoav Gimmon, Liora Shmueli
PMID 39932784WHAT IT FOUND
ChatGPT answered 70% of vestibular rehabilitation questions correctly, beating students but losing to experienced therapists.
It struggled most with clinical reasoning. Therapists should use AI for basic knowledge retrieval, not complex decision-making, and watch for outdated advice in explanations.
Key findings
01ChatGPT scored 70% and Google Gemini 60% on the vestibular knowledge test, significantly lower than experienced physical therapists but higher than students.
02Both AI models performed perfectly on clinical knowledge questions but poorly on clinical reasoning, where ChatGPT scored 50% and Gemini 25%.
03Experts rated only 45% of ChatGPT's explanations as comprehensive, with 25% considered completely incorrect, often due to outdated information or poor handling of complex scenarios.
STILL TO COME
How it was doneWhat they foundWhat it means for PTs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The study used a new, non-standardized questionnaire (VKT) developed specifically for this research. The internal consistency of the test was low for students (Cronbach alpha .37) and moderate for therapists (.68). The AI models were queried only once, reflecting a single real-world interaction but not accounting for iterative refinement. The test was originally developed in Hebrew and translated to English, which may have introduced linguistic bias despite expert review. Results are specific to the versions of ChatGPT-3.5 and Google Gemini tested in 2023/2024 and may not apply to newer models.
Declared interests
No specific funding sources or conflicts of interest are declared in the provided text.
The easy way to misread this
Do not assume AI is ready for clinical decision-making. While ChatGPT outperformed students, it scored only 50% on clinical reasoning questions and provided completely incorrect explanations for 25% of its answers, often due to outdated information. It is a tool for basic knowledge retrieval, not a substitute for professional judgment.