GPT-4 outperforms junior expert physical therapists in sports medicine rehabilitation: an evaluation of AI response quality and adaptiveness.
Eric Hamrin Senorski, Ramana Piussi, Janina Kaarre and 7 others
GPT-4's written answers to sports rehabilitation questions outscored those of three junior expert physical therapists on both quality and adaptiveness, for patient-facing and therapist-facing responses alike.
Key findings
1For patient-targeted responses, GPT-4 received a median quality score of 3 (mean 2.774) versus 2 (mean 2.182) for JEPs, and a median adaptiveness score of 4 (mean 3.296) versus 3 (mean 2.384), both with p ≤ 0.001.
2For physical therapist-targeted responses, GPT-4 received a median quality score of 3 (mean 2.805) versus 2 (mean 2.025) for JEPs, and a median adaptiveness score of 3.25 (mean 3.113) versus 2.50 (mean 2.025), both with p ≤ 0.001.
3All three experts unanimously preferred GPT-4 in 26 patient-targeted questions and 34 physiotherapist-targeted questions, while JEPs were preferred in 3 patient-targeted questions and never in physiotherapist-targeted questions.
Still to come
How it was doneWhat they foundWhat it means for PTs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
Written responses only; no real-time interaction, non-verbal cues, physical assessment, or patient feedback, all of which a PT uses in practice. Only three JEPs, all from the same clinic, all sports-medicine-specialised PhD students aged 29 to 35; results may not represent the broader PT population. No systematic safety validation of the AI-generated advice; the quality rating included an implicit safety judgment but was not a formal safety test. The AI model's knowledge has a cutoff date and it is not continuously updated; it can generate plausible but incorrect or outdated recommendations (hallucination risk). The study design favoured AI in standardisation: GPT-4 gave a consistent, well-structured written answer, while human therapists bring variable, experience-based reasoning that a text-only comparison does not capture. Inter-rater agreement was weak (Fleiss' kappa 0.20 and 0.07), meaning the three experts' individual choices varied even though the overall direction favoured GPT-4. Using GPT-4 effectively requires prompt engineering knowledge; a patient without that knowledge may not get well-adapted responses. The black-box architecture makes it difficult to trace the sources or reasoning behind GPT-4's responses, limiting transparency and accountability.
Declared interests
The authors declared that no financial support was received for the work or its publication.
The easy way to misread this
Do not read this as evidence that GPT-4 can replace a physical therapist in clinical practice. The study compared written text only, in a controlled setting, with no real-time interaction, no physical assessment, and no safety validation. The three raters' individual choices showed weak agreement (Fleiss' kappa 0.20 for patient responses, 0.07 for therapist responses), so the overall preference for GPT-4 was driven by a skewed distribution rather than strong consensus on every question. The AI's knowledge has a fixed cutoff and it can produce confident but incorrect recommendations, particularly on loading, progression, and return-to-sport decisions where patient safety is at stake.
Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →