Benchmarking large language models against human experts in rehabilitation medicine: a multidimensional evaluation.
Wenhui Cao, Mengjian Qu, Tao Zhu and 14 others
PMID 41772676WHAT IT FOUND
Top-tier AI models wrote rehabilitation plans that scored higher than textbook expert answers on safety and evidence, but missed the patient's life context.
Use AI to draft the medical checklist. You must still add the human strategy and personal goals.
Key findings
01Grok-4 and Gemini-2.5-pro generated plans with higher weighted scores than the expert benchmark, with Grok-4 scoring 4.31 versus 3.56 for experts.
02Human experts outperformed all AI models on integrating Traditional Chinese Medicine, scoring 2.31 versus 1.77 for the top AI model.
03Qualitative review showed AI plans were static checklists, while human plans prioritized core limiting factors and used patient-specific goals like stocking fruit shelves.
STILL TO COME
How it was doneWhat they foundWhat it means for PTsWhat it means for OTsWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The study compared AI to 'textbook' expert answers, not the variable and often hurried plans of clinicians in routine practice. No real patients were involved; this was a simulation using written case summaries, so it cannot predict actual health outcomes or patient safety in the real world. The expert panel was small (6 people) and recruited from high-level Chinese institutions, which may bias the preference for integrated Traditional Chinese Medicine approaches. The evaluation was text-only. It did not test the AI's ability to interpret imaging or observe patient movement, which are critical for physical and occupational therapy.
Declared interests
Funded by The First Affiliated Hospital of University of South China.
The easy way to misread this
Do not interpret the high AI scores as evidence that AI is safer or more effective than a human therapist. The AI won on 'standardized textbook' criteria like structure and evidence citation. It failed to demonstrate the clinical judgment to prioritize a patient's psychological stability or adapt goals to their specific life context, which are the skills that actually drive rehabilitation success.
Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →