Use of artificial intelligence large language models as a clinical tool in rehabilitation medicine: a comparative test case.
Liang Zhang, Syoichi Tashiro, Masahiko Mukaino and 1 others
PMID 37691497WHAT IT FOUND
ChatGPT-4 produced rehabilitation prescriptions and ICF codes for one textbook stroke case.
The 3-digit codes were reported as accurate, but one body-structure code listed the left hand instead of the right precentral gyrus, so clinician checking is needed.
Key findings
01The 3-digit ICF codes generated by ChatGPT-4 were reported as accurate, and no discrepancies were found in the ICF code inclusion criteria.
02Personal factors were included in the description even though they are not coded in the ICF.
03A body-structure code was wrong: the model listed the left upper extremity instead of the right precentral gyrus as the damaged structure.
STILL TO COME
How it was doneWhat they foundWhat it means for PTsWhat it means for OTsWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
This is one textbook case, not a real patient or a series, so it cannot show how the tool performs in ordinary clinic work. The authors state the case was straightforward and real-world acute and chronic rehabilitation is more complex and uncertain. The authors describe the rehabilitation prescriptions as general. The model made a body-structure coding error and the authors report weakness in determining chronicity and selecting approaches based on prognosis prediction. Only two PMR clinicians reviewed the output, and the paper reports comments and comparison with a textbook standard. The authors state ChatGPT-4 may generate hallucinations, its knowledge stops at 2021, and LLMs are not suitable for replacing human decision-making.
The easy way to misread this
Do not read this as evidence that ChatGPT-4 can write rehabilitation prescriptions or ICF codes without clinician review. It was one textbook case, the model made a body-structure coding error, and the authors state LLMs are not suitable for replacing human decision-making.