PTOtherJournal of physical therapy science2024

Evaluation of the accuracy of ChatGPT's responses to and references for clinical questions in physical therapy.

Shogo Sawamura, Takanobu Bito, Takahiro Ando and 3 others

PMID 38694019

WHAT IT FOUND

Overall, ChatGPT's answer content was rated almost correct, but 40.5% of references were fictitious.

Check every reference before using it.

Key findings

01Overall, ChatGPT's answer content was rated a median 3.0 on the study's four-point scale, while reference accuracy was rated a median 1.00.

02Only 18.9% (7/37) of appended PMIDs or DOIs matched the output, and 40.5% (15/37) of references were fictitious.

03In the musculoskeletal section, 0.0% (0/11) of references matched the output and 45.5% (5/11) were fictitious.

STILL TO COME

How it was doneWhat they foundWhat it means for PTs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The study used only five questions from each of three guideline sections, so it does not cover all physical therapy topics. It tested the free Japanese version of ChatGPT, not ChatGPT-4 or the paid version, so results may not apply to newer tools. Scoring used a rough four-point scale based on expert judgment, which may miss details such as answer complexity or depth. The same question may not get the same answer because ChatGPT's temperature can cause fluctuations. Raters were experienced in their sections, but the evaluation remained subjective and could be biased.

Declared interests

This study was funded by the Reiwa 5th Year President's Recommended Research Activity Grant of Heisei College of Health Sciences. The funder had no role in design, conduct, data interpretation, or manuscript preparation and approval. The authors declared no conflicts of interest.

The easy way to misread this

Do not conclude that ChatGPT is a reliable source for physical therapy references because its answer content was rated almost correct. The reference accuracy was poor: 40.5% of references were fictitious and only 18.9% of PMIDs or DOIs matched the output.

Read it on PubMed →