PTOTSLPOtherJMIR rehabilitation and assistive technologies2026

Can GPT-5 Support Licensing Examination Preparation? Analysis of Accuracy, Reasoning, and Semantic Similarity Across Rehabilitation Disciplines.

Christy Muasher-Kerwin, M Courtney Hughes, Aida Sanatizadeh

PMID 42166754

WHAT IT FOUND

GPT-5 answered 91.0% of PT, 79.0% of OT, and 83.0% of SLP board-preparation questions correctly, but its explanations were less similar to expert rationales for inferential and evaluative reasoning.

Use it as a supervised study aid, not a replacement for expert reasoning.

Key findings

01GPT-5 answered 91.0% of PT, 79.0% of OT, and 83.0% of SLP questions correctly across 100 items per discipline.

02GPT-5 rationales had mean semantic similarity 0.707 to source rationales, with deductive reasoning highest at 0.730 and evaluative and inferential reasoning lower at 0.671 and 0.685 (P = .02 across reasoning types).

03Only PT deductive reasoning reached 100% accuracy, with 22 of 22 correct.

STILL TO COME

How it was doneWhat they foundWhat it means for PTsWhat it means for OTsWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The questions came from board-preparation materials, not official licensing examinations, so results do not represent live exam performance. Difficulty indices and blueprint mappings were unavailable, so items were not stratified. The study did not measure learner outcomes. Only 100 questions per discipline were used, and the sources are not publicly available. Qualitative coding used consensus review, not formal interrater reliability statistics.

Declared interests

No funding was provided and no conflicts were declared.

The easy way to misread this

Do not conclude that GPT-5 can replace supervised exam preparation or expert rationale checking. The study used board-preparation questions, not official licensing exams, did not measure learner outcomes, and found lower alignment for inferential and evaluative reasoning.

Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →