Evaluating Large Language Model-Assisted Emergency Triage: A Comparison of Acuity Assessments by GPT-4 and Medical Experts.
Gal Ben Haim, Mor Saban, Yiftach Barash and 6 others
PMID 39610042WHAT IT FOUND
GPT-4 assigned emergency triage scores that were more urgent than senior nurses and a physician did, median 2.0 versus 3.0, and its agreement with the physician was 0.10.
This does not support replacing nurse triage judgement.
Key findings
01GPT-4 assigned a lower median ESI score of 2.0 than the human evaluators, who mostly assigned 3.0 (p < 0.001).
02Agreement with the senior physician was 0.30, 0.20 and 0.28 for the three nurses, but only 0.10 for GPT-4.
03Older patients tended to receive lower ESI scores, which corresponds to higher urgency, and this age pattern was strongest for GPT-4 (correlation -0.38).
STILL TO COME
How it was doneWhat they foundWhat it means for RNs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
Only 100 patients from one emergency department on one day were studied, so the results may not generalise to other emergency departments. GPT-4 was tested with one prompt and the OpenAI web interface, without tuning model settings, so a different prompt or model configuration could change the scores. Only one large language model, GPT-4, was evaluated, so the findings may not apply to other AI tools. The nurses' scores varied, with agreement values of 0.30, 0.20 and 0.28 against the senior physician, despite ESI as a standardised triage system. The study compared ESI assignments, not patient outcomes, so it does not show clinical benefit.
The easy way to misread this
Do not read GPT-4's lower ESI scores as evidence that AI triage is safer or more accurate. In this study, GPT-4 assigned a median score of 2.0 while human evaluators assigned 3.0 (p < 0.001), and its agreement with the senior physician was 0.10, so it did not match expert assessment.