Enhancing Adverse Event Reporting With Clinical Language Models: Inpatient Falls.
Insook Cho, Hyunchul Park, Byeong Sun Park and 1 others
PMID 39948219WHAT IT FOUND
Fine-tuned BERT models and GPT-4 with tailored prompts could detect inpatient falls from nursing notes, and note review found 30% to 91% more unreported falls.
Key findings
01Note review identified an additional 30% to 91% of unreported fall incidents across hospitals.
02Fine-tuned BERT models outperformed GPT-4 with prompt programming in both languages.
03GPT-4 with standardised prompts performed poorly in both languages and was worse for Korean than English data.
STILL TO COME
How it was doneWhat they foundWhat it means for RNs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The models were tested on retrospective notes and reports, not real-time patient care, so the authors say further validation in real-world settings is needed. The data set was supplemented with national incident reports, which were longer and contained more clinical history than nursing notes. Performance depends on nursing documentation quality, which varies across hospitals and nurses. Cases with a self-report but no matching nursing note were excluded, accounting for 4.4% to 17.1% of cases depending on the hospital. GPT-4 with standardised prompts performed poorly in both languages.
Declared interests
The authors declared no conflicts of interest. The study was funded by the National Research Foundation of Korea and Inha University.
The easy way to misread this
Do not read the model performance as proof that AI can safely replace fall assessment or incident reporting. The study used retrospective notes and national reports, not real-time patient care, and the authors state the models need further validation in real-world settings.