RNCohortJournal of nursing scholarship : an official publication of Sigma Theta Tau International Honor Society of Nursing2025

Automating sedation state assessments using natural language processing.

Aaron Conway, Jack Li, Mohammad Goudarzi Rad and 3 others

PMID 38532639

WHAT IT FOUND

Transcribed speech from 82 sedated radiology patients was sorted into pain and movement states.

The best model still made errors. This points to a documentation prompt that needs nurse confirmation, not a complete record.

Key findings

01The RoBERTa transformer pipeline had the best overall performance, with F1 0.87, precision 0.86, and recall 0.89.

02For pain, the best model had F1 0.81, precision 0.85, and recall 0.77; for movement, it had F1 0.79, precision 0.82, and recall 0.78.

03The GPT-3.5 model was less accurate than the trained pipelines, with F1 0.65, precision 0.57, and recall 0.93.

STILL TO COME

How it was doneWhat they foundWhat it means for RNs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The study was done in one interventional radiology department with 82 adults, and most were male and White, and the median age was 68, so it may not apply to other settings or patient groups. The models were tested on audio transcribed after the procedure, not on live clinical audio, so real-time accuracy is unknown. Transcription errors were not checked in detail, so imperfect transcripts may have affected the classification. The work did not test full completion of the Pediatric Sedation State Scale or another formal sedation scale. The best model still had errors for pain and movement, with pain precision 0.85 and recall 0.77 and movement precision 0.82 and recall 0.78. Participants usual language was not recorded, and non-English pain words may not have been transcribed. Most procedures used small intravenous doses of midazolam and fentanyl, so other sedation depths or drugs may produce different speech patterns.

The easy way to misread this

Do not conclude that sedation documentation can be automated now. The best model still made errors. For pain it had precision 0.85 and recall 0.77, and for movement it had precision 0.82 and recall 0.78. The study tested recorded transcripts, not real-time clinical use.

Read it on PubMed →