Automating sedation state assessments using natural language processing.
Aaron Conway, Jack Li, Mohammad Goudarzi Rad and 3 others
PMID 38532639WHAT IT FOUND
Transcribed speech from 82 sedated radiology patients was sorted into pain and movement states.
The best model still made errors. This points to a documentation prompt that needs nurse confirmation, not a complete record.
Key findings
01The RoBERTa transformer pipeline had the best overall performance, with F1 0.87, precision 0.86, and recall 0.89.
02For pain, the best model had F1 0.81, precision 0.85, and recall 0.77; for movement, it had F1 0.79, precision 0.82, and recall 0.78.
03The GPT-3.5 model was less accurate than the trained pipelines, with F1 0.65, precision 0.57, and recall 0.93.
STILL TO COME
How it was doneWhat they foundWhat it means for RNs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The study was done in one interventional radiology department with 82 adults, and most were male and White, and the median age was 68, so it may not apply to other settings or patient groups. The models were tested on audio transcribed after the procedure, not on live clinical audio, so real-time accuracy is unknown. Transcription errors were not checked in detail, so imperfect transcripts may have affected the classification. The work did not test full completion of the Pediatric Sedation State Scale or another formal sedation scale. The best model still had errors for pain and movement, with pain precision 0.85 and recall 0.77 and movement precision 0.82 and recall 0.78. Participants usual language was not recorded, and non-English pain words may not have been transcribed. Most procedures used small intravenous doses of midazolam and fentanyl, so other sedation depths or drugs may produce different speech patterns.
The easy way to misread this
Do not conclude that sedation documentation can be automated now. The best model still made errors. For pain it had precision 0.85 and recall 0.77, and for movement it had precision 0.82 and recall 0.78. The study tested recorded transcripts, not real-time clinical use.