RNCohortAmerican journal of critical care : an official publication, American Association of Critical-Care Nurses2018

Advancing In-Hospital Clinical Deterioration Prediction Models.

Alvin D Jeffery, Mary S Dietrich, Daniel Fabbri and 4 others

PMID 30173171

WHAT IT FOUND

Four models for predicting in-hospital cardiopulmonary arrest were similarly good at ranking risk, but poor at balancing true and false alarms.

Machine learning models flagged more patients and generated more alarms, and the study did not test bedside use.

Key findings

01Four models for predicting in-hospital cardiopulmonary arrest had similar discrimination (AUROC 0.847 to 0.861) but poor and variable balance of true and false alarms (F1 0.170 to 0.325).

02Machine learning models flagged more patients and had higher sensitivity than logistic and Cox regression models at cardiopulmonary arrest probability thresholds from 0.006 to 0.12.

03ICD-9 codes for respiratory, circulatory, genitourinary, endocrine, and symptom diagnoses or procedures were among the 10 most important variables in all four models.

STILL TO COME

How it was doneWhat they foundWhat it means for RNs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The total number of patients analysed is not given in the article text, so the sample size cannot be checked. The study used only first-day admission data, so it may miss information that becomes available later during hospitalization. There was a large amount of missing data, about 40% for laboratory values and 60% for vital signs, and the models were built after imputation. The outcome was identified using CPT codes, which may not perfectly capture clinical events. The models included ICD-9 codes that would not be available for a real-time clinical decision support tool. The study was retrospective and used cross-sectional first-day data, so it did not show how clinicians would use the alerts in real time. Machine learning models produced more alarms, which could burden clinicians. Validation used held-out data from the same retrospective database, not external hospitals.

Declared interests

No author conflict statement appears in the article text. The article is listed as supported by NIH extramural research.

The easy way to misread this

Do not read the AUROC values of 0.847 to 0.861 as evidence that these models are ready for bedside use. They were developed retrospectively from electronic records with large amounts of missing data, had F1 scores of 0.170 to 0.325, and were not tested prospectively for effect on patient care or outcomes.

Read it on PubMed →