RNOtherJournal of nursing scholarship : an official publication of Sigma Theta Tau International Honor Society of Nursing2025

Does synthetic data augmentation improve the performances of machine learning classifiers for identifying health problems in patient-nurse verbal communications in home healthcare settings?

Jihye Kim Scroggins, Maxim Topaz, Jiyoun Song and 1 others

PMID 38961517

WHAT IT FOUND

Synthetic patient-nurse conversations generated by GPT-4 did not clearly improve home healthcare health-problem detection models.

Some problems were detected better, others worse, and patient outcomes were not tested.

Key findings

01Adding synthetic conversations raised the average F1 score only from 0.62 to 0.63, and the gain was not shared by all classifiers.

02After synthetic data augmentation, F1 scores rose for skin and nutrition problems, which had low initial scores, but fell for medication regimen, which had a higher initial score.

03Training only on synthetic conversations gave average F1 scores similar to training on real conversations, with a slight decrease from 0.62 to 0.61.

STILL TO COME

How it was doneWhat they foundWhat it means for RNs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The real data came from 15 patients and 23 recordings at one home healthcare organization, so the classifier results may not apply elsewhere. The synthetic conversations were generated by GPT-4 and may not capture the complexity of real patient-nurse interactions. The study measured classifier scores, not whether nurses identified health problems better or whether patients avoided hospitalization or emergency visits.

Declared interests

The authors reported no conflicts of interest. The article is listed as NIH extramural research support.

The easy way to misread this

Do not read the F1 score gains as proof that synthetic conversations improve patient care. The study reported classifier scores from 15 patients and 23 recordings, not nurse-delivered care or patient outcomes.

Read it on PubMed →