Development and validation of a stroke risk prediction model using regional healthcare big data and machine learning.
Yunxia Duan, Rui Wang, Yumei Sun and 6 others
PMID 41367592WHAT IT FOUND
A machine-learning stroke risk model beat a standard model in 92,172 Chinese adults; hypertension, age, and diabetes were top factors, but it was not tested outside its original records.
Key findings
01Random Forest had the highest validation AUC at 0.980 and outperformed logistic regression.
02The final model used 13 factors, and hypertension, age, and diabetes were the top-ranked predictors.
03A predicted risk score of 0.789 or higher was proposed as the high-risk cutoff.
STILL TO COME
How it was doneWhat they foundWhat it means for RNs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The model was developed and validated only on one regional Chinese health-record dataset, so it may not work in other populations or health systems. Only 436 incident strokes were identified among 92,172 participants, and the dataset was balanced by generating synthetic stroke cases and reducing non-stroke cases. Random Forest had poorer calibration than logistic regression, decision tree and neural network models in validation, so its predicted probabilities may not match observed risk. The study reported model performance, not evidence that nurses using the model changed care or prevented strokes. The model was described as a black box, and SHAP interpretability analysis was not performed because of privacy, computational and latency constraints.
Declared interests
The study was funded by the Beijing Natural Science Foundation - Haidian Original Innovation Joint Fund (Grant No. L222103) and the National Natural Science Foundation of China (Grant No. 72174012). The authors declared no conflict of interest.
The easy way to misread this
Do not conclude that this model is ready to guide stroke prevention in nursing practice. It was not tested in an independent cohort, and the Random Forest model had poorer calibration than logistic regression, decision tree and neural network models.