Machine learning methods for functional recovery prediction and prognosis in post-stroke rehabilitation: a systematic review.
Silvia Campagnini, Chiara Arienti, Michele Patrini and 3 others
PMID 35659246WHAT IT FOUND
Most stroke prediction models rely on internal validation, not new data, so their real-world accuracy is unproven.
High risk of bias and small sample sizes limit current tools. Clinicians cannot yet use these algorithms to confidently predict patient recovery.
Key findings
01Only three of the nineteen included studies validated their models on external, new data; the rest used only internal validation.
02The majority of models had a high or unclear risk of bias, primarily due to insufficient sample sizes and unclear outcome timing.
03Linear and logistic regressions were the most common algorithms, chosen for interpretability over more complex machine learning methods.
STILL TO COME
How it was doneWhat they foundWhat it means for PTsWhat it means for OTsWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
Most models were validated only internally, meaning their ability to generalize to new patients is unproven. A large proportion of models had a high risk of bias due to small sample sizes relative to the number of predictors. The timing of outcome measurement was often unclear, making it difficult to know if the prediction was made before or after the rehabilitation effect. Heterogeneity in outcome measures and rehabilitation protocols prevented a meta-analysis and limits direct comparison between studies.
Declared interests
The study was funded by the Ministero della Salute (Italian Ministry of Health). No other conflicts of interest were declared.
The easy way to misread this
Do not interpret the reported high accuracy metrics (such as AUCs up to 0.97) as proof that these tools work in practice. These figures come from models that were mostly tested on the same data they were trained on, which inflates performance estimates and does not reflect real-world predictive ability.