PTOTSLPSystematic ReviewJournal of neuroengineering and rehabilitation2022

Machine learning methods for functional recovery prediction and prognosis in post-stroke rehabilitation: a systematic review.

Silvia Campagnini, Chiara Arienti, Michele Patrini and 3 others

PMID 35659246

WHAT IT FOUND

Most stroke prediction models rely on internal validation, not new data, so their real-world accuracy is unproven.

High risk of bias and small sample sizes limit current tools. Clinicians cannot yet use these algorithms to confidently predict patient recovery.

Key findings

01Only three of the nineteen included studies validated their models on external, new data; the rest used only internal validation.

02The majority of models had a high or unclear risk of bias, primarily due to insufficient sample sizes and unclear outcome timing.

03Linear and logistic regressions were the most common algorithms, chosen for interpretability over more complex machine learning methods.

STILL TO COME

How it was doneWhat they foundWhat it means for PTsWhat it means for OTsWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

Most models were validated only internally, meaning their ability to generalize to new patients is unproven. A large proportion of models had a high risk of bias due to small sample sizes relative to the number of predictors. The timing of outcome measurement was often unclear, making it difficult to know if the prediction was made before or after the rehabilitation effect. Heterogeneity in outcome measures and rehabilitation protocols prevented a meta-analysis and limits direct comparison between studies.

Declared interests

The study was funded by the Ministero della Salute (Italian Ministry of Health). No other conflicts of interest were declared.

The easy way to misread this

Do not interpret the reported high accuracy metrics (such as AUCs up to 0.97) as proof that these tools work in practice. These figures come from models that were mostly tested on the same data they were trained on, which inflates performance estimates and does not reflect real-world predictive ability.

Read it on PubMed →