Movement Competency Screens Can Be Reliable In Clinical Practice By A Single Rater Using The Composite Score.
Kerry J Mann, Nicholas O'Dwyer, Michaela R Bruton and 2 others
PMID 35693862WHAT IT FOUND
The total score from this movement screen is reliable whether rated by novices or experts, but individual movement scores are not.
Do not use single-movement ratings to make clinical decisions; rely only on the composite total.
Key findings
01The composite score for all nine movements showed excellent reliability for both novice and expert raters, with no improvement from multiple viewings.
02Individual movement scores had poor to moderate reliability and did not improve when raters watched the video multiple times to focus on fewer errors.
03Expert raters detected more errors than novices, but both groups improved their error detection across repeated viewing sessions.
STILL TO COME
How it was doneWhat they foundWhat it means for PTsWhat it means for OTs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The study used video ratings rather than real-time field assessment, which may not reflect the speed and pressure of clinical practice. Individual movements were scored on a narrow three-point scale, which likely reduced reliability compared to a scale with more categories. The athletes had generally poor movement competency, leading to skewed data distributions that made it difficult to distinguish between agreement and chance. No formal training was provided to raters prior to the study, which may have impacted the baseline reliability of the novice group.
Declared interests
No conflicts of interest or funding sources were reported in the provided text.
The easy way to misread this
Do not assume that watching a patient perform a movement multiple times will improve the accuracy of your individual movement scores. The study found that multiple viewings did not increase reliability for individual tasks, likely because the three-point scoring scale is too narrow to capture subtle differences.