Is this assessment any good?
A validation study asks whether an assessment measures what it claims, in the people you are using it on. It answers with a handful of numbers, and two of them decide whether the tool is usable at all.
Written and reviewed by Reza D, OTR/L, MOT, CPACC · Published 10 September 2026 · Free to read
A validation study asks whether an assessment measures what it claims, in the people you are using it on. It answers with a handful of numbers, and two of them decide whether the tool is usable at all.
What is the difference between reliability, validity and responsiveness?
Reliability asks whether the tool gives the same answer twice. Validity asks whether it measures the thing it says it measures. Responsiveness asks whether it moves when the patient actually changes.
They fail independently, and in that order. An unreliable tool cannot be valid, because a score that varies at random cannot be measuring anything. A reliable and valid tool can still be useless for tracking progress if it does not move when your patient improves.
Two kinds of reliability, and papers often report only one
Test-retest reliability is the same rater measuring the same stable patient twice. Inter-rater reliability is two different people measuring the same patient. A tool can be excellent at the first and mediocre at the second, and which one matters depends on whether you are the only person who will ever score it.
Which numbers matter, and what counts as good?
Four statistics carry most of the weight. The thresholds below are the conventional ones, and the point of the last column is that a threshold is a convention rather than a law.
| Statistic | What it is asking | Rough reading |
|---|---|---|
| ICC | Do repeated measurements of the same patient agree? | Below 0.5 is poor and the tool is not usable for tracking. 0.5 to 0.75 is moderate. 0.75 to 0.9 is good. Above 0.9 is what you want before making a decision about one patient. |
| Cronbach's alpha | Do the items on this scale hang together as one thing? | Around 0.7 to 0.9 is the usual target. Below that the items may be measuring several different things. Above about 0.95 usually means the questions are redundant. |
| Kappa | Do two raters agree more than chance would produce? | Below 0.4 is poor, 0.4 to 0.6 moderate, 0.6 to 0.8 substantial. Used where the score is a category rather than a number. |
| SEM and MDC | How much does the score wobble on its own? | The SEM is the typical wobble. The MDC is the change you need to exceed before believing something moved. Covered in full in the MCID guide. |
The number a paper does not report is often the informative one. A validation study with no test-retest data cannot tell you whether the tool is stable. One with no minimal detectable change cannot tell you what counts as a real change, which is the question you will have first.
Is a translated assessment a validated one?
No, and this is the most common misreading in this literature. Translating an assessment produces a document in another language. Whether it still measures the same thing in the people who speak it is a separate empirical question, and the answer is often that it does not.
How do you use a validation paper in practice?
- Check the population first, not the numbers. Reliability established in healthy adolescents does not transfer to your post-surgical caseload. The numbers are properties of the tool in that sample.
- Find the minimal detectable change before you track anyone. Without it you cannot tell improvement from noise, and the guide on MCID and MDC covers what to do when the paper does not give you one.
- Check whether it was the whole tool. Papers frequently validate a subset of subtests. That says nothing about the rest.
- Look for the parts that failed. Honest validation papers report what did not work, and that is usually the sentence that changes your practice.
These papers almost all rate low or very low certainty here, which is normal for the genre rather than a mark against them. The certainty guide explains why. If the paper you are holding turns out to be a first look at a new tool rather than a validation of an established one, the guide on studies that cannot prove anything covers what that can and cannot support.
Written and reviewed by Reza D, OTR/L, MOT, CPACC. Guides are the only pages on Applied Evidence written by a person. Paper summaries are generated by software and marked as such. Last reviewed 10 September 2026.