Applied Evidence
Guide six of seven

Is this assessment any good?

A validation study asks whether an assessment measures what it claims, in the people you are using it on. It answers with a handful of numbers, and two of them decide whether the tool is usable at all.

Written and reviewed by Reza D, OTR/L, MOT, CPACC · Published 10 September 2026 · Free to read

A validation study asks whether an assessment measures what it claims, in the people you are using it on. It answers with a handful of numbers, and two of them decide whether the tool is usable at all.

What is the difference between reliability, validity and responsiveness?

Reliability asks whether the tool gives the same answer twice. Validity asks whether it measures the thing it says it measures. Responsiveness asks whether it moves when the patient actually changes.

They fail independently, and in that order. An unreliable tool cannot be valid, because a score that varies at random cannot be measuring anything. A reliable and valid tool can still be useless for tracking progress if it does not move when your patient improves.

Two kinds of reliability, and papers often report only one

Test-retest reliability is the same rater measuring the same stable patient twice. Inter-rater reliability is two different people measuring the same patient. A tool can be excellent at the first and mediocre at the second, and which one matters depends on whether you are the only person who will ever score it.

Worked example
One paper, two very different numbers · free to readReliability and concurrent validity of the Likert-type stuttering severity scale and the visual analog scale in Turkish-speaking adults who stutterInternational Journal of Language & Communication Disorders, 2026 · cohort studyThe same paper reports test-retest agreement of 0.997 for the numerical scale and inter-rater agreement of 0.573. Read the first number alone and the tool looks near-perfect. Read both and the real finding appears: it is highly consistent when one person uses it, and considerably less so when the person changes. That matters if your service hands patients between clinicians.Read the summary

Which numbers matter, and what counts as good?

Four statistics carry most of the weight. The thresholds below are the conventional ones, and the point of the last column is that a threshold is a convention rather than a law.

Comparison
The four numbers a validation paper usually reports.
StatisticWhat it is askingRough reading
ICCDo repeated measurements of the same patient agree?Below 0.5 is poor and the tool is not usable for tracking. 0.5 to 0.75 is moderate. 0.75 to 0.9 is good. Above 0.9 is what you want before making a decision about one patient.
Cronbach's alphaDo the items on this scale hang together as one thing?Around 0.7 to 0.9 is the usual target. Below that the items may be measuring several different things. Above about 0.95 usually means the questions are redundant.
KappaDo two raters agree more than chance would produce?Below 0.4 is poor, 0.4 to 0.6 moderate, 0.6 to 0.8 substantial. Used where the score is a category rather than a number.
SEM and MDCHow much does the score wobble on its own?The SEM is the typical wobble. The MDC is the change you need to exceed before believing something moved. Covered in full in the MCID guide.

The number a paper does not report is often the informative one. A validation study with no test-retest data cannot tell you whether the tool is stable. One with no minimal detectable change cannot tell you what counts as a real change, which is the question you will have first.

Is a translated assessment a validated one?

No, and this is the most common misreading in this literature. Translating an assessment produces a document in another language. Whether it still measures the same thing in the people who speak it is a separate empirical question, and the answer is often that it does not.

Worked example
The test did not survive the crossing · occupational therapyPsychometric characteristics and cultural interpretation of LOTCA visual subtests among adults in TaiwanHong Kong Journal of Occupational Therapy, 2026 · cohort studyIn 101 Taiwanese adults with no cognitive impairment, the LOTCA visual subtests did not hold together as the manual describes, and the original domain structure did not reproduce. The authors' reading is the useful part: a low score may mean the person did not recognise the pictured object, not that their thinking is impaired. Scoring an unvalidated translation as though it were the original produces a diagnosis out of a cultural mismatch.Read the summary

How do you use a validation paper in practice?

  • Check the population first, not the numbers. Reliability established in healthy adolescents does not transfer to your post-surgical caseload. The numbers are properties of the tool in that sample.
  • Find the minimal detectable change before you track anyone. Without it you cannot tell improvement from noise, and the guide on MCID and MDC covers what to do when the paper does not give you one.
  • Check whether it was the whole tool. Papers frequently validate a subset of subtests. That says nothing about the rest.
  • Look for the parts that failed. Honest validation papers report what did not work, and that is usually the sentence that changes your practice.
Worked example
Half the tool worked · physical therapyShoulder internal and external rotator strength reliability and validity in healthy adolescent athletesInternational Journal of Sports Physical Therapy, 2026 · cohort studyHandheld dynamometry agreed well with the reference device for internal rotation, prone, and poorly for external rotation in every position tested. A clinician recording both and comparing them week to week is reading one real number and one unreliable one off the same sheet, with nothing on the page to say which is which.Read the summary

These papers almost all rate low or very low certainty here, which is normal for the genre rather than a mark against them. The certainty guide explains why. If the paper you are holding turns out to be a first look at a new tool rather than a validation of an established one, the guide on studies that cannot prove anything covers what that can and cannot support.


Other guides
01What the certainty rating on a research summary actually meansEvery paper on this site carries one of three words: moderate, low, or very low. Here is where each comes from, what it does not mean, and what it should change about how you read the summary underneath it.02MCID vs MDC: telling whether a change on an outcome measure is realTwo numbers decide whether a score has moved: one asks whether the change is bigger than the measurement error, the other whether it is big enough for the patient to care. They are not the same, and one is useless without the other.03Pilot, feasibility, protocol, case report: what a study that cannot prove anything still tells youA pilot, a feasibility study, a case report and a protocol each answer a real question. None of those questions is whether the treatment works — and reading them as if it were is the commonest way to misread the literature.04Scoping review, systematic review, meta-analysis: which one answers a clinical questionOnly one of them is designed to answer whether a treatment works. Telling them apart takes about ten seconds, and it changes how much weight the paper can carry.05What a qualitative study can tell you, and what it cannotIt does not measure whether a treatment works. It finds out what something is like, in enough detail that you recognise it. That is a different question, with answers a trial cannot produce.07"No significant difference" is not "no difference"It means the study did not find an effect. It does not mean there is no effect. Those are different claims, and only one of them is usually supported by the paper in front of you.

Written and reviewed by Reza D, OTR/L, MOT, CPACC. Guides are the only pages on Applied Evidence written by a person. Paper summaries are generated by software and marked as such. Last reviewed 10 September 2026.