What the certainty rating actually means
Every paper on this site carries one of three words: moderate, low, or very low. Here is where each comes from, what it does not mean, and what it should change about how you read the summary underneath it.
Updated 9 September 2026 · Free to read
Every paper summary on this site carries one word near the top: moderate, low, or very low. It answers one question — how much weight this single study can carry on its own — and it is not a judgement about whether the treatment works.
What does each certainty level mean?
Moderate certainty means the study design can support a causal claim: a randomised trial, a systematic review, or a meta-analysis. Further research could still change the picture. Low certainty means the design can show that two things occur together but cannot establish that one caused the other — cohort studies, surveys, and narrative reviews sit here. Very low certainty means the study describes something rather than testing it: one patient, a handful of patients, a feasibility pilot, a qualitative interview study, or a protocol for research not yet done.
| Level | What it rests on | What it should change |
|---|---|---|
| Moderate | Randomised trials, systematic reviews, meta-analyses | Reasonable to let it influence practice, alongside what else you know. Still read the limitations — a randomised trial can be small, short, or done in a population unlike yours. |
| Low | Cohort studies, surveys, narrative reviews, and any paper whose design was not recorded | Useful for describing what happens and generating a question. Not grounds on its own for changing what you do to a patient. |
| Very low | Case reports, case series, pilot and feasibility studies, qualitative studies, protocols | Read it for the mechanism, the idea, or the patient's experience. Treat every number in it as illustration, not evidence. |
Where does the rating come from?
It is computed from the study design, using the design label the paper already carries in PubMed. It is not a judgement made by the language model that writes the summary, and it is not made by a person.
That is a deliberate choice. On an early sample the model disagreed with PubMed’s own design label 13% of the time, and a model that decides a case report is moderate certainty does real harm to whoever reads it between patients. So the rating is worked out from the design and the model is not consulted. The trade is that the rating is blunt: it cannot tell a careful randomised trial from a sloppy one. What it can do is never be wildly wrong in the dangerous direction.
How this differs from GRADE
If you have seen a Cochrane summary-of-findings table, you have seen GRADE certainty ratings, which use the same four words. Ours are not those, and the difference matters in one specific way.
GRADE rates a body of evidence for one outcome — every trial of that treatment, considered together, with the reviewers moving the rating up or down for five reasons: risk of bias in how the studies were run, inconsistency between their results, indirectness (the studies answer a question next to the one you asked), imprecision (confidence intervals wide enough to include both benefit and harm), and publication bias (the negative trials that were never printed).
This site rates one paper at a time, because that is what a page here is. Those five reasons cannot be applied to a single study — four of them are comparisons between studies. So the rating starts and ends with what kind of study it is. When you see moderate on a paper page, read it as “this design is capable of supporting a causal claim,” not as “a review panel weighed the evidence and landed here.”
Why is no paper rated high certainty?
Because no single study earns it. GRADE reserves high certainty for a body of consistent evidence, and the strongest thing that can appear on a page here is one meta-analysis — which, if its included trials disagree with each other, is not high certainty either. At the time of writing the corpus holds 149 summarised papers: 31 moderate, 73 low, 45 very low, and no high. The rating is capped by design, so a paper page will never show one.
The largest single group is worth knowing about. Forty-two of those 73 low papers are low because PubMed recorded no design label for them at all. Where the design is unknown, the rating defaults downward rather than guessing upward. Some of those are stronger studies than their rating suggests — the label is missing, not the rigour. It is the one place the rating is systematically unfair, and it is unfair in the safe direction.
Does low certainty mean the effect was small?
No, and this is the most common way the word is misread. Certainty is about how much the study design lets you trust the finding. Effect size is about how big the finding was. They are independent: a case report can describe an enormous change and still be very low certainty, and a large randomised trial can find almost nothing and still be moderate.
Two papers in the corpus make the point better than an explanation can.
Very low certainty · enormous effect
International Journal of Sports Physical Therapy, 2026 · case report
One 41-year-old sprinter with twenty years of hamstring activation failure showed roughly a fortyfold increase in muscle activity with real-time EMG feedback. That is a dramatic number. It is also one person, with no control condition and no way to know whether anyone else would respond — which is exactly what very low certainty is telling you.
Moderate certainty · no effect
International Journal of Sports Physical Therapy, 2026 · randomised controlled trial
Adding electrical stimulation to a PNF stretching routine produced no more flexibility than the routine alone. Nothing happened, and the moderate rating is unchanged by that — it describes the design, which was capable of detecting a difference had there been one. A well-run trial that finds nothing is a useful result, not a failed one.
Why does a study without randomisation start at low?
Because without randomisation you cannot rule out that the groups differed before the treatment did anything. Patients who receive an intervention often differ from those who do not in ways nobody measured — they were healthier, more motivated, referred earlier, seen at a better-resourced clinic. Randomising distributes those unmeasured differences evenly by chance. Nothing else does.
That does not make observational research second-rate. It makes it a different tool. A cohort study can tell you what a measure looks like across hundreds of real patients, which no trial of thirty people can.
Low certainty · directly usable
Hong Kong Journal of Occupational Therapy, 2026 · cohort study
A four-point gain on the 64-point Extended Barthel Index is the smallest improvement a patient notices; six points is what a family member notices. Low certainty, and immediately useful at a bedside — because it is not claiming a treatment works. It is establishing what a number on a scale means, and a cohort is the right design for that.
How should you use the rating?
As a reading instruction, not a verdict. It tells you how hard to lean on one paper — and the honest answer for any single paper, at any level, is “less hard than you would like to.”
- Moderate — worth acting on together with what you already know. Read the limitations section before you do.
- Low — worth knowing, worth raising with a colleague, not worth changing a plan of care over on its own.
- Very low — worth reading for the idea. If you find yourself quoting its numbers, stop and find a trial.
Every summary here also carries a limitations section and the authors’ funding and conflict-of-interest declarations, both free to read on every paper, signed in or not. Those two sections will tell you more about how far to trust a specific paper than the one-word rating ever can. The rating is where to start, not where to stop.
You can browse the corpus by design at every paper we have read, or by discipline for physical therapy, occupational therapy and speech-language pathology.
Applied Evidence summarises peer-reviewed therapy research from the full paper. Browse the index.