Quality of evidence (certainty of evidence)

Studies are designed to answer important questions, but not all studies are equally reliable. Researchers therefore evaluate how much confidence they have in a studyโ€™s results. This is called the quality (or certainty) of the evidence. One of the most commonly used systems for evaluating certainty of evidence is called GRADE (Grading of Recommendations Assessment, Development, and Evaluation). Based on this system, the quality of evidence can be rated as very low, low, moderate, or high.

Very low The studyโ€™s results may be very different from what really happens
Low The studyโ€™s results may be different from what really happens
Moderate The studyโ€™s results are probably close to what really happens
High The studyโ€™s results are very likely close to what really happens

With the GRADE system, randomized controlled trials (RCTs) usually start as high-quality evidence, while observational studies usually start as low-quality evidence. The quality of evidence can then be downgraded or upgraded based on the strengths and limitations of the evidence. This helps explain why two studies can report similar results, but one is considered much more reliable than the other.

Reasons for downgrading evidence quality include:

  • Risk of Bias: Flaws in study design or execution that could introduce errors.
  • Inconsistency: Discrepancy in results across studies, indicated by varying effect sizes or significant heterogeneity.
  • Indirectness: When the study population, interventions, outcomes, or comparators do not directly match the research question.
  • Imprecision: Evidence is considered less reliable if results are based on data with wide confidence intervals or a small number of events (small sample size or small number of affected patients).
  • Publication Bias: Suspected selective publication of studies, such as only those with positive outcomes.

Reasons for upgrading evidence quality include:

  • Large Magnitude of Effect: When the intervention demonstrates a large effect.
  • Dose-Response Gradient: Finding that higher exposures to the intervention lead to increased effects.
  • All Plausible Confounding: If after considering all possible confounding factors, a change is still found, then the evidence can be upgraded.

GRADE uses a structured approach to make these ratings as objective as possible, although some judgment is still involved. As a result, different researchers may sometimes rate the same evidence differently.

Evaluating the quality of evidence is especially important in systematic reviews and meta-analyses, which combine the results of multiple studies. If many of the included studies have important limitations, the overall conclusions may be less reliable.

You can read more about GRADE in the GRADE handbook.