
How supplements are rated on the evidence strength scale (and how to read these ratings)
The hierarchy of evidence, GRADE, NNT, and letter grades explained based on source documents. Check what the rating really means for a supplement.
On one shelf of the pharmacy stands a product tested in dozens of randomized studies and one backed by a single cell culture experiment. The labels look similar. The difference is only visible when you know how to read where the number comes from and how strong the study design that produced it is. This text explains the concepts that frequently appear in supplement descriptions: the hierarchy of evidence, the difference between randomized and observational studies, the difference between systematic and narrative reviews, the four levels of certainty in the GRADE system, and the abbreviations NNT, NNTB, and NNTH. We have verified each in the document that introduced them, not in someone else’s summary, and for each, we provide what exactly it measures. You will also find two things that no rating includes: the size of the effect in minutes or number of people and information about who funded the study.
KEY INFORMATION
• The pyramid of evidence arranges research projects from case reports and basic studies at the bottom to randomized studies at the top, but its authors propose removing systematic reviews from the top and treating them as a tool for reading evidence (Murad et al., Evidence-Based Medicine, 2016).
• GRADE has four levels of certainty: high, moderate, low, and very low. They assess the entirety of evidence for a single question, not a single study (Balshem et al., J Clin Epidemiol, 2011).
• Certainty is lowered by five factors, not four: risk of systematic error, inconsistency, indirectness, low precision, and publication bias.
• NNT is the inverse of absolute risk reduction. NNTB and NNTH separate this number into benefit and harm (Cook and Sackett, BMJ, 1995).
• Studies funded by manufacturers more often end with results favorable to the sponsor, relative risk 1.27 (95% CI 1.17-1.37) in a review of 75 studies (Lundh et al., Cochrane, 2017).
What is the hierarchy of evidence and what does it not resolve?
The hierarchy of evidence is an ordering of research projects according to how well each deals with systematic error. It is most often depicted as a pyramid: at the bottom are basic studies and case reports, higher are case-control and cohort studies, even higher are randomized studies, and at the top are systematic reviews with meta-analysis. The authors of the review dedicated to the pyramid itself describe this arrangement while also undermining it: they propose removing systematic reviews from the top because they are tools for assessing and applying evidence, not another level of research design (Murad et al., Evidence-Based Medicine, 2016).
The reason lower tiers are lower is practical. If you observe people taking vitamin C and see they have fewer colds, you do not know whether it was the vitamin or whether people who reach for supplements usually also take care of their sleep and visit the doctor more often. An observational study cannot separate this and therefore shows a correlation, not a cause.
However, the hierarchy does not resolve two things that readers most often ask about. It does not say whether the effect is large or whether it applies to them. A meta-analysis can gather twenty studies and show a statistically significant effect that changes very little in practice. It also has its own limitation: the heterogeneity of combined studies, clinical, methodological, and statistical, can be described or reduced, but never eliminated. How the same scale looks when applied to specific substances is shown in our comparison of supplements with the strongest evidence.
What is the difference between a randomized study and an observational study?
Random assignment. In a randomized study, participants are assigned to either the treatment group or the control group by randomization, so both groups differ before treatment only by chance. The difference in outcome can then be attributed to what they received. In an observational study, the assignment is determined by the participants themselves, and then all the characteristics that prompted people to reach for it come into one group along with the supplement. Blinding on both sides adds a second layer: neither the participant nor the person measuring the outcome knows who received what.
The most costly example of this difference in the history of supplementation concerns beta-carotene. Observational studies linked high intake of carotenoids with lower lung cancer rates, so a randomized trial was conducted: 29,133 smoking men from Finland, 20 mg of beta-carotene daily, observation from five to eight years. The incidence of lung cancer in the beta-carotene group turned out to be 18 percent higher (95% CI from 3 to 36 percent), and overall mortality was 8 percent higher (95% CI from 1 to 16 percent) than in the others (ATBC Study Group, N Engl J Med, 1994). Correlation pointed one way, the experiment the other.
The table below compares three study designs along with what each cannot resolve.
| Study Design | What it controls | What it does not resolve | Example read at the source |
|---|---|---|---|
| Systematic review with meta-analysis | Random error of a single trial; searching is declared in advance | Heterogeneity of combined studies and quality of input material | Creatine and upper limb strength: 53 studies, effect size 0.317 (95% CI 0.185-0.449), Lanhers et al. 2017 |
| Randomized study | Differences between groups before treatment; with blinding also expectations | Whether the result will transfer to another population and another observation time | D-mannose and recurrent urinary tract infections: 308 women, three arms, no blinding and no placebo, Kranjčec et al. 2014 |
| Observational study | None of the above; describes people as they were found | Causality; differences between groups may explain the entire result | Carotenoids and lung cancer: the relationship reversed in the randomized trial, ATBC 1994 |
Systematic review, meta-analysis, and narrative review: what divides them?
A systematic review is conducted according to a protocol written before the start. The authors declare the question, databases, search terms, inclusion and exclusion criteria, and the method of assessing the risk of bias, and then report how many studies they found and how many they rejected at each stage. The standard for reporting such work is described in the PRISMA guideline, now in its 2021 version (Page et al., BMJ, 2021). If there is no protocol or description of the search in the publication, it is not a systematic review, even if the word review is in the title.
A meta-analysis is a step further, not a synonym. It is a statistical combination of results from studies gathered in a review into a single common value along with a confidence interval. A systematic review can end without a meta-analysis if the studies are too different to sum, and such a decision is a signal of caution, not a lack.
A narrative review has none of these safeguards. The author selects the works they consider relevant and describes them in their own words. It can be a great introduction to the topic and is often the only thing that exists for a little-studied substance, but it provides no guarantee that works inconvenient for the thesis were included at all. In practice, this is the most common type of source cited in supplement descriptions, simply labeled as research.
What do the four levels of certainty in the GRADE system mean?
GRADE assigns a rating of high, moderate, low, or very low, and it assigns it to the entirety of evidence gathered for a single question, not a single study. The rating indicates how much one can trust the effect estimate. Randomized studies start at a high level, observational studies at a low level, and then the rating moves down or up (Balshem et al., J Clin Epidemiol, 2011).
It is lowered by five factors: risk of systematic error in input studies, inconsistency of results among them, indirectness of data, low precision of estimates, and publication bias. The last of this five is most often omitted from guidelines, yet it is a separate point in methodology and has its own elaboration in the GRADE guideline series. The rating can also be raised by circumstances acting in the opposite direction, such as a very large effect or dose dependence. The entirety was first described in a short review article (Guyatt et al., BMJ, 2008).
The practical significance of this scale is simple. Moderate certainty means that further studies may shift the estimate. Low certainty means they may reverse it. A statement about promising results, which product descriptions replace the rating with, has no equivalent in GRADE and usually hides the lowest level. Moderate certainty is practically visible in the entry about what Cochrane says about N-acetylcysteine: at this very rating, the review states the number needed to treat is eight.
What do the abbreviations NNT, NNTB, and NNTH say?
NNT, or number needed to treat, is the inverse of absolute risk reduction. It answers the question of how many people need to be included in the intervention for one additional person to benefit, which would not have happened without it. The authors who popularized this measure argued that it conveys both statistical and practical significance to the physician, which relative risk does not do (Cook and Sackett, BMJ, 1995).
NNTB and NNTH are the same measure split into two directions. NNTB counts people for one additional benefit, NNTH for one additional harm. The separation makes sense where the intervention helps some and harms others, allowing both effects to be compared in one table. In the analysis of thrombolytic treatment of stroke, NNTB was 6.1 (95% CI 5.6-6.7), and NNTH was 37.5 (95% CI 34.6-40.5), which the authors summarized as a better outcome for about one in six and worse for one in thirty-five (Saver et al., Stroke, 2009).
With supplements, these numbers appear rarely, and that is why it is worth asking about them. The mere information that the difference compared to placebo was statistically significant does not say how many people would need to take the supplement for one to feel a change. Without this number, a high rating and a moderate rating sound identical to the reader.
How to read letter grades and what they do not say?
The Examine service assigns a grade from A to F, separately for each substance-effect pairing, so the same substance can have an A for one effect and a D for another. Contrary to popular interpretation, the grade does not measure the strength of evidence itself. It arises from an algorithm that combines the consistency of results across studies with the size of the measured effect and corresponds to the overall effectiveness. An A grade means many mostly consistent studies indicating at least a moderate effect, a D grade means few studies, divergent results, or a zero effect, and an F grade does not mean a lack of evidence, but rather evidence that the intervention may worsen the effect.
The second thing that the letter does not convey is the significance of the effect for a specific person. A meta-analysis of nineteen studies involving 1,683 participants found that melatonin shortens the time to fall asleep by an average of 7.06 minutes (95% CI from 4.37 to 9.75), extends sleep by 8.25 minutes, and improves its quality, and the authors themselves called these effects moderate and smaller than with sleeping pills (Ferracioli-Oda et al., PLOS ONE, 2013). The result is solid and simultaneously small, and these two things are not in conflict.
Additionally, there is the question of who funded the study. A methodological review by Cochrane covering 75 studies found that manufacturer-sponsored studies more often yield favorable efficacy results (RR 1.27; 95% CI 1.17-1.37) and more often end with favorable conclusions (RR 1.34; 95% CI 1.19-1.51), and the agreement of conclusions with their own results is lower in them (RR 0.83; 95% CI 0.70-0.98). The authors called this an industry bias that standard risk assessment does not capture (Lundh et al., Cochrane, 2017). Separate caution is required for results labeled by the authors as exploratory. These are calculations that were not planned as the main research question, so they weigh less than the main analysis, even if they sound more impressive. If a paper itself calls its result exploratory, and the product description presents it as a finding, a discrepancy occurred along the way, not in the study.
Frequently Asked Questions
What is the hierarchy of scientific evidence?
It is an ordering of research projects according to how well they deal with systematic error: from basic research and case reports at the bottom, through case-control and cohort studies, to randomized studies. The authors of the review on the pyramid itself propose to remove systematic reviews from the top and treat them as a tool for assessing evidence (Murad et al., 2016).
What is the difference between a randomized study and an observational study?
Random assignment. Randomization ensures that groups differ before treatment only by chance, so the difference in outcome can be attributed to the substance being studied. In an observational study, the assignment is determined by the participants themselves, and along with them come all the characteristics that prompted them to reach for the supplement. Therefore, observation shows a correlation, not a cause.
What do the four levels of certainty in GRADE mean?
High, moderate, low, and very low certainty indicate how much one can trust the effect estimate for the overall evidence on a single question, rather than for a single study. Randomized studies start high, observational studies low. The rating is lowered by five factors: risk of bias, inconsistency, indirectness, low precision, and publication bias (Balshem et al., 2011).
What do the abbreviations NNT, NNTB, and NNTH mean?
NNT is the inverse of absolute risk reduction, meaning the number of people receiving the intervention needed for one additional benefit. NNTB and NNTH separate this measure into benefit and harm, allowing both directions to be compared side by side. In the analysis of stroke treatment, NNTB was 6.1, and NNTH was 37.5 (Saver et al., 2009).
Does the letter grade for a supplement measure the strength of evidence?
Not only. The grades from A to F on Examine are generated from an algorithm that combines the consistency of results across studies with the size of the measured effect and correspond to the overall effectiveness for a single substance-effect pairing. An F grade does not mean a lack of evidence, but rather evidence that the intervention may worsen the effect.
Is the result of a manufacturer-sponsored study reliable?
It can be, but it requires more careful reading. A Cochrane review covering 75 studies found that manufacturer-funded studies more often yield favorable results (RR 1.27) and favorable conclusions (RR 1.34), and their conclusions are less likely to agree with their own results (RR 0.83). The authors called this a bias that standard risk assessment does not capture (Lundh et al., 2017).
The above tools are worth applying before you buy anything from the supplements category. A separate compilation of where the evidence is weakest has been gathered in our entry about which supplements are a waste of money.
This article is for informational and educational purposes and does not replace consultation with a doctor. If you are pregnant, breastfeeding, taking medications, or have chronic conditions, consult a specialist before using supplements or herbs.
Author: Michał Waluk · Published: 2026-08-09 · Updated: 2026-08-16







