Research and Review Platform
Evidence Guide

Understanding Supplement Evidence Labels

A positive study is not the same thing as strong evidence. The label should reflect how broad, consistent and directly relevant the human research actually is.

Direct answer: Strong, Moderate, Preliminary and Insufficient describe confidence in a claim—not a guaranteed individual result. Safety Concern is tracked separately because an ingredient can have efficacy evidence and still require caution.
Read our full Editorial Methodology
Evidence Guide research and evidence visual
Strong

A broader and reasonably consistent human evidence base supports the outcome.

Preliminary

There is an early signal, but replication, sample size, population or formulation limitations keep confidence low.

Insufficient

Current evidence does not justify a confident benefit claim; it does not prove that an effect is impossible.

Study design changes confidence

Randomization, controls, sample size, duration, replication and systematic review quality all affect how much weight a result deserves.

Population and formulation still matter

Even strong evidence may not transfer cleanly to a different population, extract, dose or finished product.

Why evidence labels exist

Supplement research rarely produces a simple yes-or-no answer. Trials can differ in dose, formulation, population, duration and outcome, and a single statistically significant result may not replicate. Evidence labels provide a compact way to express how confident the site is that a specific claim is supported by relevant human evidence.

Strong evidence

A Strong label is reserved for claims supported by a comparatively broad, consistent and directly relevant human evidence base. Replication matters. High-quality systematic reviews and meta-analyses can strengthen the case when the underlying studies are appropriate and reasonably consistent. Strong does not mean every user will benefit, and it does not turn a population-level effect into an individualized recommendation.

Moderate evidence

Moderate evidence supports a cautious conclusion but still has meaningful limitations. The number of trials may be smaller, results may show some inconsistency, the effect may depend on a particular population, or the evidence may not cover all commonly sold formulations. Moderate is intentionally different from both Strong and Preliminary: there is enough support to take the signal seriously, but not enough to present it as settled across contexts.

Preliminary evidence

Preliminary evidence means human research has produced an interesting signal, but confidence remains low. This can occur when studies are small, short, industry-specific, limited to one extract, conducted in an unusual population or not replicated. Preliminary does not mean useless or disproven. It means the appropriate language is 'may,' 'suggests' or 'early evidence' rather than a confident performance or hormone claim.

Insufficient evidence

Insufficient means the available research does not justify a confident benefit claim. There may be too little human evidence, studies may conflict, the relevant outcome may not have been measured, or controlled trials may fail to show the marketed effect. Importantly, insufficient evidence is not proof that an effect is impossible; it is a statement about what the current evidence can support.

Safety concern is a separate label

Efficacy and safety should not be compressed into one score. An ingredient can have good evidence for an outcome while still creating meaningful interaction or dose concerns. Conversely, an ingredient can appear relatively well tolerated while having little evidence of benefit. Safety Concern is therefore tracked separately from Strong, Moderate, Preliminary or Insufficient efficacy labels.

Study design and replication

Randomized controlled trials usually provide stronger causal evidence than uncontrolled observations, but study quality still matters. Sample size, allocation, blinding, missing data, outcome selection and statistical analysis can all affect confidence. Replication across independent research groups is especially valuable because one positive trial can be an outlier.

Population, dose and formulation

Evidence does not automatically transfer from one population to another. A result in deficient men may not apply to men with adequate nutrient status. A branded botanical extract may not represent every powder sold under the same ingredient name. A study dose provides research context, not a universal dose recommendation. These factors can lower the confidence assigned to a broad consumer claim.

Systematic reviews are not automatically Strong

A systematic review or meta-analysis is powerful when it combines good studies that address the same question. It can still be limited when trials are small, heterogeneous or at high risk of bias. Pooling several weak studies does not automatically create strong evidence. The underlying evidence remains important.

How labels change over time

Evidence labels are provisional in the scientific sense: they can change as better evidence appears. A promising ingredient can move from Preliminary toward Moderate if larger replicated trials support it. A once-popular claim can weaken if subsequent controlled studies fail to reproduce the effect. Updating a label is therefore a sign that the methodology is functioning, not that evidence should never change.

Examples from the current Ingredient Library

Creatine monohydrate illustrates a Strong performance evidence position because resistance-training outcomes have been studied extensively. Maca illustrates why outcomes must remain separate: human research can support a preliminary libido signal without establishing a testosterone-raising effect. D-aspartic acid shows the importance of controlled negative evidence, because popularity does not override trials that fail to show a reliable testosterone increase.

Why mechanisms and animal data have a supporting role

Mechanistic and animal research can explain why an effect is plausible and can help generate hypotheses, but it does not establish a consumer benefit in humans. The site may discuss mechanism where useful, yet a Strong human-efficacy label requires relevant human outcome evidence. This prevents plausible biology from being presented as though the clinical result has already been demonstrated.

Directness of evidence matters

Evidence can be high quality yet indirect for the exact claim being evaluated. A trial in women, a study in severely deficient participants, or an outcome limited to a laboratory biomarker may not directly answer a claim about strength or testosterone in healthy men. Directness therefore influences the label alongside study quality.

Statistical significance is not the whole decision

A statistically significant result can still be small, imprecise or of uncertain practical value. Evidence grading should consider effect size, confidence intervals and whether the measured outcome matters to the user. The site should not equate a p-value below a conventional threshold with a meaningful consumer benefit.

Frequently asked questions

Does Strong evidence guarantee a benefit?

No. It indicates stronger population-level evidence, not a guaranteed individual response.

Is Preliminary the same as ineffective?

No. It means the signal needs stronger or more consistent confirmation.

Can an ingredient have strong efficacy evidence and still raise safety concerns?

Yes. Safety is evaluated separately.

Can evidence labels change?

Yes. They should change when better evidence materially shifts the balance of support.

Sources

Authoritative sources and operational records supporting the page.