Veterinary Evidence Hierarchy: How to Evaluate Supplement Claims

Our Veterinary Editorial Board —

The supplement industry generates claims faster than the scientific literature can test them. “Clinically proven!” “Vet recommended!” “Backed by science!” — these phrases appear on labels and websites with little regard for what they actually mean. For veterinary professionals, the ability to evaluate these claims against the evidence hierarchy is not optional — it is a core clinical competency. This guide provides a practical framework for assessing supplement evidence, from the apex of the hierarchy (systematic reviews of RCTs) to its base (anecdotes and testimonials), with specific attention to the statistical and ethical dimensions that determine whether a claim deserves clinical trust.

The Evidence Hierarchy: A Visual Framework

Evidence-based veterinary medicine organizes research designs by their resistance to bias. From strongest to weakest:

  1. Systematic reviews and meta-analyses of multiple RCTs
  2. Randomized controlled trials (RCTs) — double-blind, placebo-controlled
  3. Cohort studies — prospective, observational
  4. Case-control studies — retrospective, observational
  5. Case series and case reports
  6. Expert opinion and mechanistic reasoning
  7. Anecdotes, testimonials, and tradition

Each level down the hierarchy introduces additional sources of bias that reduce confidence in the conclusion. The supplement industry overwhelmingly operates at levels 5-7, while marketing materials imply level 1-2 support.

Level 1-2: RCTs and Meta-Analyses

What Makes a Good RCT?

Not all RCTs are equal. A well-designed veterinary supplement RCT should have:

  • Randomization: Animals assigned to treatment and control groups by a random process, not by owner preference, clinician judgment, or convenience.
  • Blinding: Double-blind (neither owner nor assessing veterinarian knows group assignment) to prevent observer bias and placebo-by-proxy effects (owner perception influencing reporting).
  • Placebo control: The comparator should be an identical formulation without the active ingredient, not “no treatment” or “standard care alone.” Without placebo control, natural recovery and owner expectation effects cannot be separated from treatment effects.
  • Adequate sample size: Powered to detect a clinically meaningful effect. Underpowered studies produce unreliable results regardless of their p-values.
  • Pre-registered protocol: The primary endpoint, sample size calculation, and analysis plan should be registered before data collection begins (e.g., on ClinicalTrials.gov or a veterinary equivalent). This prevents “p-hacking” — testing multiple endpoints and reporting only the significant ones.
  • Target species: Conducted in dogs (or cats), not extrapolated from rodents, humans, or in vitro models.
  • Peer-reviewed publication: Published in a journal with genuine editorial oversight, not a pay-to-publish outlet with nominal “review.”

What Does a Meta-Analysis Add?

A meta-analysis pools data from multiple independent RCTs to increase statistical power and assess consistency. The 2025 canine probiotic meta-analysis (PMC12299376) is an example: by combining multiple small trials, it revealed that the pooled gut health effect was statistically nonsignificant — a conclusion invisible in any individual trial. Meta-analyses also expose heterogeneity: if studies disagree, the pooled result is less trustworthy.

Canine Supplement RCTs: The Current Landscape

Honesty requires acknowledging that level 1-2 evidence for canine supplements is rare. Most products on the market have zero RCTs in dogs. The postbiotic oral health trial (PMID: 40509062) is notable precisely because it is an exception — a properly designed, adequately powered, peer-reviewed RCT in the target species. Most “evidence” cited by supplement companies is:

  • Extrapolated from human studies (different species, different conditions)
  • Based on in vitro data (petri dish ≠ patient)
  • Derived from manufacturer-sponsored “trials” that lack blinding, randomization, or peer review
  • Anecdotal (customer reviews, social media posts)

Understanding P-Values (Without Being Fooled by Them)

The p-value is the most widely misunderstood statistic in medicine. Here is what it does and does not mean:

What a P-Value Is

The p-value is the probability of observing the measured result (or a more extreme result) if the null hypothesis were true — that is, if the supplement had no real effect. A p-value of 0.05 means: “If this supplement were truly inert, there is a 5% chance we would see a result this large or larger, purely by random variation.”

What a P-Value Is Not

  • It is not the probability that the supplement “works.”
  • It is not the probability that the result is due to chance.
  • It is not a measure of effect size. A p-value of 0.001 can accompany a clinically trivial effect if the sample size is large enough.
  • It is not a measure of study quality. A poorly designed, biased study can produce a low p-value.

The P-Value in Context: The Postbiotic VSC Trial

The 2025 postbiotic oral health RCT (PMID: 40509062) reported a 27% VSC reduction with p=0.004. This is a strong result because:

  • The p-value is well below the conventional 0.05 threshold
  • The effect size (27%) is clinically meaningful — perceptible to owners and close contacts
  • The study design (randomized, double-blind, placebo-controlled) minimizes bias
  • The endpoint (VSC concentration) is objective and quantifiable, not subjective

Contrast this with a hypothetical study reporting a 3% improvement in fecal score with p=0.04. Statistically significant? Yes. Clinically meaningful? Almost certainly not. The p-value alone does not distinguish these cases.

Effect Size: The Number That Actually Matters

Effect size quantifies the magnitude of the treatment effect, independent of sample size. Common measures include:

  • Absolute risk reduction (ARR): The difference in outcome rates between treatment and control groups.
  • Relative risk reduction (RRR): The proportional reduction. (A 50% RRR sounds impressive but may represent a 1% ARR if the baseline risk is 2%.)
  • Mean difference: For continuous outcomes (e.g., VSC concentration, fecal score), the absolute difference in means between groups.
  • Number needed to treat (NNT): How many animals must be treated for one to benefit. Lower is better.

When evaluating a supplement claim, always ask: “What is the effect size, and is it large enough to matter to my patient?” A statistically significant result with a trivial effect size is a scientific curiosity, not a clinical recommendation.

Conflict of Interest: Following the Money

Supplement research is expensive, and the entities with the most to gain from positive results — manufacturers — are often the ones funding the studies. This does not automatically invalidate industry-funded research, but it introduces systematic biases that require scrutiny:

Funding Bias

Meta-epidemiological studies in human medicine consistently show that industry-funded trials are more likely to report favorable conclusions than independently funded trials studying the same interventions (Lundh et al., 2017; PMID: 29231868). The mechanism is rarely outright fraud; more commonly, it involves:

  • Choosing favorable comparators (comparing to placebo rather than to an established effective treatment)
  • Selecting favorable endpoints (measuring a biomarker that responds rather than a clinical outcome that matters)
  • Choosing favorable populations (studying mild cases likely to improve regardless)
  • Suppressing unfavorable results (not publishing negative trials)

Authorship Bias

Check author affiliations. If the first author, corresponding author, or senior author is an employee or paid consultant of the manufacturer, the study requires additional scrutiny. This does not mean the data are fabricated — it means the study design, analysis choices, and interpretation may unconsciously (or consciously) favor the sponsor’s product.

Publication Bias

The “file drawer problem”: negative trials are less likely to be published than positive ones. If a company runs five trials and publishes only the two with favorable results, the literature presents a misleadingly positive picture. Pre-registration of trials (declaring the study exists before results are known) mitigates this, but pre-registration is rare in veterinary supplement research.

A Practical COI Checklist

  1. Who funded the study? (Check the acknowledgments and conflict-of-interest disclosure.)
  2. Are any authors employed by or consulting for the manufacturer?
  3. Did the manufacturer provide the product, and did they have input on study design or manuscript preparation?
  4. Is the journal peer-reviewed with a genuine editorial board, or is it a pay-to-publish outlet?
  5. Are there unpublished negative trials? (Search trial registries for registered studies with no corresponding publication.)

The Anecdote Problem: Why Testimonials Are Not Evidence

Customer testimonials are the most common “evidence” cited by supplement companies. They are also the least reliable. Multiple cognitive and methodological biases undermine their validity:

  • Placebo-by-proxy: Owners who believe a supplement works perceive improvement more readily. A dog’s subjective “energy level” or “coat shine” is filtered through the owner’s expectations.
  • Regression to the mean: Owners typically start supplements when symptoms are at their worst. Natural fluctuation means symptoms would likely improve regardless. The supplement gets credit for the natural recovery.
  • Confirmation bias: Owners notice and remember improvements, while discounting or forgetting non-improvement.
  • Selection bias: Satisfied customers post reviews. Dissatisfied customers often do not, or their reviews are buried. The visible review distribution is not representative.
  • Concurrent interventions: Owners who start a supplement often simultaneously change diet, increase exercise, or begin other treatments. The supplement gets credit for the combined effect.

Testimonials generate hypotheses. They cannot test them. No supplement should be recommended based on testimonials alone, regardless of how numerous or enthusiastic they are.

A Practical Evaluation Framework

When confronted with a supplement claim, apply this sequential filter:

  1. Is there an RCT in the target species? If yes, proceed to evaluate its quality. If no, the claim rests on extrapolation or anecdote — acknowledge this explicitly.
  2. Was the RCT properly designed? Randomized? Blinded? Placebo-controlled? Adequately powered? Pre-registered?
  3. Is the effect size clinically meaningful? Not just statistically significant — large enough to matter in practice?
  4. Who funded it? Industry-funded? Author conflicts? Published in a credible journal?
  5. Is the result replicated? One positive trial is suggestive. Multiple independent replications are convincing.
  6. Is the evidence specific to the product being sold? Or is it extrapolated from a different strain, species, dose, or formulation?

If a product clears all six filters, it has earned a place in evidence-based practice. If it fails at step 1 — as most do — the honest recommendation is: “The evidence is insufficient to recommend this product for this indication.”

Conclusion

The evidence hierarchy is not an academic exercise. It is a clinical tool for separating signal from noise in a market saturated with unsupported claims. RCTs in the target species, with adequate power, proper blinding, and clinically meaningful endpoints, represent the threshold for evidence-based recommendation. P-values without effect sizes are misleading. Industry funding requires scrutiny but not automatic dismissal. Testimonials are not evidence. The canine supplement market is vast, and the evidence base is thin. Veterinary professionals owe their patients and clients the honesty to say “we don’t know yet” when the evidence does not support a claim — and the rigor to recognize and recommend the rare products that do meet the standard.

References

  1. Salminen S, Collado MC, Endo A, et al. ISAPP consensus statement on postbiotics. Nat Rev Gastroenterol Hepatol. 2021;18(9):649-667. PMID: 33903774.
  2. Canine probiotic gut health meta-analysis. 2025. PMC12299376.
  3. Canine oral postbiotic RCT. 2025. PMID: 40509062.
  4. Lundh A, Lexchin J, Mintzes B, et al. Industry sponsorship and research outcome. Cochrane Database Syst Rev. 2017;2:MR000033. PMID: 28196072.
  5. Sackett DL, Rosenberg WM, Gray JA, et al. Evidence based medicine: what it is and what it isn’t. BMJ. 1996;312(7023):71-72. PMID: 8555924.
  6. Weese JS, Martin H. Assessment of commercial probiotic products for dogs and cats. Can Vet J. 2011;52(3):287-290. PMID: 21392016.

Frequently Asked Questions

What is the highest level of evidence for a supplement claim?

The highest level is a well-powered, randomized, double-blind, placebo-controlled trial (RCT) conducted in the target species (dogs, not rats or humans), published in a peer-reviewed journal, with a pre-registered protocol and declared funding sources. Systematic reviews and meta-analyses of multiple such RCTs represent the apex of the evidence hierarchy. Most canine supplements on the market have no evidence at this level.

What does a p-value actually tell me?

A p-value tells you the probability of observing the measured result (or more extreme) if the null hypothesis were true — i.e., if the supplement had no real effect. A p-value of 0.05 means there is a 5% probability the result occurred by chance alone. It does NOT tell you the magnitude of the effect, whether the effect is clinically meaningful, or the probability that the supplement “works.” Always consider effect size alongside statistical significance.

How do I spot conflict of interest in supplement research?

Check the funding disclosure and author affiliations. Red flags include: the study was funded by the product manufacturer, authors are employees or consultants of the manufacturer, the manufacturer had input on study design or manuscript preparation, results are published in a low-impact or pay-to-publish journal, and negative results are absent from the literature (publication bias). Industry-funded studies are not automatically invalid, but they require additional scrutiny of design and analysis choices.

Are customer testimonials evidence?

No. Testimonials are anecdotal reports subject to placebo-by-proxy (owner perception bias), regression to the mean (natural recovery attributed to the supplement), confirmation bias, and selection bias. They generate hypotheses but cannot test them. No supplement should be recommended based on testimonials alone, regardless of volume or enthusiasm.

What’s the difference between statistical and clinical significance?

Statistical significance (p < 0.05) means the result is unlikely due to chance alone. Clinical significance means the effect is large enough to matter in practice. A study might find a statistically significant 3% improvement in fecal score (p=0.02) that no owner would notice. Conversely, a clinically dramatic improvement in a small case series may not reach statistical significance due to insufficient sample size. Both dimensions must be considered when evaluating a claim.

Similar Posts