"Clinically proven" and "backed by science" sit on almost every supplement label, and almost every one of them can point to at least one study. That's a low bar. The question that actually separates a real finding from a marketing sentence isn't whether a study exists: it's whether the result inside it was large enough, well-measured enough, and honestly reported enough to matter. A p-value under 0.05 is the first checkpoint most people stop at, and it's the wrong place to stop.
A p-value answers one narrow statistical question: if this supplement actually did nothing at all, how likely is it that a result this large (or larger) would show up just from random variation between the two groups? When that probability drops under 5%, researchers call the result "statistically significant" and, by convention, treat chance as an unlikely explanation.
That is genuinely useful information. It is also frequently mistaken for something it isn't. A p-value does not tell you how big the effect was. It does not tell you whether the effect matters to a real person's life. It does not tell you the study was well designed, that the comparison group was appropriate, or that the finding will hold up in anyone else's hands. A trial with thousands of participants can produce a statistically significant result for an effect so small it would never be noticed outside a lab: a fraction of a point on a lab value, for instance. Meanwhile a small, poorly controlled trial can produce an eye-catching, "significant" number that never replicates once anyone else tries to reproduce it.
The honest way to use a p-value is as a gate, not a verdict: it tells you a result is worth looking at more closely. It does not tell you what you're looking at.
Effect size is the measurement that actually answers "how much did this help?" It comes in a few common forms, and knowing which one you're looking at changes how you should read it.
| Form | What it means | Watch for |
|---|---|---|
| Percent change vs. comparator | The most common form in supplement research, e.g. "+8% strength vs. placebo." Useful only if the comparator is stated. | A percentage with no named comparator (vs. what?) is not yet a finding: it could be vs. placebo, vs. baseline, or vs. nothing. |
| Absolute vs. relative risk | Relative risk compares two rates to each other ("cuts risk by a quarter"). Absolute risk tells you the real difference across a fixed number of people. | A 25% relative risk reduction can be a 10-point absolute change or a 0.4-point one: both are true, but only one tells you what to expect for yourself. |
| Cohen's d (standardized effect size) | Expresses the difference between groups in standard-deviation units, so results from different studies and different measurement scales can be compared. Roughly: 0.2 = small, 0.5 = medium, 0.8+ = large. | A "statistically significant" result with a Cohen's d under 0.2 is real but small: often too small to notice subjectively. |
The pattern to watch for is a percentage with the comparator quietly removed. "Improves focus by 30%" sounds like a finding. Thirty percent compared to what? A placebo group, the same people's own baseline measured twice, or nothing at all: each of those changes whether that number means anything. Baseline-only comparisons are especially common and especially misleading, because people tend to improve somewhat on a second measurement regardless of what they were given, simply from paying more attention, expecting a result, or measurement variability.
Open a random handful of supplement trials and a pattern shows up fast: a lot of them enroll somewhere between ten and thirty people, often using a crossover design where every participant receives both the supplement and the placebo, in a randomized order, with a washout period between. This isn't automatically bad science. Crossover trials are efficient: each person acts as their own comparison, which reduces the noise from person-to-person variation and lets a real effect show up with fewer total participants than a trial where each person is only measured once.
The tradeoff is precision and generalizability. A twelve-person trial has wide confidence intervals around its result: the true effect could plausibly be meaningfully larger or smaller than the number reported. A handful of unusually strong or unusually weak responders can shift the group average substantially, in a way that a few hundred participants would average out. Small trials are also, as a body of methods research on publication and replication has repeatedly found, disproportionately likely to fail to reproduce when someone runs the study again at a larger scale. The SELECT vitamin E and selenium trial, discussed below, is a well-documented example of an earlier, smaller signal not holding up.
None of that makes a small trial worthless. It's the reason "an early study found X" and "a large, replicated body of evidence supports X" are different sentences, and the honest version of a claim says which one you're looking at.
A surrogate marker is a measurable stand-in for the outcome you actually care about: a blood hormone level, an inflammatory marker, a score on a lab panel. Surrogate markers are used because they're cheap and fast to measure, and a trial that tracks them can be finished in weeks instead of years. The catch is that a surrogate marker moving in the right direction does not guarantee the outcome it's supposed to predict actually improves.
The clearest everyday version of this: a supplement that "raises testosterone" in a blood test is reporting a surrogate marker. Whether that translates into more strength, more muscle, or more energy is a separate question, answered by a separate kind of study: one that measures the outcome itself (a 1-rep max, a body-composition scan, a validated fatigue questionnaire) rather than the blood value alone. Some surrogate markers are well validated stand-ins with a strong track record of predicting the real outcome. Others are measured mainly because they're convenient, and the connection to anything you'd actually notice is thinner than the headline implies. The question worth asking every time a study touts a marker is simple: did they also measure the thing I actually care about, or just the thing that was easy to measure?
How long a trial ran changes what its result can tell you. An eight-to-twelve-week trial can establish that an effect exists over that window; it cannot tell you whether the effect holds, fades with tolerance, or is outweighed by a slow-developing side effect over a year or five years. Long-term safety and long-term efficacy are both, definitionally, only answerable by long-term studies, and the large, multi-year outcome trials that can answer them are far rarer and far more expensive than the short trials that dominate the supplement literature. A result described from a 6-week trial and a result from a 5-year trial can both be genuine and still aren't interchangeable evidence for "is this safe and effective to take for the next decade."
Who paid for a study is not, on its own, proof the result is wrong. It is a fact worth knowing before weighing the finding, because funders, like anyone, tend to fund research they expect to support their own product or position, and industry-funded nutrition research has a well-documented tendency to produce more favorable results than independently funded research on the same question. Checking funding takes about two minutes:
This guide is for informational and research-literacy purposes only. It is not medical advice and does not replace a conversation with a doctor or pharmacist about a specific supplement, condition, or medication. Always tell your healthcare provider about everything you take, and get individualized guidance before starting, stopping, or changing any supplement, especially if you're pregnant, nursing, managing a chronic condition, or taking prescription medication.
Creatine monohydrate is a useful example because it survives every question above instead of failing one of them. The strength and power findings come from meta-analyses pooling more than 20 randomized controlled trials, not a single small study, with a consistently replicated effect of roughly +8% strength and +14% power output versus placebo (Rawson & Volek, 2003, Journal of Strength and Conditioning Research). The 2017 International Society of Sports Nutrition position stand concluded 3-5g/day is safe and effective for healthy adults. The outcome measured is a real one: a 1-rep max lift, total work performed, not only a surrogate marker. And the finding has been replicated across trials run by different research groups over roughly three decades.
Contrast that with a case where an early, promising signal did not hold up once it was tested properly at scale. In the late 1990s, observational studies found that men with higher blood levels of vitamin E and selenium had lower rates of prostate cancer: a real, published association. That correlation was strong enough to justify a large randomized trial, SELECT, the Selenium and Vitamin E Cancer Prevention Trial, funded by the National Cancer Institute and registered on ClinicalTrials.gov (NCT00006392) before enrollment began, eventually following more than 32,000 healthy men. When the results came in, vitamin E supplementation did not reduce prostate cancer risk; it was associated with a 17% relative increase in cases, or roughly 1.6 more diagnoses per 1,000 person-years than placebo (Klein et al., 2011, JAMA). The real outcome, the public funding, and a broad study population all pointed the same direction, and clinical guidance on vitamin E for cancer prevention changed as a result.
| Question | Creatine (strength/power) | Vitamin E (prostate cancer, SELECT) |
|---|---|---|
| Evidence level | Meta-analysis of 20+ RCTs, replicated ~30 years | Started as observational; tested directly in one large RCT |
| Sample size | Pooled across thousands of trial participants | 32,000+ in the definitive trial |
| Outcome measured | Real outcome (1-rep max, total work) | Real outcome (cancer diagnosis) |
| Comparator | Placebo | Placebo |
| Funding | Independent sports-science research bodies | National Cancer Institute (public, no product to sell) |
| What changed | Confirmed and strengthened the case for use | Reversed the earlier observational signal |
Neither outcome makes the earlier, smaller studies dishonest: the observational vitamin E data was real data, and the early creatine trials were real trials too. The difference is that one line of research kept getting tested and kept replicating, while the other got tested properly at scale and did not hold up. That is exactly the distinction a p-value alone can never show you.
The dose printed on a supplement label tells you how much of something you are taking. It does not tell you what that amount actually did in a study, and treating a dose as if it were a proven outcome is one of the most common misreadings of a label. 500 mg is a dose; it is not a measured effect. A trial might test that exact dose and find a real result, or it might test a completely different dose, in a different form, in a different population — and the label in front of you gives you none of that context on its own. The number on the front of the bottle is a starting point for research, not a conclusion you can read straight off the shelf.
Before starting any new supplement, a baseline blood panel is cheaper than guessing for six months. Instead of trying a supplement, waiting, and hoping you notice a difference, a blood test tells you where your levels actually stand before you start, which gives you something concrete to compare against later. Guessing costs money and time either way; the difference is whether you spend it with a number to check against or without one. That single step — testing before you start rather than after — turns a subjective sense of “is this working” into something you can actually measure.
Most claims that get made in an ad, a headline, or a product page don't survive past question three. The ones that make it through all seven are worth taking seriously, and they are, not coincidentally, usually the ones the people making the claim are happy to have you check for yourself.
It means the result probably wasn't random noise, at whatever threshold the study picked (usually p<0.05). It says nothing about the size of the effect. A large trial can find a statistically significant result that's too small to notice in real life, and a small trial can find a large, meaningful-looking effect that fails to replicate. Significance is a gate, not a verdict.
Effect size measures how big the difference actually is, not just whether it's likely to be real. It's usually expressed as a percent change (e.g. +8% strength vs. placebo), an absolute difference (1.6 more cases per 1,000 person-years), or a standardized measure like Cohen's d. The p-value tells you a result probably wasn't chance; the effect size tells you whether that result is worth caring about.
Small crossover trials are cheap and fast to run, and they can still detect a real effect if that effect is large and consistent. The tradeoff is precision: a 12-person trial has wide margins of error, is more vulnerable to a few outlier responders skewing the average, and is more likely to fail to replicate than a trial with hundreds of participants. It's a legitimate starting point for research, not a settled answer.
A surrogate marker is a measurable stand-in, like a blood hormone level or an inflammatory marker, used because it's cheaper and faster to track than the outcome you actually care about. A real outcome is the thing itself: strength gained, a fracture avoided, a diagnosis made. Surrogate markers can move without the real outcome changing at all, which is why a supplement that "raises testosterone" on paper doesn't automatically make you stronger.
Look up the trial's registration on ClinicalTrials.gov, which lists the sponsor. Published papers also carry a funding statement and a conflict-of-interest disclosure, usually at the end of the article or in the methods section, and both are visible on the free PubMed abstract page even when the full paper is paywalled. Industry funding doesn't automatically invalidate a result, but it's a fact worth knowing before you weigh the finding.
The same effect-size-first framework behind this guide is what scores every compound reviewed on this site, see how the scoring actually works on the methodology page, or see it applied to the single most-studied supplement on the shelf in the complete creatine guide.
A one-variable-at-a-time fillable tracker for anything you add to your routine: log the dose, log a month of sleep and energy, and find out whether it did anything.