Every substantive claim on this site carries an evidence grade. That page explains what the four grades mean. This one explains the machinery behind them — how you would look at a study yourself and work out how much it is worth.
It exists because grading is not a service you should have to take on trust, and because the same handful of design features explain most of the distance between what a trial found and what gets said about it afterwards. Every worked example below is a study already in this site's reference list, so you can check the reasoning against the source rather than against the argument.
1. What was actually measured?
The first question, and the one that disposes of most of the supplement literature in a sentence. A hard endpoint is something that happened to a person: death, a heart attack, a stroke, a hospital admission. A surrogate endpoint is a number that is believed to predict those things: a cholesterol level, a blood pressure, an inflammatory marker, a score on a scan.
Surrogates are used because they are faster and cheaper. They are legitimate when the link between moving the surrogate and changing the outcome has itself been demonstrated. They are treacherous otherwise, because a drug can move the number beautifully and do nothing.
So: if a study reports only that something improved a marker, the short summary is that it improved a marker. That is not nothing — it is how you decide what to test next — but it is not evidence that anyone lived longer, and this site grades it Emerging accordingly.
2. If the endpoint is composite, what is driving it?
Most cardiovascular trials report a bundle: cardiovascular death, plus non-fatal heart attack, plus stroke, plus urgent revascularisation. Bundling is not a trick — it is how you get enough events to answer the question without running for twenty years. But the components are not equally serious and they are not equally objective.
Death is unambiguous. Revascularisation is a decision made by a clinician who may know which arm the patient is in, and which can be triggered by a symptom rather than an event. A composite that looks positive can turn out to be carried almost entirely by its softest component, with no movement at all in death or infarction.
What to look for: the breakdown by component, which good papers publish. If a headline result is significant but every individual component is not, that is worth knowing before quoting the headline.
The trap in the other direction. A component that fails to reach significance on its own has usually not been shown to be unaffected — it has been shown that the trial was not large enough to answer that question separately. That is why the composite existed. Reading a null component as a negative finding is the mirror-image error and is just as common.
3. How big was the effect, in absolute terms?
Covered in more detail on the clinical page, because it matters most in a consultation. The short version: a relative risk reduction with no absolute figure alongside it is uninterpretable, and the direction of that omission is rarely accidental. Halving a risk of 1 in 10,000 is halving almost nothing.
4. Hazard ratios and confidence intervals
A hazard ratio compares the rate of events between arms over time. Below 1 favours treatment, above 1 favours control, and 1 means no difference. The number quoted is the single best estimate given the data — but the trial did not find that number, it found a range, and the range is the more informative part.
- If the interval crosses 1, the trial is compatible with no effect. It is also compatible with the point estimate, which is why "no significant difference" and "no difference" are different statements.
- Width tells you how much the trial actually pinned down. An HR of 0.99 with an interval of 0.88–1.11, as in ZEUS, is a genuinely informative null: the trial was big enough to exclude anything more than a small effect in either direction. An HR of 0.99 with an interval of 0.55–1.80 would tell you almost nothing.
- A point estimate near the edge of its own interval does not exist. If someone quotes 0.72 from a trial whose interval was 0.52–0.99, the fair reading is "somewhere between a large benefit and almost none".
5. What was it compared against?
The comparator determines what the result means, and it is the most commonly skipped question.
- Placebo or active control? Beating a placebo tells you the treatment does something. Beating the current standard tells you whether to switch. These are different claims and get reported identically.
- What dose of the comparator? An active control given at a low dose is a straw man, and this is a well-documented way of engineering a favourable result.
- Was there a run-in period? Many trials give everyone the active drug briefly and exclude those who cannot tolerate it before randomising. That is a defensible design, and it means the trial's tolerability figures do not describe the population you are treating.
- Who was excluded? Trial populations are systematically younger, healthier and on fewer other medicines than the people the results get applied to. The CTT muscle-symptom analysis is a good result honestly reported, and it still under-represents the frail and the polypharmacy patient, because the trials in it did.
6. Was it blinded, and did the blinding hold?
Blinding matters most for outcomes that involve judgement — symptoms, decisions to intervene, quality-of-life scores — and least for death.
7. Subgroups, and why they are usually noise
A trial reporting twenty subgroups will typically find one that looks significant by chance alone. That is not misconduct, it is arithmetic — and it is why a headline like "worked especially well in women over 60" should be treated as a hypothesis rather than a finding, unless it was specified in advance and the analysis tested whether the effect genuinely differs between groups rather than whether it reached significance within one.
The reverse also applies: a subgroup in which the effect was not significant has usually not been shown to be a group in which the treatment does not work.
8. Was it registered, and did the outcome change?
Trials are registered before they start, with their primary outcome declared. This exists so that a study which fails on its stated endpoint cannot quietly be rewritten around whichever measure did move. Outcome switching is common enough that checking is worthwhile, and the registration is public: the trial's identifier, usually beginning NCT, will find it.
The same logic applies to trials that never report at all. A field where the positive trials are published and the negative ones are abandoned will look more convincing than it is, which is one of the arguments for weighting large pre-registered outcome trials heavily and single positive studies lightly.
9. Who paid for it?
Industry funding is not disqualifying and treating it as such would exclude most of the evidence base for most effective drugs. Commercial sponsors run the largest, best-monitored, most-scrutinised trials in medicine.
What funding should change is where you look. Sponsored trials are more likely to use favourable comparators, favourable doses, surrogate endpoints and per-protocol rather than intention-to-treat analyses — all of which are visible in the methods if you check. The question is not "can I trust these people" but "which of the design choices above went the sponsor's way".
10. Three designs, three different questions
| Design | Answers | Main weakness |
|---|---|---|
| Observational cohort | Do people who do X have more of Y? | Confounding, and reverse causation. People who do X differ in a hundred other ways, and illness changes behaviour. |
| Randomised trial | Does doing X to these people change Y? | Short, expensive, and run on a selected population. Cannot answer questions about lifetime exposure. |
| Mendelian randomisation | Does lifelong slightly-more-X change Y? | Assumes the gene variants affect the outcome only through X. Estimates lifetime exposure, not the effect of intervening in middle age. |
They are strongest in combination and most informative when they disagree. When all three point the same way, as with apoB and cardiovascular events, the claim is about as settled as this field gets. When the observational and genetic evidence conflict, as with moderate drinking, the disagreement is itself the finding — and which one you weight is a judgement that should be stated rather than smuggled.
The short version
Find the primary endpoint, and check it is something that happened to a person
If it is a marker, the study is about a marker.
If it is a composite, find the component breakdown
And notice if the softest component is carrying it.
Read the confidence interval, not the point estimate
The width is the information. A tight null is a real finding; a wide one is an unanswered question.
Check the comparator and who was excluded
Then ask whether you resemble the people in the trial.
Treat a single positive study as a reason to look for the second one
Most findings that fail to replicate looked exactly like this at the start.
None of this requires statistical training, and none of it is about catching people out. It is about reading the same paper the press release was written from.