MedSysEvidence-graded reference

MedSys / How we grade evidence

How we grade evidence

Every substantive claim on this site carries one of four grades. They exist because "there is evidence for this" is nearly meaningless on its own — the useful question is how much, of what kind, and pointing which way.

Strong

Consistent randomised trials, meta-analyses of them, or Mendelian randomisation across many independent genetic variants. Effects that hold across different populations and that scale with dose or with duration of exposure.

What to do with it. Act on these. Where a strong-grade claim conflicts with something you have read elsewhere, the other source is usually the problem.

On this siteThat lowering apoB lowers cardiovascular events — trials, genetics and epidemiology all agree, and the effect scales with both size and duration of the reduction.

Moderate

Observational studies that agree with one another, supported by a plausible mechanism and usually a dose-response relationship, but without randomised confirmation. Residual confounding cannot be excluded.

What to do with it. Probably real. Worth acting on where the action is cheap, safe and has other reasons to recommend it anyway.

On this siteThat an extra-virgin olive oil-led dietary pattern reduces cardiovascular events — large trial support exists but the trial has been reanalysed, and the observational base carries the weight.

Emerging

Early human data, small studies, surrogate endpoints rather than outcomes, or animal work with a convincing mechanism. Often exactly the stage at which findings are loudest in the press and weakest on paper.

What to do with it. Interesting, not actionable. If something here turns out to matter, the grade will change and the page will say when.

On this siteThat urolithin A meaningfully extends healthspan in humans — the mitochondrial mechanism is real, the human outcome data are not there.

Contested

Competent researchers looking at the same evidence reach different conclusions, and the disagreement is about interpretation rather than about who has read more. Trials pointing in opposite directions land here, not in a lower grade.

What to do with it. Read both positions. This site gives both and no verdict, because a verdict would be an invention rather than a summary.

On this siteThat lowering inflammation reduces cardiovascular events — CANTOS said yes, CLEAR SYNERGY and ZEUS said no, and hs-CRP remains a good predictor either way.

Why grade at all

The failure mode of health writing is a uniform confident voice applied to claims of wildly different quality. That apoB lowering reduces cardiovascular events is about as settled as cardiovascular medicine gets. That urolithin A extends healthspan in humans is a hypothesis with some encouraging mitochondrial data. Presenting both in the same tone is how readers end up unable to tell the difference — and how they end up either dismissing everything or believing everything.

What a grade is not

It is not a measure of how large an effect is. A Strong intervention can have a small effect; an Emerging one could turn out to be substantial. The grade describes how confident we can be that the effect exists, nothing more. Effect sizes are given separately wherever they are known.

It is also not a recommendation. Grades describe evidence; whether something suits you depends on your situation, your other medicines and your own judgement in conversation with your GP.

When a grade changes

Grades move in both directions, and downgrades are more common than the enthusiasm around new compounds suggests. When a grade changes, the page it sits on gets a new review date.

A worked example, from this August. Raised hs-CRP predicts cardiovascular events even when cholesterol is well controlled — that much is Strong and unchanged. Whether lowering inflammation helps was carried at Moderate here on the strength of CANTOS and two colchicine trials. Then CLEAR SYNERGY (2025) randomised 7,062 people after a heart attack and found nothing, and ZEUS (July 2026) inhibited IL-6 directly in over 6,300 people, hit every biomarker target, and returned a hazard ratio of 0.99. Two large trials pointing the opposite way to the earlier ones is the definition of Contested, so that is what the therapeutic claim now carries — while the predictive claim keeps its Strong. The same biomarker, two different questions, two different grades.

Where the guidelines sit in all this

National guidelines are not a fifth grade. They are a synthesis of the same evidence, filtered through a health system's priorities and constraints — which is why the UK, Europe and the US can read the same trials and land in different places. Where the guidelines stand sets out the current position in each, and why the numbers do not transfer between them.

How to arrive at a grade yourself

Grading should not be something you have to take on trust. The design features that decide how much a study is worth — surrogate against hard endpoints, what is driving a composite, how wide the confidence interval is, what the comparator was — are set out with worked examples from this site's own reference list on how to read a trial.

What the grades are based on

The primary sources behind the graded claims on this site are listed on one page, grouped by topic, each with what it actually supports and how far it can be pushed. The reference list.

Found something graded wrongly? That is a useful thing to hear. Grading is a judgement and judgements can be mistaken.