Hierarchy of Evidence
The hierarchy of evidence ranks study designs by how exposed they are to bias when answering a treatment question, from systematic reviews of randomised trials down to mechanistic reasoning and opinion.
The pyramid orders designs by structural vulnerability to confounding and selection effects: systematic reviews of randomised trials at the top, then individual trials, prospective cohorts, case-control studies, cross-sectional surveys, case series and case reports, and at the base mechanistic reasoning, in-vitro work, animal models and expert opinion. It ranks designs, not studies: what each design protects you from, not how competently any example was executed.
The modern refinement is that design sets a starting point, not a verdict. GRADE begins randomised trials at high certainty and observational studies at low, then moves each up or down, so a small unblinded trial with a surrogate endpoint can finish below a large well-conducted cohort. It is also question-specific, having been built for therapy questions: for prognosis the cohort is the right design, for diagnostic accuracy a comparison against a reference standard, and for rare or delayed harms no feasible trial is large enough, so registries and post-marketing surveillance carry the load.
Practically, it tells you what a citation can support before you read it. A claim resting only on in-vitro data has not been tested in an organism; one resting on an uncontrolled case series has not been separated from natural history, regression to the mean or placebo response. Moving up a level costs years and money, which is why many compounds sit permanently at the base while being discussed as though near the top.
That gap is where most peptide marketing lives: clinical studies show, attached to a cell-culture assay, an eight-person open-label series or a rodent model. The mirror error is using the hierarchy to dismiss everything below a randomised trial: an anecdote cannot establish efficacy, but a case series can raise a real safety signal, and often that is where the first one appears.
Worked examples — what a censored literature looks like
Both panels start from the same 46 simulated trials drawn at their own standard errors around a true risk ratio of 0.82. The second panel simply withholds the small studies that came out null or unfavourable — exactly what publication bias does — and the funnel goes lopsided.
Every panel is redrawn from its own equation by scripts/glossary-figures.js — no traced or stock artwork, and a rebuild is byte-identical.