Effect Size
Effect size is the magnitude of a difference or an association, expressed on a scale that does not depend on how many participants were studied, unlike a p-value.
Effect size answers how much, where the p-value answers how surprising. Unstandardised effect sizes stay in the units of the outcome, kilograms or millimetres of mercury or percentage points. Ratio measures such as the risk ratio and hazard ratio express one group's outcome as a multiple of another's. Standardised effect sizes divide a mean difference by a pooled standard deviation, giving Cohen's d, or Hedges' g with a small-sample correction; Cohen's labels of 0.2, 0.5 and 0.8 were offered for fields with no meaningful scale, and he cautioned against leaning on them.
The unstandardised version is usually the informative one. In STEP 1, 68 weeks of weekly semaglutide 2.4 mg against placebo in adults with obesity produced a mean body weight change of about minus 15 percent versus minus 2 percent, a gap of roughly 12 percentage points, or about 12 kg for a 100 kg participant. Divided by a pooled standard deviation it becomes a figure that is harder to picture and easier to misuse.
Effect size and p-value are not substitutes. With 15,000 participants a difference too small to matter will be highly significant; with 30 a large one may not reach significance. Sample size changes the precision of an estimate, not its expected magnitude, so the honest report is an effect size with its confidence interval, checked against the minimal clinically important difference.
Standardised effect sizes are inflated by narrow eligibility, because a homogeneous sample has a small standard deviation and the same clinical change becomes a larger d, so comparing d across trials compares different denominators. The other predictable distortion is the winner's curse: in a small study only a large observed effect clears significance, so the earliest published estimate is systematically the largest and should be expected to shrink on replication.
Worked example — what sample size buys
The 95% interval half-width is 1.96·σ/√n. Because n sits under a square root, precision is bought slowly: going from 50 to 200 participants per arm halves the interval, and you need 800 to halve it again.
Every panel is redrawn from its own equation by scripts/glossary-figures.js — no traced or stock artwork, and a rebuild is byte-identical.