What a 20 percent risk reduction actually tells you
It tells you the ratio between two event rates, and nothing else. An untreated rate of 8 percent against a treated rate of 6.4 percent is a 20 percent relative reduction. An untreated rate of 0.5 percent against a treated rate of 0.4 percent is also a 20 percent relative reduction. The two situations are not remotely comparable for someone deciding whether to take a drug, yet the headline number is identical.
That is the whole problem in one sentence. Relative measures are baseline-independent by construction, and the baseline is the single most important thing you need to know. A ratio deliberately divides it out, so the figure that travels furthest through press releases and news copy is precisely the one stripped of the decision-relevant information.
The honest translation is the absolute risk reduction: control rate minus treated rate, in percentage points. One divided by that difference gives the number needed to treat, which answers the question a reader is actually asking. How many people like me have to take this, for how long, for one of us to avoid the outcome? Every example below is that same two-step calculation on a real trial.
The arithmetic worked on a real cardiovascular trial
The cleanest recent example is the semaglutide cardiovascular outcomes trial, which randomised roughly 17,600 adults with established cardiovascular disease and a body mass index of 27 or above, none with diabetes, to semaglutide 2.4 mg weekly or placebo. The primary endpoint was a composite of cardiovascular death, non-fatal myocardial infarction and non-fatal stroke; mean follow-up was about 39.8 months, so roughly 3.3 years.
The reported hazard ratio was 0.80, with a 95 percent confidence interval of about 0.72 to 0.90. That is where the 20 percent comes from. The event rates behind it were approximately 6.5 percent on semaglutide and 8.0 percent on placebo. Subtract: 1.5 percentage points over 3.3 years. Now invert. One divided by 0.015 is 66.7, so about 67 people had to take semaglutide for three and a third years for one to avoid a cardiovascular death, heart attack or stroke. Multiply 67 by 3.3 and that is roughly 220 person-years of treatment per event prevented.
None of that makes the result unimpressive. A 1.5 point reduction on a hard composite endpoint, in a population already taking statins and antiplatelets, is a real finding, and in March 2024 the FDA added a cardiovascular risk reduction indication to the semaglutide 2.4 mg label on the strength of it. The point is that 20 percent and 1.5 percentage points are the same result, and only one of them can be weighed against cost, burden and side effects.
Why the same 20 percent is worth wildly different amounts
Hold the relative reduction at 20 percent and vary only the baseline. From 40 percent you get 32 percent: an 8 point absolute reduction, a number needed to treat of about 13. From 20 percent you get 16 percent, 4 points, and 25. From 8 percent, 6.4 percent, 1.6 points, and 63. From 2 percent, 1.6 percent, 0.4 points, and 250. From 0.5 percent, 0.4 percent, a tenth of a point, and 1,000.
Same drug, same effect, same headline, and the number who must be treated to help one of them varies by a factor of nearly eighty. This is why indications are restricted by risk stratum rather than by mechanism: a treatment can be excellent in secondary prevention and close to worthless in low-risk primary prevention while producing the identical relative figure in both.
It is also the error made when a result is carried to someone unlike the people studied. The semaglutide trial enrolled people who had already had a heart attack, a stroke or symptomatic peripheral arterial disease. Someone with excess weight and no cardiovascular history has a far lower three-year event rate, so the correct operation is to multiply the 20 percent against their baseline, not the trial's. The reverse error exists too: a modest average figure understates what the drug does for the sickest enrolled subgroup, who get the largest absolute benefit from the same unchanged ratio.
Number needed to treat only means something with a clock attached
A number needed to treat is not a property of a drug. It is a property of a drug, a population, an endpoint and a time horizon, and dropping any one makes it meaningless. Sixty-seven is not semaglutide's number needed to treat; it is the figure for that composite endpoint, in that population, over about 3.3 years. Quote it without the horizon and you have invented a statistic.
Time matters because absolute risk accumulates while relative risk usually does not. If the hazard is roughly constant and events stay uncommon, doubling follow-up roughly doubles the absolute reduction and halves the number needed to treat. Over one year the semaglutide figure would be nearer 220 than 67; over ten years, if the effect persisted, far smaller. Nobody has run that trial, so the ten-year number is an extrapolation, not a finding.
This is why person-years of treatment is often the more useful unit. Roughly 220 person-years per event prevented can be set against the injections, side effects, monitoring and expenditure those same 220 years generate.
Hazard ratio, risk ratio and odds ratio are three different 20 percents
Compute all three from that trial's own numbers. The crude risk ratio is 6.5 divided by 8.0, about 0.81, so a 19 percent reduction. The odds ratio, comparing 6.5 against 93.5 with 8.0 against 92.0, comes out near 0.80, a 20 percent reduction. The reported hazard ratio, which uses the timing of every event rather than final counts alone, was 0.80. All three land within about a percentage point here.
They agree because events were uncommon. When the outcome is rare the odds ratio approximates the risk ratio well; when it is common it does not, and always looks like the bigger effect. A control rate of 40 percent against a treated rate of 32 percent is a risk ratio of 0.80 but an odds ratio of about 0.71, which a careless write-up renders as a 29 percent reduction. That gap is pure artefact.
The hazard ratio has its own subtlety: it averages the instantaneous event rate among those still at risk and assumes that ratio is roughly constant over time. When curves separate late, as they often do in cardiovascular prevention, it compresses a story in which early benefit was near zero and later benefit larger. None of the three is wrong; only the absolute event rates let you check which one a headline used.
Composite endpoints hide which component actually moved
That trial's primary endpoint bundled three events together. This is standard practice because it raises the event count and therefore the power, but it means the headline reduction is an average across components that may not have behaved alike.
Here the reduction was driven substantially by non-fatal myocardial infarction, hazard ratio in the region of 0.72. Non-fatal stroke showed essentially no separation, its confidence interval comfortably spanning 1.0. Cardiovascular death moved favourably but its interval touched or crossed 1.0, so no mortality benefit was established on that component alone. A reader assuming the 20 percent applied uniformly to all three would be wrong on two of them.
This matters across trials, because different programmes bundle different things. A composite including hospitalisation for unstable angina or revascularisation accrues more events than one restricted to death, infarction and stroke, and those softer components are more vulnerable to ascertainment differences between arms when side effects can unblind participants. Find out what went into the composite before comparing relative figures.
The same arithmetic applies to harms, and usually is not done
Number needed to harm is the mirror image: one divided by the absolute increase in risk. It is rarely reported, producing an asymmetry in which benefits arrive as inflated relative figures and harms as reassuring absolute ones, or as nothing at all.
In the semaglutide cardiovascular trial, discontinuation for adverse events ran at roughly 17 percent on drug against roughly 8 percent on placebo: an absolute difference near 9 percentage points, a number needed to harm of about twelve. Set that beside a number needed to treat of 67 over the same period and the trade-off becomes visible. For every person who avoided a cardiovascular event, roughly five stopped the drug because they could not tolerate it. The counterweight is that serious adverse events were not more common on semaglutide; they were slightly less, around 33 percent against 36 percent.
The classic demonstration outside this field is the Women's Health Initiative combined hormone therapy arm, where a 26 percent relative increase in invasive breast cancer was, in absolute terms, about 38 cases per 10,000 women per year against about 30. Eight extra cases per 10,000 per year is a number needed to harm near 1,250 a year. Both are accurate; the relative one dominated a decade of coverage.
Converting the incretin outcome trials into comparable numbers
The liraglutide cardiovascular outcomes trial randomised about 9,340 people with type 2 diabetes at high cardiovascular risk over a median of roughly 3.8 years. The primary composite occurred in about 13.0 percent against 14.9 percent on placebo, a hazard ratio of 0.87. Subtract: 1.9 percentage points, a number needed to treat of about 53. Cardiovascular death was about 4.7 percent against 6.0 percent, 1.3 points, and a number needed to treat near 77. That trial supported adding a cardiovascular indication to the liraglutide 1.8 mg label in 2017.
The earlier semaglutide trial in type 2 diabetes, a two-year pre-approval safety study of roughly 3,300 participants, reported the composite in about 6.6 percent against 8.9 percent: 2.3 points, a number needed to treat near 44 over two years, from a hazard ratio around 0.74. It reads as the strongest of the three and is also the smallest and shortest, designed to rule out harm rather than demonstrate benefit, with a correspondingly wide confidence interval. Precision and effect size are different things.
Tirzepatide is where comparison gets genuinely hard. Its cardiovascular outcomes trial in type 2 diabetes used dulaglutide as an active comparator rather than placebo and reported a hazard ratio near 0.92 with an interval reaching about 1.0. No placebo-relative absolute benefit can be read off that design, because the control arm was itself receiving a drug with an established cardiovascular effect. The placebo-controlled obesity morbidity and mortality trial is still running. Lined up, the lesson is not that one drug beats another: these trials differ in population, comparator, endpoint and duration, and converting each to an absolute reduction mainly makes visible why the comparison should not be made.
Why relative measures get reported at all
They are not a conspiracy. Relative effects are often more stable across populations than absolute ones, which makes them the sensible unit for pooling studies in a meta-analysis and for checking consistency across subgroups; a forest plot of absolute differences across trials with different baseline risks would be mostly noise. The scale also has a legitimate biological reading: if a mechanism removes a fixed proportion of a pathway's contribution to events, constancy across risk strata is what you would expect, and departures from it are informative.
The failure is omission rather than the measure itself. A trial report giving event counts in both arms lets any reader reconstruct everything: the absolute difference, the number needed to treat, both relative measures. A press release giving only the percentage lets a reader reconstruct nothing. When the control group event rate appears nowhere in a piece of coverage, that absence is itself the finding.
Reading a risk headline in ninety seconds
Find the control group event rate. If it is not in the article, look in the abstract, then the results table. If it exists nowhere, no number in the headline can be interpreted. Then subtract the treated rate from the control rate for the absolute difference in percentage points, and divide one by that decimal for the number needed to treat.
Then ask four questions. Over what period, because the number needed to treat is meaningless without one. In whom, because the enrolled population's baseline risk is what the ratio was multiplied against and may be nothing like yours. On what endpoint, and if it is a composite, which component moved. And what is the number needed to harm, computed the same way from the adverse event rates in the same table.
Doing this to the semaglutide trial turns a 20 percent headline into a usable sentence: in adults who already had cardiovascular disease and excess weight, treating about 67 for a bit over three years prevented one cardiovascular death, heart attack or stroke, while about one in twelve stopped for side effects. Longer, less quotable, and the only version that supports a decision.