Researched and fact-checked in-house against primary literature and regulator records. Not reviewed by a named clinician — how we work.
Evidence-rated reference Updated August 2026
We sell nothing. No vendor sponsorship. Editorial policy
pepteyes .com

Evidence & Statistics

Mean Difference and Standardised Mean Difference

A mean difference is the between-group difference in means on the original scale, while a standardised mean difference divides it by a pooled standard deviation so different instruments can be combined.

A mean difference is subtraction: the treated group's mean outcome minus the control group's, in whatever units it was measured in. Pooled across trials using the same instrument it becomes a weighted mean difference, each trial weighted by the inverse of its variance. When trials measured one construct with different instruments, the standardised mean difference divides each trial's mean difference by a pooled within-group standard deviation, converting everything into standard deviation units. Cohen's d uses that pooled deviation directly; Hedges' g adds a small-sample correction.

The choice is usually forced by the outcome. Body weight in kilograms, or as a percentage of baseline, needs no standardising, which is why obesity meta-analyses report plain mean differences. Pain, physical function and quality of life are measured with several non-interchangeable instruments, so the standardised form is the only way to pool them, and reviews often back-transform the result using the standard deviation of one familiar instrument.

Report on the native scale wherever one exists. A mean difference of 12 kilograms can be checked against a minimal clinically important difference immediately; a standardised difference of 0.6 cannot, until someone supplies the denominator. The trade is comparability against interpretability, and standardising is a last resort.

Standardised values get compared across reviews as though the units were shared. They are not: the denominator is the standard deviation of that review's populations, so a trial with narrow eligibility reports a larger standardised effect for an identical clinical change and restrictive trials look more effective than pragmatic ones. Two further traps are mixing change-from-baseline with endpoint standard deviations in one denominator, and reading a group mean as what an individual gained.

← All 572 glossary terms