Multiplicity and Subgroup Analysis
Multiplicity is the inflation of false-positive risk created by testing many outcomes, subgroups or time points, and subgroup analysis is the form of it most often mistaken for a finding.
Every extra hypothesis test is another chance to cross a threshold by luck: twenty independent tests at a two-sided alpha of 0.05 leave roughly a 64 percent chance that at least one comes out significant when nothing is happening. Bonferroni splits alpha evenly across the family; fixed-sequence testing spends it all on the primary endpoint and moves down a ranked list only while each test succeeds; alpha spending allocates a share of the total error rate to each interim look. False discovery rate procedures instead control the expected proportion of false positives among the results called significant.
Subgroups multiply fast: a trial reporting effects by sex, age band, baseline severity, region and background therapy has run a dozen comparisons before anyone calls it an analysis. The canonical demonstration is ISIS-2, whose investigators subdivided their aspirin result by astrological birth sign and found no benefit under Gemini or Libra, while the overall effect on vascular mortality was large and real.
The right question is not whether an effect is significant within a subgroup but whether it differs between subgroups, which is a test of interaction. Within-subgroup p-values disagree routinely, since each subgroup carries a fraction of the sample and of the power. Interaction tests are themselves underpowered in a trial sized for an overall effect, so a null one is weak evidence of uniformity.
The claim to recognise rests on a nominal p-value: unadjusted, outside the hierarchy that protected the primary endpoint, and treated by regulators as descriptive for exactly that reason. When a page reports that a compound worked especially well in women, ask whether the subgroup was prespecified, whether an interaction test was reported, and how many subgroups went unmentioned.