The short answer, and what it rests on
Across the two randomised trials that gave the same patients one drug or the other, tirzepatide produced the larger average effect on the primary endpoint. In type 2 diabetes over 40 weeks, every tirzepatide dose beat semaglutide 1 mg on HbA1c and on weight. In obesity without diabetes over 72 weeks, tirzepatide at its maximum tolerated dose produced about a fifth of body weight lost against about a seventh for semaglutide at its maximum tolerated dose. That is the finding, and it is consistent in direction across both settings.
What those two trials do not establish matters as much. Neither was powered for hard clinical outcomes, neither compared doses chosen to be pharmacologically equivalent, and one of them was open-label. Neither tells you why the difference exists, which matters because the mechanistic explanation people reach for first is the one the preclinical literature is least sure about.
Almost everything else circulating as a comparison is cross-trial arithmetic: a headline number from a semaglutide trial set beside one from a tirzepatide trial. That exercise carries far more uncertainty than it appears to, for reasons set out below.
SURPASS-2: the head-to-head in type 2 diabetes
SURPASS-2 was a 40-week, randomised, double-blind trial in adults with type 2 diabetes inadequately controlled on metformin, published in 2021. Roughly 1,900 participants were assigned to once-weekly tirzepatide at 5, 10 or 15 mg, or to once-weekly semaglutide at 1 mg. The primary endpoint was change in HbA1c, with tirzepatide tested first for non-inferiority and then for superiority.
All three tirzepatide doses were superior. Mean HbA1c fell by roughly 2.0, 2.2 and 2.3 percentage points across the ascending tirzepatide doses against roughly 1.9 for semaglutide, from a baseline near 8.3 percent. The gaps are small in isolation but were consistent, and the proportion reaching an HbA1c below 5.7 percent, essentially a non-diabetic range, was several times higher on the two higher tirzepatide doses.
Body weight was a secondary endpoint and the separation there was wider: about 7.6, 9.3 and 11.2 kg lost across the tirzepatide doses against about 5.7 kg on semaglutide. In a diabetes trial that is a large difference, since weight loss in type 2 diabetes is typically blunted relative to what the same agents achieve in people without diabetes.
The design objection to SURPASS-2 is the comparator dose. Semaglutide 1 mg was the highest dose approved for type 2 diabetes when the trial ran, so the comparison was fair against the label of the day, but a 2 mg dose was subsequently approved and the 2.4 mg obesity dose already existed. SURPASS-2 therefore answers what was true of the approved diabetes doses in 2021, not what is true of the two molecules at their present ceilings in diabetes. No trial has repaired that gap.
SURMOUNT-5: the head-to-head in obesity
SURMOUNT-5 is the trial most people mean when they ask about a head-to-head. It randomised roughly 750 adults with obesity, or with overweight plus a weight-related complication, and without type 2 diabetes, to 72 weeks of once-weekly tirzepatide or once-weekly semaglutide. Crucially, both arms were titrated to a maximum tolerated dose rather than a fixed one: 10 or 15 mg of tirzepatide, 1.7 or 2.4 mg of semaglutide. Results were reported in 2025.
Mean weight change was about 20 percent with tirzepatide against about 14 percent with semaglutide, a difference on the order of six percentage points of starting body weight. The categorical endpoints moved in the same direction: the share of participants reaching at least a 15 percent reduction was substantially higher on tirzepatide, and waist circumference fell by several centimetres more.
The design detail that matters most is that SURMOUNT-5 was open-label: participants and investigators knew which drug was which. That is a real but bounded concern here, because weight on a calibrated scale is not a subjective rating, so knowledge of assignment cannot directly inflate it. It can influence adherence, titration decisions, eating behaviour and who stays in the trial, and none of those can be ruled out from the published data.
The second limitation is scale. Seven hundred and fifty participants is ample to detect a six-point difference in mean weight change and far too few to say anything about uncommon adverse events or clinical outcomes. SURMOUNT-5 is a strong answer to a narrow question.
Why STEP versus SURMOUNT is not a head-to-head
The most common comparison in circulation sets STEP 1, which reported about 15 percent mean weight loss on semaglutide 2.4 mg over 68 weeks, against SURMOUNT-1, which reported roughly 15, 20 and 21 percent across ascending tirzepatide doses over 72 weeks. Read side by side that looks like a six-point gap at the top dose, which is close to what SURMOUNT-5 later found. The agreement is partly coincidence, and treating it as confirmation is a mistake.
The trials differ in duration by four weeks, which is not nothing on a curve that has not fully plateaued. They differ in eligibility, in the intensity of the accompanying lifestyle programme, in geography and calendar year, and in the placebo response, which was around 2 percent in one and around 3 percent in the other. They also differ in how missing data and treatment discontinuation were handled, and the choice between an efficacy estimand and a treatment-regimen estimand routinely moves a reported mean by one to three percentage points in this drug class.
None of those differences is a scandal; they are ordinary trial-design variation. The point is that they are large relative to the effect being inferred. A cross-trial gap of six points carries an uncertainty band wide enough to contain three or nine. We can speak about a six-point difference only because a randomised trial measured it directly.
The GIP question: an extra receptor, an unsettled mechanism
Semaglutide is a GLP-1 receptor agonist. Tirzepatide is a single peptide that activates both the GLP-1 receptor and the receptor for glucose-dependent insulinotropic polypeptide, the other major incretin hormone. The obvious inference is that the second receptor arm explains the larger effect. That inference is plausible, widely repeated, and not established.
The complication is that tirzepatide's two activities are not balanced. It engages the GIP receptor with an affinity in the range of the native hormone, while its affinity at the GLP-1 receptor is substantially lower than native GLP-1's, and its signalling at that receptor is biased, favouring the cyclic AMP pathway over beta-arrestin recruitment and the receptor internalisation that follows. A drug that is a weaker but less desensitising GLP-1 receptor agonist is not simply semaglutide with something added, and some of the difference in effect could come from that altered GLP-1 signalling rather than from GIP at all.
The deeper problem is directional. If GIP receptor agonism drives weight loss, then GIP receptor blockade should oppose it. It does not. Antagonising the GIP receptor also reduces body weight in animal models, and a GIP receptor antagonist combined with a GLP-1 receptor agonist has produced substantial weight loss in human phase 2 work. Both pushing and blocking the same receptor apparently help. The leading reconciliations are that sustained agonism functionally desensitises the receptor and so ends up resembling blockade, or that the relevant GIP receptor populations differ between brain and adipose tissue and the two strategies act at different sites. Neither has been settled in humans.
The honest formulation is that tirzepatide has an additional receptor activity and a larger measured effect, and those two facts are associated. The causal chain between them is a live research question, and anyone presenting it as settled is ahead of the data.
Dose is not a neutral variable in either trial
Both head-to-head trials compared maximum approved or maximum tolerated doses. That is the right comparison for a clinical question, because those are the doses that exist. It is the wrong comparison for a mechanistic one, because there is no reason to think 15 mg of tirzepatide and 2.4 mg of semaglutide sit at equivalent points on their respective dose-response curves.
Milligrams are not comparable across molecules with different receptor affinities, clearance and steady-state exposure, and semaglutide's dose-response for weight is not flat at 2.4 mg either. If one drug's approved ceiling sits further up its own curve than the other's, part of the observed gap is a regulatory and tolerability artefact rather than a property of the molecules.
That is not a reason to dismiss the result. A patient can only take an approved dose, so the pragmatic comparison is the one that governs real decisions. It is a reason to be careful about the sentence people extract from it, which is usually about which molecule is stronger rather than which regimen produced more weight loss in a specific trial.
Tolerability: similar in kind, different in degree by less than the efficacy gap
In both head-to-head trials the adverse-event profile was dominated by gastrointestinal effects in every arm: nausea, vomiting, diarrhoea and constipation, mostly mild to moderate, mostly concentrated during dose escalation and mostly declining thereafter. This is what would be expected from two drugs that slow gastric emptying and act on the same brainstem and hypothalamic circuits.
The frequencies were broadly comparable between the drugs, discontinuation for adverse events sat in the single-digit percentages in both arms of both trials, and neither trial demonstrated a clinically meaningful tolerability advantage. The claim that tirzepatide is better tolerated because GIP receptor agonism suppresses nausea has preclinical support, but the head-to-head trials do not show a difference large enough to build a recommendation on.
One asymmetry affects interpretation: in an open-label trial, a participant who knows they are on the drug with the bigger reputation may report and tolerate symptoms differently. SURPASS-2 was double-blind and shows the same broad picture, but it ran in a different population, for less time, at a lower comparator dose.
Outcomes beyond weight and HbA1c
Weight and HbA1c are surrogate endpoints. The question that eventually matters is whether people have fewer heart attacks, strokes, kidney failures and deaths, and here the two molecules are not in the same evidentiary position.
Semaglutide 2.4 mg was tested against placebo in a cardiovascular outcome trial of more than 17,000 people with established cardiovascular disease and overweight or obesity but without diabetes, and reduced major adverse cardiovascular events by about a fifth. It also has a positive kidney outcome trial in type 2 diabetes with chronic kidney disease, trial evidence in heart failure with preserved ejection fraction, and a regulatory approval in metabolic dysfunction-associated steatohepatitis.
Tirzepatide's cardiovascular programme took a different shape. Its large cardiovascular trial in type 2 diabetes used an active comparator, dulaglutide, rather than placebo, and was designed around non-inferiority. A result showing tirzepatide is not worse than another GLP-1 receptor agonist establishes cardiovascular safety; it does not establish a placebo-controlled benefit, and it cannot be read as equivalent to the semaglutide result. Tirzepatide does have positive randomised outcome data in obstructive sleep apnoea with obesity, which supported a regulatory approval, and in heart failure with preserved ejection fraction. The point is not that one drug has outcomes and the other does not; it is that the outcome evidence is not parallel, so a weight-loss ranking is not an outcomes ranking.
What the observational data adds, and what it cannot
Large electronic health record analyses comparing people dispensed tirzepatide with people dispensed semaglutide for overweight or obesity have found a difference in the same direction and of a broadly similar size to the trials, with tirzepatide users losing more weight over a year of follow-up. Consistency between a randomised finding and a real-world one is genuinely worth something, particularly on external validity, since trial participants are not a random sample of patients.
The limits are severe. Prescribing was not randomised, so the groups differ systematically in insurance status, baseline weight, comorbidity and prescriber, and adherence and actual dose reached are poorly captured. Confounding by indication runs in unpredictable directions when shortages, cost and coverage determine who receives which drug in a given month.
Real-world data is best read as a check that the randomised result did not evaporate outside the trial setting, not as an independent estimate of the size of the difference.
Where this leaves a decision
On average weight loss over roughly a year and a half, in people with obesity and without diabetes, the direct randomised evidence favours tirzepatide by about six percentage points of body weight. On glycaemic control in type 2 diabetes on metformin, the direct evidence favours tirzepatide by a smaller margin against a comparator dose that is now below the current ceiling. Those are the two things the head-to-head literature supports.
Averages are not individuals. Both trials show wide distributions: some participants on semaglutide lost more than the tirzepatide mean, and a substantial minority on either drug lost little. A six-point difference in group means is a poor predictor of any single person's response, and nothing in either trial identifies in advance who will fall where.
Beyond weight the comparison stops being one-dimensional. Duration of outcomes evidence, the specific comorbidity being treated, tolerability during escalation and continuity of supply all enter, and they do not all point the same way. Whether either drug suits a given person is a clinical question that belongs with their clinician.