Another non-weight indication that can change an insurance conversation, and evidence that the benefit is not only cosmetic.
| Arm | Cardiovascular death or worsening heart-failure event; KCCQ-CSS change |
|---|---|
| Worsening HF event or CV death | HR 0.62 |
| KCCQ-CSS symptom score | +6.9 points vs placebo |
What else it found
- Worsening heart-failure events fell from 15.9% to 9.9%.
- Six-minute walk distance improved by 18.3 m versus placebo.
Safety and tolerability
GI adverse events more frequent on tirzepatide; discontinuations 6.3% vs 1.4%.
Trial record
| Question | Tirzepatide in heart failure with preserved ejection fraction and obesity. |
|---|---|
| Design | Phase 3, randomized, double-blind, placebo-controlled; median 104 weeks |
| Population | HFpEF (EF ≥50%) with BMI ≥30 |
| Participants | 731 adults |
| Arms | Tirzepatide max tolerated (15 mg) vs placebo |
| Primary endpoint | Cardiovascular death or worsening heart-failure event; KCCQ-CSS change |
| Registration | NCT04847557 |
| Publication | N Engl J Med 2025;392:427-437 |
| DOI | 10.1056/NEJMoa2410027 |
How to read this trial critically
Design: Double-blind. Comparator: Placebo. Endpoint type: clinical events. Funding: Eli Lilly and Company.
Estimands — which question the headline number answers
A composite endpoint combining clinical events and a symptom score. Composite endpoints can be driven by their softest component, so the breakdown matters more than the headline.
Modern trials report the same result under more than one estimand — a precise statement of the question being answered. The treatment-regardless-of-adherence estimand (often called the treatment policy estimand) asks what happened to everyone randomised, including people who stopped the drug early or started another one. The efficacy estimand asks what would have happened if everyone had stayed on treatment as assigned. The efficacy estimand almost always produces a larger number, because it removes the people for whom the drug did not work well enough to keep taking. Marketing tends to quote the efficacy figure; a patient deciding whether to start should generally weigh the treatment-regardless figure, because it includes the possibility of being someone who stops.
Discontinuation and what it means here
Heart-failure populations have high competing-risk mortality, which complicates interpretation of any non-fatal endpoint.
Missing data
Symptom scores are patient-reported and missing-data assumptions carry more weight than for an objective measure.
Sponsorship and conflicts of interest
Sponsor-funded trial in obesity-related heart failure with preserved ejection fraction.
Who this result applies to
Adults with HFpEF and obesity. A narrow and clinically important population, not a general weight-management one.
| Criterion | Met | Points |
|---|---|---|
| Randomised allocation | Yes | 20/20 |
| Participants and investigators blinded | Yes | 15/15 |
| Active comparator rather than placebo | No | 0/10 |
| 1,000 or more participants | No | 0/10 |
| 52 weeks or longer | Yes | 15/15 |
| Clinical events rather than a surrogate endpoint | Yes | 15/15 |
| Prospectively registered | Yes | 10/10 |
| Funding independent of the manufacturer | No | 0/5 |
What a placebo-controlled result does and does not establish
A placebo comparison answers one question: does this drug beat nothing, under trial conditions, in this population. It does not tell you how the drug performs against an active alternative, how it performs in people who would have been excluded from the trial, or what happens after the last study visit. Every one of those is a separate question requiring separate evidence.
What this trial cannot tell you
Every limitation below applies to SUMMIT as it applies to the rest of the program. None of them makes the result unreliable; all of them constrain how far it can be generalised to you.
- Trial participants receive the medication free, with structured follow-up and lifestyle support that most patients do not get. Real-world results are consistently smaller.
- Enrollment criteria exclude many people who would take the drug in practice, including some with significant comorbidity.
- Mean results conceal the distribution. A trial mean contains both strong responders and people who lost very little.
- Adverse events are collected under trial conditions, which detects some events better than routine care and others worse.
- These results were generated with the FDA-approved product. No compounded preparation has been tested in a trial of this kind.
How we use this trial on the rest of the site
Where a page on this site quotes a figure from SUMMIT, it names the trial, gives the dose the figure came from, and states the population. We do not present a single arm's result as though it applied to everyone, and we do not round a range into one dramatic number. If you find a page that does, tell us through the corrections process and we will fix it.
Trial populations are selected and trial participants receive free medication and structured follow-up. Real-world results are consistently smaller, mostly because people stop early — and cost is the most common reason they stop. See our evidence policy for how we handle that gap.
How to read the SUMMIT trial
SUMMIT is a Phase 3, randomized, double-blind, placebo-controlled; median 104 weeks study. That design matters for how much weight the result carries: randomisation is what allows the difference between arms to be attributed to the drug rather than to the kind of person who chose it, and blinding is what stops expectation from moving subjective endpoints. Where a trial is open-label, treat patient-reported outcomes with more caution than objectively measured ones such as body weight or HbA1c.
What the primary endpoint does and does not tell you
The primary endpoint here was Cardiovascular death or worsening heart-failure event; KCCQ-CSS change. A primary endpoint is chosen before the trial starts and is the only result the study is properly powered to detect; everything else reported alongside it is secondary or exploratory and carries less weight, however striking the number looks. When a marketing page quotes a figure from a trial, the first question worth asking is whether that figure was the primary endpoint or something found further down the results table.
Who SUMMIT actually studied
HFpEF (EF ≥50%) with BMI ≥30 — 731 adults participants. Trial populations are selected: they exclude many of the comorbidities, medications and circumstances that ordinary patients bring, and participants receive structured follow-up and free medication that real-world patients do not. That is the main reason real-world weight outcomes across this drug class run consistently below trial averages, and it is a reason to read the mean effect as a ceiling rather than a forecast.
SUMMIT reported harms and adverse events
GI adverse events more frequent on tirzepatide; discontinuations 6.3% vs 1.4%. Adverse-event rates in trials are collected systematically, which makes them more reliable than anecdote but also means they capture mild events that patients might not otherwise report. Our side-effect reference puts these figures alongside the rest of the program.
Why SUMMIT matters
Another non-weight indication that can change an insurance conversation, and evidence that the benefit is not only cosmetic.
Reading the SUMMIT primary source
We summarize; we do not substitute. The registration record carries the pre-specified protocol, endpoints and eligibility criteria, and the published report carries the full results and limitations sections that summaries like this one necessarily compress. Both are linked in the sources below, and if our summary and the primary report disagree, the primary report is right — tell us through corrections and we will fix it.
SUMMIT: what a result like this can and cannot establish
A randomized trial establishes that a difference between arms is attributable to the intervention rather than to who chose it. It does not establish that you will experience the mean, that the effect persists beyond the observation window, or that it generalizes to people who would have been excluded from enrolment. Those three limits apply to every figure quoted from SUMMIT, including the ones quoted on this site.
Duration is the limit worth holding onto. A 72-week trial tells you what happened over 72 weeks. For a therapy that current evidence says should continue indefinitely, that is a short window, and honest reading treats the long-term profile as still accumulating rather than settled.
SUMMIT in the context of the wider program
No single trial carries a drug. SUMMIT sits alongside the rest of the program, and the weight of evidence comes from consistency across them rather than from any one headline. The most directly related studies indexed here:
- SURMOUNT-1 (2022) — Does tirzepatide cause weight loss in adults with obesity and without diabetes?
- SURMOUNT-5 (2025) — Head-to-head: tirzepatide vs semaglutide for weight loss in adults with obesity and without diabetes.
- SURMOUNT-4 (2024) — What happens if you stop tirzepatide after reaching a maintenance dose?
How marketing distorts SUMMIT
Three distortions recur. Quoting a secondary or exploratory endpoint as though it were the primary result. Quoting the completer or per-protocol figure rather than the intention-to-treat figure, which is higher and less representative. And quoting a trial mean as an individual forecast. When a program's landing page cites a percentage without naming the trial, the dose and the timepoint, all three of those are available to it.
What SUMMIT means for what you pay
Efficacy evidence attaches to the molecule, not to the vial it arrives in. The same tirzepatide studied here is dispensed by compounding pharmacies from $133 per month, and by the manufacturer as an FDA-approved product at a considerably higher price. What the higher price buys is regulatory review of that specific product, metered dosing and a manufacturer's quality system — not a different result in a trial like this one. Anyone claiming clinical superiority for a compounded preparation on the strength of SUMMIT is over-reading it, and so is anyone claiming a compounded preparation was proven equivalent by it.
SUMMIT: reading the record yourself
The registration record carries the pre-specified protocol, endpoints and eligibility criteria; the published report carries the full results and the limitations section that any summary compresses. Both are linked in the sources below. If our summary and the primary report conflict, the primary report governs and we would like to hear about it through corrections.
SUMMIT: statistics worth understanding before quoting it
Two conventions decide how large a trial result looks. The intention-to-treat analysis counts everyone randomized, including those who stopped early, and is the conservative and more honest figure. The completer or efficacy analysis counts only those who finished on treatment, and always looks better. Weight-loss trials in this class report both, and marketing predictably quotes the second. When a percentage appears without a stated analysis population, assume the flattering one.
The second convention is the confidence interval. A mean of twenty per cent with a wide interval and a mean of twenty per cent with a narrow one are different results, and only the interval tells you how precisely the effect was estimated. Neither appears in an advertisement.
SUMMIT: who is missing from it
Trial eligibility criteria exclude people, and the exclusions are systematic rather than random: severe comorbidity, certain concurrent medications, recent cardiovascular events, pregnancy, and frequently significant psychiatric or eating-disorder history. If you would have been excluded from this trial, its mean is a weaker guide for you than it is for someone who would have qualified — and your prescriber is the person who can say which of those you are.
SUMMIT: what would change the conclusion
Longer follow-up showing the effect attenuating. Outcome data showing the weight change does not translate into the clinical benefits it is assumed to produce. Or head-to-head evidence against a newer comparator reversing the ranking. None of those has happened, but naming them is what distinguishes reading evidence from repeating it — and it is why our evidence policy commits to updating rather than defending what we published.
SUMMIT: what patients get wrong most often
Four mistakes recur, and all four are expensive. Judging the drug on the first month, when the starter dose is not a treatment dose and a flat scale predicts nothing. Comparing programs on the advertised price, which describes four weeks of a twelve-month course. Treating a plateau as failure rather than as the expected shape of the trial curves. And stopping for cost without pricing the market first, when the spread between the cheapest and dearest route to the identical molecule runs from $133 to $399 a month at entry.
The fifth, less common but more serious, is buying outside the licensed channel because a price looked unbeatable. Research-labelled peptide sold without a prescription is a different market with no clinical oversight and no accountability if something goes wrong, and the disclaimer that it is not for human consumption exists so the seller can say you were told.
SUMMIT: the three-year view
Current evidence treats obesity pharmacotherapy as ongoing while it remains effective, tolerated and appropriate, and the withdrawal data show substantial regain after stopping. That makes three years a more honest planning horizon than twelve months. At the cheapest tracked route, holding a 10 mg maintenance dose for three years costs roughly $7,164; at the dearest compounded route it is several times that for the same molecule. Deciding on a first-month price is deciding on the least representative number available.
SUMMIT: what would change this page
New trial evidence, a change in the labelling, a shift in the compounding framework, or a correction from a reader. Clinical statements here follow the prescribing information and the published trial program; price figures carry a verification date and sit in a downloadable dataset. Where the evidence is genuinely limited — long-term data beyond the trial horizon, bioavailability of compounded oral preparations, the human relevance of the rodent thyroid finding — this site says so rather than filling the gap with confident language, and updates rather than defends when better evidence arrives.
Primary sources
Clinical, dosing and regulatory statements on this page rest on the documents below. Links go to the publisher, not to a summary of it. Prices are not sourced here — they carry a verification date instead, and the reason is set out in the source ledger.
- ZEPBOUND (tirzepatide) full prescribing information — DailyMed, US National Library of Medicine. Dose ladder, contraindications, warnings, storage.
- FDA’s concerns with unapproved GLP-1 drugs used for weight loss — US Food and Drug Administration. Compounded GLP-1 risks, API import alert, cold-chain complaints.
- FD&C Act provisions that apply to human drug compounding — US Food and Drug Administration. Why a compounded preparation is lawful without being FDA-approved.
- Tirzepatide once weekly for the treatment of obesity (SURMOUNT-1) — N Engl J Med 2022;387:205-216. Registration NCT04184622. The weight-reduction and adverse-event figures used across this site.
- The full source ledger — every primary source, what it supports, the date we read it, and the claims we deliberately do not source.