This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
SURMOUNT-OSA was a randomized, double-blind, parallel phase 3 study of tirzepatide in participants with obstructive sleep apnea and obesity. The registry reports 469 enrolled participants, 4 arms, 1 registered primary endpoint, 9 posted outcome measures, and 20 posted statistical analyses.
| Feature | SURMOUNT-OSA |
|---|---|
| Trial name | SURMOUNT-OSA |
| Brief title | Obstructive Sleep Apnea Master Protocol GPIF: A Study of Tirzepatide (LY3298176) in Participants With Obstructive Sleep Apnea |
| Phase | Phase 3 |
| Status | COMPLETED |
| Conditions | Obstructive Sleep Apnea; Obesity |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | DOUBLE |
| Primary purpose | TREATMENT |
| Enrollment | 469 |
| Interventions | Tirzepatide; Placebo |
| Lead sponsor | Eli Lilly and Company |
| Sponsor type | INDUSTRY |
| ClinicalTrials.gov | NCT05412004 |
2. Clinical Question
The central statistical question was whether treatment with tirzepatide produced a different change from baseline in the Apnea-Hypopnea Index at Week 52 compared with placebo in the two reported comparison pairs.
Population
Participants with the registered conditions of obstructive sleep apnea and obesity.
Intervention
Tirzepatide.
Comparator
Placebo.
Primary question
Does tirzepatide produce a different change from baseline in AHI at Week 52 compared with placebo?
The registry reports a superiority hypothesis for the primary analyses. The primary endpoint is listed as a binary endpoint in the registered endpoint information, while the posted statistical-analysis records describe the AHI outcome as a count/rate measured in events per hour and analyse its change from baseline with ANCOVA.
3. Trial Design
Tirzepatide MTD_GPI1 vs Placebo_GPI1
- Primary AHI analysis at Baseline, Week 52
- Secondary AHI percentage-change analysis
- Secondary binary AHI-reduction analyses
- Secondary sleep-related, body-weight, hsCRP, and SBP analyses
Tirzepatide MTD_GPI2 vs Placebo_GPI2
- Primary AHI analysis at Baseline, Week 52
- Secondary AHI percentage-change analysis
- Secondary binary AHI-reduction analyses
- Secondary sleep-related, body-weight, hsCRP, and SBP analyses
4. Trial Timeline
Trial start
The registry lists 2022-06-21 as the study start date.
Primary completion
The registry lists 2024-03-12 as the primary completion date.
Registry status
The trial status is recorded as COMPLETED, with results posted on ClinicalTrials.gov.
5. Primary Endpoint
| Endpoint | Registered definition / time frame | Posted analysis |
|---|---|---|
| Change From Baseline in Apnea-Hypopnea Index (AHI) | Baseline, Week 52. AHI is the number of apneas or hypopneas recorded via polysomnography during the study per hour of sleep. Apnea is defined as a cessation of airflow lasting at least 10 seconds; hypopnea as a decrease in airflow by at least 30% from baseline for at least 10 seconds occurring with a drop in oxygen saturation (SpO₂) by at least 4%. AHI values are categorized as 5-15 events/hr. | ANCOVA; LS Mean Change difference; superiority |
The primary statistical analyses use the modified intent-to-treat (mITT) population, defined in the registry as all randomized participants who are exposed to at least 1 dose of study intervention and have evaluable data for the outcome.
6. Primary Results: Change From Baseline in AHI
The registry reports two formal primary analyses for the same endpoint, corresponding to the two posted tirzepatide-versus-placebo comparison pairs. Both analyses use ANCOVA and report a least-squares mean change difference with a two-sided 95% confidence interval.
MTD_GPI1 comparison
LS Mean Change difference in AHI
95% CI: -25.82 to -14.20 · P < 0.001
Tirzepatide MTD_GPI1 vs Placebo_GPI1 at Baseline, Week 52
| Feature | Primary analysis |
|---|---|
| Analysis population | Modified intent-to-treat (mITT) |
| Comparison | Tirzepatide MTD_GPI1 vs Placebo_GPI1 |
| Outcome | Change From Baseline in Apnea-Hypopnea Index (AHI) |
| Time frame | Baseline, Week 52 |
| Effect measure | LS Mean Change difference |
| Estimate | -20.01 |
| 95% CI | -25.82 to -14.20 |
| P-value | <0.001 |
| Hypothesis | Superiority |
The estimate of -20.01 events per hour is the reported difference in covariate-adjusted least-squares mean change between the MTD_GPI1 tirzepatide and placebo groups. The negative direction means the estimated change in AHI was lower in the tirzepatide comparison group than in the placebo comparison group.
The 95% CI of -25.82 to -14.20 describes the statistical uncertainty around that estimated difference under the reported ANCOVA framework. Because the entire interval is below zero, the interval is consistent with a lower adjusted mean change in the tirzepatide group relative to placebo.
The P-value < 0.001 addresses evidence against the null hypothesis used for the superiority comparison. It does not measure the size or clinical importance of the treatment effect. Effect size is described by the estimate and its confidence interval, not by the P-value alone.
This analysis is based on an mITT population rather than every randomized participant regardless of exposure and evaluability. The ClinicalTrials.gov record also do not provide enough information to reconstruct the individual-level AHI distribution or to assess the assumptions of the fitted ANCOVA directly.
MTD_GPI2 comparison
LS Mean Change difference in AHI
95% CI: -29.61 to -17.93 · P < 0.001
Tirzepatide MTD_GPI2 vs Placebo_GPI2 at Baseline, Week 52
| Feature | Primary analysis |
|---|---|
| Analysis population | Modified intent-to-treat (mITT) |
| Comparison | Tirzepatide MTD_GPI2 vs Placebo_GPI2 |
| Outcome | Change From Baseline in Apnea-Hypopnea Index (AHI) |
| Time frame | Baseline, Week 52 |
| Effect measure | LS Mean Change difference |
| Estimate | -23.77 |
| 95% CI | -29.61 to -17.93 |
| P-value | <0.001 |
| Hypothesis | Superiority |
The reported estimate of -23.77 events per hour represents the difference in covariate-adjusted least-squares mean change between the MTD_GPI2 tirzepatide and placebo groups.
The 95% CI of -29.61 to -17.93 gives the corresponding interval estimate of uncertainty. It remains entirely below zero, so the reported data are consistent with a lower adjusted change in AHI for the tirzepatide comparison group.
The P-value < 0.001 provides evidence against the null hypothesis of no difference under the reported superiority analysis. It is not a measure of how large the effect is, how important it is clinically, or the probability that the treatment hypothesis is true.
As with the first primary analysis, interpretation is tied to the mITT population and the ANCOVA model specified in the registry. The ClinicalTrials.gov record does not provide individual observations, residual diagnostics, or a complete statistical analysis plan from which model assumptions could be independently checked.
7. Statistical Methodology
ANCOVA for the primary endpoint
The primary AHI comparisons were analysed using analysis of covariance (ANCOVA). The reported model included baseline, geographic region, sex, and treatment as covariates, with Type III sum of squares.
The important statistical idea is that the comparison is adjusted for prespecified covariates rather than relying only on an unadjusted difference between observed group means.
For the primary AHI endpoint, the reported effect measure is the LS Mean Change difference. Least-squares means are model-based adjusted means. Thus, the reported estimate is not simply the arithmetic difference between two raw observed means.
Covariate adjustment
Baseline AHI and the other reported covariates can improve the precision of a treatment comparison when they explain outcome variability. Adjustment does not change the randomized treatment assignment; instead, it uses the fitted model to estimate the treatment contrast after accounting for the specified covariates.
Logistic regression for binary secondary endpoints
The registry reports logistic regression for the endpoint measuring the percentage of participants with ≥50% AHI reduction from baseline. The effect measure reported for these analyses is the risk difference (RD).
This is an important distinction: logistic regression models a binary outcome, but the reported treatment effect here is a risk difference rather than an odds ratio. The risk difference describes the difference in the estimated probability of meeting the binary criterion between treatment groups.
Median difference for SASHB
For change from baseline in Sleep Apnea-Specific Hypoxic Burden (SASHB), the registry reports a Median Difference (Net) as the effect measure. The ClinicalTrials.gov record does not report a normalized statistical method for these two analyses, so this page does not assign an unreported method to them.
Modified intent-to-treat analysis
The mITT population includes randomized participants who received at least 1 dose of study intervention and had evaluable data for the relevant outcome. This preserves the randomized origin of the comparison while applying the registry's exposure and evaluability criteria.
Two-sided confidence intervals
All statistical analyses reported here report 95% two-sided confidence intervals. A confidence interval gives a range of parameter values compatible with the observed data and statistical model under the stated confidence procedure. It is not a probability statement about the individual treatment effect.
8. Secondary Endpoint Results
Percent Change From Baseline in AHI
| Comparison | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Tirzepatide MTD_GPI1 vs Placebo_GPI1 | -47.65 | -65.76 to -29.55 | <0.001 | ANCOVA |
| Tirzepatide MTD_GPI2 vs Placebo_GPI2 | -56.21 | -73.73 to -38.70 | <0.001 | ANCOVA |
These estimates are reported as LS Mean Change differences for percentage change in AHI from baseline to Week 52. The negative estimates indicate a lower adjusted percentage change in the tirzepatide comparison groups relative to placebo under the reported model.
Participants With ≥50% AHI Reduction From Baseline
| Comparison | Risk difference | 95% CI | P-value | Method |
|---|---|---|---|---|
| Tirzepatide MTD_GPI1 vs Placebo_GPI1 | 42.77 | 30.76 to 54.79 | <0.001 | Logistic regression |
| Tirzepatide MTD_GPI2 vs Placebo_GPI2 | 48.60 | 36.55 to 60.65 | <0.001 | Logistic regression |
A risk difference of 42.77 means that the model-based difference in the percentage meeting the ≥50% AHI-reduction criterion was 42.77 percentage points for the MTD_GPI1 comparison. Similarly, the reported MTD_GPI2 risk difference was 48.60 percentage points. These are absolute differences in the probability of meeting the specified binary endpoint, not relative risks or odds ratios.
AHI <5 or AHI 5-14 With ESS ≤10
| Comparison | Risk difference | 95% CI | P-value | Method |
|---|---|---|---|---|
| Tirzepatide MTD_GPI1 vs Placebo_GPI1 | 28.74 | 18.27 to 39.22 | <.001 | Logistic regression |
| Tirzepatide MTD_GPI2 vs Placebo_GPI2 | 33.22 | 22.12 to 44.31 | <0.001 | Logistic regression |
The endpoint combines an AHI threshold with an Epworth Sleepiness Scale criterion. The reported risk differences quantify the adjusted absolute separation between the two comparison groups for achieving that composite definition at Week 52.
Sleep Apnea-Specific Hypoxic Burden
| Comparison | Median difference (Net) | 95% CI |
|---|---|---|
| Tirzepatide MTD_GPI1 vs Placebo_GPI1 | -70.13 | -90.94 to -49.31 |
| Tirzepatide MTD_GPI2 vs Placebo_GPI2 | -61.29 | -84.66 to -37.93 |
The registry reports these effects as median differences in change from baseline in SASHB, measured in %.min/hr. A formal normalized analysis method is not reported in the ClinicalTrials.gov record, so interpretation should remain at the level of the reported median contrast and confidence interval rather than assigning an unreported model.
PROMIS Sleep-Related Outcomes
| Comparison | Reported measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| MTD_GPI1 vs Placebo_GPI1 | PROMIS SD | -2.03 | -3.95 to -0.12 | 0.037 |
| MTD_GPI1 vs Placebo_GPI1 | PROMIS SRI | -3.43 | -5.69 to -1.17 | 0.003 |
| MTD_GPI2 vs Placebo_GPI2 | PROMIS SD | -3.90 | -6.21 to -1.58 | <0.001 |
| MTD_GPI2 vs Placebo_GPI2 | PROMIS SRI | -4.26 | -6.97 to -1.56 | 0.002 |
These analyses used ANCOVA and report LS Mean Change differences in T scores. The MTD_GPI1 Sleep Disturbance model adjusted for baseline, geographic region, sex, baseline OSA severity Group, and treatment. The MTD_GPI1 Sleep-Related Impairment model used the same covariate structure. The registry-reported MTD_GPI2 analyses likewise report ANCOVA with these covariates.
Percent Change From Baseline in Body Weight
| Comparison | Estimate | 95% CI | P-value |
|---|---|---|---|
| Tirzepatide MTD_GPI1 vs Placebo_GPI1 | -16.09 | -17.99 to -14.19 | <0.001 |
| Tirzepatide MTD_GPI2 vs Placebo_GPI2 | -17.28 | -19.29 to -15.28 | <.001 |
The body-weight endpoint was analysed with ANCOVA. The reported effect is the LS Mean Change difference in percent change from baseline at Week 52. The negative direction indicates a lower adjusted percentage change in the tirzepatide comparison groups.
High Sensitivity C Reactive Protein
| Comparison | Estimate | 95% CI | P-value |
|---|---|---|---|
| Tirzepatide MTD_GPI1 vs Placebo_GPI1 | -0.71 | -1.21 to -0.22 | 0.752 |
| Tirzepatide MTD_GPI2 vs Placebo_GPI2 | -1.04 | -1.57 to -0.51 | 0.350 |
These results illustrate why effect estimates and P-values should be read together with the confidence interval. The ClinicalTrials.gov record reports the estimates, intervals, and P-values above; it does not provide additional information here that would justify reconciling the P-values with the corresponding confidence intervals or reconstructing an alternative analysis.
Systolic Blood Pressure
| Comparison | Time frame | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Tirzepatide MTD_GPI1 vs Placebo_GPI1 | Baseline, Week 48 | -7.62 | -10.48 to -4.77 | <0.001 |
| Tirzepatide MTD_GPI2 vs Placebo_GPI2 | Baseline, Week 48 | -3.70 | -6.75 to -0.65 | 0.017 |
The SBP endpoint was analysed with ANCOVA. Unlike most of the reported secondary analyses in this record, its time frame is Baseline, Week 48 rather than Week 52. The reported effect measure is a mean difference, with the MTD_GPI2 record specifically describing it as an LS Mean difference.
9. Safety Results
The ClinicalTrials.gov record reports serious adverse events by the four arms as affected participants divided by participants at risk. Because the underlying four-arm sample sizes and further safety definitions are not provided, the safety analysis below reports the registry values exactly rather than calculating additional rates.
| Arm | Serious adverse events affected / at risk |
|---|---|
| Tirzepatide MTD_GPI1 | 9/114 |
| Placebo_GPI1 | 7/120 |
| Tirzepatide MTD_GPI2 | 7/119 |
| Placebo_GPI2 | 12/114 |
These figures describe serious adverse events by randomized arm as reported in the ClinicalTrials.gov record. They should be kept separate from the efficacy analyses: an efficacy estimate such as the AHI LS Mean Change difference and a safety count such as serious adverse events answer different statistical questions.
10. Statistical Methods Explained
Why was ANCOVA used for the primary AHI endpoint?
ANCOVA is useful when the outcome is measured at follow-up and a baseline measurement is available. Instead of comparing only raw Week 52 means, the model incorporates baseline AHI together with geographic region, sex, and treatment. This can account for baseline variation and produce an adjusted treatment contrast.
What does an LS Mean Change difference of -20.01 mean?
It means that the reported model-based adjusted mean change in AHI differed by -20.01 events per hour between MTD_GPI1 tirzepatide and placebo. The negative sign identifies the direction of the contrast. It does not mean that every participant experienced exactly a 20.01-events-per-hour reduction.
Why is the confidence interval important?
The 95% confidence interval provides an uncertainty range around the estimated treatment contrast. For the MTD_GPI1 primary analysis, the interval extends from -25.82 to -14.20. For MTD_GPI2, it extends from -29.61 to -17.93. The interval is therefore important for understanding the precision of the estimate rather than relying on the P-value alone.
Why doesn't the P-value measure effect size?
A P-value describes the compatibility of the observed data with the null hypothesis under the specified statistical framework. It is affected by both the size of an observed difference and the amount of information available. The treatment-effect estimate and confidence interval provide the more direct description of magnitude and precision.
Why use logistic regression for ≥50% AHI reduction?
The endpoint classifies participants according to whether they achieved a specified binary response: at least 50% AHI reduction from baseline. Logistic regression is designed for binary outcomes. In this registry record, the reported effect measure is a risk difference, which expresses the absolute difference in estimated probabilities between treatment groups.
What does a risk difference of 42.77 mean?
A risk difference of 42.77 represents a 42.77-percentage-point difference in the probability of meeting the specified ≥50% AHI-reduction criterion between the MTD_GPI1 tirzepatide and placebo groups under the reported analysis. It is not a hazard ratio, relative risk, or odds ratio.
Why does the mITT population matter?
The mITT population is narrower than an all-randomized intention-to-treat population because the registry definition requires exposure to at least 1 dose and evaluable data for the outcome. The estimate therefore describes the randomized comparison within the specified analysis population rather than automatically representing every randomized participant regardless of exposure or evaluability.
11. Interpreting the Primary AHI Results
Both primary analyses report negative LS Mean Change differences: -20.01 for MTD_GPI1 versus placebo and -23.77 for MTD_GPI2 versus placebo. Under the reported change-from-baseline definition, the negative direction corresponds to a lower adjusted AHI change in the tirzepatide comparison groups.
The 95% confidence intervals are -25.82 to -14.20 and -29.61 to -17.93, respectively. Both intervals exclude zero, so the reported analyses provide interval evidence for a nonzero difference in the prespecified direction.
The AHI estimate does not by itself quantify individual treatment response, establish that every participant improved, or provide a complete characterization of clinical benefit. It is a group-level model-based comparison of change from baseline at the specified time point.
Because the primary endpoint was analysed with ANCOVA, the reported estimate depends on the specified covariate model and the mITT population. The ClinicalTrials.gov record does not provide residual diagnostics, missing-data assumptions, or individual-level observations, so those aspects cannot be independently evaluated from the registry summary alone.
12. Confidence Intervals and Statistical Significance
The primary analyses illustrate the difference between an effect estimate, a confidence interval, and a P-value.
| Comparison | Estimate | 95% CI | P-value |
|---|---|---|---|
| MTD_GPI1 vs Placebo_GPI1 | -20.01 | -25.82 to -14.20 | <0.001 |
| MTD_GPI2 vs Placebo_GPI2 | -23.77 | -29.61 to -17.93 | <0.001 |
The estimate is the central numerical description of the treatment contrast. The confidence interval adds information about precision. The P-value quantifies evidence against the null hypothesis within the specified testing framework. None of these quantities should be interpreted as the probability that the treatment works for an individual participant.
13. Multiplicity and Multiple Comparisons
The ClinicalTrials.gov record contains 20 statistical analyses, including two primary-analyses records and multiple secondary analyses. The trial data identify superiority hypotheses for the posted analyses, but they do not provide an alpha-allocation scheme, hierarchical testing sequence, multiplicity-adjustment procedure, or interim-analysis plan.
| Analysis family | Reported information | Interpretive implication |
|---|---|---|
| Primary AHI | 2 formal analyses | Both are identified as primary and superiority analyses |
| Secondary continuous outcomes | ANCOVA-based analyses | Each reported estimate should be interpreted in its own endpoint context |
| Secondary binary outcomes | Logistic regression with risk difference | Binary endpoint comparisons require attention to multiplicity across endpoints |
| Additional secondary outcomes | Median difference or method not reported | Interpret only the reported effect measure and uncertainty |
14. Missing Data, Imputation, and Censoring
The ClinicalTrials.gov record defines the primary analysis population as randomized participants who were exposed to at least 1 dose and had evaluable data for the outcome. They do not specify a missing-data imputation method, a pattern-mixture model, multiple imputation strategy, or other explicit missing-data procedure.
For that reason, this page does not attribute a particular imputation method to SURMOUNT-OSA. The distinction matters because an ANCOVA estimate can depend on how missing Week 52 observations are handled. The registry summary provided here is sufficient to identify the analysis population, but not sufficient to reconstruct the full missing-data strategy.
15. Stratification and Covariate Adjustment
The primary AHI ANCOVA included baseline, geographic region, sex, and treatment as covariates. Several secondary ANCOVA analyses additionally included baseline OSA severity Group.
| Analysis | Reported covariates |
|---|---|
| Primary AHI | Baseline; geographic region; sex; treatment |
| Percent change in AHI | Baseline; geographic region; sex; treatment |
| PROMIS SD / SRI | Baseline; geographic region; sex; baseline OSA severity Group; treatment |
| Body weight | Baseline; geographic region; sex; baseline OSA severity Group; treatment |
| SBP | Baseline; geographic region; sex; baseline OSA severity Group; treatment |
The use of baseline adjustment is especially important for change-from-baseline outcomes. A covariate-adjusted estimate can differ from a simple observed change difference because the model estimates the treatment contrast after accounting for the specified covariates.
16. Why the Binary AHI Endpoints Are Different From the Continuous AHI Endpoint
The primary endpoint evaluates change in AHI, expressed in events per hour in the statistical-analysis records. The secondary ≥50% AHI-reduction endpoint instead asks whether each participant crossed a predefined response threshold.
Continuous-style comparison
The primary analysis estimates an adjusted difference in mean change. Participants contribute information according to their measured AHI change.
Binary comparison
The ≥50% endpoint classifies each participant as meeting or not meeting the response criterion and uses logistic regression with risk difference as the reported effect measure.
These endpoints answer related but distinct questions. A continuous change analysis preserves more information about the magnitude of change, whereas a binary threshold focuses on whether a prespecified level of improvement was achieved.
17. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The two posted primary analyses report negative adjusted AHI change differences with 95% confidence intervals entirely below zero and P-values <0.001.
Clinical interpretation
The ClinicalTrials.gov record shows differences in AHI and several related secondary endpoints, including binary AHI response, sleep-related outcomes, body weight, hsCRP, and SBP. Clinical meaning should be considered endpoint by endpoint rather than reduced to a single statistical number.
18. Important Limitations and Interpretation Issues
- Registry-level detail: the ClinicalTrials.gov record does not contain the complete statistical analysis plan, individual participant data, full baseline characteristics, or complete missing-data methodology.
- mITT population: the primary analyses require exposure to at least 1 dose and evaluable outcome data, so the analysis population is not simply all randomized participants.
- Four-arm design: the registry reports 4 arms, but the ClinicalTrials.gov record does not provide detailed arm descriptions or arm-specific sample sizes beyond the serious-adverse-event denominators.
- Multiplicity: 20 statistical analyses are posted, but the ClinicalTrials.gov record does not specify a familywise-error strategy or hierarchy for the full set of analyses.
- Secondary endpoints: several secondary analyses have different scales and effect measures, so they should not be treated as interchangeable measures of the same outcome.
- Unreported methods: the SASHB analyses report a median difference but no normalized statistical method in the ClinicalTrials.gov record. A method should not be inferred merely from the effect measure.
- Missing data: the ClinicalTrials.gov record identifies the mITT/evaluable-data definition but does not specify an imputation strategy.
- Confidence intervals: a confidence interval describes uncertainty in the estimated population contrast; it does not describe the range of individual participant responses.
- P-values: P-values should not be interpreted as measures of effect magnitude or clinical importance.
19. Why This Trial Matters Statistically
SURMOUNT-OSA is a useful teaching case because its registry results combine randomized treatment comparisons with several distinct statistical estimands. The primary endpoint uses ANCOVA and a model-based mean-change contrast, while important secondary endpoints use logistic regression and risk differences. Other outcomes use median differences or additional ANCOVA models.
| Concept | How it appears in SURMOUNT-OSA |
|---|---|
| Randomization | The study is registered as RANDOMIZED. |
| Blinding | The study is registered as DOUBLE masked. |
| Parallel design | The design model is PARALLEL. |
| ANCOVA | Used for the primary AHI analysis and multiple secondary continuous outcomes. |
| Covariate adjustment | Baseline, geographic region, sex, treatment, and in selected secondary analyses baseline OSA severity Group. |
| Logistic regression | Used for binary AHI-response endpoints. |
| Risk difference | Reported for the ≥50% AHI-reduction and AHI/ESS composite endpoints. |
| Confidence intervals | 95% two-sided intervals are reported for the posted estimates. |
| Modified intent-to-treat | Primary and secondary analyses use an mITT population defined by exposure and evaluability. |
| Multiple endpoints | 20 statistical analyses are posted across primary and secondary outcomes. |
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Calculators
22. Sources
- ClinicalTrials.gov: NCT05412004 — SURMOUNT-OSA.
- Linked publication record: PubMed PMID 42675225.
- Linked publication record: PubMed PMID 41793578.
- Linked publication record: PubMed PMID 41540105.
- Linked publication record: PubMed PMID 41135142.
- Linked publication record: PubMed PMID 40774158.
Continue through the Clinical Biostats statistical pathway
Use the methods in this trial as a starting point for deeper study of ANCOVA, logistic regression, confidence intervals, covariate adjustment, randomization, and risk differences.
23. Record Summary
SURMOUNT-OSA provides a clear example of how a modern randomized trial can use different statistical estimands for different clinical questions. The primary endpoint, Change From Baseline in Apnea-Hypopnea Index at Baseline and Week 52, was analysed with ANCOVA in an mITT population. The reported LS Mean Change differences were -20.01 for Tirzepatide MTD_GPI1 versus Placebo_GPI1 and -23.77 for Tirzepatide MTD_GPI2 versus Placebo_GPI2, with 95% confidence intervals of -25.82 to -14.20 and -29.61 to -17.93, respectively, and P-values <0.001 for both comparisons.
The secondary analyses broaden the statistical picture. AHI percentage change was evaluated with ANCOVA; binary AHI-response endpoints were evaluated with logistic regression and reported as risk differences; SASHB was reported using median differences; PROMIS, body-weight, hsCRP, and SBP outcomes were evaluated with ANCOVA. The ClinicalTrials.gov record reports serious adverse events separately for all four arms.
The most important statistical lesson is that the treatment effect cannot be reduced to a single P-value. Proper interpretation requires the effect estimate, confidence interval, analysis population, covariate structure, endpoint definition, and multiplicity context to be considered together.