← Clinical Trials
Obstructive Sleep Apnea Phase 3 Completed NCT05412004

SURMOUNT-OSA: Complete Statistical Analysis of Tirzepatide in Obstructive Sleep Apnea

An independent statistical analysis of the randomized, double-blind phase 3 SURMOUNT-OSA trial evaluating tirzepatide versus placebo in participants with obstructive sleep apnea and obesity, with emphasis on the Apnea-Hypopnea Index, covariate-adjusted ANCOVA, logistic regression, and risk differences.

Trial start: 2022-06-21  ·  Primary completion: 2024-03-12  ·  Enrollment: 469
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SURMOUNT-OSA was a randomized, double-blind, parallel phase 3 study of tirzepatide in participants with obstructive sleep apnea and obesity. The registry reports 469 enrolled participants, 4 arms, 1 registered primary endpoint, 9 posted outcome measures, and 20 posted statistical analyses.

469
Enrolled
Phase 3
4
Arms
Parallel design
1
Primary endpoint
AHI change
20
Statistical analyses
Posted in registry
FeatureSURMOUNT-OSA
Trial nameSURMOUNT-OSA
Brief titleObstructive Sleep Apnea Master Protocol GPIF: A Study of Tirzepatide (LY3298176) in Participants With Obstructive Sleep Apnea
PhasePhase 3
StatusCOMPLETED
ConditionsObstructive Sleep Apnea; Obesity
AllocationRANDOMIZED
Design modelPARALLEL
MaskingDOUBLE
Primary purposeTREATMENT
Enrollment469
InterventionsTirzepatide; Placebo
Lead sponsorEli Lilly and Company
Sponsor typeINDUSTRY
ClinicalTrials.govNCT05412004

2. Clinical Question

The central statistical question was whether treatment with tirzepatide produced a different change from baseline in the Apnea-Hypopnea Index at Week 52 compared with placebo in the two reported comparison pairs.

Population

Participants with the registered conditions of obstructive sleep apnea and obesity.

Intervention

Tirzepatide.

Comparator

Placebo.

Primary question

Does tirzepatide produce a different change from baseline in AHI at Week 52 compared with placebo?

The registry reports a superiority hypothesis for the primary analyses. The primary endpoint is listed as a binary endpoint in the registered endpoint information, while the posted statistical-analysis records describe the AHI outcome as a count/rate measured in events per hour and analyse its change from baseline with ANCOVA.

3. Trial Design

01
Randomized469 enrolled
02
Double-blind4-arm study
03
ParallelTirzepatide or placebo
04
Week 52Primary AHI assessment
05
AnalysisANCOVA / logistic regression
Allocation
RANDOMIZED
Design model
PARALLEL
Masking
DOUBLE
Primary purpose
TREATMENT
COMPARISON 1

Tirzepatide MTD_GPI1 vs Placebo_GPI1

  • Primary AHI analysis at Baseline, Week 52
  • Secondary AHI percentage-change analysis
  • Secondary binary AHI-reduction analyses
  • Secondary sleep-related, body-weight, hsCRP, and SBP analyses
COMPARISON 2

Tirzepatide MTD_GPI2 vs Placebo_GPI2

  • Primary AHI analysis at Baseline, Week 52
  • Secondary AHI percentage-change analysis
  • Secondary binary AHI-reduction analyses
  • Secondary sleep-related, body-weight, hsCRP, and SBP analyses
Important design distinction: the registry reports 4 arms, but the posted statistical analyses are presented as two tirzepatide-versus-placebo comparison pairs: MTD_GPI1 versus Placebo_GPI1 and MTD_GPI2 versus Placebo_GPI2. The ClinicalTrials.gov record does not provide further arm descriptions, doses, or sample sizes for those four arms, so this page does not infer them.

4. Trial Timeline

2022-06-21

Trial start

The registry lists 2022-06-21 as the study start date.

2024-03-12

Primary completion

The registry lists 2024-03-12 as the primary completion date.

Completed

Registry status

The trial status is recorded as COMPLETED, with results posted on ClinicalTrials.gov.

5. Primary Endpoint

EndpointRegistered definition / time framePosted analysis
Change From Baseline in Apnea-Hypopnea Index (AHI) Baseline, Week 52. AHI is the number of apneas or hypopneas recorded via polysomnography during the study per hour of sleep. Apnea is defined as a cessation of airflow lasting at least 10 seconds; hypopnea as a decrease in airflow by at least 30% from baseline for at least 10 seconds occurring with a drop in oxygen saturation (SpO₂) by at least 4%. AHI values are categorized as 5-15 events/hr. ANCOVA; LS Mean Change difference; superiority

The primary statistical analyses use the modified intent-to-treat (mITT) population, defined in the registry as all randomized participants who are exposed to at least 1 dose of study intervention and have evaluable data for the outcome.

6. Primary Results: Change From Baseline in AHI

The registry reports two formal primary analyses for the same endpoint, corresponding to the two posted tirzepatide-versus-placebo comparison pairs. Both analyses use ANCOVA and report a least-squares mean change difference with a two-sided 95% confidence interval.

MTD_GPI1 comparison

LS Mean Change difference in AHI

-20.01 events/hour

95% CI: -25.82 to -14.20   ·   P < 0.001

Tirzepatide MTD_GPI1 vs Placebo_GPI1 at Baseline, Week 52

FeaturePrimary analysis
Analysis populationModified intent-to-treat (mITT)
ComparisonTirzepatide MTD_GPI1 vs Placebo_GPI1
OutcomeChange From Baseline in Apnea-Hypopnea Index (AHI)
Time frameBaseline, Week 52
Effect measureLS Mean Change difference
Estimate-20.01
95% CI-25.82 to -14.20
P-value<0.001
HypothesisSuperiority
Clinical Biostats interpretation

The estimate of -20.01 events per hour is the reported difference in covariate-adjusted least-squares mean change between the MTD_GPI1 tirzepatide and placebo groups. The negative direction means the estimated change in AHI was lower in the tirzepatide comparison group than in the placebo comparison group.

The 95% CI of -25.82 to -14.20 describes the statistical uncertainty around that estimated difference under the reported ANCOVA framework. Because the entire interval is below zero, the interval is consistent with a lower adjusted mean change in the tirzepatide group relative to placebo.

The P-value < 0.001 addresses evidence against the null hypothesis used for the superiority comparison. It does not measure the size or clinical importance of the treatment effect. Effect size is described by the estimate and its confidence interval, not by the P-value alone.

This analysis is based on an mITT population rather than every randomized participant regardless of exposure and evaluability. The ClinicalTrials.gov record also do not provide enough information to reconstruct the individual-level AHI distribution or to assess the assumptions of the fitted ANCOVA directly.

MTD_GPI2 comparison

LS Mean Change difference in AHI

-23.77 events/hour

95% CI: -29.61 to -17.93   ·   P < 0.001

Tirzepatide MTD_GPI2 vs Placebo_GPI2 at Baseline, Week 52

FeaturePrimary analysis
Analysis populationModified intent-to-treat (mITT)
ComparisonTirzepatide MTD_GPI2 vs Placebo_GPI2
OutcomeChange From Baseline in Apnea-Hypopnea Index (AHI)
Time frameBaseline, Week 52
Effect measureLS Mean Change difference
Estimate-23.77
95% CI-29.61 to -17.93
P-value<0.001
HypothesisSuperiority
Clinical Biostats interpretation

The reported estimate of -23.77 events per hour represents the difference in covariate-adjusted least-squares mean change between the MTD_GPI2 tirzepatide and placebo groups.

The 95% CI of -29.61 to -17.93 gives the corresponding interval estimate of uncertainty. It remains entirely below zero, so the reported data are consistent with a lower adjusted change in AHI for the tirzepatide comparison group.

The P-value < 0.001 provides evidence against the null hypothesis of no difference under the reported superiority analysis. It is not a measure of how large the effect is, how important it is clinically, or the probability that the treatment hypothesis is true.

As with the first primary analysis, interpretation is tied to the mITT population and the ANCOVA model specified in the registry. The ClinicalTrials.gov record does not provide individual observations, residual diagnostics, or a complete statistical analysis plan from which model assumptions could be independently checked.

7. Statistical Methodology

ANCOVA for the primary endpoint

The primary AHI comparisons were analysed using analysis of covariance (ANCOVA). The reported model included baseline, geographic region, sex, and treatment as covariates, with Type III sum of squares.

Conceptual model
Outcome at Week 52 = treatment effect + baseline + geographic region + sex + error

The important statistical idea is that the comparison is adjusted for prespecified covariates rather than relying only on an unadjusted difference between observed group means.

For the primary AHI endpoint, the reported effect measure is the LS Mean Change difference. Least-squares means are model-based adjusted means. Thus, the reported estimate is not simply the arithmetic difference between two raw observed means.

Covariate adjustment

Baseline AHI and the other reported covariates can improve the precision of a treatment comparison when they explain outcome variability. Adjustment does not change the randomized treatment assignment; instead, it uses the fitted model to estimate the treatment contrast after accounting for the specified covariates.

Logistic regression for binary secondary endpoints

The registry reports logistic regression for the endpoint measuring the percentage of participants with ≥50% AHI reduction from baseline. The effect measure reported for these analyses is the risk difference (RD).

This is an important distinction: logistic regression models a binary outcome, but the reported treatment effect here is a risk difference rather than an odds ratio. The risk difference describes the difference in the estimated probability of meeting the binary criterion between treatment groups.

Median difference for SASHB

For change from baseline in Sleep Apnea-Specific Hypoxic Burden (SASHB), the registry reports a Median Difference (Net) as the effect measure. The ClinicalTrials.gov record does not report a normalized statistical method for these two analyses, so this page does not assign an unreported method to them.

Modified intent-to-treat analysis

The mITT population includes randomized participants who received at least 1 dose of study intervention and had evaluable data for the relevant outcome. This preserves the randomized origin of the comparison while applying the registry's exposure and evaluability criteria.

Two-sided confidence intervals

All statistical analyses reported here report 95% two-sided confidence intervals. A confidence interval gives a range of parameter values compatible with the observed data and statistical model under the stated confidence procedure. It is not a probability statement about the individual treatment effect.

8. Secondary Endpoint Results

Percent Change From Baseline in AHI

ComparisonEstimate95% CIP-valueMethod
Tirzepatide MTD_GPI1 vs Placebo_GPI1-47.65-65.76 to -29.55<0.001ANCOVA
Tirzepatide MTD_GPI2 vs Placebo_GPI2-56.21-73.73 to -38.70<0.001ANCOVA

These estimates are reported as LS Mean Change differences for percentage change in AHI from baseline to Week 52. The negative estimates indicate a lower adjusted percentage change in the tirzepatide comparison groups relative to placebo under the reported model.

Participants With ≥50% AHI Reduction From Baseline

ComparisonRisk difference95% CIP-valueMethod
Tirzepatide MTD_GPI1 vs Placebo_GPI142.7730.76 to 54.79<0.001Logistic regression
Tirzepatide MTD_GPI2 vs Placebo_GPI248.6036.55 to 60.65<0.001Logistic regression

A risk difference of 42.77 means that the model-based difference in the percentage meeting the ≥50% AHI-reduction criterion was 42.77 percentage points for the MTD_GPI1 comparison. Similarly, the reported MTD_GPI2 risk difference was 48.60 percentage points. These are absolute differences in the probability of meeting the specified binary endpoint, not relative risks or odds ratios.

AHI <5 or AHI 5-14 With ESS ≤10

ComparisonRisk difference95% CIP-valueMethod
Tirzepatide MTD_GPI1 vs Placebo_GPI128.7418.27 to 39.22<.001Logistic regression
Tirzepatide MTD_GPI2 vs Placebo_GPI233.2222.12 to 44.31<0.001Logistic regression

The endpoint combines an AHI threshold with an Epworth Sleepiness Scale criterion. The reported risk differences quantify the adjusted absolute separation between the two comparison groups for achieving that composite definition at Week 52.

Sleep Apnea-Specific Hypoxic Burden

ComparisonMedian difference (Net)95% CI
Tirzepatide MTD_GPI1 vs Placebo_GPI1-70.13-90.94 to -49.31
Tirzepatide MTD_GPI2 vs Placebo_GPI2-61.29-84.66 to -37.93

The registry reports these effects as median differences in change from baseline in SASHB, measured in %.min/hr. A formal normalized analysis method is not reported in the ClinicalTrials.gov record, so interpretation should remain at the level of the reported median contrast and confidence interval rather than assigning an unreported model.

PROMIS Sleep-Related Outcomes

ComparisonReported measureEstimate95% CIP-value
MTD_GPI1 vs Placebo_GPI1PROMIS SD-2.03-3.95 to -0.120.037
MTD_GPI1 vs Placebo_GPI1PROMIS SRI-3.43-5.69 to -1.170.003
MTD_GPI2 vs Placebo_GPI2PROMIS SD-3.90-6.21 to -1.58<0.001
MTD_GPI2 vs Placebo_GPI2PROMIS SRI-4.26-6.97 to -1.560.002

These analyses used ANCOVA and report LS Mean Change differences in T scores. The MTD_GPI1 Sleep Disturbance model adjusted for baseline, geographic region, sex, baseline OSA severity Group, and treatment. The MTD_GPI1 Sleep-Related Impairment model used the same covariate structure. The registry-reported MTD_GPI2 analyses likewise report ANCOVA with these covariates.

Percent Change From Baseline in Body Weight

ComparisonEstimate95% CIP-value
Tirzepatide MTD_GPI1 vs Placebo_GPI1-16.09-17.99 to -14.19<0.001
Tirzepatide MTD_GPI2 vs Placebo_GPI2-17.28-19.29 to -15.28<.001

The body-weight endpoint was analysed with ANCOVA. The reported effect is the LS Mean Change difference in percent change from baseline at Week 52. The negative direction indicates a lower adjusted percentage change in the tirzepatide comparison groups.

High Sensitivity C Reactive Protein

ComparisonEstimate95% CIP-value
Tirzepatide MTD_GPI1 vs Placebo_GPI1-0.71-1.21 to -0.220.752
Tirzepatide MTD_GPI2 vs Placebo_GPI2-1.04-1.57 to -0.510.350

These results illustrate why effect estimates and P-values should be read together with the confidence interval. The ClinicalTrials.gov record reports the estimates, intervals, and P-values above; it does not provide additional information here that would justify reconciling the P-values with the corresponding confidence intervals or reconstructing an alternative analysis.

Systolic Blood Pressure

ComparisonTime frameEstimate95% CIP-value
Tirzepatide MTD_GPI1 vs Placebo_GPI1Baseline, Week 48-7.62-10.48 to -4.77<0.001
Tirzepatide MTD_GPI2 vs Placebo_GPI2Baseline, Week 48-3.70-6.75 to -0.650.017

The SBP endpoint was analysed with ANCOVA. Unlike most of the reported secondary analyses in this record, its time frame is Baseline, Week 48 rather than Week 52. The reported effect measure is a mean difference, with the MTD_GPI2 record specifically describing it as an LS Mean difference.

9. Safety Results

The ClinicalTrials.gov record reports serious adverse events by the four arms as affected participants divided by participants at risk. Because the underlying four-arm sample sizes and further safety definitions are not provided, the safety analysis below reports the registry values exactly rather than calculating additional rates.

ArmSerious adverse events affected / at risk
Tirzepatide MTD_GPI19/114
Placebo_GPI17/120
Tirzepatide MTD_GPI27/119
Placebo_GPI212/114

These figures describe serious adverse events by randomized arm as reported in the ClinicalTrials.gov record. They should be kept separate from the efficacy analyses: an efficacy estimate such as the AHI LS Mean Change difference and a safety count such as serious adverse events answer different statistical questions.

Safety interpretation: The registry values identify affected participants and participants at risk, but the ClinicalTrials.gov record does not provide event definitions, exposure duration, seriousness criteria, treatment-emergent conventions, or statistical comparisons for these serious adverse events. No additional safety inference is therefore made here.

10. Statistical Methods Explained

Why was ANCOVA used for the primary AHI endpoint?

ANCOVA is useful when the outcome is measured at follow-up and a baseline measurement is available. Instead of comparing only raw Week 52 means, the model incorporates baseline AHI together with geographic region, sex, and treatment. This can account for baseline variation and produce an adjusted treatment contrast.

What does an LS Mean Change difference of -20.01 mean?

It means that the reported model-based adjusted mean change in AHI differed by -20.01 events per hour between MTD_GPI1 tirzepatide and placebo. The negative sign identifies the direction of the contrast. It does not mean that every participant experienced exactly a 20.01-events-per-hour reduction.

Why is the confidence interval important?

The 95% confidence interval provides an uncertainty range around the estimated treatment contrast. For the MTD_GPI1 primary analysis, the interval extends from -25.82 to -14.20. For MTD_GPI2, it extends from -29.61 to -17.93. The interval is therefore important for understanding the precision of the estimate rather than relying on the P-value alone.

Why doesn't the P-value measure effect size?

A P-value describes the compatibility of the observed data with the null hypothesis under the specified statistical framework. It is affected by both the size of an observed difference and the amount of information available. The treatment-effect estimate and confidence interval provide the more direct description of magnitude and precision.

Why use logistic regression for ≥50% AHI reduction?

The endpoint classifies participants according to whether they achieved a specified binary response: at least 50% AHI reduction from baseline. Logistic regression is designed for binary outcomes. In this registry record, the reported effect measure is a risk difference, which expresses the absolute difference in estimated probabilities between treatment groups.

What does a risk difference of 42.77 mean?

A risk difference of 42.77 represents a 42.77-percentage-point difference in the probability of meeting the specified ≥50% AHI-reduction criterion between the MTD_GPI1 tirzepatide and placebo groups under the reported analysis. It is not a hazard ratio, relative risk, or odds ratio.

Why does the mITT population matter?

The mITT population is narrower than an all-randomized intention-to-treat population because the registry definition requires exposure to at least 1 dose and evaluable data for the outcome. The estimate therefore describes the randomized comparison within the specified analysis population rather than automatically representing every randomized participant regardless of exposure or evaluability.

11. Interpreting the Primary AHI Results

Direction of effect

Both primary analyses report negative LS Mean Change differences: -20.01 for MTD_GPI1 versus placebo and -23.77 for MTD_GPI2 versus placebo. Under the reported change-from-baseline definition, the negative direction corresponds to a lower adjusted AHI change in the tirzepatide comparison groups.

Precision

The 95% confidence intervals are -25.82 to -14.20 and -29.61 to -17.93, respectively. Both intervals exclude zero, so the reported analyses provide interval evidence for a nonzero difference in the prespecified direction.

What the analysis does not establish by itself

The AHI estimate does not by itself quantify individual treatment response, establish that every participant improved, or provide a complete characterization of clinical benefit. It is a group-level model-based comparison of change from baseline at the specified time point.

Model dependence

Because the primary endpoint was analysed with ANCOVA, the reported estimate depends on the specified covariate model and the mITT population. The ClinicalTrials.gov record does not provide residual diagnostics, missing-data assumptions, or individual-level observations, so those aspects cannot be independently evaluated from the registry summary alone.

12. Confidence Intervals and Statistical Significance

The primary analyses illustrate the difference between an effect estimate, a confidence interval, and a P-value.

ComparisonEstimate95% CIP-value
MTD_GPI1 vs Placebo_GPI1-20.01-25.82 to -14.20<0.001
MTD_GPI2 vs Placebo_GPI2-23.77-29.61 to -17.93<0.001

The estimate is the central numerical description of the treatment contrast. The confidence interval adds information about precision. The P-value quantifies evidence against the null hypothesis within the specified testing framework. None of these quantities should be interpreted as the probability that the treatment works for an individual participant.

13. Multiplicity and Multiple Comparisons

The ClinicalTrials.gov record contains 20 statistical analyses, including two primary-analyses records and multiple secondary analyses. The trial data identify superiority hypotheses for the posted analyses, but they do not provide an alpha-allocation scheme, hierarchical testing sequence, multiplicity-adjustment procedure, or interim-analysis plan.

Analysis familyReported informationInterpretive implication
Primary AHI2 formal analysesBoth are identified as primary and superiority analyses
Secondary continuous outcomesANCOVA-based analysesEach reported estimate should be interpreted in its own endpoint context
Secondary binary outcomesLogistic regression with risk differenceBinary endpoint comparisons require attention to multiplicity across endpoints
Additional secondary outcomesMedian difference or method not reportedInterpret only the reported effect measure and uncertainty
Multiplicity caution: A collection of P-values from several secondary endpoints should not automatically be interpreted as if every endpoint were an independently confirmatory hypothesis test. The ClinicalTrials.gov record does not state the familywise-error strategy, so this page does not infer one.

14. Missing Data, Imputation, and Censoring

The ClinicalTrials.gov record defines the primary analysis population as randomized participants who were exposed to at least 1 dose and had evaluable data for the outcome. They do not specify a missing-data imputation method, a pattern-mixture model, multiple imputation strategy, or other explicit missing-data procedure.

For that reason, this page does not attribute a particular imputation method to SURMOUNT-OSA. The distinction matters because an ANCOVA estimate can depend on how missing Week 52 observations are handled. The registry summary provided here is sufficient to identify the analysis population, but not sufficient to reconstruct the full missing-data strategy.

15. Stratification and Covariate Adjustment

The primary AHI ANCOVA included baseline, geographic region, sex, and treatment as covariates. Several secondary ANCOVA analyses additionally included baseline OSA severity Group.

AnalysisReported covariates
Primary AHIBaseline; geographic region; sex; treatment
Percent change in AHIBaseline; geographic region; sex; treatment
PROMIS SD / SRIBaseline; geographic region; sex; baseline OSA severity Group; treatment
Body weightBaseline; geographic region; sex; baseline OSA severity Group; treatment
SBPBaseline; geographic region; sex; baseline OSA severity Group; treatment

The use of baseline adjustment is especially important for change-from-baseline outcomes. A covariate-adjusted estimate can differ from a simple observed change difference because the model estimates the treatment contrast after accounting for the specified covariates.

16. Why the Binary AHI Endpoints Are Different From the Continuous AHI Endpoint

The primary endpoint evaluates change in AHI, expressed in events per hour in the statistical-analysis records. The secondary ≥50% AHI-reduction endpoint instead asks whether each participant crossed a predefined response threshold.

Continuous-style comparison

The primary analysis estimates an adjusted difference in mean change. Participants contribute information according to their measured AHI change.

Binary comparison

The ≥50% endpoint classifies each participant as meeting or not meeting the response criterion and uses logistic regression with risk difference as the reported effect measure.

These endpoints answer related but distinct questions. A continuous change analysis preserves more information about the magnitude of change, whereas a binary threshold focuses on whether a prespecified level of improvement was achieved.

17. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The two posted primary analyses report negative adjusted AHI change differences with 95% confidence intervals entirely below zero and P-values <0.001.

Clinical interpretation

The ClinicalTrials.gov record shows differences in AHI and several related secondary endpoints, including binary AHI response, sleep-related outcomes, body weight, hsCRP, and SBP. Clinical meaning should be considered endpoint by endpoint rather than reduced to a single statistical number.

18. Important Limitations and Interpretation Issues

19. Why This Trial Matters Statistically

SURMOUNT-OSA is a useful teaching case because its registry results combine randomized treatment comparisons with several distinct statistical estimands. The primary endpoint uses ANCOVA and a model-based mean-change contrast, while important secondary endpoints use logistic regression and risk differences. Other outcomes use median differences or additional ANCOVA models.

ConceptHow it appears in SURMOUNT-OSA
RandomizationThe study is registered as RANDOMIZED.
BlindingThe study is registered as DOUBLE masked.
Parallel designThe design model is PARALLEL.
ANCOVAUsed for the primary AHI analysis and multiple secondary continuous outcomes.
Covariate adjustmentBaseline, geographic region, sex, treatment, and in selected secondary analyses baseline OSA severity Group.
Logistic regressionUsed for binary AHI-response endpoints.
Risk differenceReported for the ≥50% AHI-reduction and AHI/ESS composite endpoints.
Confidence intervals95% two-sided intervals are reported for the posted estimates.
Modified intent-to-treatPrimary and secondary analyses use an mITT population defined by exposure and evaluability.
Multiple endpoints20 statistical analyses are posted across primary and secondary outcomes.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Use the methods in this trial as a starting point for deeper study of ANCOVA, logistic regression, confidence intervals, covariate adjustment, randomization, and risk differences.

23. Record Summary

SURMOUNT-OSA provides a clear example of how a modern randomized trial can use different statistical estimands for different clinical questions. The primary endpoint, Change From Baseline in Apnea-Hypopnea Index at Baseline and Week 52, was analysed with ANCOVA in an mITT population. The reported LS Mean Change differences were -20.01 for Tirzepatide MTD_GPI1 versus Placebo_GPI1 and -23.77 for Tirzepatide MTD_GPI2 versus Placebo_GPI2, with 95% confidence intervals of -25.82 to -14.20 and -29.61 to -17.93, respectively, and P-values <0.001 for both comparisons.

The secondary analyses broaden the statistical picture. AHI percentage change was evaluated with ANCOVA; binary AHI-response endpoints were evaluated with logistic regression and reported as risk differences; SASHB was reported using median differences; PROMIS, body-weight, hsCRP, and SBP outcomes were evaluated with ANCOVA. The ClinicalTrials.gov record reports serious adverse events separately for all four arms.

The most important statistical lesson is that the treatment effect cannot be reduced to a single P-value. Proper interpretation requires the effect estimate, confidence interval, analysis population, covariate structure, endpoint definition, and multiplicity context to be considered together.

Clinical Biostats methodology: This page separates the numerical results reported in the ClinicalTrials.gov record from statistical interpretation. Where the registry does not provide enough information to identify a method, imputation strategy, multiplicity procedure, or arm-level characteristic, no additional method or result is inferred.