← Clinical Trials
COPD Phase 3 12-Week Treatment NCT02347761

GOLDEN-3: Complete Statistical Analysis of Nebulized SUN-101 in COPD

An independent statistical analysis of the randomized phase 3 GOLDEN-3 trial evaluating 12 weeks of treatment with nebulized SUN-101 in patients with COPD, focusing on change from baseline in trough FEV1 at Week 12 and the trial's longitudinal mixed-model methodology.

Trial status: COMPLETED  ·  Start: 2015-02  ·  Primary completion: 2015-11
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the GOLDEN-3 trial data from ClinicalTrials.gov. The registry record is the official source for the trial's design and posted results.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

GOLDEN-3 was a randomized, parallel, quadruple-masked, phase 3 trial evaluating two doses of nebulized SUN-101 against placebo in patients with COPD. The registry reports an enrollment of 653 subjects and two registered primary endpoint entries concerning change from baseline in trough FEV1 at Week 12.

653
Enrollment
Patients
3
Treatment arms
2 SUN-101 doses + placebo
12 wk
Primary time point
Baseline and Week 12
<0.0001
Primary p-values
Reported for all 4 analyses
FeatureGOLDEN-3
PhasePhase 3
ConditionCOPD
Brief titleEfficacy and Safety Trial of 12 Weeks of Treatment With Nebulized SUN-101 in Patients With COPD
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment653
Lead sponsorSunovion Respiratory Development Inc.
Sponsor typeIndustry
ClinicalTrials.govNCT02347761

2. Clinical Question

The central statistical question was whether treatment with either dose of SUN-101 produced a greater improvement in change from baseline in trough FEV1 at Week 12 than placebo in patients with COPD.

Population

Patients with COPD enrolled in the phase 3 GOLDEN-3 trial. The ClinicalTrials.gov record reports 653 enrolled subjects.

Intervention

SUN-101 50 mcg BID delivered by eFlow closed-system nebulizer, or SUN-101 25 mcg BID delivered by the same nebulizer.

Comparator

Placebo BID delivered by an eFlow closed-system nebulizer.

Primary question

Does either SUN-101 dose produce a greater change from baseline in trough FEV1 at Week 12 than placebo?

3. Trial Design

01
Randomize653 enrolled
02
3 armsTwo SUN-101 doses + placebo
03
12 weeksTreatment period
04
SpirometryTrough FEV1 assessment
05
CompareChange from baseline
ARM A · 50 MCG

SUN-101 50 mcg BID

  • SUN-101 50 mcg BID
  • eFlow closed-system nebulizer
  • Drug intervention
ARM B · 25 MCG

SUN-101 25 mcg BID

  • SUN-101 25 mcg BID
  • eFlow closed-system nebulizer
  • Drug intervention
ARM C · CONTROL

Placebo BID

  • Placebo BID
  • eFlow closed-system nebulizer
  • Placebo intervention
Allocation
Randomized
Randomization reduces systematic differences between treatment groups at the point of assignment and provides the basis for an intention-to-treat treatment comparison.
Masking
Quadruple
The registry identifies the study as quadruple-masked. This reduces opportunities for knowledge of assignment to influence trial conduct or assessment.
Design model
Parallel
The registry classifies GOLDEN-3 as a parallel-group study rather than a crossover or factorial design.
Primary purpose
Treatment
The registered primary purpose was treatment, consistent with the comparison of active SUN-101 doses against placebo.

4. Endpoints

The registry lists two primary endpoint entries. Both concern the same clinical construct: change from baseline in trough FEV1 at Week 12. The ClinicalTrials.gov record distinguishes one entry as “at Week 12” and another as “Week 12,” with both using the time frame “Baseline and Week 12.”

Registered endpointTime frameRegistry description reported
Change From Baseline in Trough Forced Expiratory Volume in 1 Second (FEV1) at Week 12 Baseline and Week 12 All collected. Spirometry was performed according to internationally accepted standards. Trough FEV1 at Week 12 was defined as the mean of the values collected at two time points 30 minutes apart at approximately 24 hours (± 1 hour) after the previous morning dose. Change from baseline in trough FEV1 was calculated as the trough FEV1 value at Week 12 minus the morning trough FEV1 at baseline.
Change From Baseline in Trough Forced Expiratory Volume in 1 Second (FEV1) Week 12 Baseline and Week 12 On treatment. Spirometry was performed according to internationally accepted standards. Trough FEV1 at Week 12 was defined as the mean of the values collected at two time points 30 minutes apart at approximately 24 hours (± 1 hour) after the previous morning dose. Change from baseline in trough FEV1 was calculated as the trough FEV1 value at Week 12 minus the morning trough FEV1 at baseline.

The ClinicalTrials.gov record classifies the posted endpoint analyses as continuous analyses even though the normalized registry endpoint-type field is inconsistent across the two registered primary endpoint entries. For the statistical results, the posted analyses use continuous-effect estimates expressed as least-squares mean differences.

Endpoint interpretation: FEV1 is measured in liters, while the analysis compares the modeled change from baseline between a SUN-101 dose and placebo. A positive treatment difference therefore indicates greater modeled improvement in trough FEV1 for SUN-101 relative to placebo.

5. Analysis Populations

The posted statistical analyses use the Intent to Treat (ITT) Population. The registry definition states that this population consisted of all subjects who were randomized to treatment and received at least one dose of study medication, with subjects analyzed according to the treatment to which they were randomized.

PopulationRole in the posted analysis
Intent to Treat (ITT) Primary efficacy analysis population. Subjects were analyzed according to randomized treatment assignment.
Safety population The ClinicalTrials.gov record provides serious adverse-event counts by treatment arm, but do not provide a separately named safety-population definition.

The ITT principle is particularly important in a randomized comparison because the treatment effect is defined by assignment rather than by an attempt to reconstruct what would have happened had patients remained perfectly adherent. The registry analysis population also requires receipt of at least one dose, so the exact population is more specifically a randomized-and-treated ITT population rather than an unrestricted set of every randomized subject.

6. Statistical Methodology

Mixed Model for Repeated Measures

The principal analysis method reported for the primary FEV1 endpoint was a mixed model for repeated measures (MMRM). This is appropriate when the same endpoint is measured repeatedly over visits and the analysis needs to account for the correlation among observations from the same patient.

Conceptual model
FEV1 change = treatment + cardiovascular risk + background LABA use + visit + visit × treatment + baseline FEV1 + repeated-measure error

The registry-reported analysis notes identify treatment, cardiovascular risk, background LABA use, visit, visit-by-treatment interaction, and baseline FEV1 as model terms or covariates. An unstructured covariance matrix was used.

Covariate adjustment

The model adjusted for baseline FEV1 and included cardiovascular risk and background LABA use. Baseline adjustment can improve precision by accounting for pre-treatment differences in the outcome measure, while prespecified clinical factors can account for systematic sources of variation.

Visit and treatment-by-visit interaction

The inclusion of visit and the visit-by-treatment interaction allows the estimated treatment difference to vary over time rather than forcing one common treatment effect across every measurement occasion. This is an important distinction from a simple comparison of one unadjusted mean at one time point.

Unstructured covariance

The registry-reported analysis notes specify an unstructured covariance matrix. This approach allows the variances and covariances among repeated measurements to be estimated without imposing a simpler correlation pattern. The trade-off is that more covariance parameters must be estimated.

Linear regression / least-squares means

One posted analysis is normalized as a linear regression analysis and reports a least-squares mean difference. In an adjusted linear model, the least-squares mean is a model-based adjusted mean rather than simply the arithmetic average of the observed values. This makes the treatment comparison conditional on the covariates and model structure used in the analysis.

Multiplicity control

The registry-reported analysis notes state that the Hochberg procedure, described in the registry text as a tree-structured gatekeeping procedure, was used to control the family-wise Type I error rate for comparisons involving the primary efficacy endpoints and key secondary efficacy endpoints.

Why this matters: GOLDEN-3 tested more than one dose against placebo. Without an appropriate multiplicity strategy, performing several hypothesis tests at the same nominal significance level can increase the probability of at least one false-positive finding. The registry specifically identifies a Hochberg-based procedure for family-wise error control.

7. Primary Results

The registry contains four posted statistical analyses for the primary endpoint construct. Two analyses compare SUN-101 50 mcg BID with placebo and two compare SUN-101 25 mcg BID with placebo. The estimates are all expressed as mean differences in liters, with two-sided 95% confidence intervals and P-values reported as <0.0001.

7.1 SUN-101 50 mcg BID vs Placebo — MMRM Analysis

Least-squares mean difference in change from baseline

0.1036 L

95% CI: 0.0663 to 0.1409   ·   P < 0.0001

MMRM; ITT population; superiority comparison

FeatureReported result
ComparisonSUN-101 50 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer
EndpointChange From Baseline in Trough FEV1 at Week 12
UnitLiters
Analysis populationIntent to Treat (ITT)
MethodMMRM (mixed model for repeated measures)
Effect measureLeast Squares Mean Difference (SE), normalized as mean difference
Estimate0.1036
95% CI0.0663 to 0.1409
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The estimated treatment difference was 0.1036 liters, meaning the model estimated a 0.1036-liter greater change from baseline in trough FEV1 for the 50 mcg SUN-101 group than for placebo at the Week 12 endpoint under the specified repeated-measures model.

The estimate is not a statement that every patient experienced a 0.1036-liter improvement. It is a population-level, model-based treatment contrast. Individual changes can be substantially larger, smaller, or in the opposite direction.

The two-sided 95% confidence interval of 0.0663 to 0.1409 liters describes uncertainty around the estimated treatment contrast under the model and sampling framework. It does not describe the range of individual patient responses.

The P-value of <0.0001 addresses the evidence against the null hypothesis under the specified testing framework. It does not measure the magnitude or clinical importance of the effect. Effect size is communicated by the estimate and its confidence interval.

Because the analysis is longitudinal and adjusted, interpretation depends on the specified model, including its covariance structure and covariates. The reported result should also be interpreted within the trial's multiplicity-control procedure because more than one primary efficacy comparison was made.

7.2 SUN-101 25 mcg BID vs Placebo — MMRM Analysis

Least-squares mean difference in change from baseline

0.0961 L

95% CI: 0.0589 to 0.1334   ·   P < 0.0001

MMRM; ITT population; superiority comparison

FeatureReported result
ComparisonSUN-101 25 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer
EndpointChange From Baseline in Trough FEV1 at Week 12
UnitLiters
Analysis populationIntent to Treat (ITT)
MethodMMRM (mixed model for repeated measures)
Effect measureLeast Squares Mean Difference (SE), normalized as mean difference
Estimate0.0961
95% CI0.0589 to 0.1334
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The estimated treatment difference was 0.0961 liters. In the fitted model, the 25 mcg SUN-101 group therefore had an estimated change from baseline in trough FEV1 that was 0.0961 liters greater than the placebo group at the Week 12 endpoint.

Again, this is a mean treatment contrast, not a prediction of an individual patient's response. The estimate summarizes the adjusted group comparison produced by the statistical model.

The 95% confidence interval of 0.0589 to 0.1334 liters provides the reported precision around the estimate. Its relatively narrow range indicates that the registry's estimated treatment contrast was not represented by a single highly uncertain point estimate.

The P-value of <0.0001 indicates strong evidence against the null hypothesis under the prespecified superiority framework, but it does not tell us that the treatment effect is “<0.0001” in size. P-values and effect estimates answer different questions.

Because this is an MMRM analysis, the interpretation depends on the longitudinal model and its assumptions. The ClinicalTrials.gov record specifies an unstructured covariance matrix and adjustment for baseline FEV1 and other listed factors. Multiplicity is also relevant because the trial compared two active doses with placebo.

7.3 SUN-101 50 mcg BID vs Placebo — Least-Squares Mean Analysis

Least-squares mean difference

0.1264 L

95% CI: 0.0856 to 0.1672   ·   P < 0.0001

Linear-model analysis; ITT population; superiority comparison

FeatureReported result
ComparisonSUN-101 50 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer
EndpointChange From Baseline in Trough FEV1 Week 12
UnitLiters
Analysis populationIntent to Treat (ITT)
MethodLinear regression / least-squares mean
Effect measureLeast Squares Mean (SE), normalized as mean difference
Estimate0.1264
95% CI0.0856 to 0.1672
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

This analysis gives an estimated treatment difference of 0.1264 liters for the 50 mcg dose versus placebo. The estimate is a model-based least-squares mean contrast rather than a raw difference between two unadjusted sample means.

The 95% confidence interval extends from 0.0856 to 0.1672 liters. The interval quantifies uncertainty around the estimated adjusted difference; it should not be read as a prediction interval for individual patients.

The P-value of <0.0001 provides evidence against the null hypothesis in the reported superiority analysis. It does not quantify the clinical magnitude of the improvement and should not be used as a substitute for examining the estimate and confidence interval.

The registry also describes repeated-measures model terms in the analysis notes, including treatment, cardiovascular risk, background LABA use, visit, visit-by-treatment interaction, and baseline FEV1. This makes the model specification more informative than simply labeling the result a “least-squares mean.”

7.4 SUN-101 25 mcg BID vs Placebo — MMRM Analysis

Least-squares mean difference in change from baseline

0.1052 L

95% CI: 0.0647 to 0.1457   ·   P < 0.0001

MMRM; ITT population; superiority comparison

FeatureReported result
ComparisonSUN-101 25 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer
EndpointChange From Baseline in Trough FEV1 Week 12
UnitLiters
Analysis populationIntent to Treat (ITT)
MethodMMRM (mixed model for repeated measures)
Effect measureLS mean (SE), normalized as mean difference
Estimate0.1052
95% CI0.0647 to 0.1457
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The reported estimate was 0.1052 liters, representing the model-based difference in change from baseline in trough FEV1 between the 25 mcg SUN-101 group and placebo.

The 95% confidence interval of 0.0647 to 0.1457 liters describes uncertainty around the treatment contrast. The interval remains above zero, consistent with the reported superiority P-value of <0.0001.

The P-value does not tell us how large the treatment effect is. The clinically interpretable quantities are the estimated difference and its uncertainty, together with the endpoint definition and measurement schedule.

The analysis is also not a simple two-sample comparison: the registry notes identify repeated measures, covariate adjustment, an unstructured covariance matrix, and visit-by-treatment interaction. These features should be considered when interpreting the estimate.

8. Primary Results Side by Side

ComparisonMethodEstimate95% CIP-value
SUN-101 50 mcg vs placebo MMRM 0.1036 0.0663 to 0.1409 <0.0001
SUN-101 25 mcg vs placebo MMRM 0.0961 0.0589 to 0.1334 <0.0001
SUN-101 50 mcg vs placebo Linear regression / least-squares mean 0.1264 0.0856 to 0.1672 <0.0001
SUN-101 25 mcg vs placebo MMRM 0.1052 0.0647 to 0.1457 <0.0001

All four posted primary analyses report positive mean differences favoring the corresponding SUN-101 dose over placebo. The estimates are not identical because the registry contains separate statistical-analysis records with different reported methods or analysis descriptions. They should therefore be presented as reported analyses rather than combined into a single derived estimate.

Do not average the estimates. The four values arise from separate posted analyses. Averaging them would create a new statistic that is not reported by the registry and would ignore differences in the analysis specifications.

9. Understanding the FEV1 Effect Measure

Mean-difference interpretation
Treatment effect = adjusted mean change in SUN-101 group − adjusted mean change in placebo group

For the reported analyses, a positive value means the modeled change from baseline in trough FEV1 was greater in the SUN-101 group than in the placebo group.

The key feature of the effect measure is that it is expressed in liters. This is fundamentally different from a hazard ratio or odds ratio. A mean difference preserves the measurement scale of FEV1 and therefore describes an absolute difference in the endpoint itself.

The term least-squares mean signals that the reported value comes from an adjusted statistical model. It should not automatically be interpreted as the ordinary observed mean for a treatment arm. The model estimates what the treatment-group mean would be after accounting for the covariates and repeated-measure structure specified by the analysis.

This distinction matters because the trial was not analyzed as a single unadjusted Week 12 comparison. The MMRM framework uses information from the repeated-measure structure and models treatment differences over visits through the treatment-by-visit interaction.

10. Why an MMRM Was Used

FEV1 is a longitudinal outcome: measurements can be collected repeatedly during the study. A mixed model for repeated measures is designed for precisely this setting because it models the relationship among repeated observations from the same patient rather than treating them as independent observations.

Repeated observations

Measurements from the same subject are correlated. An MMRM explicitly models that within-subject structure.

Time dependence

The model includes visit and visit-by-treatment interaction, allowing treatment differences to be evaluated within a longitudinal framework.

Baseline adjustment

Baseline FEV1 is included as a covariate, improving the precision of the estimated treatment comparison when baseline values explain outcome variability.

Unstructured covariance

The registry-reported analysis specifies an unstructured covariance matrix, allowing the repeated measurements to have flexible variances and covariances.

11. Covariate Adjustment

The registry-reported analysis notes identify four important classes of model terms: treatment, cardiovascular risk, background LABA use, visit, and visit-by-treatment interaction, together with baseline FEV1 as a covariate.

Model componentStatistical role
TreatmentDefines the principal SUN-101 versus placebo comparison.
Cardiovascular riskAdjustment factor included in the repeated-measures model.
Background LABA useAdjustment factor included in the repeated-measures model.
VisitRepresents the longitudinal measurement structure.
Visit × treatmentAllows treatment differences to vary across visits.
Baseline FEV1Covariate accounting for baseline lung-function level.

Covariate adjustment does not mean that randomization has been replaced by regression. Randomization remains the foundation for causal comparison. The regression model is used to obtain a more precise adjusted estimate and to account for clinically relevant sources of variation identified in the statistical analysis.

12. Multiplicity and the Hochberg Procedure

The trial included two SUN-101 doses and placebo, producing more than one active-versus-control efficacy comparison. The registry analysis notes state that the Hochberg procedure, described there as a tree-structured gatekeeping procedure, was used to control the family-wise Type I error rate for comparisons of the primary efficacy endpoints and key secondary efficacy endpoints.

Why multiplicity exists

When several hypotheses are tested, the chance of obtaining at least one apparently positive result can increase if no multiplicity strategy is used.

What family-wise control means

The objective is to control the probability of making one or more Type I errors across the prespecified family of hypotheses.

The important point is that the individual P-values should not be interpreted in isolation from the trial's multiplicity framework. A P-value is a property of a hypothesis test, while the decision about whether a collection of tests preserves the desired family-wise error rate depends on the prespecified testing procedure.

13. Statistical Methods Explained

Why was an MMRM used for FEV1?

Because FEV1 is measured longitudinally, observations from the same patient are related. MMRM accounts for that repeated-measure structure and permits treatment effects to be modeled over visits. The registry-reported analysis also includes a visit-by-treatment interaction, which allows the treatment contrast to vary with time.

What does a mean difference of 0.1036 mean?

It means the estimated adjusted change from baseline in trough FEV1 was 0.1036 liters greater with SUN-101 50 mcg BID than with placebo in the corresponding posted MMRM analysis. It does not mean that every individual patient's FEV1 increased by 0.1036 liters.

Why use a least-squares mean instead of a raw mean?

A least-squares mean is derived from the fitted statistical model. It incorporates the covariate adjustment and model structure specified for the analysis. A raw arithmetic mean would not reflect those adjustments.

What does the 95% confidence interval tell us?

The 95% confidence interval describes statistical uncertainty around the estimated treatment difference under the specified model and sampling framework. For the 50 mcg MMRM analysis, the interval is 0.0663 to 0.1409 liters. It is not a range containing 95% of individual patient responses.

Why doesn't the P-value measure treatment effect size?

The P-value measures evidence against a null hypothesis under a specified statistical model. It is affected by both effect magnitude and information in the data. The estimated mean difference tells us the size of the treatment contrast, while the confidence interval describes its precision.

Why does multiplicity matter when both doses have small P-values?

Because multiple hypotheses were evaluated. A prespecified multiplicity procedure such as the reported Hochberg procedure is used to maintain the desired family-wise Type I error control across the relevant family of comparisons.

What does an unstructured covariance matrix contribute?

It gives the repeated-measures model flexibility to estimate different variances and covariances among measurement occasions rather than imposing a simpler fixed correlation pattern. This flexibility can be useful when the longitudinal covariance structure is not known in advance.

14. What the Confidence Intervals Say

ComparisonEstimate95% confidence intervalInterpretive focus
50 mcg vs placebo 0.1036 0.0663 to 0.1409 Uncertainty around the MMRM treatment contrast
25 mcg vs placebo 0.0961 0.0589 to 0.1334 Uncertainty around the MMRM treatment contrast
50 mcg vs placebo 0.1264 0.0856 to 0.1672 Uncertainty around the least-squares mean contrast
25 mcg vs placebo 0.1052 0.0647 to 0.1457 Uncertainty around the MMRM treatment contrast

All four reported confidence intervals lie above zero. This is consistent with the corresponding reported superiority P-values of <0.0001. The intervals also provide information that a P-value alone cannot: they show the range of treatment-effect values that remain compatible with the statistical uncertainty represented by the model.

Importantly, the confidence intervals do not establish that the treatment effect is clinically important. Statistical precision and clinical importance are separate questions. A precise estimate can still represent an effect whose practical importance requires clinical context, and the ClinicalTrials.gov record does not provide a clinical-importance threshold.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm. The available safety information is presented as affected subjects over subjects at risk.

Treatment armSerious adverse eventsAffected / at risk
SUN-101 50 mcg BID eFlow (CS) Nebulizer Serious adverse events 10 / 218
SUN-101 25 mcg BID eFlow (CS) Nebulizer Serious adverse events 8 / 217
Placebo BID eFlow (CS) Nebulizer Serious adverse events 11 / 218

These figures describe the number of subjects affected and the corresponding number at risk reported in the ClinicalTrials.gov record. They should not be converted into an additional statistical comparison here because the ClinicalTrials.gov record does not provide a formal statistical analysis, confidence interval, or hypothesis test for serious adverse events.

Safety interpretation: a serious-adverse-event count is not directly comparable with the continuous FEV1 treatment effect. Efficacy and safety address different outcomes and require different statistical summaries. The ClinicalTrials.gov record supports reporting the serious-adverse-event counts by arm, but not a formal between-arm safety-effect estimate.

16. Design Topics Not Supported by the Supplied Registry Data

TopicWhat can be established from the ClinicalTrials.gov record
Non-inferiority marginNot applicable to the reported hypothesis type. The posted analyses are identified as superiority analyses.
CrossoverThe ClinicalTrials.gov record does not report a crossover design or crossover analysis.
Factorial designThe design model is reported as parallel; the ClinicalTrials.gov record does not identify a factorial design.
Interim analysisThe ClinicalTrials.gov record does not report an interim efficacy analysis or interim stopping boundary.
Missing-data imputationThe ClinicalTrials.gov record does not specify an imputation method. The MMRM framework is reported, but no additional missing-data procedure is reported.
Bayesian methodsNo Bayesian analysis is reported in the ClinicalTrials.gov record.

This distinction is important for a trial-results page. Absence of a registry-reported registry detail should not be converted into an invented methodological claim. The statistical analysis can be described confidently where the registry provides the method, while unsupported topics should remain limited to what the record establishes.

17. Limitations

18. Why This Trial Matters Statistically

GOLDEN-3 is a useful teaching case because its primary endpoint illustrates a common clinical-trial situation: a continuous physiologic outcome measured repeatedly over time, analyzed using a model that adjusts for baseline and other clinical factors while accounting for within-patient correlation.

ConceptHow it appears in GOLDEN-3
RandomizationThe trial is randomized, providing the foundation for treatment-group comparison.
BlindingThe registry classifies the trial as quadruple-masked.
Parallel designThe trial is classified as a parallel-group study.
ITT analysisPrimary posted analyses use the Intent to Treat population as defined in the registry.
Continuous endpointChange from baseline in trough FEV1 is analyzed in liters.
MMRMThe principal statistical method is a mixed model for repeated measures.
Covariate adjustmentBaseline FEV1 and specified clinical factors are incorporated into the model.
Visit-by-treatment interactionThe model permits treatment differences to vary across visits.
Unstructured covarianceThe repeated-measure covariance structure is modeled flexibly.
Least-squares meanReported effects are adjusted model-based treatment contrasts.
Confidence intervalEach posted primary analysis includes a two-sided 95% CI.
P-valueAll four posted primary analyses report P < 0.0001.
MultiplicityThe registry reports the Hochberg procedure for family-wise Type I error control.
Superiority testingAll four statistical analyses are classified as superiority analyses.

19. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The posted analyses estimate positive treatment differences in change from baseline in trough FEV1 for both SUN-101 doses versus placebo, with two-sided 95% confidence intervals and P-values <0.0001. The analyses use adjusted models rather than simple unadjusted comparisons.

Clinical interpretation

The ClinicalTrials.gov record establishes the direction and magnitude of the reported FEV1 treatment contrasts, but they do not supply a clinical threshold for determining how important a particular FEV1 difference is for an individual patient.

This distinction is central to responsible interpretation. A statistically precise difference can establish that the randomized groups differ under the prespecified model without, by itself, establishing how meaningful that difference is to a patient's symptoms, daily function, or long-term outcomes.

20. A Closer Look at the Two SUN-101 Doses

The registry contains estimates for both active doses. The reported MMRM estimates are 0.1036 for the 50 mcg comparison and 0.0961 for the 25 mcg comparison in their respective posted analyses.

Reported MMRM mean differences
SUN-101 50 mcg
0.1036 L
SUN-101 25 mcg
0.0961 L

The visual comparison is descriptive only. It should not be interpreted as a formal test that the two SUN-101 doses differ from each other. The statistical analyses posted on ClinicalTrials.gov compare each dose with placebo; they do not provide a formal SUN-101 50 mcg versus SUN-101 25 mcg hypothesis test.

This is an important general statistical principle: two treatment effects being statistically significant versus control does not automatically imply that the two treatment effects are statistically different from one another. A direct dose-to-dose comparison would require its own estimated contrast and uncertainty measure.

21. P-values, Estimates, and Confidence Intervals

Estimate

The estimate answers: How large is the modeled treatment difference? For example, 0.1036 is the reported mean difference in change from baseline in trough FEV1 for one of the 50 mcg analyses.

Confidence interval

The confidence interval answers: How precisely has the treatment difference been estimated? The interval incorporates sampling uncertainty around the point estimate under the specified analysis.

P-value

The P-value answers a different question: How compatible are the observed data with the null hypothesis under the specified test? It is not an effect-size measure and should not be used to rank the four estimates by clinical importance.

For GOLDEN-3, these three quantities should be read together. The estimates establish the direction and scale of the reported treatment contrasts, the confidence intervals show their precision, and the P-values quantify evidence against the corresponding null hypotheses within the statistical testing framework.

22. Statistical Analysis Workflow

01
RandomizationAssign patients to treatment
02
BaselineMeasure FEV1
03
Repeated visitsCollect longitudinal data
04
MMRMAdjust and model correlation
05
ContrastSUN-101 vs placebo

The statistical workflow illustrates why the final estimate should not be reduced to a simple subtraction of two observed means. The model uses treatment assignment, baseline FEV1, specified clinical covariates, visit, the treatment-by-visit interaction, and an unstructured covariance matrix to estimate the treatment contrast.

The final inferential step is then performed within the trial's multiplicity framework. This is the point at which the reported P-values and confidence intervals become part of the broader confirmatory testing strategy rather than isolated numerical findings.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Calculators

25. Sources

Continue through the Clinical Biostats statistical learning pathway

Explore the methods behind randomized trials, longitudinal models, covariate adjustment, confidence intervals, and multiplicity through tutorials and statistical tools.

26. Record Summary

GOLDEN-3 provides a clear example of longitudinal clinical-trial analysis in which a continuous physiologic endpoint is measured repeatedly and analyzed using adjusted mixed models. The trial was randomized, parallel, and quadruple-masked, with three treatment arms and 653 enrolled subjects. Its primary efficacy analyses used an ITT population and reported mean differences in change from baseline in trough FEV1 at Week 12.

The statistical story is more informative than the P-values alone. The reported treatment differences ranged from 0.0961 to 0.1264 liters across the four posted primary analyses, with two-sided 95% confidence intervals posted on ClinicalTrials.gov for every analysis and P-values of <0.0001. The principal MMRM specification incorporated treatment, cardiovascular risk, background LABA use, visit, visit-by-treatment interaction, and baseline FEV1, with an unstructured covariance matrix.

The trial also illustrates why multiplicity must be considered whenever several hypotheses are evaluated. The registry reports use of a Hochberg procedure to control the family-wise Type I error rate for the primary efficacy endpoints and key secondary efficacy endpoints. Consequently, the inferential results should be understood as components of a prespecified testing strategy rather than as four unrelated P-values.

Finally, the available safety data show serious adverse events in 10/218 subjects receiving SUN-101 50 mcg, 8/217 receiving SUN-101 25 mcg, and 11/218 receiving placebo. These figures complement the efficacy analysis but do not constitute a formal comparative safety analysis because the ClinicalTrials.gov record does not provide a corresponding statistical test or confidence interval.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical analysis from additional interpretation. For GOLDEN-3, that means preserving the registry's endpoint definitions, analysis populations, model specifications, effect estimates, confidence intervals, P-values, multiplicity information, and available safety counts while avoiding unsupported claims about clinical importance or unreported statistical procedures.