This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the GOLDEN-3 trial data from ClinicalTrials.gov. The registry record is the official source for the trial's design and posted results.
1. Trial at a Glance
GOLDEN-3 was a randomized, parallel, quadruple-masked, phase 3 trial evaluating two doses of nebulized SUN-101 against placebo in patients with COPD. The registry reports an enrollment of 653 subjects and two registered primary endpoint entries concerning change from baseline in trough FEV1 at Week 12.
| Feature | GOLDEN-3 |
|---|---|
| Phase | Phase 3 |
| Condition | COPD |
| Brief title | Efficacy and Safety Trial of 12 Weeks of Treatment With Nebulized SUN-101 in Patients With COPD |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 653 |
| Lead sponsor | Sunovion Respiratory Development Inc. |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT02347761 |
2. Clinical Question
The central statistical question was whether treatment with either dose of SUN-101 produced a greater improvement in change from baseline in trough FEV1 at Week 12 than placebo in patients with COPD.
Population
Patients with COPD enrolled in the phase 3 GOLDEN-3 trial. The ClinicalTrials.gov record reports 653 enrolled subjects.
Intervention
SUN-101 50 mcg BID delivered by eFlow closed-system nebulizer, or SUN-101 25 mcg BID delivered by the same nebulizer.
Comparator
Placebo BID delivered by an eFlow closed-system nebulizer.
Primary question
Does either SUN-101 dose produce a greater change from baseline in trough FEV1 at Week 12 than placebo?
3. Trial Design
SUN-101 50 mcg BID
- SUN-101 50 mcg BID
- eFlow closed-system nebulizer
- Drug intervention
SUN-101 25 mcg BID
- SUN-101 25 mcg BID
- eFlow closed-system nebulizer
- Drug intervention
Placebo BID
- Placebo BID
- eFlow closed-system nebulizer
- Placebo intervention
Randomization reduces systematic differences between treatment groups at the point of assignment and provides the basis for an intention-to-treat treatment comparison.
The registry identifies the study as quadruple-masked. This reduces opportunities for knowledge of assignment to influence trial conduct or assessment.
The registry classifies GOLDEN-3 as a parallel-group study rather than a crossover or factorial design.
The registered primary purpose was treatment, consistent with the comparison of active SUN-101 doses against placebo.
4. Endpoints
The registry lists two primary endpoint entries. Both concern the same clinical construct: change from baseline in trough FEV1 at Week 12. The ClinicalTrials.gov record distinguishes one entry as “at Week 12” and another as “Week 12,” with both using the time frame “Baseline and Week 12.”
| Registered endpoint | Time frame | Registry description reported |
|---|---|---|
| Change From Baseline in Trough Forced Expiratory Volume in 1 Second (FEV1) at Week 12 | Baseline and Week 12 | All collected. Spirometry was performed according to internationally accepted standards. Trough FEV1 at Week 12 was defined as the mean of the values collected at two time points 30 minutes apart at approximately 24 hours (± 1 hour) after the previous morning dose. Change from baseline in trough FEV1 was calculated as the trough FEV1 value at Week 12 minus the morning trough FEV1 at baseline. |
| Change From Baseline in Trough Forced Expiratory Volume in 1 Second (FEV1) Week 12 | Baseline and Week 12 | On treatment. Spirometry was performed according to internationally accepted standards. Trough FEV1 at Week 12 was defined as the mean of the values collected at two time points 30 minutes apart at approximately 24 hours (± 1 hour) after the previous morning dose. Change from baseline in trough FEV1 was calculated as the trough FEV1 value at Week 12 minus the morning trough FEV1 at baseline. |
The ClinicalTrials.gov record classifies the posted endpoint analyses as continuous analyses even though the normalized registry endpoint-type field is inconsistent across the two registered primary endpoint entries. For the statistical results, the posted analyses use continuous-effect estimates expressed as least-squares mean differences.
5. Analysis Populations
The posted statistical analyses use the Intent to Treat (ITT) Population. The registry definition states that this population consisted of all subjects who were randomized to treatment and received at least one dose of study medication, with subjects analyzed according to the treatment to which they were randomized.
| Population | Role in the posted analysis |
|---|---|
| Intent to Treat (ITT) | Primary efficacy analysis population. Subjects were analyzed according to randomized treatment assignment. |
| Safety population | The ClinicalTrials.gov record provides serious adverse-event counts by treatment arm, but do not provide a separately named safety-population definition. |
The ITT principle is particularly important in a randomized comparison because the treatment effect is defined by assignment rather than by an attempt to reconstruct what would have happened had patients remained perfectly adherent. The registry analysis population also requires receipt of at least one dose, so the exact population is more specifically a randomized-and-treated ITT population rather than an unrestricted set of every randomized subject.
6. Statistical Methodology
Mixed Model for Repeated Measures
The principal analysis method reported for the primary FEV1 endpoint was a mixed model for repeated measures (MMRM). This is appropriate when the same endpoint is measured repeatedly over visits and the analysis needs to account for the correlation among observations from the same patient.
The registry-reported analysis notes identify treatment, cardiovascular risk, background LABA use, visit, visit-by-treatment interaction, and baseline FEV1 as model terms or covariates. An unstructured covariance matrix was used.
Covariate adjustment
The model adjusted for baseline FEV1 and included cardiovascular risk and background LABA use. Baseline adjustment can improve precision by accounting for pre-treatment differences in the outcome measure, while prespecified clinical factors can account for systematic sources of variation.
Visit and treatment-by-visit interaction
The inclusion of visit and the visit-by-treatment interaction allows the estimated treatment difference to vary over time rather than forcing one common treatment effect across every measurement occasion. This is an important distinction from a simple comparison of one unadjusted mean at one time point.
Unstructured covariance
The registry-reported analysis notes specify an unstructured covariance matrix. This approach allows the variances and covariances among repeated measurements to be estimated without imposing a simpler correlation pattern. The trade-off is that more covariance parameters must be estimated.
Linear regression / least-squares means
One posted analysis is normalized as a linear regression analysis and reports a least-squares mean difference. In an adjusted linear model, the least-squares mean is a model-based adjusted mean rather than simply the arithmetic average of the observed values. This makes the treatment comparison conditional on the covariates and model structure used in the analysis.
Multiplicity control
The registry-reported analysis notes state that the Hochberg procedure, described in the registry text as a tree-structured gatekeeping procedure, was used to control the family-wise Type I error rate for comparisons involving the primary efficacy endpoints and key secondary efficacy endpoints.
7. Primary Results
The registry contains four posted statistical analyses for the primary endpoint construct. Two analyses compare SUN-101 50 mcg BID with placebo and two compare SUN-101 25 mcg BID with placebo. The estimates are all expressed as mean differences in liters, with two-sided 95% confidence intervals and P-values reported as <0.0001.
7.1 SUN-101 50 mcg BID vs Placebo — MMRM Analysis
Least-squares mean difference in change from baseline
95% CI: 0.0663 to 0.1409 · P < 0.0001
MMRM; ITT population; superiority comparison
| Feature | Reported result |
|---|---|
| Comparison | SUN-101 50 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer |
| Endpoint | Change From Baseline in Trough FEV1 at Week 12 |
| Unit | Liters |
| Analysis population | Intent to Treat (ITT) |
| Method | MMRM (mixed model for repeated measures) |
| Effect measure | Least Squares Mean Difference (SE), normalized as mean difference |
| Estimate | 0.1036 |
| 95% CI | 0.0663 to 0.1409 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The estimated treatment difference was 0.1036 liters, meaning the model estimated a 0.1036-liter greater change from baseline in trough FEV1 for the 50 mcg SUN-101 group than for placebo at the Week 12 endpoint under the specified repeated-measures model.
The estimate is not a statement that every patient experienced a 0.1036-liter improvement. It is a population-level, model-based treatment contrast. Individual changes can be substantially larger, smaller, or in the opposite direction.
The two-sided 95% confidence interval of 0.0663 to 0.1409 liters describes uncertainty around the estimated treatment contrast under the model and sampling framework. It does not describe the range of individual patient responses.
The P-value of <0.0001 addresses the evidence against the null hypothesis under the specified testing framework. It does not measure the magnitude or clinical importance of the effect. Effect size is communicated by the estimate and its confidence interval.
Because the analysis is longitudinal and adjusted, interpretation depends on the specified model, including its covariance structure and covariates. The reported result should also be interpreted within the trial's multiplicity-control procedure because more than one primary efficacy comparison was made.
7.2 SUN-101 25 mcg BID vs Placebo — MMRM Analysis
Least-squares mean difference in change from baseline
95% CI: 0.0589 to 0.1334 · P < 0.0001
MMRM; ITT population; superiority comparison
| Feature | Reported result |
|---|---|
| Comparison | SUN-101 25 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer |
| Endpoint | Change From Baseline in Trough FEV1 at Week 12 |
| Unit | Liters |
| Analysis population | Intent to Treat (ITT) |
| Method | MMRM (mixed model for repeated measures) |
| Effect measure | Least Squares Mean Difference (SE), normalized as mean difference |
| Estimate | 0.0961 |
| 95% CI | 0.0589 to 0.1334 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The estimated treatment difference was 0.0961 liters. In the fitted model, the 25 mcg SUN-101 group therefore had an estimated change from baseline in trough FEV1 that was 0.0961 liters greater than the placebo group at the Week 12 endpoint.
Again, this is a mean treatment contrast, not a prediction of an individual patient's response. The estimate summarizes the adjusted group comparison produced by the statistical model.
The 95% confidence interval of 0.0589 to 0.1334 liters provides the reported precision around the estimate. Its relatively narrow range indicates that the registry's estimated treatment contrast was not represented by a single highly uncertain point estimate.
The P-value of <0.0001 indicates strong evidence against the null hypothesis under the prespecified superiority framework, but it does not tell us that the treatment effect is “<0.0001” in size. P-values and effect estimates answer different questions.
Because this is an MMRM analysis, the interpretation depends on the longitudinal model and its assumptions. The ClinicalTrials.gov record specifies an unstructured covariance matrix and adjustment for baseline FEV1 and other listed factors. Multiplicity is also relevant because the trial compared two active doses with placebo.
7.3 SUN-101 50 mcg BID vs Placebo — Least-Squares Mean Analysis
Least-squares mean difference
95% CI: 0.0856 to 0.1672 · P < 0.0001
Linear-model analysis; ITT population; superiority comparison
| Feature | Reported result |
|---|---|
| Comparison | SUN-101 50 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer |
| Endpoint | Change From Baseline in Trough FEV1 Week 12 |
| Unit | Liters |
| Analysis population | Intent to Treat (ITT) |
| Method | Linear regression / least-squares mean |
| Effect measure | Least Squares Mean (SE), normalized as mean difference |
| Estimate | 0.1264 |
| 95% CI | 0.0856 to 0.1672 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
This analysis gives an estimated treatment difference of 0.1264 liters for the 50 mcg dose versus placebo. The estimate is a model-based least-squares mean contrast rather than a raw difference between two unadjusted sample means.
The 95% confidence interval extends from 0.0856 to 0.1672 liters. The interval quantifies uncertainty around the estimated adjusted difference; it should not be read as a prediction interval for individual patients.
The P-value of <0.0001 provides evidence against the null hypothesis in the reported superiority analysis. It does not quantify the clinical magnitude of the improvement and should not be used as a substitute for examining the estimate and confidence interval.
The registry also describes repeated-measures model terms in the analysis notes, including treatment, cardiovascular risk, background LABA use, visit, visit-by-treatment interaction, and baseline FEV1. This makes the model specification more informative than simply labeling the result a “least-squares mean.”
7.4 SUN-101 25 mcg BID vs Placebo — MMRM Analysis
Least-squares mean difference in change from baseline
95% CI: 0.0647 to 0.1457 · P < 0.0001
MMRM; ITT population; superiority comparison
| Feature | Reported result |
|---|---|
| Comparison | SUN-101 25 mcg BID eFlow (CS) Nebulizer vs Placebo BID eFlow (CS) Nebulizer |
| Endpoint | Change From Baseline in Trough FEV1 Week 12 |
| Unit | Liters |
| Analysis population | Intent to Treat (ITT) |
| Method | MMRM (mixed model for repeated measures) |
| Effect measure | LS mean (SE), normalized as mean difference |
| Estimate | 0.1052 |
| 95% CI | 0.0647 to 0.1457 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The reported estimate was 0.1052 liters, representing the model-based difference in change from baseline in trough FEV1 between the 25 mcg SUN-101 group and placebo.
The 95% confidence interval of 0.0647 to 0.1457 liters describes uncertainty around the treatment contrast. The interval remains above zero, consistent with the reported superiority P-value of <0.0001.
The P-value does not tell us how large the treatment effect is. The clinically interpretable quantities are the estimated difference and its uncertainty, together with the endpoint definition and measurement schedule.
The analysis is also not a simple two-sample comparison: the registry notes identify repeated measures, covariate adjustment, an unstructured covariance matrix, and visit-by-treatment interaction. These features should be considered when interpreting the estimate.
8. Primary Results Side by Side
| Comparison | Method | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| SUN-101 50 mcg vs placebo | MMRM | 0.1036 | 0.0663 to 0.1409 | <0.0001 |
| SUN-101 25 mcg vs placebo | MMRM | 0.0961 | 0.0589 to 0.1334 | <0.0001 |
| SUN-101 50 mcg vs placebo | Linear regression / least-squares mean | 0.1264 | 0.0856 to 0.1672 | <0.0001 |
| SUN-101 25 mcg vs placebo | MMRM | 0.1052 | 0.0647 to 0.1457 | <0.0001 |
All four posted primary analyses report positive mean differences favoring the corresponding SUN-101 dose over placebo. The estimates are not identical because the registry contains separate statistical-analysis records with different reported methods or analysis descriptions. They should therefore be presented as reported analyses rather than combined into a single derived estimate.
9. Understanding the FEV1 Effect Measure
For the reported analyses, a positive value means the modeled change from baseline in trough FEV1 was greater in the SUN-101 group than in the placebo group.
The key feature of the effect measure is that it is expressed in liters. This is fundamentally different from a hazard ratio or odds ratio. A mean difference preserves the measurement scale of FEV1 and therefore describes an absolute difference in the endpoint itself.
The term least-squares mean signals that the reported value comes from an adjusted statistical model. It should not automatically be interpreted as the ordinary observed mean for a treatment arm. The model estimates what the treatment-group mean would be after accounting for the covariates and repeated-measure structure specified by the analysis.
This distinction matters because the trial was not analyzed as a single unadjusted Week 12 comparison. The MMRM framework uses information from the repeated-measure structure and models treatment differences over visits through the treatment-by-visit interaction.
10. Why an MMRM Was Used
FEV1 is a longitudinal outcome: measurements can be collected repeatedly during the study. A mixed model for repeated measures is designed for precisely this setting because it models the relationship among repeated observations from the same patient rather than treating them as independent observations.
Repeated observations
Measurements from the same subject are correlated. An MMRM explicitly models that within-subject structure.
Time dependence
The model includes visit and visit-by-treatment interaction, allowing treatment differences to be evaluated within a longitudinal framework.
Baseline adjustment
Baseline FEV1 is included as a covariate, improving the precision of the estimated treatment comparison when baseline values explain outcome variability.
Unstructured covariance
The registry-reported analysis specifies an unstructured covariance matrix, allowing the repeated measurements to have flexible variances and covariances.
11. Covariate Adjustment
The registry-reported analysis notes identify four important classes of model terms: treatment, cardiovascular risk, background LABA use, visit, and visit-by-treatment interaction, together with baseline FEV1 as a covariate.
| Model component | Statistical role |
|---|---|
| Treatment | Defines the principal SUN-101 versus placebo comparison. |
| Cardiovascular risk | Adjustment factor included in the repeated-measures model. |
| Background LABA use | Adjustment factor included in the repeated-measures model. |
| Visit | Represents the longitudinal measurement structure. |
| Visit × treatment | Allows treatment differences to vary across visits. |
| Baseline FEV1 | Covariate accounting for baseline lung-function level. |
Covariate adjustment does not mean that randomization has been replaced by regression. Randomization remains the foundation for causal comparison. The regression model is used to obtain a more precise adjusted estimate and to account for clinically relevant sources of variation identified in the statistical analysis.
12. Multiplicity and the Hochberg Procedure
The trial included two SUN-101 doses and placebo, producing more than one active-versus-control efficacy comparison. The registry analysis notes state that the Hochberg procedure, described there as a tree-structured gatekeeping procedure, was used to control the family-wise Type I error rate for comparisons of the primary efficacy endpoints and key secondary efficacy endpoints.
Why multiplicity exists
When several hypotheses are tested, the chance of obtaining at least one apparently positive result can increase if no multiplicity strategy is used.
What family-wise control means
The objective is to control the probability of making one or more Type I errors across the prespecified family of hypotheses.
The important point is that the individual P-values should not be interpreted in isolation from the trial's multiplicity framework. A P-value is a property of a hypothesis test, while the decision about whether a collection of tests preserves the desired family-wise error rate depends on the prespecified testing procedure.
13. Statistical Methods Explained
Why was an MMRM used for FEV1?
Because FEV1 is measured longitudinally, observations from the same patient are related. MMRM accounts for that repeated-measure structure and permits treatment effects to be modeled over visits. The registry-reported analysis also includes a visit-by-treatment interaction, which allows the treatment contrast to vary with time.
What does a mean difference of 0.1036 mean?
It means the estimated adjusted change from baseline in trough FEV1 was 0.1036 liters greater with SUN-101 50 mcg BID than with placebo in the corresponding posted MMRM analysis. It does not mean that every individual patient's FEV1 increased by 0.1036 liters.
Why use a least-squares mean instead of a raw mean?
A least-squares mean is derived from the fitted statistical model. It incorporates the covariate adjustment and model structure specified for the analysis. A raw arithmetic mean would not reflect those adjustments.
What does the 95% confidence interval tell us?
The 95% confidence interval describes statistical uncertainty around the estimated treatment difference under the specified model and sampling framework. For the 50 mcg MMRM analysis, the interval is 0.0663 to 0.1409 liters. It is not a range containing 95% of individual patient responses.
Why doesn't the P-value measure treatment effect size?
The P-value measures evidence against a null hypothesis under a specified statistical model. It is affected by both effect magnitude and information in the data. The estimated mean difference tells us the size of the treatment contrast, while the confidence interval describes its precision.
Why does multiplicity matter when both doses have small P-values?
Because multiple hypotheses were evaluated. A prespecified multiplicity procedure such as the reported Hochberg procedure is used to maintain the desired family-wise Type I error control across the relevant family of comparisons.
What does an unstructured covariance matrix contribute?
It gives the repeated-measures model flexibility to estimate different variances and covariances among measurement occasions rather than imposing a simpler fixed correlation pattern. This flexibility can be useful when the longitudinal covariance structure is not known in advance.
14. What the Confidence Intervals Say
| Comparison | Estimate | 95% confidence interval | Interpretive focus |
|---|---|---|---|
| 50 mcg vs placebo | 0.1036 | 0.0663 to 0.1409 | Uncertainty around the MMRM treatment contrast |
| 25 mcg vs placebo | 0.0961 | 0.0589 to 0.1334 | Uncertainty around the MMRM treatment contrast |
| 50 mcg vs placebo | 0.1264 | 0.0856 to 0.1672 | Uncertainty around the least-squares mean contrast |
| 25 mcg vs placebo | 0.1052 | 0.0647 to 0.1457 | Uncertainty around the MMRM treatment contrast |
All four reported confidence intervals lie above zero. This is consistent with the corresponding reported superiority P-values of <0.0001. The intervals also provide information that a P-value alone cannot: they show the range of treatment-effect values that remain compatible with the statistical uncertainty represented by the model.
Importantly, the confidence intervals do not establish that the treatment effect is clinically important. Statistical precision and clinical importance are separate questions. A precise estimate can still represent an effect whose practical importance requires clinical context, and the ClinicalTrials.gov record does not provide a clinical-importance threshold.
15. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm. The available safety information is presented as affected subjects over subjects at risk.
| Treatment arm | Serious adverse events | Affected / at risk |
|---|---|---|
| SUN-101 50 mcg BID eFlow (CS) Nebulizer | Serious adverse events | 10 / 218 |
| SUN-101 25 mcg BID eFlow (CS) Nebulizer | Serious adverse events | 8 / 217 |
| Placebo BID eFlow (CS) Nebulizer | Serious adverse events | 11 / 218 |
These figures describe the number of subjects affected and the corresponding number at risk reported in the ClinicalTrials.gov record. They should not be converted into an additional statistical comparison here because the ClinicalTrials.gov record does not provide a formal statistical analysis, confidence interval, or hypothesis test for serious adverse events.
16. Design Topics Not Supported by the Supplied Registry Data
| Topic | What can be established from the ClinicalTrials.gov record |
|---|---|
| Non-inferiority margin | Not applicable to the reported hypothesis type. The posted analyses are identified as superiority analyses. |
| Crossover | The ClinicalTrials.gov record does not report a crossover design or crossover analysis. |
| Factorial design | The design model is reported as parallel; the ClinicalTrials.gov record does not identify a factorial design. |
| Interim analysis | The ClinicalTrials.gov record does not report an interim efficacy analysis or interim stopping boundary. |
| Missing-data imputation | The ClinicalTrials.gov record does not specify an imputation method. The MMRM framework is reported, but no additional missing-data procedure is reported. |
| Bayesian methods | No Bayesian analysis is reported in the ClinicalTrials.gov record. |
This distinction is important for a trial-results page. Absence of a registry-reported registry detail should not be converted into an invented methodological claim. The statistical analysis can be described confidently where the registry provides the method, while unsupported topics should remain limited to what the record establishes.
17. Limitations
- Registry-level reporting: the ClinicalTrials.gov record provides selected statistical-analysis records rather than a complete statistical analysis plan.
- Endpoint duplication: the registry contains two primary endpoint entries that describe the same FEV1 construct with slightly different wording, and the posted statistical analyses contain overlapping comparisons.
- Analysis-record heterogeneity: the four posted analyses are not identical. Three are normalized as MMRM analyses and one as a linear-model / least-squares mean analysis, so the estimates should not be combined into one derived result.
- Incomplete population detail: the registry-reported ITT definition specifies randomized subjects who received at least one dose, but the ClinicalTrials.gov record does not provide detailed counts of analyzed subjects by endpoint.
- Missing-data detail: the ClinicalTrials.gov record does not specify an imputation procedure or sensitivity-analysis framework for missing FEV1 observations.
- Model assumptions: MMRM results depend on the specified covariance structure, model terms, and assumptions concerning the repeated observations and missing data.
- Multiplicity: the registry reports a Hochberg-based procedure for family-wise error control, so the primary P-values should be interpreted within that prespecified multiple-testing framework rather than as four unrelated tests.
- Safety detail: only serious-adverse-event counts are reported here. The available data do not support a comprehensive safety-profile analysis.
- Clinical importance: the registry supplies statistical effect estimates but does not provide a clinical threshold for interpreting the magnitude of FEV1 change.
18. Why This Trial Matters Statistically
GOLDEN-3 is a useful teaching case because its primary endpoint illustrates a common clinical-trial situation: a continuous physiologic outcome measured repeatedly over time, analyzed using a model that adjusts for baseline and other clinical factors while accounting for within-patient correlation.
| Concept | How it appears in GOLDEN-3 |
|---|---|
| Randomization | The trial is randomized, providing the foundation for treatment-group comparison. |
| Blinding | The registry classifies the trial as quadruple-masked. |
| Parallel design | The trial is classified as a parallel-group study. |
| ITT analysis | Primary posted analyses use the Intent to Treat population as defined in the registry. |
| Continuous endpoint | Change from baseline in trough FEV1 is analyzed in liters. |
| MMRM | The principal statistical method is a mixed model for repeated measures. |
| Covariate adjustment | Baseline FEV1 and specified clinical factors are incorporated into the model. |
| Visit-by-treatment interaction | The model permits treatment differences to vary across visits. |
| Unstructured covariance | The repeated-measure covariance structure is modeled flexibly. |
| Least-squares mean | Reported effects are adjusted model-based treatment contrasts. |
| Confidence interval | Each posted primary analysis includes a two-sided 95% CI. |
| P-value | All four posted primary analyses report P < 0.0001. |
| Multiplicity | The registry reports the Hochberg procedure for family-wise Type I error control. |
| Superiority testing | All four statistical analyses are classified as superiority analyses. |
19. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The posted analyses estimate positive treatment differences in change from baseline in trough FEV1 for both SUN-101 doses versus placebo, with two-sided 95% confidence intervals and P-values <0.0001. The analyses use adjusted models rather than simple unadjusted comparisons.
Clinical interpretation
The ClinicalTrials.gov record establishes the direction and magnitude of the reported FEV1 treatment contrasts, but they do not supply a clinical threshold for determining how important a particular FEV1 difference is for an individual patient.
This distinction is central to responsible interpretation. A statistically precise difference can establish that the randomized groups differ under the prespecified model without, by itself, establishing how meaningful that difference is to a patient's symptoms, daily function, or long-term outcomes.
20. A Closer Look at the Two SUN-101 Doses
The registry contains estimates for both active doses. The reported MMRM estimates are 0.1036 for the 50 mcg comparison and 0.0961 for the 25 mcg comparison in their respective posted analyses.
The visual comparison is descriptive only. It should not be interpreted as a formal test that the two SUN-101 doses differ from each other. The statistical analyses posted on ClinicalTrials.gov compare each dose with placebo; they do not provide a formal SUN-101 50 mcg versus SUN-101 25 mcg hypothesis test.
This is an important general statistical principle: two treatment effects being statistically significant versus control does not automatically imply that the two treatment effects are statistically different from one another. A direct dose-to-dose comparison would require its own estimated contrast and uncertainty measure.
21. P-values, Estimates, and Confidence Intervals
The estimate answers: How large is the modeled treatment difference? For example, 0.1036 is the reported mean difference in change from baseline in trough FEV1 for one of the 50 mcg analyses.
The confidence interval answers: How precisely has the treatment difference been estimated? The interval incorporates sampling uncertainty around the point estimate under the specified analysis.
The P-value answers a different question: How compatible are the observed data with the null hypothesis under the specified test? It is not an effect-size measure and should not be used to rank the four estimates by clinical importance.
For GOLDEN-3, these three quantities should be read together. The estimates establish the direction and scale of the reported treatment contrasts, the confidence intervals show their precision, and the P-values quantify evidence against the corresponding null hypotheses within the statistical testing framework.
22. Statistical Analysis Workflow
The statistical workflow illustrates why the final estimate should not be reduced to a simple subtraction of two observed means. The model uses treatment assignment, baseline FEV1, specified clinical covariates, visit, the treatment-by-visit interaction, and an unstructured covariance matrix to estimate the treatment contrast.
The final inferential step is then performed within the trial's multiplicity framework. This is the point at which the reported P-values and confidence intervals become part of the broader confirmatory testing strategy rather than isolated numerical findings.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Calculators
25. Sources
- ClinicalTrials.gov: GOLDEN-3, NCT02347761.
- Linked publication: PubMed record for PMID 28838685.
- Linked publication: PubMed record for PMID 30587959.
- Linked publication: PubMed record for PMID 30584583.
Continue through the Clinical Biostats statistical learning pathway
Explore the methods behind randomized trials, longitudinal models, covariate adjustment, confidence intervals, and multiplicity through tutorials and statistical tools.
26. Record Summary
GOLDEN-3 provides a clear example of longitudinal clinical-trial analysis in which a continuous physiologic endpoint is measured repeatedly and analyzed using adjusted mixed models. The trial was randomized, parallel, and quadruple-masked, with three treatment arms and 653 enrolled subjects. Its primary efficacy analyses used an ITT population and reported mean differences in change from baseline in trough FEV1 at Week 12.
The statistical story is more informative than the P-values alone. The reported treatment differences ranged from 0.0961 to 0.1264 liters across the four posted primary analyses, with two-sided 95% confidence intervals posted on ClinicalTrials.gov for every analysis and P-values of <0.0001. The principal MMRM specification incorporated treatment, cardiovascular risk, background LABA use, visit, visit-by-treatment interaction, and baseline FEV1, with an unstructured covariance matrix.
The trial also illustrates why multiplicity must be considered whenever several hypotheses are evaluated. The registry reports use of a Hochberg procedure to control the family-wise Type I error rate for the primary efficacy endpoints and key secondary efficacy endpoints. Consequently, the inferential results should be understood as components of a prespecified testing strategy rather than as four unrelated P-values.
Finally, the available safety data show serious adverse events in 10/218 subjects receiving SUN-101 50 mcg, 8/217 receiving SUN-101 25 mcg, and 11/218 receiving placebo. These figures complement the efficacy analysis but do not constitute a formal comparative safety analysis because the ClinicalTrials.gov record does not provide a corresponding statistical test or confidence interval.