This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are limited to the information reported in the ClinicalTrials.gov record for GOLDEN-4.
1. Trial at a Glance
GOLDEN-4 was a randomized, phase 3, parallel-group, quadruple-masked trial of nebulized SUN-101 in patients with COPD. The trial enrolled 641 participants and compared SUN-101 50 mcg twice daily, SUN-101 25 mcg twice daily, and placebo, with trough FEV1 at Week 12 serving as the registered primary endpoint framework.
| Feature | GOLDEN-4 |
|---|---|
| Trial name | GOLDEN-4 |
| NCT identifier | NCT02347774 |
| Phase | Phase 3 |
| Therapeutic area | Pulmonology |
| Condition | COPD; Chronic Obstructive Pulmonary Disease |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 641 |
| Lead sponsor | Sunovion Respiratory Development Inc. |
| Sponsor type | Industry |
| Trial status | Completed |
| Start | 2015-02 |
| Primary completion | 2015-12 |
2. Clinical Question
The central statistical question was whether either dose of SUN-101 produced a greater improvement in change from baseline in trough FEV1 at Week 12 than placebo in patients with COPD.
Population
Patients with COPD enrolled in the GOLDEN-4 phase 3 trial.
Intervention
SUN-101 50 mcg BID administered by eFlow (CS) nebulizer or SUN-101 25 mcg BID administered by eFlow (CS) nebulizer.
Comparator
Placebo administered by eFlow (CS) nebulizer.
Primary question
Does either SUN-101 dose improve the registered trough FEV1 endpoint relative to placebo?
3. Trial Design
50 mcg eFlow (CS) nebulizer
- SUN-101 50 mcg twice daily
- eFlow (CS) nebulizer
- Compared directly with placebo
- Primary analysis used the ITT population
25 mcg eFlow (CS) nebulizer
- SUN-101 25 mcg twice daily
- eFlow (CS) nebulizer
- Compared directly with placebo
- Primary analysis used the ITT population
Placebo eFlow (CS) nebulizer
- Placebo administered twice daily
- eFlow (CS) nebulizer
- Comparator for both SUN-101 dose groups
Randomized and quadruple-masked
- Randomized allocation
- Parallel design
- Quadruple masking
- Primary purpose: treatment
The registry data identify three parallel treatment groups but do not provide a randomized allocation ratio. The safety data identify 214 participants at risk in each SUN-101 group and 212 in the placebo group for the serious-adverse-event summaries.
4. Endpoints
ClinicalTrials.gov lists two registered primary endpoint entries. Both concern change from baseline in trough FEV1 at Week 12, but they differ in how the spirometry observations are defined for analysis.
| Registered primary endpoint | Time frame | Registry definition reported |
|---|---|---|
| Change From Baseline in Trough Forced Expiratory Volume in 1 Second (FEV1) at Week 12 | baseline and Week 12 | All collected spirometry was performed according to internationally accepted standards. Trough FEV1 at Week 12 was defined as the mean of the values collected at two time points 30 minutes apart at approximately 24 hours (± 1 hour) after the previous morning dose. The registry definition states that all collected values were used regardless of whether the subject remained on randomized treatment. |
| Change From Baseline in Trough Forced Expiratory Volume in 1 Second (FEV1) Week 12 | Week 12 | On-treatment spirometry was performed according to internationally accepted standards. Trough FEV1 at Week 12 was defined as the mean of the values collected at two time points 30 minutes apart at approximately 24 hours (± 1 hour) after the previous morning dose. Only on-treatment values were used for this analysis. |
5. Statistical Analysis Populations
The posted primary analyses use the Intent to Treat (ITT) Population. The registry definition describes this as all subjects randomized to treatment who received at least one dose of study medication, with subjects analyzed according to the treatment to which they were randomized.
| Population / analysis concept | Role in GOLDEN-4 |
|---|---|
| Intent to Treat (ITT) | Primary efficacy analysis population. Subjects were analyzed according to randomized treatment assignment. |
| All collected spirometry analysis | Uses collected trough FEV1 values regardless of whether the subject remained on randomized treatment. |
| On-treatment spirometry analysis | Uses only spirometry values collected while subjects were taking study drug. |
This distinction illustrates an important clinical-trial principle: the same nominal endpoint can produce different statistical targets depending on which observations are included. The all-collected analysis is closer to an assignment-based treatment comparison, whereas restricting measurements to periods during which participants are taking study drug changes the population of observations contributing to the estimate.
6. Results: Trough FEV1 Using All Collected Values
The first primary endpoint was change from baseline in trough FEV1 at Week 12 using all collected spirometry values. The analysis was conducted in the ITT population and compared each SUN-101 dose with placebo.
SUN-101 50 mcg BID vs Placebo
Adjusted mean difference in change from baseline
95% CI: 0.0346 to 0.1127 L · P = 0.0002
Method: least mean squared estimate from the reported MMRM analysis.
The estimated treatment difference was 0.0736 liters, meaning the model-estimated change from baseline in trough FEV1 was higher for SUN-101 50 mcg BID than for placebo by 0.0736 L under this analysis.
The estimate does not mean that every participant improved by exactly 0.0736 L, nor does it represent the raw difference between two simple arithmetic means. It is a least mean squared treatment difference from a longitudinal model that adjusted for prespecified covariates and repeated measurements.
The 95% confidence interval of 0.0346 to 0.1127 L describes the statistical uncertainty around the estimated treatment difference under the model. It does not describe the range of individual patient responses.
The P-value of 0.0002 addresses evidence against the relevant null hypothesis under the prespecified testing framework. It is not a measure of the size or clinical importance of the treatment effect. Effect magnitude is better understood from the estimate itself and its confidence interval.
The analysis notes also state that a Hochberg procedure was used to control the family-wise Type I error rate. That matters because two dose comparisons are being considered rather than treating each comparison as an isolated test.
SUN-101 25 mcg BID vs Placebo
Adjusted mean difference in change from baseline
95% CI: 0.0416 to 0.1204 L · P = 0.0001
Method: least mean squared estimate from the reported MMRM analysis.
The estimated treatment difference was 0.0810 liters, indicating a higher model-estimated change from baseline in trough FEV1 for SUN-101 25 mcg BID than for placebo under the all-collected-values analysis.
Again, this is a model-adjusted mean difference, not a statement that every patient experienced an 0.0810 L improvement. The MMRM incorporates the longitudinal structure of the data and the covariates specified in the analysis.
The 95% confidence interval of 0.0416 to 0.1204 L quantifies uncertainty around the estimated group difference. A narrower interval would indicate greater precision; the interval itself does not describe individual variability.
The P-value of 0.0001 describes the strength of evidence against the null hypothesis within the specified testing framework. It should not be interpreted as the probability that the treatment effect is real, nor as a measure of effect magnitude.
Because the trial compared two SUN-101 dose groups with placebo, multiplicity is relevant. The registry-reported analysis notes identify the Hochberg procedure as the family-wise Type I error-control method for the relevant analysis.
7. Results: On-Treatment Trough FEV1 at Week 12
The second registered primary endpoint restricted the spirometry data to measurements collected while participants were taking study drug. The same general MMRM framework was used, but the analysis population of observations differs from the all-collected-values analysis.
SUN-101 50 mcg BID vs Placebo
On-treatment adjusted mean difference
95% CI: 0.0417 to 0.1224 L · P < 0.0001
Method: least mean squared estimate from the reported MMRM analysis.
For the on-treatment analysis, the estimated difference in change from baseline was 0.0820 liters in favor of SUN-101 50 mcg BID relative to placebo.
This estimate should not be treated as interchangeable with the 0.0736 L estimate from the all-collected-values analysis. The distinction comes from the observations included in the model: the on-treatment analysis excludes measurements obtained when subjects were no longer taking study drug.
The 95% CI of 0.0417 to 0.1224 L gives the uncertainty associated with the estimated treatment difference. It does not indicate that 95% of individual patients would experience an effect within this interval.
The P-value of less than 0.0001 indicates strong statistical evidence against the relevant null hypothesis under the stated analysis. It does not tell us how large the treatment effect is; that information comes from the estimated difference and its confidence interval.
SUN-101 25 mcg BID vs Placebo
On-treatment adjusted mean difference
95% CI: 0.0433 to 0.1246 L · P = 0.0001
Method: least mean squared estimate from the reported MMRM analysis.
The on-treatment analysis estimated a 0.0840 L difference in change from baseline between SUN-101 25 mcg BID and placebo.
The estimate represents the treatment contrast produced by the longitudinal model among the on-treatment observations. It is therefore not a simple unadjusted comparison and should not be interpreted as a universal effect for every participant.
The 95% CI of 0.0433 to 0.1246 L expresses uncertainty around the estimated mean difference. Confidence intervals are particularly useful here because they show both the direction and the precision of the estimated treatment contrast rather than reducing the result to a binary P-value.
The P-value of 0.0001 is evidence against the null hypothesis under the reported testing procedure. It does not quantify treatment benefit, does not give the probability that the null hypothesis is true, and should not be used as a substitute for examining the estimated effect and confidence interval.
8. Primary Results Side by Side
| Primary analysis | Comparison | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| All collected values | SUN-101 50 mcg BID vs placebo | 0.0736 L | 0.0346–0.1127 L | 0.0002 |
| All collected values | SUN-101 25 mcg BID vs placebo | 0.0810 L | 0.0416–0.1204 L | 0.0001 |
| On-treatment values | SUN-101 50 mcg BID vs placebo | 0.0820 L | 0.0417–0.1224 L | <0.0001 |
| On-treatment values | SUN-101 25 mcg BID vs placebo | 0.0840 L | 0.0433–0.1246 L | 0.0001 |
The four posted primary analyses point in the same direction: each SUN-101 dose had a positive model-estimated difference relative to placebo. The important statistical distinction is not merely the numerical size of the four estimates, but the definition of the observations contributing to each analysis.
9. Statistical Methodology
Mixed model for repeated measures
The central analytical method was a mixed model for repeated measures (MMRM). This is a longitudinal modeling approach designed for outcomes measured repeatedly over time. Rather than analyzing only a single endpoint value in isolation, the model uses the repeated measurements and their covariance structure to estimate treatment differences while accounting for within-subject correlation.
The reported model included terms for:
- treatment;
- cardiovascular risk;
- background LABA use;
- visit week;
- visit week by treatment interaction; and
- baseline FEV1 as a covariate.
The analysis used an unstructured covariance matrix. This allows the repeated measurements to have a flexible covariance relationship rather than imposing a simple correlation pattern across visits.
The exact fitted model is more detailed than this schematic expression. The purpose of the representation is to show why the reported treatment effect is a model-adjusted longitudinal contrast rather than a simple difference between two observed means.
Least mean squared treatment differences
The registry reports the effect estimate as a least mean squared estimate with its standard error, while the effect measure is a mean difference. In an adjusted longitudinal model, this estimate represents the model-based treatment contrast after accounting for the covariates and repeated-measure structure specified in the analysis.
The posted estimates are expressed in liters because trough FEV1 is measured in liters.
Covariate adjustment
Baseline FEV1 was included as a covariate, along with cardiovascular risk and background LABA use. Covariate adjustment can improve precision when the included variables explain some of the outcome variation. It also makes the reported treatment estimate conditional on the structure of the prespecified model rather than being a raw comparison alone.
Visit and treatment-by-visit interaction
Visit week and the visit week-by-treatment interaction are important in a repeated-measures model because treatment differences can vary across visits. The interaction allows the treatment contrast to be represented as a function of visit rather than forcing a single common difference across all repeated measurements.
Unstructured covariance
An unstructured covariance matrix permits the model to estimate a separate variance for each repeated measurement and a separate covariance for each pair of measurements, subject to the data supporting those estimates. This is more flexible than imposing a fixed correlation pattern.
10. Why Was an MMRM Used?
FEV1 is a continuous outcome measured repeatedly during a clinical trial. An MMRM is therefore a natural method for using the longitudinal information rather than reducing the dataset to a single unadjusted comparison.
Repeated observations
Multiple measurements from the same participant are correlated. Treating them as independent would misrepresent the information structure.
Covariate adjustment
The model incorporates baseline FEV1 and other specified factors so the treatment contrast is adjusted for those variables.
Time dependence
The visit term and treatment-by-visit interaction allow the modeled treatment effect to vary over the repeated-measure schedule.
Model-based estimate
The least mean squared treatment difference is obtained from the fitted model rather than from a simple arithmetic comparison of raw endpoint means.
11. Statistical Methods Explained
Why was baseline FEV1 included as a covariate?
Baseline FEV1 was explicitly included in the MMRM. Adjustment for a baseline measurement can improve precision because the analysis accounts for where participants started rather than relying only on their follow-up values. The resulting treatment estimate is therefore a covariate-adjusted contrast.
What does a mean difference of 0.0736 L mean?
It means that the model estimated a 0.0736-liter greater change from baseline in trough FEV1 for SUN-101 50 mcg BID than placebo in the all-collected-values analysis. It is not a statement that each participant experienced that exact change, and it is not the same as a percentage improvement.
Why are there two primary FEV1 analyses?
The registry identifies two primary endpoint entries that differ in the observations included. One uses all collected values, while the other uses only on-treatment values. The distinction changes the statistical estimand because the second analysis excludes measurements collected after subjects were no longer taking study drug.
Why use an unstructured covariance matrix?
Repeated FEV1 measurements from the same participant are correlated, but the strength of that correlation need not follow a simple pattern. An unstructured covariance matrix gives the model flexibility to represent the covariance relationships among the repeated measurements.
Why does multiplicity matter when there are two SUN-101 doses?
The trial compares both SUN-101 dose groups with placebo. If each hypothesis were tested independently at the same nominal significance level, the probability of at least one false-positive result across the family could exceed the intended level. The registry-reported analysis notes state that a Hochberg procedure was used to control the family-wise Type I error rate for the relevant primary analysis.
What does a P-value of 0.0001 tell us?
A P-value of 0.0001 quantifies the compatibility of the observed result with the null hypothesis under the specified statistical model and testing framework. It does not measure the magnitude of the treatment effect, the clinical importance of the effect, or the probability that the treatment is effective.
What does the 95% confidence interval tell us?
The confidence interval communicates uncertainty around the estimated treatment difference. For example, the 0.0736 L estimate has a 95% CI of 0.0346 to 0.1127 L. The interval is about uncertainty in the estimated population treatment contrast, not about the range in which 95% of individual patients' responses will fall.
12. Sample Size and Statistical Power
The registry-reported analysis notes state that a sample size of 215 subjects per treatment would provide approximately 90% power to detect a treatment difference of 80 mL in the change from baseline in trough FEV1 at Week 12 between each of the two SUN-101 dose groups and placebo.
| Design assumption | Registry-reported value |
|---|---|
| Sample size | 215 subjects per treatment |
| Target power | ~90% |
| Target treatment difference | 80 mL |
| Assumed standard deviation | 255 mL |
| Alpha | 0.05 |
| Test | 2-sided |
These are design assumptions, not estimates of the treatment effect actually observed in the trial. In particular, the 80 mL value should not be substituted for the reported treatment estimates of 0.0736 L, 0.0810 L, 0.0820 L, or 0.0840 L.
13. Multiplicity and Hypothesis Testing
The trial used superiority hypotheses for the SUN-101 comparisons with placebo. Because there were two SUN-101 dose groups, the family of comparisons introduces a multiplicity issue.
| Feature | GOLDEN-4 information |
|---|---|
| Hypothesis type | Superiority |
| Primary treatment comparisons | SUN-101 50 mcg BID vs placebo; SUN-101 25 mcg BID vs placebo |
| Family-wise error control | Hochberg procedure stated in the registry-reported analysis notes |
| Design alpha | 0.05 |
| Test direction | 2-sided |
The important statistical point is that a P-value should be interpreted within the testing procedure that generated it. The ClinicalTrials.gov record explicitly identify multiplicity adjustment for the first two primary analyses, so the interpretation should not treat those two dose comparisons as though they were completely unrelated statistical tests.
14. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk. These figures are presented exactly as reported rather than converting them into separately calculated percentages.
| Treatment arm | Serious adverse events | At risk |
|---|---|---|
| SUN-101 50 mcg BID eFlow (CS) Nebulizer | 8/214 | 214 |
| SUN-101 25 mcg BID e-Flow (CS) Nebulizer | 5/214 | 214 |
| Placebo BID Eflow (CS) Nebulizer | 13/212 | 212 |
The serious-adverse-event summaries provide an important safety context, but they are not directly comparable to the primary efficacy estimates because they answer a different statistical question. The FEV1 analyses estimate a continuous treatment difference with a confidence interval and P-value; the serious-adverse-event data shown here are event counts among participants at risk.
The registry reports serious adverse events in all three treatment groups. The ClinicalTrials.gov record does not provide additional information about the individual serious adverse-event types, severity distribution, timing, treatment relatedness, or formal statistical comparison of these safety counts. Those characteristics therefore should not be inferred from the affected/at-risk counts alone.
15. What the Primary Estimates Do — and Do Not — Mean
An estimate such as 0.0810 L represents the model-based difference in change from baseline in trough FEV1 between SUN-101 25 mcg BID and placebo under the specified all-collected-values MMRM. It is a population-level treatment contrast, not an individual patient prediction.
The 95% CI of 0.0416 to 0.1204 L describes uncertainty around that estimated treatment contrast. It does not say that 95% of patients will have treatment effects within those limits.
The P-value of 0.0001 measures evidence against the relevant null hypothesis within the specified testing framework. It is not an effect-size measure and should not be interpreted as a probability that the treatment hypothesis is true.
The reported least mean squared differences come from an MMRM containing treatment, cardiovascular risk, background LABA use, visit week, treatment-by-visit interaction, and baseline FEV1. Consequently, these are adjusted model-based results rather than simple raw differences.
16. Why the All-Collected and On-Treatment Analyses Differ
The two registered primary analyses are especially useful for understanding how analysis populations and observation rules can affect an endpoint.
| Feature | All collected values | On-treatment values |
|---|---|---|
| Endpoint | Change from baseline in trough FEV1 at Week 12 | Change from baseline in trough FEV1 at Week 12 |
| Observations | All collected spirometry values | Only values collected while subjects were taking study drug |
| 50 mcg estimate | 0.0736 L | 0.0820 L |
| 25 mcg estimate | 0.0810 L | 0.0840 L |
| Model | MMRM | MMRM |
The differences between the estimates are not evidence by themselves of inconsistency. The analyses answer slightly different questions because they use different sets of observations. The all-collected analysis retains measurements regardless of whether participants remained on randomized treatment, whereas the on-treatment analysis restricts the measurement set to periods during study-drug exposure.
This is an important reason to avoid reading the four estimates as four independent replications of exactly the same estimand. They are related analyses of the same trial, but the registry definitions distinguish their observation rules.
17. Blinding and Randomization
The trial was randomized and quadruple-masked. Randomization is the design mechanism that supports a causal comparison between treatment assignments by balancing measured and unmeasured characteristics in expectation. Masking reduces opportunities for knowledge of treatment assignment to influence participant behavior, investigator assessment, or other aspects of trial conduct.
Randomization
Participants were assigned through a randomized parallel-group design, supporting comparison of treatment assignments under the trial protocol.
Quadruple masking
The registry identifies the study as quadruple-masked, limiting awareness of treatment assignment among the groups covered by the masking procedure.
Objective outcome
FEV1 is a quantitative spirometric outcome, which provides a continuous measurement for the MMRM rather than a binary responder classification.
Placebo comparator
Both SUN-101 dose groups were compared with placebo delivered through the eFlow (CS) nebulizer.
18. Missing Data and Model Interpretation
The ClinicalTrials.gov record specifies an MMRM and distinguish all-collected from on-treatment observations, but they do not provide the complete missing-data or imputation specification. The page therefore does not attribute a particular imputation method to GOLDEN-4.
This limitation matters because repeated-measures analyses depend on the observed longitudinal data and the assumptions used to relate observed and unobserved outcomes. An MMRM does not make missing observations disappear; rather, its validity depends on the statistical assumptions underlying the model and the pattern of missingness.
19. Proportional-Hazards and Other Model Issues
The primary GOLDEN-4 endpoint is a continuous FEV1 outcome analyzed with an MMRM, not a time-to-event endpoint analyzed with a Cox proportional-hazards model. Therefore, the proportional-hazards assumption is not a relevant assumption for the reported primary analysis.
This distinction is useful because many clinical-trial results pages involve hazard ratios, but the appropriate statistical assumptions depend on the endpoint and model. GOLDEN-4 illustrates a different class of analysis: a longitudinal continuous endpoint analyzed using repeated-measures modeling.
| Statistical concept | Relevant to GOLDEN-4 primary analysis? |
|---|---|
| MMRM | Yes |
| Covariate adjustment | Yes |
| Repeated-measure covariance | Yes; unstructured covariance matrix |
| Multiplicity adjustment | Yes; Hochberg procedure stated in registry-reported analysis notes |
| Two-sided testing | Yes |
| Cox proportional-hazards assumption | No; not the primary endpoint model |
| Kaplan-Meier estimation | No; not reported for the primary endpoint |
| Hazard ratio | No; the reported effect measure is a mean difference |
20. Primary Analysis Assumptions and Interpretation
The statistical interpretation of the MMRM results depends on several layers of the analysis. First, the treatment contrast is estimated within a longitudinal model. Second, baseline FEV1 and other prespecified covariates enter the model. Third, the repeated observations are linked through an unstructured covariance matrix. Finally, the two-sided testing framework and multiplicity procedure determine how the reported P-values should be interpreted.
That structure means the result should not be reduced to the statement that "the P-value was small." The more informative statistical description is that the model estimated a positive treatment difference, registry-reported a confidence interval describing uncertainty around that difference, and evaluated the hypothesis within a multiplicity-controlled framework.
- Effect: What was the estimated treatment difference?
- Precision: How wide is the confidence interval?
- Hypothesis test: What does the P-value indicate under the prespecified testing framework?
- Estimand: Which observations and treatment assignment rules generated the estimate?
- Model: Which covariates and covariance assumptions produced the adjusted result?
21. Limitations
- Two related primary analyses: The all-collected and on-treatment analyses use different observation rules, so their estimates should not be treated as identical estimands.
- Incomplete baseline information: The ClinicalTrials.gov record does not provide a baseline characteristics table, so baseline group comparisons are not presented.
- Secondary endpoints: Sixteen outcome measures were posted in the registry, but the ClinicalTrials.gov record does not provide secondary endpoint statistical analyses or estimates. No secondary efficacy results are therefore inferred.
- Missing-data specification: The ClinicalTrials.gov record identifies MMRM but do not provide the complete missing-data or imputation strategy.
- No raw distributions: The ClinicalTrials.gov record does not provide arm-level raw Week 12 FEV1 means, standard deviations, or individual-level observations, so the reported model estimates cannot be independently reconstructed from those quantities.
- Safety detail: Serious adverse events are reported as affected/at-risk counts, but the ClinicalTrials.gov record does not include detailed event categories or formal statistical comparisons.
- Generalizability: The ClinicalTrials.gov record identifies the COPD population and trial design but does not provide the full eligibility and baseline-characteristic information needed to characterize external validity in detail.
22. Why This Trial Matters Statistically
GOLDEN-4 is a useful teaching example because it demonstrates how a randomized clinical trial with a continuous physiological endpoint can require a substantially different statistical framework from a survival trial. The central question is not a hazard ratio or median time to event; it is an adjusted longitudinal difference in FEV1.
| Concept | How it appears in GOLDEN-4 |
|---|---|
| Randomization | Randomized parallel-group phase 3 design |
| Blinding | Quadruple-masked study |
| Continuous endpoint | Change from baseline in trough FEV1 |
| Repeated measures | FEV1 observations collected longitudinally |
| MMRM | Primary statistical method |
| Covariate adjustment | Baseline FEV1, cardiovascular risk, and background LABA use included in the model |
| Treatment-by-visit interaction | Allows the treatment contrast to depend on visit week |
| Unstructured covariance | Used to model within-subject repeated-measure covariance |
| Multiplicity | Hochberg procedure stated for family-wise Type I error control |
| Superiority testing | Two-sided superiority framework |
| ITT analysis | Primary analyses use randomized treatment assignment in the ITT population |
| Analysis estimand | All-collected versus on-treatment observation rules provide related but distinct treatment contrasts |
23. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The reported MMRM analyses produced positive adjusted mean differences for both SUN-101 dose groups relative to placebo, with 95% confidence intervals and P-values reported for each comparison. The analyses were conducted under a two-sided superiority framework.
Clinical interpretation
The numerical clinical meaning of an FEV1 difference requires context beyond the P-value. The ClinicalTrials.gov record establishes the estimated treatment differences and their statistical uncertainty, but do not provide a broader clinical-threshold analysis for this page.
This distinction is important. Statistical evidence answers whether the observed treatment contrast is compatible with a null hypothesis under the specified model and testing framework. Clinical interpretation asks what the magnitude of that contrast means for patients and clinical practice. Those are related but different questions.
24. What the Four Primary Analyses Show Together
The four posted primary analyses provide a useful consistency check because both dose groups are evaluated under both observation rules.
The visualization is only a display of the four registry-reported estimates; it does not replace the confidence intervals or the formal model. In particular, the closeness of the point estimates should not be interpreted as a formal dose-response analysis because the ClinicalTrials.gov record does not provide such a statistical comparison.
25. Related Tutorials
Learn more about the methods used in this trial:
26. Related Calculators
27. Sources
- ClinicalTrials.gov: NCT02347774 — GOLDEN-4.
- PubMed: PMID 28838685.
- PubMed: PMID 30587959.
- PubMed: PMID 30584583.
The ClinicalTrials.gov record is the controlling source for the trial facts and numerical results presented on this page. The linked PubMed records are provided as publication references associated with the trial; no additional numerical results from those publications are incorporated here.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts that connect randomized trial design, repeated-measures modeling, confidence intervals, multiplicity, and clinical-trial interpretation.
28. Record Summary
GOLDEN-4 provides a clear example of how a randomized phase 3 trial with a continuous physiological endpoint can be analyzed using a longitudinal mixed model rather than a time-to-event framework. The primary outcome was change from baseline in trough FEV1 at Week 12, and the ClinicalTrials.gov record reports four primary analyses: two using all collected spirometry values and two using only on-treatment values.
The statistical story is defined by the combination of randomization, quadruple masking, an ITT analysis population, MMRM modeling, baseline and clinical covariate adjustment, an unstructured covariance matrix, two-sided superiority testing, and a stated Hochberg procedure for family-wise Type I error control. The four reported treatment differences were 0.0736 L, 0.0810 L, 0.0820 L, and 0.0840 L according to the respective analysis definitions, with the corresponding confidence intervals and P-values reported above.
The most important interpretive lesson is that the estimate, confidence interval, and P-value answer different questions. The estimate describes the magnitude of the adjusted treatment contrast; the confidence interval describes uncertainty around that contrast; and the P-value evaluates the observed evidence against a null hypothesis within the specified testing framework. None of these measures, by itself, describes the response of an individual patient.