This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record. The registry provides the official trial record.
1. Trial at a Glance
DUO-E is a randomized, parallel, quadruple-masked phase 3 trial evaluating maintenance strategies after first-line treatment of advanced and recurrent endometrial cancer. The trial has three arms and an enrollment of 805 participants.
| Feature | DUO-E |
|---|---|
| Trial name | DUO-E |
| NCT identifier | NCT04269200 |
| Phase | Phase 3 |
| Condition | Endometrial Neoplasms |
| Brief title | Durvalumab With or Without Olaparib as Maintenance Therapy After First-Line Treatment of Advanced and Recurrent Endometrial Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 805 |
| Primary endpoint type | Time-to-event |
| Primary endpoints registered | 1 |
| Results posted | Yes |
| Outcome measures posted | 12 |
| Statistical analyses posted | 14 |
| Lead sponsor | AstraZeneca |
| Sponsor type | Industry |
2. Clinical Question
The central statistical question is whether adding durvalumab to platinum-based chemotherapy, followed by maintenance durvalumab, and whether additionally adding olaparib to the maintenance strategy, improves progression-free survival compared with platinum-based chemotherapy alone in patients with advanced and recurrent endometrial cancer.
Population
Patients with advanced and recurrent endometrial cancer represented by the registered condition Endometrial Neoplasms.
Interventions
Durvalumab, olaparib, durvalumab placebo, olaparib placebo, carboplatin, and paclitaxel were the registered interventions.
Comparator
The primary comparisons reported in the statistical analyses use standard of care (SoC) as the reference group.
Primary question
Does the addition of durvalumab, with or without olaparib, improve investigator-assessed progression-free survival?
3. Trial Design
Standard-of-care comparison
- Carboplatin
- Paclitaxel
- Serves as the reference group in the reported global comparisons
Durvalumab strategy
- Carboplatin
- Paclitaxel
- Durvalumab
- Compared directly with SoC for the primary PFS analysis
Durvalumab plus olaparib strategy
- Carboplatin
- Paclitaxel
- Durvalumab
- Olaparib
- Compared directly with SoC for the primary PFS analysis
The ClinicalTrials.gov record also identify a China cohort. The China analyses compare SoC with SoC + durvalumab and SoC with SoC + durvalumab + olaparib, respectively.
4. Endpoints
Primary Endpoint
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Progression-free Survival (PFS) According to RECIST 1.1, Based on Investigator Assessments | At baseline, every 9 weeks (wks) up to 18 wks, then every 12 wks until objective radiological disease progression. Assessed until 12 Apr 2023 DCO (08 Jul 2024 DCO for China cohort), up to 50 months | Time-to-event |
The registry defines this endpoint as the basis for assessing the efficacy of durvalumab in combination with platinum-based chemotherapy followed by maintenance durvalumab or durvalumab with olaparib compared with platinum-based chemotherapy.
Secondary Endpoints With Reported Statistical Analyses
| Endpoint | Time frame | Analysis method | Effect measure |
|---|---|---|---|
| Time From Randomisation to Second Progression or Death (PFS2) Based on Local Standard Clinical Practice | At baseline, every 9 to 18 wks, then every 12 wks until objective radiological disease progression. Assessments then per local practice every 12 wks until second progression. Assessed until 12Apr2023 DCO (08Jul2024 DCO for China cohort), up to 50 months | Cox proportional-hazards model | Hazard ratio |
| Objective Response Rate (ORR) Based on Investigator Assessment | At baseline, every 9 to 18 wks, then every 12 wks until objective radiological disease progression. Assessed until 12 Apr 2023 DCO (08 Jul 2024 DCO for China cohort), up to 50 months | Logistic regression | Odds ratio |
| Change From Baseline in Physical Functioning Score of the EORTC QLQ-C30 | At baseline, every 3 wks until Wk 18, and every 4 wks until the second progression. Assessed until 12 Apr 2023 DCO, up to 35 months. The average treatment effect over the first 12 months after randomisation is presented. | MMRM | Mean difference |
| Change From Baseline in Global Health Status/QoL Score of the EORTC QLQ-C30 | At baseline, every 3 wks until Wk 18, and every 4 wks until the second progression. Assessed until 12 Apr 2023 DCO, up to 35 months. The average treatment effect over the first 12 months after randomisation is presented. | MMRM | Mean difference |
5. Analysis Populations and Stratification
The primary global PFS analyses used a Full Analysis Set (FAS) consisting of all patients randomized as part of global enrollment, including patients from China who were randomized before global recruitment was closed. The ClinicalTrials.gov record describes this population as the basis for the global primary analyses.
The reported Cox models for the global primary PFS comparisons were stratified by MMR status (proficient versus deficient) and disease status (recurrent versus newly diagnosed). The China primary PFS analyses were stratified by disease status.
| Analysis | Population | Stratification |
|---|---|---|
| Global PFS: SoC vs SoC + Durvalumab | Full Analysis Set | MMR status; disease status |
| Global PFS: SoC vs SoC + Durvalumab + Olaparib | Full Analysis Set | MMR status; disease status |
| China PFS: SoC vs SoC + Durvalumab | Full Analysis Set for the China cohort analysis | Disease status |
| China PFS: SoC vs SoC + Durvalumab + Olaparib | Full Analysis Set for the China cohort analysis | Disease status |
This structure illustrates an important distinction between randomization and stratified analysis. Randomization creates the treatment groups, while stratification in the analysis adjusts the Cox comparison according to prespecified factors identified in the registry analysis record.
6. Primary Results: Progression-Free Survival
The registry contains four formal primary statistical analyses for the investigator-assessed RECIST 1.1 PFS endpoint. Two compare the global cohort with SoC, and two repeat the comparison structure within the China cohort.
Global Cohort: SoC vs SoC + Durvalumab
Hazard ratio for progression or death
95% CI: 0.57–0.89 · P = 0.003
Cox proportional-hazards model; two-sided 95% confidence interval.
A hazard ratio of 0.71 means that, within the fitted Cox model and over the analyzed follow-up, the estimated instantaneous rate of progression or death was 29% lower with SoC + durvalumab than with SoC. The 29% figure is a direct interpretation of the relative hazard estimate, not an estimate of the proportion of patients who avoid progression.
The 95% confidence interval of 0.57–0.89 describes uncertainty around the estimated hazard ratio. Because the interval lies below 1, the registry result is consistent with a lower estimated hazard in the durvalumab group under the specified model.
The P = 0.003 is evidence against the null hypothesis under the stated statistical test. It does not measure the size or clinical importance of the treatment effect. The effect size is represented by the hazard ratio, while the confidence interval describes its statistical precision.
As with any Cox analysis, interpretation depends on the model assumptions and the handling of censoring. A single hazard ratio also should not be interpreted as a statement that every patient experiences the same relative reduction in risk.
Global Cohort: SoC vs SoC + Durvalumab + Olaparib
Hazard ratio for progression or death
95% CI: 0.43–0.69 · P < 0.0001
Cox proportional-hazards model; two-sided 95% confidence interval.
A hazard ratio of 0.55 corresponds to an estimated 45% lower instantaneous rate of progression or death for SoC + durvalumab + olaparib relative to SoC under the fitted Cox model. It does not mean that 45% of participants benefited, nor does it mean that individual patients have exactly a 45% reduction in their personal risk.
The 95% CI of 0.43–0.69 is entirely below 1, indicating that the estimated treatment effect is consistently below the null value within this confidence interval. The width of the interval provides information about precision; it should not be interpreted as a range in which individual treatment effects must fall.
The P < 0.0001 indicates strong evidence against the null hypothesis under the reported two-sided analysis. A small P-value is not a measure of effect size. The hazard ratio and confidence interval are needed to understand the magnitude and uncertainty of the result.
The comparison is a superiority analysis. Because the analysis is time-to-event based, censoring and the proportional-hazards framework remain important considerations when translating the estimate into a clinical interpretation.
China Cohort: SoC vs SoC + Durvalumab
Hazard ratio for progression or death
95% CI: 0.42–1.25
Cox proportional-hazards model; two-sided 95% confidence interval.
A hazard ratio of 0.73 is an estimated 27% lower instantaneous rate of progression or death for SoC + durvalumab relative to SoC under the fitted model. That relative estimate should not be converted into an absolute probability or interpreted as the percentage of patients who benefit.
The 95% CI of 0.42–1.25 is substantially wider than the corresponding global estimate and crosses 1. This indicates considerably more statistical uncertainty around the China-cohort estimate. The interval includes values compatible with both a lower and a higher estimated hazard relative to the null value.
No P-value is provided in the ClinicalTrials.gov record for this comparison. The absence of a reported P-value does not permit a separate significance conclusion to be substituted for the confidence-interval interpretation.
This analysis is also a useful reminder that subgroup or regional estimates can be less precise than the corresponding overall randomized comparison. Precision is determined by the information contributing to the analysis, not simply by the existence of a point estimate.
China Cohort: SoC vs SoC + Durvalumab + Olaparib
Hazard ratio for progression or death
95% CI: 0.57–1.61
Cox proportional-hazards model; two-sided 95% confidence interval.
A hazard ratio of 0.96 is close to the null value of 1. Under the fitted model, it corresponds to an estimated instantaneous progression-or-death rate approximately 4% lower with SoC + durvalumab + olaparib than with SoC.
The 95% CI of 0.57–1.61 is wide and crosses 1. The estimate therefore has substantial uncertainty, and the interval encompasses materially different possible relative hazard values. The confidence interval is more informative than the point estimate alone.
No P-value is provided in the ClinicalTrials.gov record. A P-value should not be reverse-engineered from the confidence interval, and the confidence interval itself should not be converted into a categorical judgment beyond what it directly communicates about uncertainty.
This China-cohort result should also not be used by itself to infer that the global treatment effect is heterogeneous. Establishing treatment-effect modification requires an appropriate interaction or heterogeneity analysis, not merely comparison of separate point estimates.
7. Primary PFS Results Side by Side
| Comparison | HR | 95% CI | P-value | Stratification |
|---|---|---|---|---|
| Global SoC vs SoC + Durvalumab | 0.71 | 0.57–0.89 | 0.003 | MMR status; disease status |
| Global SoC vs SoC + Durvalumab + Olaparib | 0.55 | 0.43–0.69 | <0.0001 | MMR status; disease status |
| China SoC vs SoC + Durvalumab | 0.73 | 0.42–1.25 | Not provided | Disease status |
| China SoC vs SoC + Durvalumab + Olaparib | 0.96 | 0.57–1.61 | Not provided | Disease status |
8. Secondary Results: PFS2
PFS2 was defined as Time From Randomisation to Second Progression or Death (PFS2) Based on Local Standard Clinical Practice. The analyses posted on ClinicalTrials.gov use Cox proportional-hazards models and hazard ratios.
Global Cohort: SoC vs SoC + Durvalumab
PFS2 hazard ratio
95% CI: 0.59–1.07
Model: unstratified Cox proportional-hazards model.
Global Cohort: SoC vs SoC + Durvalumab + Olaparib
PFS2 hazard ratio
95% CI: 0.40–0.76
Model: unstratified Cox proportional-hazards model.
China Cohort: SoC vs SoC + Durvalumab
PFS2 hazard ratio
95% CI: 0.48–2.21
Model: Cox proportional-hazards model stratified by disease status.
China Cohort: SoC vs SoC + Durvalumab + Olaparib
PFS2 hazard ratio
95% CI: 0.68–2.90
Model: Cox proportional-hazards model stratified by disease status.
| Comparison | HR | 95% CI | Model |
|---|---|---|---|
| Global SoC vs SoC + Durvalumab | 0.80 | 0.59–1.07 | Unstratified Cox |
| Global SoC vs SoC + Durvalumab + Olaparib | 0.55 | 0.40–0.76 | Unstratified Cox |
| China SoC vs SoC + Durvalumab | 1.03 | 0.48–2.21 | Cox stratified by disease status |
| China SoC vs SoC + Durvalumab + Olaparib | 1.38 | 0.68–2.90 | Cox stratified by disease status |
The PFS2 analyses illustrate why later time-to-event endpoints should be kept conceptually distinct from the primary PFS endpoint. PFS2 incorporates an event occurring after the first progression and therefore captures a later point in the treatment pathway.
9. Secondary Results: Objective Response Rate
Objective Response Rate was assessed by investigator assessment. The analysis population included all participants with measurable disease at baseline, within the broader FAS framework described in the registry.
Global Cohort: SoC vs SoC + Durvalumab
Odds ratio for objective response
95% CI: 0.89–1.98
Logistic regression stratified by disease status.
Global Cohort: SoC vs SoC + Durvalumab + Olaparib
Odds ratio for objective response
95% CI: 0.95–2.18
Logistic regression stratified by disease status.
| Comparison | Odds ratio | 95% CI | Analysis |
|---|---|---|---|
| Global SoC vs SoC + Durvalumab | 1.32 | 0.89–1.98 | Logistic regression, stratified by disease status |
| Global SoC vs SoC + Durvalumab + Olaparib | 1.44 | 0.95–2.18 | Logistic regression, stratified by disease status |
An odds ratio greater than 1 favors the durvalumab-containing treatment according to the registry analysis notes. An odds ratio of 1.44, for example, describes the estimated ratio of the odds of response between the treatment and reference groups; it is not the same as saying that response was 44 percentage points higher or that the probability of response increased by 44%.
The confidence intervals for both global ORR comparisons include 1. This means the point estimates should be interpreted together with their uncertainty rather than treated as definitive measures of the absolute response difference.
10. Secondary Results: Patient-Reported Quality of Life
The registry contains MMRM analyses of change from baseline in physical functioning and global health status/QoL scores of the EORTC QLQ-C30. The models included fixed effects for treatment, visit, and baseline score, together with treatment-by-visit and baseline-score-by-visit interactions and a random patient effect.
Physical Functioning
| Comparison | Least-squares mean difference | 95% CI |
|---|---|---|
| SoC + Durvalumab vs SoC | 1.7 | -1.2 to 4.5 |
| SoC + Durvalumab + Olaparib vs SoC | -0.6 | -3.4 to 2.2 |
Global Health Status / QoL
| Comparison | Least-squares mean difference | 95% CI |
|---|---|---|
| SoC + Durvalumab vs SoC | 0.0 | -2.5 to 2.6 |
| SoC + Durvalumab + Olaparib vs SoC | -0.9 | -3.4 to 1.6 |
The time frame for these assessments was baseline, every 3 weeks until Week 18, and every 4 weeks until the second progression, with assessment through the 12 Apr 2023 data cutoff as specified in the ClinicalTrials.gov record.
An MMRM least-squares mean difference describes a model-based difference between treatment groups over the repeated-measures framework specified in the analysis. It is not a hazard ratio and should not be interpreted as a relative risk or odds ratio.
The confidence intervals provide the principal information about uncertainty around the estimated mean differences. All four registry-reported intervals span zero, the null value for a mean difference. No P-values are provided in the ClinicalTrials.gov record for these MMRM analyses.
MMRM is particularly useful when the same participant contributes repeated measurements over time because it models the longitudinal structure rather than reducing each participant to a single observation.
11. Safety Results
The ClinicalTrials.gov record reports serious adverse events by cohort and treatment arm as affected participants divided by participants at risk. These figures are presented exactly as provided.
| Cohort / treatment group | Serious adverse events | Format |
|---|---|---|
| Global Cohort - SoC | 73/236 | Affected / at risk |
| Global Cohort - SoC + Durvalumab | 73/235 | Affected / at risk |
| Global Cohort - SoC + Durvalumab + Olapa | 85/238 | Affected / at risk |
| China Cohort - SoC | 14/44 | Affected / at risk |
| China Cohort - SoC + Durvalumab | 8/41 | Affected / at risk |
| China Cohort - SoC + Durvalumab + Olapar | 19/41 | Affected / at risk |
The serious-adverse-event data are not converted into percentages here because the registry field specifies the affected/at-risk counts directly and the task rules require numbers to be reported exactly as provided.
12. Statistical Methodology
Cox Proportional-Hazards Model
The primary PFS analyses and the PFS2 analyses use Cox proportional-hazards models. The Cox model relates the instantaneous event rate to treatment and other covariates or stratification factors without requiring the baseline hazard function to take a particular parametric form.
The hazard ratio for a treatment comparison is obtained from the corresponding regression coefficient. In this trial, the reported effect measure is the hazard ratio.
The global primary models were stratified by MMR status and disease status. The China primary analyses were stratified by disease status. The PFS2 global analyses were unstratified according to the registry-reported analysis notes, while the China PFS2 analyses were stratified by disease status.
Hazard Ratio
The hazard ratio compares estimated instantaneous event rates between groups under the fitted time-to-event model. A value below 1 favors the durvalumab-containing group in the registry-reported DUO-E analyses.
The hazard ratio is not an absolute risk difference, a probability of response, or a statement that every individual patient experiences the same relative effect.
Logistic Regression
Objective response rate is a binary endpoint: participants are classified according to whether the prespecified response criterion was met. The registry analyses use logistic regression and report an odds ratio.
An odds ratio of 1 represents equal odds under the model. The DUO-E registry analysis notes specify that an odds ratio greater than 1 favors the corresponding durvalumab-containing treatment.
MMRM
The quality-of-life analyses use a mixed model for repeated measures. The registry-reported model specification includes treatment, visit, and baseline score as fixed effects; treatment-by-visit and baseline-score-by-visit interactions; and a random patient effect.
This structure recognizes that repeated measurements from the same participant are correlated. Instead of treating each measurement as independent, the model accounts for the repeated-measures nature of the observations.
Stratified Analysis
Stratification is used in the Cox analyses to account for specified factors while estimating the treatment comparison. For the global primary PFS analyses, the reported strata are MMR status and disease status. For the China primary analyses, disease status is the reported stratification factor.
Intention-to-Treat Analysis
The registry-reported primary analysis population is described as a Full Analysis Set consisting of all patients randomized as part of global enrollment, including eligible China participants randomized before global recruitment closed. This preserves the connection between the analysis population and randomized treatment assignment.
13. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
PFS and PFS2 are time-to-event endpoints. Participants can experience progression or death at different times, while others may remain event-free at their last assessment and therefore be censored. Cox regression is designed for this type of data and provides a hazard ratio as a relative treatment-effect measure.
What does an HR of 0.55 mean?
Under the fitted Cox model, an HR of 0.55 means the estimated instantaneous rate of progression or death is 55% of the reference rate, corresponding to an estimated 45% lower instantaneous rate. It does not mean that 45% of participants are guaranteed to benefit or that survival probabilities are reduced by 45 percentage points.
Why are confidence intervals important?
A point estimate is only one estimate of the underlying treatment effect. The 95% confidence interval communicates statistical uncertainty around that estimate. For example, the global SoC versus SoC + durvalumab + olaparib PFS HR is 0.55 with a 95% CI of 0.43–0.69. The interval is therefore an essential part of the result, not an optional addition to the hazard ratio.
Why doesn't the P-value measure effect size?
A P-value quantifies evidence against a specified null hypothesis under the statistical model. It depends not only on the observed effect but also on the amount of information in the analysis. The hazard ratio describes the estimated relative effect, while the confidence interval communicates its uncertainty.
Why use logistic regression for ORR?
ORR is binary rather than time-to-event. Logistic regression models the probability of a binary outcome through the odds scale. The resulting odds ratio is therefore appropriate to the reported analysis, but it should not be interpreted as a percentage-point difference in response probability.
Why use MMRM for physical functioning and QoL?
Participants have repeated questionnaire measurements over time. MMRM can incorporate the longitudinal structure of these observations and the specified treatment-by-visit and baseline-score-by-visit interactions. The random patient effect recognizes that observations from the same participant are related.
Why is a subgroup result not automatically evidence of treatment-effect modification?
A treatment estimate within a subgroup answers whether the treatment comparison can be estimated in that subgroup. It does not by itself test whether the treatment effect differs from another subgroup. A formal interaction or heterogeneity analysis is needed to make that comparison.
14. Confidence Intervals, P-values, and the Null Value
The different statistical methods in DUO-E use different natural null values.
| Effect measure | Null value | Examples in DUO-E |
|---|---|---|
| Hazard ratio | 1 | PFS, PFS2 |
| Odds ratio | 1 | Objective response rate |
| Mean difference | 0 | Physical functioning; global health status/QoL |
This distinction is important when reading confidence intervals. A hazard-ratio or odds-ratio interval that crosses 1 includes the null value for that measure. A mean-difference interval that crosses 0 includes its null value.
The primary global PFS comparisons have reported P-values of 0.003 and <0.0001. The China PFS comparisons, PFS2 analyses, ORR analyses, and MMRM analyses in the ClinicalTrials.gov record does not include P-values, so none are added here.
15. Primary Endpoint Analysis Structure
| Question | DUO-E analysis |
|---|---|
| Endpoint | Investigator-assessed PFS according to RECIST 1.1 |
| Endpoint type | Time-to-event |
| Primary hypothesis type | Superiority |
| Primary model | Cox proportional-hazards model |
| Global stratification | MMR status and disease status |
| China stratification | Disease status |
| Effect measure | Hazard ratio |
| Confidence interval | Two-sided 95% |
This design creates a coherent statistical pathway: randomization establishes the treatment comparison, RECIST 1.1 investigator assessments define the event endpoint, time-to-event methods accommodate differing follow-up and censoring, and the Cox model estimates the relative treatment effect.
16. Multiplicity and Multiple Comparisons
The ClinicalTrials.gov record identifies a single registered primary endpoint but four primary statistical analyses: two global-cohort comparisons and two China-cohort comparisons. The ClinicalTrials.gov record does not specify an alpha-spending procedure, multiplicity-adjustment procedure, or hierarchical testing strategy.
The same principle applies to secondary endpoints. The presence of multiple PFS2, ORR, and quality-of-life analyses means that individual estimates should be interpreted in the context of the overall analysis program rather than treating every reported comparison as a separate confirmatory hypothesis test.
17. Missing Data, Censoring, and Analysis Assumptions
Censoring in time-to-event analyses
PFS and PFS2 are time-to-event endpoints. A participant may not experience the specified event during the period of observation, creating a censored observation. Cox analysis and Kaplan-Meier estimation are designed to use the information available up to censoring.
Proportional-hazards assumption
The Cox model is called a proportional-hazards model because its conventional interpretation assumes that the relative hazard between groups is adequately represented by a treatment hazard ratio over the relevant analysis period. A single hazard ratio is therefore a model-based summary rather than a complete description of how treatment effects evolve at every time point.
Missing repeated-measures data
The quality-of-life analyses use MMRM, which is specifically suited to repeated observations. The ClinicalTrials.gov record specifies the model structure but do not provide a separate missing-data or imputation strategy. No additional imputation procedure is therefore described here.
18. Timeline and Trial Status
Trial start
The registered trial start date is May 5, 2020.
Global enrollment cutoff referenced in analyses
The registry-reported FAS descriptions for several analyses refer to patients randomized through and including April 20, 2022.
Quality-of-life data cutoff
The registry-reported MMRM analyses reference a 12 Apr 2023 data cutoff for the EORTC QLQ-C30 outcomes.
Primary completion
The registered primary completion date is July 8, 2024.
Active, not recruiting
the ClinicalTrials.gov record lists the study status as ACTIVE_NOT_RECRUITING.
19. Limitations
- Primary endpoint scope: The registry provides one registered primary endpoint, PFS according to RECIST 1.1 based on investigator assessments. Other outcomes should therefore not be silently promoted to primary endpoints.
- Regional estimates: The China-cohort estimates have wider confidence intervals than the corresponding global estimates in the ClinicalTrials.gov record. They should be interpreted with attention to precision.
- No inferred P-values: Several statistical analyses provide estimates and confidence intervals without P-values. P-values are not reconstructed from the registry-reported confidence intervals.
- Multiplicity information: The ClinicalTrials.gov record does not specify an alpha-allocation or multiplicity-adjustment procedure. No such procedure is assumed.
- Proportional hazards: Hazard ratios are model-based summaries and rely on the appropriateness of the Cox proportional-hazards framework.
- Endpoint differences: PFS, PFS2, ORR, physical functioning, and global health status/QoL measure different aspects of the treatment experience and cannot be reduced to a single effect measure.
- Analysis populations: Different endpoints use different analysis-population descriptions, such as the FAS and participants with measurable disease for ORR. Estimates should therefore not be compared without considering their population definitions.
- Safety denominators: Serious adverse-event results are reported as affected/at-risk counts for specified cohorts and arms. These values are reported without deriving additional rates.
- Registry-level detail: The ClinicalTrials.gov recordset does not include baseline characteristics, median PFS, median PFS2, response percentages, subgroup forest plots, crossover information, or a detailed missing-data plan. Those items are not added from outside sources.
20. Why This Trial Matters Statistically
DUO-E is a useful statistical teaching case because several major clinical-trial methods appear in one randomized phase 3 analysis program. The primary endpoint is a time-to-event outcome analyzed with Cox regression, while secondary endpoints demonstrate logistic regression and mixed models for repeated measures.
| Concept | How it appears in DUO-E |
|---|---|
| Randomization | The trial uses randomized allocation in a parallel-group phase 3 design. |
| Blinding | The registered masking is quadruple. |
| Time-to-event endpoint | Primary PFS is defined according to RECIST 1.1 and investigator assessment. |
| Cox regression | Used for the primary PFS and secondary PFS2 analyses. |
| Hazard ratio | Reports the relative treatment effect for PFS and PFS2. |
| Stratified analysis | Global primary PFS models are stratified by MMR status and disease status. |
| Logistic regression | Used for investigator-assessed objective response rate. |
| Odds ratio | Used as the effect measure for ORR. |
| MMRM | Used for repeated EORTC QLQ-C30 physical-functioning and global-health-status/QoL scores. |
| Confidence intervals | Two-sided 95% confidence intervals accompany the reported effect estimates. |
| Multiple analyses | The registry reports 14 statistical analyses across primary and secondary outcomes. |
The statistical structure is especially instructive because the same randomized trial requires different models depending on the endpoint. Time-to-event outcomes require methods that handle event timing and censoring; binary outcomes use logistic regression; and longitudinal questionnaire outcomes use repeated-measures modeling.
21. What the Primary Hazard Ratios Do — and Do Not — Mean
The primary PFS HR of 0.71 for SoC + durvalumab versus SoC corresponds to an estimated 29% lower instantaneous rate of progression or death under the reported Cox model.
It does not mean that PFS was 29 percentage points higher, that 29% of participants avoided progression, or that every individual participant experienced a 29% reduction in risk.
The primary PFS HR of 0.55 corresponds to an estimated 45% lower instantaneous rate of progression or death under the fitted model relative to SoC.
The associated 95% CI of 0.43–0.69 indicates uncertainty around that relative estimate. It is not a range describing the treatment effect experienced by individual patients.
The China-cohort HRs of 0.73 and 0.96 have wider confidence intervals of 0.42–1.25 and 0.57–1.61, respectively. This illustrates why a subgroup estimate should be interpreted together with its precision rather than by its point estimate alone.
22. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The reported global primary PFS analyses produced hazard ratios below 1, with two-sided 95% confidence intervals and reported P-values of 0.003 and <0.0001. The corresponding China-cohort estimates were less precise.
Endpoint interpretation
PFS, PFS2, ORR, physical functioning, and global health status/QoL describe different outcomes. Each requires interpretation using its own statistical scale and null value.
The statistical evidence should therefore be read as a collection of endpoint-specific estimates rather than as one universal numerical measure of treatment effect.
23. A Practical Reading of the DUO-E Results
A useful way to read the registry results is to move from the design to the endpoint and then to the model:
- Start with randomization. The study is randomized, parallel, phase 3, and quadruple-masked.
- Identify the primary endpoint. The registered primary endpoint is investigator-assessed PFS according to RECIST 1.1.
- Identify the analysis population. The global primary analysis uses the Full Analysis Set described in the registry.
- Identify the statistical model. PFS is analyzed using Cox proportional-hazards regression.
- Read the effect estimate. The hazard ratio communicates the relative event rate under the model.
- Read the confidence interval. The interval communicates statistical uncertainty and whether the null value is contained within the reported interval.
- Read the P-value only after reading the effect size. The P-value describes evidence against the null hypothesis; it does not replace the hazard ratio or its confidence interval.
- Then examine secondary endpoints. PFS2, ORR, and quality-of-life analyses answer different questions and use different models.
This sequence helps prevent a common clinical-trial interpretation error: treating a P-value as the primary description of an effect instead of first understanding the endpoint, effect measure, confidence interval, and analysis population.
24. Related Tutorials
Learn more about the methods used in this trial:
25. Related Statistical Calculators
26. Sources
- ClinicalTrials.gov: NCT04269200 — DUO-E.
- PubMed: PMID 41943281.
- PubMed: PMID 40590327.
- PubMed: PMID 38431043.
- PubMed: PMID 37864337.
Continue through the Clinical Biostats statistical pathway
Use the trial's endpoints as a starting point for deeper study of survival analysis, regression, confidence intervals, repeated-measures models, and clinical-trial methodology.
27. Record Summary
DUO-E provides a useful example of a modern randomized phase 3 statistical program in which multiple endpoint types require different analytical methods. The registered primary endpoint is investigator-assessed progression-free survival according to RECIST 1.1, analyzed with Cox proportional-hazards regression. The global PFS comparisons produced hazard ratios of 0.71 for SoC + durvalumab versus SoC and 0.55 for SoC + durvalumab + olaparib versus SoC, with corresponding two-sided 95% confidence intervals of 0.57–0.89 and 0.43–0.69 and reported P-values of 0.003 and <0.0001.
The secondary analyses extend the statistical framework to PFS2, objective response rate, and longitudinal quality-of-life outcomes. PFS2 uses Cox regression, ORR uses logistic regression, and EORTC QLQ-C30 outcomes use MMRM. Together, these analyses demonstrate why the interpretation of a clinical trial depends on matching each endpoint to its appropriate effect measure, model, confidence interval, and analysis population.