This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
ZUMA-7 was a randomized, parallel, open-label phase 3 treatment trial in relapsed/refractory diffuse large B-cell lymphoma (DLBCL). The trial compared axicabtagene ciloleucel with standard of care therapy, with Event Free Survival (EFS) per blinded central assessment registered as the primary endpoint.
| Feature | ZUMA-7 |
|---|---|
| Trial name | ZUMA-7 |
| NCT identifier | NCT03391466 |
| Phase | Phase 3 |
| Population | Relapsed/Refractory Diffuse Large B-Cell Lymphoma (DLBCL) |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 359 |
| Primary endpoint | Event Free Survival (EFS) Per Blinded Central Assessment |
| Results posted | Yes |
| Lead sponsor | Kite, A Gilead Company |
| Sponsor type | Industry |
2. Clinical Question
The primary statistical question was whether axicabtagene ciloleucel improved event-free survival relative to standard of care therapy in patients with relapsed/refractory DLBCL. The registered hypothesis type was superiority.
Population
Patients with relapsed/refractory diffuse large B-cell lymphoma (DLBCL).
Intervention
Axicabtagene ciloleucel, with cyclophosphamide and fludarabine listed among the trial interventions.
Comparator
Standard of care therapy, including platinum-containing salvage chemotherapy.
Primary question
Does axicabtagene ciloleucel improve EFS compared with standard of care therapy?
3. Trial Design
Randomization was the basis for the primary comparative analysis.
The registry identifies the study as unmasked.
Participants were compared across two treatment groups.
The registry classifies the primary purpose as treatment.
Axicabtagene Ciloleucel
- Axicabtagene ciloleucel
- Cyclophosphamide
- Fludarabine
Standard of Care Therapy
- Platinum-containing salvage chemotherapy
- Standard of care therapy
4. Endpoints
Primary Endpoint: Event Free Survival
Registered endpoint: Event Free Survival (EFS) Per Blinded Central Assessment.
Time frame: From randomization date up to a median follow-up: 24.9 months.
Registry definition: EFS is time from randomization to disease progression (PD), best response of stable disease (SD) up to Day 150, start of new anti-lymphoma therapy including stem cell transplant, or death from any cause.
The endpoint is a composite time-to-event measure. Its event definition combines disease progression, an insufficient response condition at a specified early time point, initiation of new anti-lymphoma therapy including stem cell transplant, and death from any cause.
The registered primary analysis used the Full Analysis Set and compared axicabtagene ciloleucel with standard of care therapy.
Secondary Endpoints With Posted Analyses
| Endpoint | Time frame | Method |
|---|---|---|
| Objective Response Rate (ORR) Per Blinded Central Assessment | From randomization date up to a median follow-up: 24.9 months | Cochran-Mantel-Haenszel test |
| Overall Survival (OS) | Up to 74.9 months | Stratified log-rank test; stratified Cox regression |
| Duration of Response (DOR) Per Blinded Central Assessments | From first confirmed objective response to disease progression or death regardless of cause | Log-rank test; stratified Cox regression |
| Modified Event Free Survival (mEFS) Per Blinded Central Assessment | From randomization date up to a median follow-up: 24.9 months | Log-rank test; stratified Cox regression |
| EFS Per Investigator Disease Assessments | From randomization date up to a median follow-up: 47.2 months | Stratified Cox regression |
| Progression-Free Survival (PFS) Per Investigator Disease Assessments | From randomization date up to a median follow-up: 47.2 months | Stratified Cox regression |
| Modified Event Free Survival (mEFS) Per Investigator Assessment | From randomization date up to a median follow-up: 47.2 months | Stratified Cox regression |
| Change From Baseline in Global Health Status Scores | Baseline, Days 50, 100, and 150; Months 9, 12, 15, 18, 21 and 24 | MMRM |
| Change From Baseline in EORTC QLQ-C30 Physical Functioning Score | Baseline, Days 50, 100, 150, Months 9, 12, 15, 18, 21 and 24 | MMRM |
| Changes From Baseline in the European Quality of Life Five Dimensions Five Levels Scale Index Score | Baseline, Days 50, 100, 150; Months 9, 12, 15, 18, 21 and 24 | MMRM |
| Change From Baseline in EQ-5D-5L VAS Scale Score | Baseline, Days 50, 100, 150; Months 9, 12, 15, 18, 21 and 24 | MMRM |
5. Statistical Methodology
The registry identifies four principal statistical methods for ZUMA-7: the log-rank test, Cox proportional-hazards model, Cochran-Mantel-Haenszel test, and mixed model with repeated measures (MMRM). Together, these methods address time-to-event outcomes, categorical response, and repeated quality-of-life measurements.
| Method | Role in ZUMA-7 | Statistical question |
|---|---|---|
| Stratified log-rank test | Primary EFS comparison and other time-to-event analyses | Are the event-time distributions different between treatment groups? |
| Stratified Cox regression | Hazard-ratio estimation for time-to-event outcomes | What is the estimated relative hazard between treatment groups? |
| Cochran-Mantel-Haenszel test | ORR comparison | Does response differ between groups after stratification? |
| MMRM | Repeated quality-of-life measurements | How does mean change from baseline differ between groups over prespecified time points? |
Stratification
The primary EFS analysis was a stratified log-rank test. The registry states that the one-sided p-value was based on a log-rank test stratified by response to first-line therapy and second-line age-adjusted International Prognostic Index (IPI), as data were collected on case report forms.
The analysis notes also describe the time-to-event analyses as stratified using the randomization stratification factors. Stratification is important because it incorporates prespecified prognostic information into the comparison without estimating a separate treatment coefficient for every stratum.
Analysis Population
For the primary EFS analysis, participants in the Full Analysis Set were analyzed. The registry also identifies intention-to-treat analysis as a concept in the primary and several secondary analyses.
Confidence Intervals and Hypothesis Testing
The primary EFS effect measure was a hazard ratio with a two-sided 95% confidence interval. The superiority hypothesis was evaluated using a one-sided p-value based on the stratified log-rank test.
These are different statistical quantities. The hazard ratio describes the estimated relative hazard, the confidence interval describes uncertainty around that estimate, and the p-value quantifies the evidence against the specified null hypothesis under the test procedure. A p-value is not a measure of the magnitude or clinical importance of an effect.
Cox Model Details
For several secondary time-to-event endpoints, the registry reports stratified Cox regression models to estimate hazard ratios and two-sided 95% confidence intervals. The Breslow method was used to handle ties for the Cox regression models.
Longitudinal Analysis
Quality-of-life endpoints were analyzed with a mixed model with repeated measures. The registry describes these analyses as differences in mean change from baseline at specified time points. This approach accounts for the repeated nature of measurements within participants rather than treating every observation as an independent data point.
Overall Survival Interim Monitoring
The OS analysis includes the statistical concept of interim analysis / alpha spending. The registry states that a Rho family spending function with parameter Rho = 6 was used to allocate alpha between the interim and primary OS analyses.
6. Primary Result: Event Free Survival
The primary endpoint was EFS per blinded central assessment, measured from randomization through the prespecified EFS events. Participants in the Full Analysis Set were analyzed.
Primary EFS effect
95% two-sided CI: 0.308 to 0.514
One-sided p-value: <0.0001
Stratified log-rank test; superiority hypothesis.
| Primary EFS element | Reported result |
|---|---|
| Outcome | Event Free Survival (EFS) Per Blinded Central Assessment |
| Analysis population | Participants in Full Analysis Set |
| Comparison | Axicabtagene Ciloleucel vs Standard of Care Therapy |
| Method | Stratified log-rank test |
| Hazard ratio | 0.398 |
| 95% CI | 0.308 to 0.514 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
| Follow-up frame | From randomization date up to a median follow-up: 24.9 months |
An HR of 0.398 means that, under the proportional-hazards interpretation of the model, the estimated instantaneous risk of an EFS event in the axicabtagene ciloleucel group was about 60.2% lower than in the standard-of-care group. This is a relative hazard interpretation, not a statement that 60.2% of patients avoided an event or that survival time was increased by 60.2%.
The two-sided 95% confidence interval of 0.308 to 0.514 quantifies uncertainty around the estimated hazard ratio. The interval remains below 1, so the reported interval is consistent with a lower estimated event hazard for axicabtagene ciloleucel across the values in that interval.
The p-value of <0.0001 addresses the statistical test of the prespecified superiority hypothesis; it does not quantify effect size. The size of the estimated effect is communicated by the HR itself and its confidence interval.
The interpretation also depends on the time-to-event framework: patients can be censored, and a hazard ratio is not automatically equivalent to a fixed-time risk ratio or a difference in median survival. The registry specifies a stratified log-rank analysis and a two-sided confidence interval, while the reported superiority p-value is one-sided.
7. Secondary Efficacy Results
Objective Response Rate
ORR difference
95% two-sided CI: 23.2 to 42.1
P-value: <0.0001
Stratified Cochran-Mantel-Haenszel test; superiority hypothesis.
The registry reports a 33.1 difference in Objective Response Rate between axicabtagene ciloleucel and standard of care therapy. The 95% confidence interval was calculated using the Wilson score method with continuity correction, while the treatment comparison used a stratified Cochran-Mantel-Haenszel test.
The estimate of 33.1 is a difference in response rates expressed in percentage-point terms. It should not be interpreted as a relative risk or odds ratio. The 95% CI of 23.2 to 42.1 describes the precision of that difference under the reported method.
The p-value of <0.0001 concerns the statistical comparison, not the size of the response difference. The stratification used for the CMH test matters because the analysis compares treatment groups while accounting for the prespecified stratification factors.
Overall Survival
OS effect
95% two-sided CI: 0.530 to 1.007
P-value: 0.027
Stratified log-rank test with stratified Cox regression for HR estimation.
| Element | Reported result |
|---|---|
| Time frame | Up to 74.9 months |
| Analysis population | Participants in Full Analysis Set |
| Hazard ratio | 0.730 |
| 95% CI | 0.530 to 1.007 |
| P-value | 0.027 |
| Hypothesis | Superiority |
| Model details | Stratified Cox regression; Breslow method for ties |
| Interim monitoring | Rho family spending function with Rho = 6 |
An HR of 0.730 corresponds to an estimated instantaneous hazard that was about 27.0% lower for axicabtagene ciloleucel under the proportional-hazards interpretation. It does not mean that the probability of death was exactly 27.0% lower at every follow-up time.
The 95% CI of 0.530 to 1.007 is relatively close to the null value of 1 at its upper boundary. That interval communicates more about precision than the p-value alone does.
The reported p-value is 0.027. Because the registry identifies interim analysis / alpha spending as part of the OS analysis, interpretation should account for the fact that alpha was allocated across the interim and primary OS analyses using the reported Rho family spending function.
Duration of Response
DOR effect
95% two-sided CI: 0.488 to 1.108
P-value: 0.0695
Stratified log-rank test; stratified Cox regression.
Duration of Response was defined from the date of first confirmed objective response (CR or PR) to disease progression or death regardless of cause. Participants in the Full Analysis Set with objective response were analyzed; participants not meeting the criteria by the analysis data cut-off date were censored at their last evaluable disease assessment.
The HR of 0.736 is a relative hazard measure among participants included in the response analysis. The 95% CI of 0.488 to 1.108 spans 1, indicating substantial uncertainty around the estimated relative hazard.
The p-value of 0.0695 is a hypothesis-testing quantity, not an estimate of the probability that one treatment is better. It should also not be used to convert the HR into a clinical magnitude. The analysis is conditional on having an objective response and therefore addresses a different population from the primary Full Analysis Set EFS analysis.
Modified Event Free Survival
mEFS per blinded central assessment
95% two-sided CI: 0.290 to 0.487
P-value: <0.0001
The reported HR of 0.376 corresponds to an estimated instantaneous event hazard about 62.4% lower for axicabtagene ciloleucel under the proportional-hazards interpretation. The 95% CI of 0.290 to 0.487 remains below 1.
The registry identifies the analysis as a stratified log-rank comparison with stratified Cox regression used to estimate the HR and two-sided 95% CI. As with EFS, the HR is a relative time-to-event measure and should not be interpreted as a fixed percentage reduction in cumulative risk at every time point.
Longer-Follow-Up Investigator-Assessed Time-to-Event Outcomes
| Endpoint | HR | 95% CI | Follow-up |
|---|---|---|---|
| EFS Per Investigator Disease Assessments | 0.422 | 0.327 to 0.545 | 47.2 months median follow-up |
| Progression-Free Survival Per Investigator Disease Assessments | 0.506 | 0.383 to 0.669 | 47.2 months median follow-up |
| Modified Event Free Survival Per Investigator Assessment | 0.412 | 0.318 to 0.532 | 47.2 months median follow-up |
All three analyses used participants in the Full Analysis Set and used stratified Cox regression models to estimate hazard ratios and two-sided 95% confidence intervals. The Breslow method was used to handle ties in the Cox regression models.
8. Quality-of-Life Results
The registry reports repeated-measures analyses for global health status, physical functioning, the EQ-5D-5L index, and the EQ-5D-5L VAS. These analyses were based on the QoL analysis set with data available at the relevant time point and used MMRM.
Global Health Status
| Time point | Difference in mean change | 95% CI | P-value |
|---|---|---|---|
| Day 100 | 18.1 | 12.3 to 23.9 | <0.0001 |
| Day 150 | 9.8 | 2.6 to 17.0 | 0.0124 |
| Month 9 | 4.4 | -3.3 to 12.0 | 0.2655 |
These estimates represent differences in mean change from baseline at the specified time point. The registry specifies MMRM as the analysis method.
EORTC QLQ-C30 Physical Functioning
| Time point | Difference in mean change | 95% CI | P-value |
|---|---|---|---|
| Day 100 | 13.1 | 8.0 to 18.2 | <0.0001 |
| Day 150 | 5.1 | -0.9 to 11.0 | 0.1253 |
EQ-5D-5L Index
| Time point | Difference in mean change | 95% CI | P-value |
|---|---|---|---|
| Day 100 | 0.081 | 0.024 to 0.138 | 0.0112 |
| Day 150 | 0.028 | -0.034 to 0.091 | 0.3703 |
EQ-5D-5L VAS
| Time point | Difference in mean change | 95% CI | P-value |
|---|---|---|---|
| Day 100 | 13.7 | 8.5 to 18.8 | <0.0001 |
| Day 150 | 11.3 | 5.4 to 17.1 | 0.0004 |
| Month 9 | 3.8 | -2.3 to 10.0 | 0.2549 |
The MMRM estimates answer a different question from the survival hazard ratios. They describe differences in mean change from baseline at specified time points. For example, the Day 100 global health status estimate of 18.1 is a between-group difference in mean change, not an HR and not a percentage change in survival.
The pattern across time points also illustrates why repeated measurements should not be reduced to a single p-value without considering timing. Several Day 100 estimates have confidence intervals entirely above zero, whereas some later estimates have confidence intervals that cross zero.
The registry's QoL analysis set is a subset of the Full Analysis Set with available data at the given time point. Consequently, these results should not automatically be interpreted as though every randomized participant contributed observations at every assessment time.
9. Safety
The registry provides serious adverse-event counts by treatment group. These are reported as affected participants over participants at risk.
| Group | Participants with serious AEs / at risk |
|---|---|
| Axicabtagene Ciloleucel Treatment | 96 / 170 |
| Standard of Care Therapy | 78 / 168 |
| Retreatment Axicabtagene Ciloleucel | 2 / 10 |
The serious-adverse-event data are presented descriptively here. The ClinicalTrials.gov record does not provide a formal statistical comparison of these serious adverse-event proportions, so no comparative p-value or confidence interval is assigned to them.
10. Statistical Methods Explained
Why was a stratified log-rank test used for EFS?
EFS is a time-to-event endpoint, so the analysis must account for both whether an event occurred and when it occurred, while also accommodating censoring. The log-rank test compares the event-time experience of the randomized groups over follow-up. In ZUMA-7, the test was stratified using the specified randomization stratification factors, allowing the comparison to account for those factors.
What does an HR of 0.398 mean?
An HR of 0.398 is a relative hazard estimate. Under the proportional-hazards interpretation, it indicates an estimated event hazard in the axicabtagene ciloleucel group that is 39.8% of the corresponding hazard in the standard-of-care group. Equivalently, the estimated hazard is about 60.2% lower. It does not mean that exactly 39.8% of participants experienced events or that every participant had a 60.2% reduction in absolute risk.
Why report both a hazard ratio and a confidence interval?
The HR provides the estimated relative treatment effect, while the confidence interval shows how precisely that effect was estimated. For ZUMA-7 EFS, the HR was 0.398 and the 95% CI was 0.308 to 0.514. Looking at both quantities gives more information than the p-value alone.
Why was the Cochran-Mantel-Haenszel test used for ORR?
ORR is a binary response endpoint rather than a time-to-event endpoint. The Cochran-Mantel-Haenszel approach provides a stratified comparison of response between randomized treatment groups. In ZUMA-7, the registry identifies the analysis as stratified by the randomization factors.
Why use MMRM for quality-of-life measurements?
Quality-of-life measurements were collected repeatedly from the same participants. MMRM is designed for longitudinal data in which observations within a participant are related. It estimates treatment-group differences in mean change at specified time points while using the repeated-measures structure rather than treating each measurement as independent.
Why does the OS analysis mention alpha spending?
When an outcome is examined at more than one planned analysis, repeated testing can increase the overall type I error if each analysis uses the full nominal alpha. An alpha-spending approach allocates the available error probability across analyses. The ZUMA-7 registry notes a Rho family spending function with Rho = 6 for allocation between the interim and primary OS analyses.
11. Understanding the EFS Hazard Ratio More Carefully
A hazard ratio compares instantaneous event rates within the survival-analysis framework. It is not an absolute risk difference, risk ratio, or direct estimate of the proportion of patients who remain event-free at a particular time.
This distinction is especially important for composite endpoints such as EFS. Because the endpoint can be triggered by several different events, the hazard ratio represents the combined time-to-first-event process defined by the registry.
The proportional-hazards interpretation also deserves caution. A Cox hazard ratio is most straightforwardly interpreted as a constant relative hazard when proportional hazards are a reasonable approximation. The registry provides the Cox model and its hazard ratio, but the ClinicalTrials.gov record does not provide a separate diagnostic assessment of the proportional-hazards assumption.
12. Confidence Intervals Versus P-values
The ZUMA-7 results illustrate why confidence intervals should accompany p-values. Consider three posted time-to-event results:
| Endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Primary EFS | 0.398 | 0.308 to 0.514 | <0.0001 |
| OS | 0.730 | 0.530 to 1.007 | 0.027 |
| DOR | 0.736 | 0.488 to 1.108 | 0.0695 |
The primary EFS interval is relatively tight and entirely below 1. The OS interval extends to 1.007, while the DOR interval extends above 1. Those interval widths and positions provide information about precision and the range of effects compatible with the analysis that a p-value by itself does not convey.
It is also important not to treat the numerical relationship among these p-values as an effect-size ranking. Statistical significance depends on both the magnitude of an estimate and its uncertainty, as well as the prespecified testing framework.
13. Analysis Populations and What They Change
The registry identifies the Full Analysis Set for the principal efficacy analyses. Quality-of-life analyses instead use the QoL Analysis Set, defined as a subset of the Full Analysis Set with data available at the relevant time point. DOR is analyzed among participants in the Full Analysis Set with objective response.
| Endpoint family | Analysis population | Interpretive consequence |
|---|---|---|
| Primary EFS | Full Analysis Set | Reflects the randomized efficacy comparison as specified in the registry. |
| ORR | Full Analysis Set | Response comparison is made in the broader efficacy population. |
| OS | Full Analysis Set | Long-term survival remains tied to the randomized comparison. |
| DOR | Full Analysis Set with objective response | Analysis is conditional on achieving objective response. |
| QoL | QoL Analysis Set with data available at the given time point | Interpretation depends on the available longitudinal observations. |
This distinction prevents an important statistical error: results from different analysis populations should not be treated as though they answer precisely the same question.
14. Missing Data and Longitudinal Measurements
The registry explicitly states that the QoL analyses used participants with data available at the given time point and that the QoL Analysis Set was a subset of the Full Analysis Set. The ClinicalTrials.gov record identifies MMRM as the method but do not provide a separate description of the missing-data mechanism or a specific imputation procedure.
15. Interim Analysis and Alpha Spending
Overall survival is the clearest example in the ClinicalTrials.gov record of a more complex inferential structure. The registry identifies interim analysis / alpha spending as an analysis concept and reports that a Rho family spending function with parameter Rho = 6 was used to allocate alpha between the interim and primary OS analyses.
Alpha is allocated according to the spending function
The available type I error is divided between the planned interim and primary OS analyses according to the reported Rho family spending function.
Final inferential assessment
The registry reports HR 0.730, a two-sided 95% CI of 0.530 to 1.007, and p = 0.027 for the OS analysis.
The important lesson is that the nominal numerical p-value should be read in the context of the trial's prespecified sequential testing structure. Alpha spending is part of the design, not an after-the-fact adjustment.
16. Stratified Analysis
Stratification appears repeatedly in the ZUMA-7 statistical methods. The primary EFS analysis used a stratified log-rank test, while the ORR analysis used a stratified Cochran-Mantel-Haenszel test. Stratified Cox regression was used to estimate hazard ratios and confidence intervals for several time-to-event endpoints.
The registry specifies the EFS and ORR p-value stratification factors as response to first-line therapy and second-line age-adjusted IPI, as recorded on case report forms. This approach recognizes that these variables can provide prognostic information and incorporates them into the treatment comparison.
Stratification is not adjustment for every baseline variable. A stratified analysis uses specified strata to structure the comparison. The treatment effect is then estimated across those strata according to the relevant statistical method.
17. Secondary Endpoint Results at a Glance
| Endpoint | Effect | 95% CI | P-value |
|---|---|---|---|
| ORR | Difference 33.1 | 23.2 to 42.1 | <0.0001 |
| OS | HR 0.730 | 0.530 to 1.007 | 0.027 |
| DOR | HR 0.736 | 0.488 to 1.108 | 0.0695 |
| mEFS, blinded central | HR 0.376 | 0.290 to 0.487 | <0.0001 |
| EFS, investigator | HR 0.422 | 0.327 to 0.545 | Not reported |
| PFS, investigator | HR 0.506 | 0.383 to 0.669 | Not reported |
| mEFS, investigator | HR 0.412 | 0.318 to 0.532 | Not reported |
The ClinicalTrials.gov record provides confidence intervals for the investigator-assessed EFS, PFS and mEFS results but do not provide p-values for those three analyses. No p-value has therefore been added.
18. Quality-of-Life Endpoint Interpretation
The longitudinal results are particularly useful for understanding how a clinical-trial analysis can move beyond a single primary endpoint. The registry reports several repeated measurements using MMRM and evaluates differences in mean change from baseline at specific time points.
Day 100
Global health status: difference in mean change 18.1, 95% CI 12.3 to 23.9, p <0.0001.
Day 150
Global health status: difference in mean change 9.8, 95% CI 2.6 to 17.0, p = 0.0124.
Physical functioning
Day 100 difference in mean change 13.1, 95% CI 8.0 to 18.2, p <0.0001.
EQ-5D-5L VAS
Day 100 difference in mean change 13.7, 95% CI 8.5 to 18.8, p <0.0001.
These values should be interpreted on their respective scales. A score difference of 13.7 on the EQ-5D-5L VAS is not commensurate with an HR of 0.398, because the two statistics measure fundamentally different quantities.
19. Limitations
- Open-label design: The registry classifies masking as none. This is important when interpreting endpoints that may be susceptible to knowledge of treatment assignment, although the primary EFS endpoint was based on blinded central assessment.
- Composite primary endpoint: EFS incorporates several distinct event types. A treatment effect on the composite does not identify which individual component is primarily responsible without separate component analyses.
- Hazard-ratio interpretation: HRs are relative time-to-event measures and should not automatically be translated into fixed-time risk reductions or differences in median event times.
- Proportional-hazards assumption: The ClinicalTrials.gov record reports Cox regression but do not provide a separate diagnostic assessment of proportional hazards.
- Different analysis populations: DOR is restricted to participants with objective response, while QoL analyses use the QoL Analysis Set with available observations.
- Longitudinal missingness: The ClinicalTrials.gov record identifies MMRM and the QoL analysis set but do not document a specific imputation strategy or missing-data mechanism.
- Multiple analyses: The trial contains a primary endpoint and numerous secondary endpoints. The ClinicalTrials.gov record explicitly identify interim alpha spending for OS, but they do not provide a complete multiplicity-adjustment hierarchy for every secondary endpoint.
- Safety comparison: Serious adverse-event counts are reported, but the ClinicalTrials.gov record does not provide formal comparative statistical testing for those counts.
- Follow-up differs by endpoint: The posted analyses use different follow-up frames, including 24.9 months, 47.2 months, and 74.9 months. These estimates should not be treated as though they came from one common analysis cutoff.
20. Why This Trial Matters Statistically
ZUMA-7 is a useful statistical teaching case because several important clinical-trial methods appear within a single randomized comparison. The primary endpoint is a time-to-event composite requiring survival-analysis methods, while secondary endpoints introduce categorical, longitudinal, and longer-follow-up analyses.
| Concept | How it appears in ZUMA-7 |
|---|---|
| Randomization | Randomized parallel phase 3 treatment comparison |
| Full Analysis Set | Used for the primary EFS and several secondary efficacy analyses |
| Time-to-event endpoint | EFS, mEFS, OS, DOR and investigator-assessed PFS |
| Hazard ratio | Relative treatment effect for time-to-event endpoints |
| Confidence interval | Quantifies uncertainty around HR and ORR-difference estimates |
| Stratified log-rank test | Primary EFS comparison and other survival analyses |
| Cox proportional-hazards model | Estimation of HRs and two-sided confidence intervals |
| Breslow ties method | Used for ties in Cox regression models |
| Cochran-Mantel-Haenszel test | Stratified comparison of ORR |
| MMRM | Repeated quality-of-life measurements |
| Interim analysis | Reported as part of the OS analysis |
| Alpha spending | Rho family spending function with Rho = 6 for interim and primary OS analyses |
21. A Practical Statistical Reading of ZUMA-7
A disciplined reading of the trial results starts with the design rather than the p-values. The study was randomized, which establishes the fundamental comparative framework. The primary endpoint was defined prospectively as EFS, and the registry identifies a stratified log-rank test as the primary analysis.
The primary estimate, HR 0.398, describes a substantial relative difference in the hazard of an EFS event under the model interpretation. The accompanying confidence interval, 0.308 to 0.514, provides the uncertainty range around that estimate. The p-value, <0.0001, supplies evidence against the specified superiority null hypothesis but should not be treated as a measure of clinical magnitude.
The secondary results then add different dimensions. ORR uses a categorical-data method, OS extends the time-to-event framework, DOR conditions on objective response, investigator-assessed endpoints provide longer follow-up, and MMRM addresses repeated quality-of-life measurements.
This layered interpretation is preferable to reducing the entire trial to a single number. Each statistical estimate answers a particular question defined by its endpoint, analysis population, follow-up period, and model.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: ZUMA-7, NCT03391466.
- PubMed: PMID 38635762.
- PubMed: PMID 35114155.
- PubMed: PMID 34922648.
- PubMed: PMID 34515338.
- PubMed: PMID 33288485.
Continue through the Clinical Biostats knowledge graph
Connect the endpoints and statistical methods in ZUMA-7 to deeper tutorials, calculators, and clinical-trial analysis workflows.
25. Record Summary
ZUMA-7 provides a compact example of several core clinical-trial statistical principles: randomized comparison, a composite time-to-event primary endpoint, stratified log-rank testing, hazard-ratio estimation with Cox regression, categorical response analysis using the Cochran-Mantel-Haenszel test, longitudinal quality-of-life analysis using MMRM, and interim monitoring with alpha spending for overall survival.
The primary EFS analysis reported an HR of 0.398 with a two-sided 95% CI of 0.308 to 0.514 and a one-sided p-value of <0.0001. The most important statistical lesson is not simply the numerical result but the framework around it: the endpoint definition, Full Analysis Set, stratification, censoring, hazard-ratio interpretation, confidence interval, and prespecified hypothesis-testing procedure all contribute to what the result actually means.
The secondary analyses illustrate why clinical-trial interpretation should remain endpoint-specific. ORR is a categorical outcome, OS and DOR are time-to-event outcomes with different analysis populations or inferential structures, and QoL outcomes are repeated continuous measurements. These estimates should therefore be interpreted on their own statistical scales rather than collapsed into a single summary statistic.