This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
CALLA was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating durvalumab added to standard-of-care chemoradiotherapy in women with locally advanced cervical cancer. The registry reports 770 participants, one primary time-to-event endpoint, seven posted outcome measures, and five posted statistical analyses.
| Feature | CALLA |
|---|---|
| Trial name | CALLA |
| ClinicalTrials.gov identifier | NCT03830866 |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Locally Advanced Cervical Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 770 |
| Lead sponsor | AstraZeneca |
| Sponsor type | Industry |
2. Clinical Question
The central statistical question was whether adding durvalumab to standard-of-care chemoradiotherapy changed progression-free survival compared with placebo plus standard-of-care chemoradiotherapy in women with locally advanced cervical cancer.
Population
Women with locally advanced cervical cancer, as defined by the CALLA trial registry record.
Intervention
Durvalumab with standard-of-care chemoradiotherapy, consisting of the registered chemotherapy and radiation components.
Comparator
Placebo with standard-of-care chemoradiotherapy.
Primary question
Does durvalumab plus standard-of-care chemoradiotherapy improve progression-free survival relative to placebo plus standard-of-care chemoradiotherapy?
3. Trial Design
Durvalumab + SoC CCRT
- Durvalumab
- Cisplatin or carboplatin
- External beam radiation therapy (EBRT)
- Brachytherapy
Placebo + SoC CCRT
- Placebo
- Cisplatin or carboplatin
- External beam radiation therapy (EBRT)
- Brachytherapy
The statistical design is important because randomization establishes the primary comparison between the two treatment strategies, while quadruple masking is intended to reduce the influence of treatment knowledge on trial conduct and assessment.
4. Endpoints
| Endpoint | Definition / time frame | Analysis |
|---|---|---|
| Primary: Progression-free Survival (PFS) | Time from date of randomisation until date of tumour progression or death by any cause, regardless of whether the patient withdrew from randomized therapy or received another anticancer therapy prior to progression. Tumor assessments start 20 weeks after randomisation then every 12 weeks up to 164 weeks, then every 24 weeks until date. | Log-rank test; hazard ratio |
| Secondary: PFS, PD-L1 Expression ≥ 1% | Tumor assessments start 20 weeks after randomisation then every 12 weeks up to 164 weeks, then every 24 weeks until date. | Log-rank test; hazard ratio |
| Secondary: Overall Survival (Duration) | Time from date of randomisation until date of death by any cause, assessed up to the data cut-off date (3rd July 2023). | Log-rank test; hazard ratio |
| Secondary: Objective Response Rate (ORR) | Tumor assessments start 20 weeks after randomisation then every 12 weeks up to 164 weeks, then every 24 weeks until date. | Logistic regression; odds ratio |
| Secondary: Complete Response Rate | Tumor assessments start 20 weeks after randomisation then every 12 weeks up to 164 weeks, then every 24 weeks until date. | Logistic regression; odds ratio |
The primary endpoint is explicitly a time-to-event endpoint. Its event definition includes either tumor progression or death, and the registry specifies that the clock begins at randomization. This matters because censoring, treatment discontinuation, and subsequent anticancer therapy do not redefine the registered PFS event itself.
5. Statistical Methodology
Log-rank testing for time-to-event endpoints
The registry reports the log-rank test for the primary PFS analysis and for the secondary PFS and overall-survival analyses. The log-rank test compares the observed pattern of event occurrence between randomized groups over follow-up rather than reducing the outcome to a single binary status at one arbitrary time point.
For CALLA, the reported formal method for the primary PFS comparison is the log-rank test, with the treatment effect summarized using a hazard ratio.
Hazard ratio
The primary PFS effect measure was a hazard ratio. A hazard ratio compares the estimated instantaneous event rates between the randomized groups over the analyzed follow-up. An HR below 1 indicates a lower estimated instantaneous event rate in the durvalumab group relative to the placebo group under the analysis framework.
An HR of 0.84 can be described as an estimated 16% lower instantaneous event hazard, because 1 − 0.84 = 0.16. It is not a statement that 16% of patients avoided progression or death.
Logistic regression for binary endpoints
The registry reports logistic regression for objective response rate and complete response rate. Logistic regression is appropriate for an endpoint represented as a binary outcome, such as whether a participant met the prespecified response definition.
An OR above 1 indicates higher estimated odds of the binary outcome in the durvalumab group. Odds are not the same as probabilities, so an odds ratio should not automatically be described as a percentage increase in response probability.
Intention-to-treat analysis
The registry analysis text identifies the intention-to-treat concept for the primary PFS analysis, overall survival analysis, objective response rate, and complete response rate. The primary PFS and overall-survival analyses use the Full Analysis Set. The intention-to-treat principle preserves the randomized treatment comparison by analyzing participants according to their assigned treatment rather than allowing later treatment changes to redefine the original randomized groups.
Kaplan-Meier estimation
Kaplan-Meier estimation is the standard descriptive framework for displaying and estimating time-to-event distributions such as PFS and overall survival. It accommodates right censoring by allowing a participant to contribute follow-up information until the participant experiences the event or reaches a censoring time. The registry's reported formal comparison for the CALLA time-to-event analyses is the log-rank test.
Confidence intervals
The reported treatment effects include two-sided 95% confidence intervals. A confidence interval communicates statistical uncertainty around the estimated treatment effect. It is not a range containing 95% of individual patient outcomes, nor does it describe the probability that the true effect lies inside the particular interval after the data have been observed.
6. Primary Result: Progression-Free Survival
The primary endpoint was progression-free survival based on investigator assessment according to RECIST 1.1 or histopathologic confirmation of local tumour progression. The analysis used the Full Analysis Set and compared durvalumab plus standard-of-care chemoradiotherapy with placebo plus standard-of-care chemoradiotherapy.
Primary PFS hazard ratio
95% CI: 0.65–1.08 · P = 0.174
Two-sided log-rank analysis; Full Analysis Set
The estimated hazard ratio of 0.84 is below 1. In relative terms, this corresponds to a 16% lower estimated instantaneous hazard of progression or death in the durvalumab group under the reported analysis.
The HR does not mean that 16% of participants avoided progression or death, and it does not mean that every participant experienced a 16% reduction in individual risk. A hazard ratio is a relative time-to-event measure describing the comparison between the two randomized groups.
The 95% confidence interval of 0.65–1.08 indicates uncertainty around the estimated HR. Because the interval includes 1, the interval is compatible with a range of relative treatment effects that includes no difference in hazard.
The P-value of 0.174 is a measure of the compatibility of the observed result with the statistical null hypothesis under the specified testing framework. It is not a measure of the size, importance, or probability of the treatment effect. The effect estimate and its confidence interval are therefore essential to interpreting the result.
Because this is a time-to-event analysis, interpretation also depends on the handling of censoring and on the assumptions underlying hazard-based summaries. A single HR is most straightforward when the relative hazards are reasonably stable over time; it should not be interpreted as an absolute risk difference or as a universal effect for every patient.
7. Secondary Time-to-Event Results
PFS in the PD-L1 Analysis Set
The registry reports a secondary PFS analysis restricted to the PD-L1 Analysis Set and defined for participants with PD-L1 expression ≥ 1%. The same tumor-assessment schedule was used: assessments start 20 weeks after randomisation, then occur every 12 weeks up to 164 weeks, followed by every 24 weeks until date.
Secondary PFS hazard ratio
95% CI: 0.64–1.10 · P = 0.203
Two-sided log-rank analysis; PD-L1 Analysis Set
The estimated HR of 0.84 corresponds to a 16% lower estimated instantaneous hazard of progression or death in the durvalumab group relative to the placebo group under this analysis. This is an effect estimate, not a statement that 16% of participants benefited.
The 95% CI of 0.64–1.10 spans 1, so the reported interval includes the possibility of no difference in hazard. The P-value of 0.203 does not measure the magnitude of the estimated effect; it quantifies evidence against the null hypothesis within the specified testing framework.
This analysis is based on the PD-L1 Analysis Set rather than the Full Analysis Set used for the primary PFS analysis. The distinction in analysis population is therefore important when comparing the two estimates.
Overall Survival
Overall survival was a secondary time-to-event endpoint defined as the time from randomisation until death by any cause. The registry specifies assessment up to the data cut-off date of 3rd July 2023.
Overall-survival hazard ratio
95% CI: 0.60–1.04 · P = 0.091
Two-sided log-rank analysis; Full Analysis Set
An HR of 0.79 corresponds to a 21% lower estimated instantaneous hazard of death in the durvalumab group relative to the placebo group under the reported model-free comparison and effect-measure framework.
The HR does not mean that 21% of patients lived longer, nor does it mean that each patient experienced exactly a 21% reduction in mortality risk. It is a relative time-to-event measure across follow-up.
The 95% CI of 0.60–1.04 includes 1, so the interval includes no difference in hazard. The P-value of 0.091 is not an effect-size measure. It should be interpreted together with the HR, confidence interval, endpoint definition, analysis population, and follow-up framework.
Overall survival is also a distinct endpoint from PFS: death from any cause is the event, regardless of the mechanism of death. Consequently, the statistical question answered by OS is not identical to the question answered by PFS.
8. Secondary Binary Endpoint Results
Objective Response Rate
Objective response rate was analyzed in the Full analysis set using logistic regression. The treatment effect was expressed as an odds ratio.
Objective response rate odds ratio
95% CI: 0.794–1.657 · P = 0.465
Two-sided logistic regression; Full analysis set
An odds ratio of 1.15 means that the estimated odds of the binary response outcome were 1.15 times those in the placebo group under the reported logistic-regression analysis. This is an odds comparison, not a 15% increase in response probability.
The 95% CI of 0.794–1.657 describes uncertainty around the odds-ratio estimate and includes 1. The P-value of 0.465 is not a measure of the size of the observed odds ratio and should not be used as a substitute for the confidence interval.
Because ORR is a binary endpoint rather than a time-to-event endpoint, the relevant statistical framework is different from the log-rank analysis used for PFS and OS.
Complete Response Rate
Complete response rate was also analyzed in the Full analysis set using logistic regression.
Complete response rate odds ratio
95% CI: 0.833–1.487 · P = 0.469
Two-sided logistic regression; Full analysis set
The estimated odds ratio of 1.11 indicates estimated odds of complete response that were 1.11 times those in the placebo group under the reported analysis. It does not mean that the probability of complete response was 11% higher.
The 95% CI of 0.833–1.487 includes 1, indicating that the uncertainty interval includes no difference in odds. The P-value of 0.469 describes statistical evidence under the specified hypothesis test; it does not quantify the clinical magnitude of the treatment effect.
9. Comparing the Different Effect Measures
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Primary PFS | Hazard ratio | 0.84 | 0.65–1.08 | 0.174 |
| PFS, PD-L1 Expression ≥ 1% | Hazard ratio | 0.84 | 0.64–1.10 | 0.203 |
| Overall Survival | Hazard ratio | 0.79 | 0.60–1.04 | 0.091 |
| Objective Response Rate | Odds ratio | 1.15 | 0.794–1.657 | 0.465 |
| Complete Response Rate | Odds ratio | 1.11 | 0.833–1.487 | 0.469 |
The table illustrates why treatment effects from different endpoint types should not be placed on a single numerical scale. Hazard ratios describe relative event hazards over time, whereas odds ratios compare odds for binary outcomes. An HR of 0.84 and an OR of 1.15 are therefore not directly comparable statements about treatment magnitude.
10. Safety Results
The registry reports serious adverse events by randomized treatment arm using affected participants over participants at risk.
| Safety measure | Durva + SoC CCRT | Placebo + SoC CCRT |
|---|---|---|
| Serious adverse events, affected / at risk | 113 / 385 | 90 / 384 |
The registry supplies affected and at-risk counts for serious adverse events rather than a formal comparative effect estimate. The denominator differs from the overall enrollment figure, so the safety counts should not be interpreted as though 385 and 384 were the original randomized sample sizes.
11. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS records the timing of progression or death, not simply whether an event occurred. The log-rank test uses the ordering of event times throughout follow-up and compares the survival experience of the randomized groups. This makes it suitable for a time-to-event endpoint such as the registered primary PFS outcome.
What does a hazard ratio of 0.84 mean?
An HR of 0.84 indicates a lower estimated instantaneous event hazard in the durvalumab group relative to the placebo group. Numerically, 1 − 0.84 = 0.16, so it can be described as a 16% lower estimated hazard. It does not mean a 16-percentage-point reduction in progression, and it does not mean every patient experienced a 16% reduction in risk.
Why does the confidence interval matter?
The point estimate alone does not communicate how precisely the treatment effect was estimated. The primary PFS confidence interval is 0.65–1.08. Because it includes 1, the interval encompasses the null value for a hazard ratio. The interval therefore gives information that cannot be obtained from the point estimate alone.
Why does the P-value not measure effect size?
A P-value summarizes evidence against a specified null hypothesis under a particular statistical model and testing framework. It is affected by factors including sample size and variability. It therefore should not be interpreted as the magnitude or clinical importance of an effect. The HR or OR and its confidence interval provide the effect-size information.
Why is logistic regression used for ORR?
ORR is a binary endpoint: a participant either meets the response definition or does not. Logistic regression models the odds of that binary outcome and can express the treatment comparison as an odds ratio. This differs fundamentally from the time-to-event framework used for PFS and OS.
Why is the Full Analysis Set important?
The primary PFS and secondary OS analyses in the registry use the Full Analysis Set. An analysis population defines which participants contribute to the treatment comparison. Using the randomized-analysis population helps preserve the comparison established by randomization and reduces the risk that post-randomization treatment decisions determine who is included in the efficacy analysis.
12. Confidence Intervals, Null Values, and Interpretation
Hazard ratio null value
For an HR, the null value is 1.00 because an HR of 1 represents equal hazards between the compared groups.
Odds ratio null value
For an OR, the null value is also 1.00 because an OR of 1 represents equal odds between the groups.
Primary PFS interval
The 95% CI is 0.65–1.08, so it spans the HR null value of 1.
ORR interval
The 95% CI is 0.794–1.657, so it spans the OR null value of 1.
Confidence intervals are especially useful for avoiding overinterpretation of isolated P-values. For CALLA, the primary PFS estimate is below 1, but its confidence interval extends above 1. The same basic interpretive principle applies to the secondary PFS, overall-survival, ORR, and complete-response estimates.
13. The Role of Analysis Populations
| Endpoint | Analysis population | Why it matters |
|---|---|---|
| Primary PFS | Full Analysis Set | Primary efficacy comparison based on the registered analysis population. |
| PFS, PD-L1 Expression ≥ 1% | PD-L1 Analysis Set | Restricted population; its estimate should not be treated as identical to the primary Full Analysis Set analysis. |
| Overall Survival | Full Analysis Set | Uses the randomized efficacy analysis population. |
| Objective Response Rate | Full analysis set | Binary response analysis based on the registered analysis population. |
| Complete Response Rate | Full analysis set | Binary response analysis based on the registered analysis population. |
The difference between the Full Analysis Set and PD-L1 Analysis Set is statistically important. A subgroup or biomarker-defined analysis can address a more specific question, but its result should not be silently substituted for the overall randomized analysis.
14. Statistical Interpretation of the Primary Endpoint
The primary PFS HR of 0.84 is below 1, corresponding to a 16% lower estimated instantaneous hazard of progression or death for durvalumab plus standard-of-care chemoradiotherapy relative to placebo plus standard-of-care chemoradiotherapy.
The 95% CI of 0.65–1.08 indicates that the estimate has substantial uncertainty. The interval includes 1, the null value for a hazard ratio, so the reported data are compatible with no difference in hazard as well as with effects in either direction within the interval.
The P-value of 0.174 does not represent the probability that the treatment works or fails, and it does not quantify the clinical size of the effect. It is a hypothesis-testing quantity that should be considered alongside the effect estimate and confidence interval.
PFS includes tumor progression or death by any cause and begins at randomisation. That composite definition is important: the reported HR concerns the registered PFS event, not tumor shrinkage alone and not mortality alone.
15. Secondary Endpoint Interpretation
The secondary analyses show the same general statistical pattern of effect estimates being close enough to the null that their confidence intervals include 1.
| Secondary endpoint | Estimate | 95% CI | P-value | Interpretive point |
|---|---|---|---|---|
| PFS, PD-L1 Expression ≥ 1% | HR 0.84 | 0.64–1.10 | 0.203 | CI includes the null HR of 1. |
| Overall Survival | HR 0.79 | 0.60–1.04 | 0.091 | CI includes the null HR of 1. |
| Objective Response Rate | OR 1.15 | 0.794–1.657 | 0.465 | CI includes the null OR of 1. |
| Complete Response Rate | OR 1.11 | 0.833–1.487 | 0.469 | CI includes the null OR of 1. |
These estimates should be interpreted endpoint by endpoint rather than combined into a single overall conclusion. PFS, OS, ORR, and complete response rate measure different aspects of the clinical course and use different statistical effect measures.
16. Limitations and Interpretation Issues
- Registry-level detail: The available trial data identify the formal statistical methods and reported estimates but do not provide every detail of the statistical analysis plan.
- Different analysis populations: The primary PFS and overall-survival analyses use the Full Analysis Set, whereas the PD-L1-defined PFS analysis uses the PD-L1 Analysis Set.
- Time-to-event assumptions: Hazard ratios are relative measures of event hazard. Their interpretation should not be converted into an absolute risk difference or individual-level treatment probability.
- Censoring: PFS and OS are time-to-event endpoints and therefore involve follow-up and censoring considerations. The registry definition of PFS specifies the event rule but does not provide the full censoring algorithm in the ClinicalTrials.gov record.
- Binary endpoints: Odds ratios for ORR and complete response rate are not directly comparable with hazard ratios for PFS and OS.
- Multiplicity: The ClinicalTrials.gov record identifies a primary endpoint and multiple secondary analyses, but do not provide a detailed multiplicity-adjustment strategy. The posted P-values should therefore be interpreted in the context of the registered endpoint hierarchy rather than assuming an unstated multiplicity procedure.
- Interim analysis: The ClinicalTrials.gov record does not report an interim-analysis or alpha-spending strategy, so no such procedure is inferred here.
- Missing-data and imputation methods: The ClinicalTrials.gov record does not report a specific missing-data or imputation strategy. No method is assumed.
- Stratification: The ClinicalTrials.gov record does not report randomization strata or stratified analysis factors. None are inferred.
- Bayesian methods: No Bayesian method is reported in the ClinicalTrials.gov record.
17. Why This Trial Matters Statistically
CALLA is a useful teaching case because the registry connects a randomized phase 3 design to several major families of clinical-trial statistics: time-to-event analysis, binary-response analysis, effect measures, confidence intervals, and intention-to-treat principles.
| Concept | How it appears in CALLA |
|---|---|
| Randomization | The trial uses randomized allocation in a parallel-group design. |
| Blinding | The registry records quadruple masking. |
| Intention-to-treat analysis | The analysis text identifies the intention-to-treat concept for several efficacy endpoints. |
| Time-to-event analysis | PFS is the registered primary endpoint and OS is a secondary endpoint. |
| Log-rank test | Used for the reported PFS and OS analyses. |
| Hazard ratio | Used to summarize the treatment comparison for PFS and OS. |
| Logistic regression | Used for ORR and complete response rate. |
| Odds ratio | Used to summarize the binary endpoint comparisons. |
| Confidence intervals | Reported as two-sided 95% intervals for all five posted statistical analyses. |
| Analysis populations | Full Analysis Set and PD-L1 Analysis Set are explicitly identified. |
The trial is therefore particularly useful for learning how the same randomized comparison can generate different statistical questions. PFS and OS ask when an event occurs, whereas ORR and complete response rate ask whether a participant meets a binary outcome definition.
18. Trial Timeline
Trial start
The registry records the start of CALLA on 2019-02-15.
Primary completion
The registry records primary completion on 2022-01-20.
Overall-survival data cut-off
The registered overall-survival endpoint specifies assessment up to the data cut-off date of 3rd July 2023.
Trial status
CALLA is recorded as completed, with results posted on ClinicalTrials.gov.
19. What the Hazard Ratio Does — and Does Not — Mean
A hazard ratio of 0.84 means that the estimated instantaneous hazard of progression or death was 0.84 times that in the comparator group under the reported analysis. Equivalently, it corresponds to a 16% lower estimated hazard.
It does not mean that 16% of participants were protected from progression or death, that survival increased by 16%, or that every participant experienced a 16% reduction in risk.
The primary PFS 95% CI of 0.65–1.08 describes uncertainty around the estimated HR. It does not describe the range of individual treatment responses. Because the interval includes 1, the null value remains within the interval.
The primary PFS P-value of 0.174 should not be read as a 17.4% probability that the treatment effect is absent or present. A P-value is conditional on a specified null hypothesis and statistical testing framework; it is not a probability assigned to the competing scientific hypotheses.
20. What the Odds Ratio Does — and Does Not — Mean
ORR: OR 1.15
The estimated odds of objective response were 1.15 times the comparator odds. This is not equivalent to saying the response probability increased by 15%.
Complete response: OR 1.11
The estimated odds of complete response were 1.11 times the comparator odds. Again, odds and probability are different quantities.
For common outcomes, an odds ratio can differ substantially from a risk ratio. Even when an OR is numerically close to 1, its interpretation depends on the underlying response probabilities. Without the underlying response counts or percentages in the ClinicalTrials.gov record, this page does not convert the reported odds ratios into response probabilities.
21. Overall Statistical Reading of the Evidence
The primary analysis reports a PFS hazard ratio of 0.84 with a two-sided 95% confidence interval of 0.65–1.08 and P-value of 0.174. The secondary PFS analysis in the PD-L1 Analysis Set reports an HR of 0.84 with a 95% CI of 0.64–1.10 and P-value of 0.203. The secondary OS analysis reports an HR of 0.79 with a 95% CI of 0.60–1.04 and P-value of 0.091.
The two binary secondary endpoints use a different statistical framework. ORR has an OR of 1.15 with a 95% CI of 0.794–1.657 and P-value of 0.465. Complete response rate has an OR of 1.11 with a 95% CI of 0.833–1.487 and P-value of 0.469.
Across these analyses, the confidence intervals for the reported HRs and ORs include their respective null value of 1. This is the most direct statistical feature to recognize when interpreting the posted effect estimates. It does not eliminate the importance of the point estimates; rather, it emphasizes that the estimates should be read together with their uncertainty intervals and analysis populations.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: CALLA, NCT03830866.
- PubMed: PubMed record for PMID 38039991.
- PubMed: PubMed record for PMID 32447296.
Continue through the Clinical Biostats statistical methods library
Use the related tutorials and calculators to explore the time-to-event and binary-outcome methods represented in CALLA.
25. Record Summary
CALLA provides a compact example of how a randomized phase 3 oncology trial can combine multiple statistical frameworks. Its primary endpoint is a time-to-event measure analyzed with a log-rank test and summarized by a hazard ratio, while secondary response endpoints use logistic regression and odds ratios. The posted results also illustrate why the analysis population, confidence interval, P-value, and endpoint definition must all be considered together rather than treating a single numerical result as a complete statistical conclusion.
The primary PFS analysis reports an HR of 0.84 (95% CI 0.65–1.08; P = 0.174). Secondary analyses report HRs of 0.84 for PFS in the PD-L1 Analysis Set and 0.79 for overall survival, alongside ORs of 1.15 for objective response rate and 1.11 for complete response rate. Each estimate has its own uncertainty interval and statistical context.