This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the ClinicalTrials.gov record. The registry provides the official trial record.
1. Trial at a Glance
CONTACT-01 was a completed, randomized phase 3 trial evaluating atezolizumab in combination with cabozantinib versus docetaxel monotherapy in patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy.
| Feature | CONTACT-01 |
|---|---|
| Phase | Phase 3 |
| Status | COMPLETED |
| Condition | Carcinoma, Non-Small-Cell Lung |
| Population | Patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy |
| Design | Randomized, parallel-group |
| Masking | NONE |
| Allocation | RANDOMIZED |
| Primary purpose | TREATMENT |
| Enrollment | 366 |
| Primary endpoint | Overall Survival (OS) |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Statistical analyses posted | 15 |
| ClinicalTrials.gov | NCT04471428 |
2. Clinical Question
The central statistical question was whether treatment with atezolizumab plus cabozantinib differed from docetaxel monotherapy with respect to overall survival in patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy.
Population
Patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy.
Intervention
Atezolizumab in combination with cabozantinib.
Comparator
Docetaxel monotherapy.
Primary question
Does the randomized treatment comparison produce a difference in overall survival?
3. Trial Design
Atezolizumab + Cabozantinib
- Atezolizumab
- Cabozantinib
- Combination treatment arm
Docetaxel Monotherapy
- Docetaxel
- Monotherapy control arm
The ClinicalTrials.gov record describes a parallel, randomized, unmasked phase 3 treatment comparison. The primary efficacy analysis used the intention-to-treat population, which included all randomized participants whether or not they received the assigned treatment.
4. Endpoints
The registry lists one registered primary endpoint: overall survival. The statistical analyses posted on ClinicalTrials.gov also report secondary analyses for progression-free survival, objective response, patient-reported deterioration, PFS rates, and OS rates.
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Overall Survival (OS) | OS was defined as the time from randomization to death from any cause. Participants alive at the time of the analysis were censored at the date when they were last known to be alive as documented by the investigator. Time frame: up to approximately 24 months. | Time-to-event |
| Progression-Free Survival (PFS) as Determined by Investigator | Time frame: up to approximately 24 months. | Time-to-event |
| Confirmed Objective Response Rate (ORR) as Determined by Investigator | Time frame: up to approximately 24 months. | Binary |
| Time to Confirmed Deterioration (TTCD) in Patient-reported Physical Functioning (PF) | Time frame: up to approximately 24 months. Participants without a confirmed deterioration at the time of analysis were censored according to the registry analysis description. | Time-to-event |
| TTCD in Patient-reported Global Health Status (GHS) | Time frame: up to approximately 24 months. Participants without a confirmed deterioration at the time of analysis were censored at the last time they were known to have not deteriorated. | Time-to-event |
| PFS Rates Assessed by Investigator | 6 months and 1 year. | Binary |
| OS Rates | 1 and 2 years. | Binary |
5. Analysis Populations and Stratification
The reported efficacy analyses use the intention-to-treat (ITT) population. The registry defines this population as all randomized participants, whether or not the participant received the assigned treatment. This is important because the treatment comparison remains anchored to randomization rather than being redefined after treatment exposure or discontinuation.
| Analysis population | Definition in the ClinicalTrials.gov record | Role |
|---|---|---|
| ITT | All randomized participants, whether or not the participant received the assigned treatment. | Primary efficacy and reported secondary efficacy analyses |
| Safety analysis at risk | The registry-reported serious-adverse-event summary reports 58/167 for docetaxel monotherapy and 76/185 for atezolizumab + cabozantinib. | Safety description |
Several analyses were explicitly described as stratified. The ClinicalTrials.gov record identifies histology and prior NSCLC treatment regimens as stratification factors. The same factors appear in the descriptions of the stratified hazard-ratio and response analyses.
Why stratification matters
Stratified analysis allows the comparison to account for prespecified factors used to organize the randomized comparison. It can improve alignment between the analysis and the trial's randomization structure.
Stratified vs unstratified
The registry reports both stratified and unstratified analyses for several endpoints. These are different statistical analyses of the same randomized comparison, not separate treatment effects.
6. Statistical Methodology
Kaplan-Meier estimation
The registered OS endpoint is a time-to-event outcome. The registry states that the Kaplan-Meier method was used to estimate the median, with the 95% confidence interval for the median computed using the method of Brookmeyer and Crowley.
The Kaplan-Meier framework is appropriate when follow-up times differ among participants and some participants remain alive at the time of analysis. Rather than treating those participants as if they had experienced the event, their observations are censored at the last date they were known to be alive.
Here, di represents the number of events at time ti, while ni is the number at risk immediately before that time.
Log-rank test
The primary OS comparison was evaluated using a log-rank test. The same method was reported for PFS and the two patient-reported time-to-deterioration endpoints. The log-rank test compares the observed and expected event patterns between randomized treatment groups over follow-up.
Cox regression and hazard ratios
The registry states that the hazard ratio was estimated using a Cox regression model. For the stratified analyses, the registry-reported analysis notes identify histology and prior NSCLC treatment regimens as stratification factors.
An HR below 1 indicates a lower estimated instantaneous event rate in the first-named treatment group relative to the comparator under the fitted model. It is not an absolute risk difference, a probability of survival, or a statement that every patient experiences the same proportional change.
Cochran-Mantel-Haenszel analysis
Confirmed objective response rate was analyzed using the Cochran-Mantel-Haenszel test in the stratified analysis. This method provides a way to compare categorical outcomes across treatment groups while accounting for prespecified strata.
Logistic regression
Odds ratios for confirmed objective response were estimated using logistic regression. The registry reports both stratified and unstratified logistic-regression analyses.
Wald / z-test methods
The reported PFS-rate and OS-rate comparisons used a z-test. The registry describes these as Wald / z-test analyses and states that confidence intervals for differences in rates were estimated using the normal approximation, with standard errors computed using the Greenwood method for the time-to-event-derived rates.
7. Primary Result: Overall Survival
Overall survival was the sole registered primary endpoint. The ClinicalTrials.gov record contains two formal analyses: a stratified analysis and an unstratified analysis. Both use the ITT population and compare docetaxel monotherapy with atezolizumab plus cabozantinib.
Stratified Overall Survival Analysis
Hazard ratio for overall survival
95% CI: 0.676–1.156 · P = 0.3668
Stratified log-rank analysis; hazard ratio estimated by Cox regression.
| Primary OS analysis | Reported result |
|---|---|
| Population | ITT population |
| Comparison | Docetaxel Monotherapy vs Atezolizumab + Cabozantinib |
| Method | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.884 |
| 95% CI | 0.676–1.156 |
| P-value | 0.3668 |
| Hypothesis type | Superiority |
| Stratification factors | Histology and prior NSCLC treatment regimens |
The estimated hazard ratio of 0.884 indicates that the fitted Cox model estimated a lower instantaneous rate of death for the first-named group relative to the comparator, with the estimate corresponding to approximately a 11.6% lower estimated hazard because 1 − 0.884 = 0.116.
That statement is about the estimated relative hazard; it does not mean that 11.6% fewer participants died, that an individual patient's probability of death fell by 11.6%, or that survival time increased by 11.6%.
The 95% confidence interval, 0.676–1.156, expresses uncertainty around the estimated hazard ratio. Because the interval includes 1, the reported data are compatible with both a lower and a higher hazard under the model. The width of the interval also shows that the estimate is not precise enough to be interpreted as a narrowly defined treatment effect.
The P-value of 0.3668 addresses the statistical evidence against the relevant null hypothesis under the reported testing framework. It does not measure the size or clinical importance of the treatment effect. A P-value should therefore be interpreted alongside the hazard ratio and its confidence interval.
Finally, a Cox hazard ratio relies on a proportional-hazards framework for its usual interpretation as a common relative hazard over time. The ClinicalTrials.gov record does not provide a separate assessment of the proportional-hazards assumption, so the single HR should not be interpreted as proof that the relative hazard was constant throughout follow-up.
Unstratified Overall Survival Analysis
Unstratified hazard ratio for overall survival
95% CI: 0.696–1.182 · P = 0.4709
Unstratified log-rank analysis; hazard ratio estimated by Cox regression.
| Primary OS analysis | Reported result |
|---|---|
| Population | ITT population |
| Comparison | Docetaxel Monotherapy vs Atezolizumab + Cabozantinib |
| Method | Unstratified log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.907 |
| 95% CI | 0.696–1.182 |
| P-value | 0.4709 |
| Hypothesis type | Superiority |
The unstratified hazard ratio of 0.907 corresponds to an estimated instantaneous hazard approximately 9.3% lower for the first-named group relative to the comparator, using the same simple interpretation of 1 − HR.
This does not mean that the probability of death was reduced by 9.3%, nor does it translate directly into a difference in median survival or a fixed difference in survival probability at a particular time point.
The 95% CI of 0.696–1.182 includes 1 and spans a range from a lower estimated hazard to a higher estimated hazard. The interval therefore communicates more information than the point estimate alone: considerable uncertainty remains about the magnitude and even the direction of the relative hazard compatible with the data.
The P-value of 0.4709 is not an effect-size measure. A larger or smaller P-value does not tell us how clinically large the hazard ratio is. Here, the appropriate interpretation is to consider the P-value together with the estimated HR, its confidence interval, the ITT population, censoring, and the prespecified superiority framework.
8. Secondary Result: Progression-Free Survival
Progression-free survival as determined by investigator was a secondary time-to-event endpoint evaluated up to approximately 24 months. The registry reports both stratified and unstratified analyses using the ITT population.
Stratified PFS Analysis
Hazard ratio for progression-free survival
95% CI: 0.585–0.923 · P = 0.0079
Stratified log-rank analysis; hazard ratio estimated by Cox regression.
| Measure | Result |
|---|---|
| Endpoint | Progression-Free Survival (PFS) as Determined by Investigator |
| Time frame | Up to approximately 24 months |
| Population | ITT |
| Method | Stratified log-rank test |
| Hazard ratio | 0.735 |
| 95% CI | 0.585–0.923 |
| P-value | 0.0079 |
An HR of 0.735 corresponds to approximately a 26.5% lower estimated instantaneous hazard for progression or death for the first-named group relative to the comparator, because 1 − 0.735 = 0.265.
The HR does not mean that 26.5% of participants avoided progression, nor does it specify how many additional months an individual participant remained progression-free. Those questions require absolute survival estimates or median time-to-event estimates, which are not included in the ClinicalTrials.gov record.
The 95% CI of 0.585–0.923 lies below 1. It provides a range of plausible relative hazard estimates under the statistical model and sampling framework, while still showing uncertainty about the exact magnitude of the effect.
The P-value of 0.0079 describes the evidence against the relevant superiority null hypothesis under the reported analysis. It does not say that there is a 0.79% probability that the null hypothesis is true, and it does not measure the size of the effect.
Because this is a Cox-model result, interpretation also depends on the proportional-hazards framework. In addition, PFS is subject to right censoring and depends on how progression and censoring are defined and observed. The ClinicalTrials.gov record does not provide a separate diagnostic of proportional hazards or a detailed missing-data assessment.
Unstratified PFS Analysis
Unstratified hazard ratio for progression-free survival
95% CI: 0.583–0.915 · P = 0.0061
Unstratified log-rank analysis; hazard ratio estimated by Cox regression.
The unstratified PFS estimate of 0.731 corresponds to approximately a 26.9% lower estimated instantaneous hazard for progression or death for the first-named group relative to the comparator.
The confidence interval, 0.583–0.915, is entirely below 1 and is relatively narrow compared with the range of possible values represented by the point estimate alone. It therefore conveys both direction and uncertainty more completely than the HR by itself.
The P-value of 0.0061 provides a measure of statistical evidence under the specified test; it should not be read as the probability that the observed treatment effect is real or as a measure of how important the effect is clinically.
9. Secondary Result: Confirmed Objective Response Rate
Confirmed objective response rate as determined by investigator was a binary endpoint assessed up to approximately 24 months. The registry reports stratified and unstratified comparisons using both difference in response rates and odds-ratio measures.
Stratified Response-Rate Difference
Difference in response rates
95% CI: −8.85 to 5.84 · P = 0.6846
Cochran-Mantel-Haenszel analysis with Wald confidence interval.
| Measure | Result |
|---|---|
| Endpoint | Confirmed Objective Response Rate (ORR) as Determined by Investigator |
| Time frame | Up to approximately 24 months |
| Population | ITT |
| Method | Cochran-Mantel-Haenszel test |
| Effect measure | Difference in Response Rates |
| Estimate | −1.51 |
| 95% CI | −8.85 to 5.84 |
| P-value | 0.6846 |
The reported difference in response rates is −1.51, with the direction defined by the registry's comparison of docetaxel monotherapy versus atezolizumab plus cabozantinib. The estimate is therefore a difference in percentage points, not a relative percentage change.
The 95% CI extends from −8.85 to 5.84. Because the interval includes 0, the data are compatible with either a lower or higher response rate for the first-named group under this analysis. The interval also illustrates why the point estimate alone should not be treated as a precise estimate of the underlying response-rate difference.
The P-value of 0.6846 measures statistical evidence under the reported test; it does not quantify the size of the difference. A response-rate comparison also answers a different question from a time-to-event analysis: ORR concerns whether a participant achieved a confirmed response, whereas PFS incorporates the timing of progression or death.
Stratified Odds Ratio for Response
Odds ratio for confirmed response
95% CI: 0.47–1.63
Odds ratio estimated by logistic regression; 95% CI computed using the Wald method.
An odds ratio of 0.88 means that the estimated odds of confirmed response for the first-named group were 0.88 times the odds for the comparator in the reported logistic-regression analysis. This is not the same as saying that the probability of response was 12% lower. Odds and probabilities are related but are not numerically interchangeable.
Unstratified Response Analyses
| Analysis | Estimate | 95% CI | P-value |
|---|---|---|---|
| Chi-squared test; difference in response rates | −1.51 | −8.85 to 5.84 | 0.7216 |
| Logistic regression; odds ratio | 0.87 | 0.47–1.62 | Not reported in the ClinicalTrials.gov record |
The unstratified logistic-regression estimate of 0.87 has a 95% CI of 0.47–1.62. The ClinicalTrials.gov record does not provide a P-value for this particular logistic-regression result, so no additional formal significance statement is made here.
10. Patient-Reported Physical Functioning: Time to Confirmed Deterioration
Time to confirmed deterioration in patient-reported physical functioning was analyzed as a secondary time-to-event endpoint up to approximately 24 months.
Stratified TTCD in Physical Functioning
95% CI: 0.59–1.16 · P = 0.2700
Stratified log-rank analysis; hazard ratio estimated by Cox regression.
The stratified estimate of 0.82 corresponds to an estimated instantaneous deterioration hazard approximately 18% lower for the first-named group relative to the comparator. The confidence interval of 0.59–1.16 includes 1, so the exact magnitude and direction of the underlying relative hazard remain uncertain within the reported interval.
The P-value of 0.2700 is a hypothesis-test quantity rather than an effect-size measure. It does not establish that the groups have identical patient-reported functioning, nor does it quantify the clinical importance of any difference.
Unstratified TTCD in Physical Functioning
Unstratified hazard ratio
95% CI: 0.60–1.17 · P = 0.3031
Unstratified log-rank analysis; hazard ratio estimated by Cox regression.
The stratified and unstratified estimates, 0.82 and 0.84, are similar in direction and magnitude, but neither confidence interval excludes 1. These results should therefore be described as estimates with uncertainty rather than as evidence of a precisely quantified reduction in deterioration hazard.
Because deterioration is defined through patient-reported functioning, the endpoint also differs conceptually from OS and PFS. A time-to-deterioration endpoint asks when a prespecified deterioration occurs and requires appropriate handling of participants who have not deteriorated by the analysis time.
11. Patient-Reported Global Health Status: Time to Confirmed Deterioration
Time to confirmed deterioration in patient-reported global health status was another secondary time-to-event endpoint evaluated up to approximately 24 months.
Stratified TTCD in Global Health Status
95% CI: 0.86–1.79 · P = 0.2408
Stratified log-rank analysis; hazard ratio estimated by Cox regression.
An HR of 1.24 corresponds to an estimated instantaneous deterioration hazard approximately 24% higher for the first-named group relative to the comparator under the reported model. That does not mean that 24% more participants deteriorated, because a hazard ratio is a relative time-to-event measure rather than an absolute event-rate difference.
The 95% CI of 0.86–1.79 includes 1 and spans both a possible lower and higher hazard. The P-value of 0.2408 does not measure the clinical importance of the estimate.
Unstratified TTCD in Global Health Status
Unstratified hazard ratio
95% CI: 0.88–1.81 · P = 0.1992
Unstratified log-rank analysis; hazard ratio estimated by Cox regression.
The stratified and unstratified estimates of 1.24 and 1.26 are close to one another. Both confidence intervals include 1. This is useful descriptively because the two analysis approaches produce similar point estimates, but it does not turn the secondary endpoint into a confirmatory result.
The distinction between a hazard ratio above 1 and an odds ratio above 1 is also important. Here, the HR describes the relative instantaneous rate of confirmed deterioration over time; it is not an odds ratio for experiencing deterioration.
12. Landmark PFS Rates
The registry reports PFS rates assessed by investigator at 6 months and 1 year. These are binary, time-specific summaries derived from the time-to-event framework. The reported comparisons used a z-test, with confidence intervals based on the normal approximation and standard errors computed using the Greenwood method.
| Time point | Difference in event-free rate | 95% CI | P-value |
|---|---|---|---|
| 6 months | 15.85 | 6.12–25.59 | 0.0014 |
| 1 year | 6.32 | −0.56 to 13.21 | 0.0719 |
At 6 Months
Difference in PFS rate
95% CI: 6.12–25.59 · P = 0.0014
The reported difference in event-free rate at 6 months was 15.85 percentage points under the registry's comparison. The confidence interval ranges from 6.12 to 25.59 percentage points. Unlike a hazard ratio, this measure is directly expressed as a difference between time-specific rates.
This result illustrates why absolute landmark estimates complement hazard ratios. A hazard ratio summarizes relative event hazards over follow-up, while a 6-month event-free-rate difference describes the estimated separation at one specific time point.
The confidence interval quantifies uncertainty around that time-specific difference. The P-value of 0.0014 addresses the corresponding hypothesis test; it does not mean that the estimated difference is 0.14% likely to have arisen by chance, nor does it measure clinical magnitude.
At 1 Year
Difference in PFS rate
95% CI: −0.56 to 13.21 · P = 0.0719
At 1 year, the reported difference in event-free rate was 6.32 percentage points, with a 95% CI from −0.56 to 13.21 percentage points. The interval includes 0, in contrast with the 6-month confidence interval.
The difference between the two time points is a useful statistical teaching point: treatment-group separation can vary over time, and a statistically supported difference at one landmark does not automatically imply the same degree of separation at another landmark.
13. Landmark Overall Survival Rates
OS rates were reported at 1 and 2 years. The registry-reported statistical analysis reports a z-test comparison of event-free rates, with the 95% confidence interval for the difference estimated using the normal approximation and Greenwood standard errors.
Reported 1-year OS-rate difference
95% CI: −11.63 to 9.92 · P = 0.8767
| Endpoint | Time frame | Difference | 95% CI | P-value |
|---|---|---|---|---|
| OS Rates | 1 and 2 years | −0.85 | −11.63 to 9.92 | 0.8767 |
The reported difference of −0.85 percentage points is accompanied by a broad 95% CI extending from −11.63 to 9.92 percentage points. The interval includes 0, and the P-value is 0.8767.
The landmark OS result should not be confused with the primary OS hazard-ratio analysis. A landmark survival-rate difference is a time-specific absolute comparison, whereas the hazard ratio summarizes relative event hazards within a Cox-model framework.
The wide interval indicates substantial uncertainty around the reported time-specific difference. It is therefore inappropriate to infer a precise equivalence of survival probabilities from the point estimate alone.
14. Secondary Results Summary
| Endpoint | Analysis | Effect | 95% CI | P-value |
|---|---|---|---|---|
| PFS | Stratified | HR 0.735 | 0.585–0.923 | 0.0079 |
| PFS | Unstratified | HR 0.731 | 0.583–0.915 | 0.0061 |
| Confirmed ORR | Stratified CMH | Difference −1.51 | −8.85 to 5.84 | 0.6846 |
| Confirmed ORR | Stratified logistic regression | OR 0.88 | 0.47–1.63 | Not reported |
| Confirmed ORR | Unstratified chi-squared | Difference −1.51 | −8.85 to 5.84 | 0.7216 |
| Confirmed ORR | Unstratified logistic regression | OR 0.87 | 0.47–1.62 | Not reported |
| TTCD Physical Functioning | Stratified | HR 0.82 | 0.59–1.16 | 0.2700 |
| TTCD Physical Functioning | Unstratified | HR 0.84 | 0.60–1.17 | 0.3031 |
| TTCD Global Health Status | Stratified | HR 1.24 | 0.86–1.79 | 0.2408 |
| TTCD Global Health Status | Unstratified | HR 1.26 | 0.88–1.81 | 0.1992 |
| PFS rate | 6 months | Difference 15.85 | 6.12–25.59 | 0.0014 |
| PFS rate | 1 year | Difference 6.32 | −0.56 to 13.21 | 0.0719 |
| OS rate | 1 year | Difference −0.85 | −11.63 to 9.92 | 0.8767 |
15. Safety Results
The ClinicalTrials.gov record includes serious adverse events by treatment arm. The reported affected/at-risk counts were 58/167 for docetaxel monotherapy and 76/185 for atezolizumab plus cabozantinib.
| Treatment arm | Participants with serious adverse events | Participants at risk |
|---|---|---|
| Docetaxel Monotherapy | 58 | 167 |
| Atezolizumab + Cabozantinib | 76 | 185 |
The affected/at-risk figures should be kept distinct from efficacy event counts. A serious adverse event is a safety outcome, whereas OS and PFS are efficacy time-to-event endpoints. The ClinicalTrials.gov record does not provide a formal between-arm statistical test for serious adverse events, so no comparative P-value or relative-effect estimate is assigned to these counts here.
16. Statistical Methods Explained
Why was the intention-to-treat population used?
The ITT population includes all randomized participants, whether or not they received the assigned treatment. This preserves the treatment comparison created by randomization. If participants were instead moved between analysis groups after randomization according to treatment received, prognostic differences could become entangled with treatment exposure.
Why use both stratified and unstratified analyses?
The stratified analysis accounts for the registry-specified stratification factors, while the unstratified analysis does not. Comparing the two can show how much the estimated treatment effect changes when those factors are incorporated. They should not be interpreted as two separate randomized experiments.
What does a hazard ratio of 0.735 mean?
For the reported PFS analysis, an HR of 0.735 means that the estimated instantaneous rate of progression or death for the first-named group was 0.735 times that of the comparator under the fitted Cox model. Equivalently, using the simple complement interpretation, the estimate corresponds to approximately a 26.5% lower hazard. It does not mean a 26.5% increase in median PFS or that 26.5% of participants were protected from progression.
Why is the confidence interval more informative than the point estimate alone?
A point estimate is only one estimate of an underlying treatment effect. The 95% confidence interval shows the uncertainty associated with that estimate under the model and sampling framework. For example, the stratified OS HR of 0.884 has a 95% CI of 0.676–1.156, demonstrating that the plausible range is substantially wider than the single point estimate.
Why is an odds ratio not the same as a risk ratio?
An odds ratio compares odds rather than probabilities. An OR of 0.88 means the estimated odds of response were 0.88 times the comparator's odds in the reported model. It does not mean that the response probability was exactly 12% lower. The difference becomes especially important when response probabilities are not small.
Why use a log-rank test for OS and PFS?
OS and PFS are time-to-event outcomes with potentially different follow-up times and right censoring. The log-rank test compares the event experience of the randomized groups over follow-up rather than reducing every participant to a simple event/no-event classification at one arbitrary time point.
Why report both hazard ratios and landmark rates?
A hazard ratio summarizes a relative time-to-event comparison, whereas a landmark rate provides an absolute estimate at a particular time. The registry-reported PFS analyses illustrate this distinction: the PFS HR is 0.735 in the stratified analysis, while the reported difference in PFS rate is 15.85 at 6 months and 6.32 at 1 year. These quantities describe different aspects of the same survival experience.
17. Multiplicity and Interpretation of Multiple Analyses
The ClinicalTrials.gov record contains 15 statistical analyses, including two analyses of the registered primary endpoint and multiple secondary endpoint analyses. The primary hypothesis type is identified as superiority.
Multiple reported analyses create an important interpretive distinction. A P-value from an individual secondary analysis should not automatically be interpreted as though it were the sole prespecified hypothesis test in the trial. Formal multiplicity control depends on the prespecified testing hierarchy, alpha allocation, and statistical analysis plan. Those details are not provided in the ClinicalTrials.gov record.
| Analysis family | Examples in the ClinicalTrials.gov record | Interpretive issue |
|---|---|---|
| Primary endpoint | Stratified and unstratified OS | One registered primary endpoint with two reported analysis approaches |
| Time-to-event secondary endpoints | PFS; TTCD in PF; TTCD in GHS | Multiple event-time comparisons |
| Binary response endpoint | ORR | Multiple methods and effect measures |
| Landmark rates | PFS at 6 months and 1 year; OS rates at 1 and 2 years | Multiple time-specific comparisons |
Accordingly, the numerical P-values should be read in the context of their individual analyses rather than assembled into a single informal list of "significant" and "nonsignificant" findings. The ClinicalTrials.gov record does not provide enough information to reconstruct a full multiplicity-adjustment hierarchy.
18. Censoring and Time-to-Event Interpretation
Time-to-event analysis depends on appropriate handling of participants whose event has not occurred by the analysis cutoff. For OS, the registry explicitly states that participants alive at the time of analysis were censored at the date when they were last known to be alive as documented by the investigator.
For the patient-reported deterioration endpoints, participants without confirmed deterioration were censored according to their last known non-deteriorated status. This preserves information up to the participant's last evaluable time rather than treating an unobserved future deterioration as if it had already occurred.
Censoring does not mean that a participant's outcome is known after the censoring time. It means that the analysis uses the information available through the last appropriate observation.
The Cox hazard ratio and Kaplan-Meier estimates therefore depend not only on the observed events but also on the timing and handling of censoring. A time-to-event estimate cannot be reconstructed accurately from an event count alone.
19. Non-Inferiority, Crossover, and Factorial Design
Superiority
The statistical analyses posted on ClinicalTrials.gov identify the hypothesis type as superiority. No non-inferiority margin is provided because the reported analyses are framed as superiority comparisons.
Crossover
The registry-reported trial design is described as a parallel-group randomized trial. The ClinicalTrials.gov record do not report a treatment crossover analysis.
Factorial design
The design model is PARALLEL rather than factorial. The ClinicalTrials.gov record therefore do not support a factorial treatment-interaction analysis.
Bayesian methods
No Bayesian method is listed among the normalized statistical methods in the ClinicalTrials.gov record.
These distinctions matter because the statistical logic of a superiority trial differs from that of a non-inferiority trial. In a non-inferiority analysis, the prespecified margin is central to the conclusion. Here, the ClinicalTrials.gov record identifies superiority as the hypothesis type and do not provide a non-inferiority margin.
20. Wald Confidence Intervals and z-Tests
The ORR analyses used Wald confidence intervals, and the landmark PFS and OS rate analyses used z-tests with confidence intervals based on normal approximation. These methods illustrate a common distinction between effect estimation and hypothesis testing.
The confidence interval combines an estimated effect with its standard error. The interval width therefore reflects both the variability of the estimate and the amount of information available for the analysis.
For the ORR difference, the reported estimate is −1.51 with a 95% CI of −8.85 to 5.84. For the 6-month PFS-rate difference, the estimate is 15.85 with a 95% CI of 6.12–25.59. These intervals are expressed on the natural scale of the reported difference rather than on the hazard-ratio scale.
The choice of effect scale matters. A confidence interval around an odds ratio is interpreted relative to the null value 1, while a confidence interval around a difference is interpreted relative to the null value 0.
21. Missing Data and Imputation
The ClinicalTrials.gov record provides censoring descriptions for the time-to-event endpoints, but they do not provide a detailed missing-data or imputation strategy for the full set of outcome measures.
For time-to-event endpoints, censoring is part of the primary analysis framework and should not automatically be equated with ordinary missing-data imputation. For patient-reported outcomes, missing questionnaires can raise additional questions about whether observations are missing independently of underlying health status.
22. What the Primary Hazard Ratio Does — and Does Not — Mean
The stratified OS HR of 0.884 means that the fitted model estimated the instantaneous rate of death for the first-named group at 0.884 times the rate for the comparator, under the reported Cox analysis.
The HR does not state how many additional participants were alive at a particular time, how many deaths were prevented, or how much longer an individual participant lived. Those are absolute or individual-level quantities and require different summaries.
The 95% CI of 0.676–1.156 indicates uncertainty around the estimated HR. Because it crosses the null value of 1, the reported interval includes both lower and higher hazards for the first-named group.
The P-value of 0.3668 measures statistical evidence under the reported superiority test. It is not the probability that the null hypothesis is true and is not a measure of the size or practical importance of the estimated HR.
The Cox interpretation depends on the time-to-event model and its assumptions. The ClinicalTrials.gov record does not report a separate proportional-hazards diagnostic, so the HR should be treated as a model-based summary rather than as proof of a constant relative hazard at every point in time.
23. Why This Trial Matters Statistically
CONTACT-01 is a useful teaching example because a single randomized trial contains several distinct statistical estimands and analysis frameworks. The registered primary endpoint is time-to-event, but the posted secondary analyses move between survival analysis, categorical-data methods, logistic regression, and landmark rate comparisons.
| Concept | How it appears in CONTACT-01 |
|---|---|
| Randomization | Parallel-group randomized phase 3 design with 366 enrolled participants |
| ITT analysis | Efficacy analyses include all randomized participants whether or not assigned treatment was received |
| Kaplan-Meier estimation | Used for OS estimation and median estimation |
| Log-rank test | Primary OS and secondary time-to-event comparisons |
| Hazard ratio | Primary OS and secondary PFS and patient-reported time-to-deterioration effects |
| Cox regression | Used to estimate hazard ratios |
| Stratified analysis | Histology and prior NSCLC treatment regimens appear as stratification factors |
| Cochran-Mantel-Haenszel test | Stratified analysis of confirmed objective response rate |
| Logistic regression | Odds-ratio estimation for confirmed objective response rate |
| Wald method | Confidence intervals for response-rate differences and odds ratios |
| z-test | Landmark PFS-rate and OS-rate comparisons |
| Greenwood method | Standard errors for reported time-specific PFS and OS rate differences |
| Multiple analyses | 15 statistical analyses are posted across primary and secondary endpoints |
The most important statistical lesson is that these methods are not interchangeable. A hazard ratio, odds ratio, response-rate difference, and landmark survival-rate difference each answer a different question. Correct interpretation requires preserving the effect scale, endpoint definition, analysis population, and time frame.
24. Important Limitations and Interpretation Issues
- Primary endpoint scope: the ClinicalTrials.gov record identifies OS as the single registered primary endpoint. Other reported analyses are secondary rather than additional registered primary endpoints.
- Stratified versus unstratified analyses: the registry reports both approaches for several endpoints. They should be understood as alternative statistical analyses of the same randomized comparison, not independent trials.
- Hazard-ratio assumptions: Cox-model HRs are model-based summaries. The ClinicalTrials.gov record does not report a separate proportional-hazards assessment.
- Censoring: time-to-event estimates depend on the timing and handling of censoring. The ClinicalTrials.gov record provides specific censoring definitions for OS and patient-reported deterioration but do not provide the full underlying event/censoring dataset.
- Multiplicity: multiple primary and secondary analyses are reported, but the ClinicalTrials.gov record does not provide a complete multiplicity-control hierarchy. Individual P-values should therefore be interpreted in their stated analytical context.
- Secondary endpoints: PFS, ORR, patient-reported deterioration, and landmark rates provide complementary information but do not answer the same question as the primary OS endpoint.
- Safety comparison: serious-adverse-event counts are reported by arm, but the ClinicalTrials.gov record does not provide a formal comparative statistical analysis for those safety counts.
- Missing data: a complete imputation and sensitivity-analysis strategy is not provided in the ClinicalTrials.gov record.
- Results granularity: the ClinicalTrials.gov record does not include median OS, median PFS, subgroup forest plots, baseline characteristic tables, or detailed event counts. Those quantities are therefore not presented here.
25. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The registered OS endpoint produced a stratified HR of 0.884 (95% CI 0.676–1.156; P = 0.3668). The secondary PFS analysis produced a stratified HR of 0.735 (95% CI 0.585–0.923; P = 0.0079). Other endpoints produced estimates on different scales with their own uncertainty.
Clinical interpretation
Clinical interpretation requires considering what each endpoint measures, the magnitude and precision of its effect estimate, the time frame, safety information, and the prespecified role of the endpoint. The ClinicalTrials.gov record alone do not justify reducing the entire trial to one numerical measure.
This distinction is especially important in a trial containing both OS and PFS results. A PFS hazard ratio and an OS hazard ratio do not represent interchangeable outcomes. OS incorporates survival through the entire follow-up period, while PFS captures the earlier occurrence of progression or death.
26. Overall Statistical Synthesis
CONTACT-01 provides a compact example of how clinical-trial evidence can change depending on the statistical lens used. The primary OS analysis estimates a hazard ratio of 0.884 in the stratified analysis, with a 95% CI of 0.676–1.156 and P = 0.3668. The corresponding unstratified estimate is 0.907, with a 95% CI of 0.696–1.182 and P = 0.4709.
For PFS, the stratified HR is 0.735 with a 95% CI of 0.585–0.923 and P = 0.0079, while the unstratified HR is 0.731 with a 95% CI of 0.583–0.915 and P = 0.0061. The landmark PFS analysis adds a time-specific perspective: the reported difference in event-free rate is 15.85 at 6 months and 6.32 at 1 year.
Other secondary endpoints illustrate why a single P-value cannot summarize the entire statistical record. Confirmed ORR has a reported stratified response-rate difference of −1.51 with a 95% CI of −8.85 to 5.84, while logistic regression gives an odds ratio of 0.88 with a 95% CI of 0.47–1.63. Patient-reported time-to-deterioration analyses similarly use hazard ratios but address different clinical constructs.
The statistically disciplined interpretation is therefore to retain the endpoint-specific effect measure, its confidence interval, its P-value, the analysis population, and the method used. The trial's results cannot be faithfully represented by treating hazard ratios, odds ratios, percentage-point differences, and landmark rates as though they were the same kind of quantity.
27. Related Tutorials
Learn more about the methods used in this trial:
28. Related Statistical Calculators
29. Sources
- ClinicalTrials.gov: CONTACT-01, NCT04471428. Official trial registry record and source for the trial data summarized on this page.
- PubMed: PubMed record associated with CONTACT-01.
- PubMed: PubMed record associated with CONTACT-01.
Continue through the Clinical Biostats statistical library
Explore the underlying statistical methods through tutorials and practical calculators for survival analysis, categorical data, regression, confidence intervals, and clinical-trial analysis.
30. Record Summary
CONTACT-01 is a randomized phase 3 parallel-group trial with 366 enrolled participants and a single registered primary endpoint, overall survival. The primary OS endpoint was analyzed in the ITT population using a log-rank test, with the hazard ratio estimated by Cox regression. The stratified OS analysis reported an HR of 0.884 (95% CI 0.676–1.156; P = 0.3668), while the unstratified analysis reported an HR of 0.907 (95% CI 0.696–1.182; P = 0.4709).
The secondary analyses demonstrate the breadth of modern clinical-trial statistics. PFS was evaluated with stratified and unstratified log-rank/Cox analyses; confirmed ORR used Cochran-Mantel-Haenszel, chi-squared, and logistic-regression approaches; patient-reported deterioration used time-to-event methods; and landmark PFS and OS rates used z-tests with Greenwood-based standard errors. The serious-adverse-event summary reports 58/167 for docetaxel monotherapy and 76/185 for atezolizumab plus cabozantinib.
The principal educational lesson is that a clinical trial does not have one universal "result." Each endpoint has an estimand, an analysis population, an effect measure, an uncertainty interval, and a statistical test. Reading CONTACT-01 rigorously therefore means keeping those elements together rather than collapsing the trial into a single hazard ratio or P-value.