← Clinical Trials
Metastatic NSCLC Phase 3 Completed NCT04471428

CONTACT-01: Complete Statistical Analysis of Atezolizumab Plus Cabozantinib in Metastatic NSCLC

An independent statistical review of the randomized phase 3 CONTACT-01 trial comparing atezolizumab plus cabozantinib with docetaxel in patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy.

Phase 3  ·  Randomized parallel design  ·  366 participants  ·  Primary endpoint: Overall Survival
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the ClinicalTrials.gov record. The registry provides the official trial record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

CONTACT-01 was a completed, randomized phase 3 trial evaluating atezolizumab in combination with cabozantinib versus docetaxel monotherapy in patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy.

366
Enrolled
Randomized participants
2
Treatment arms
Parallel-group design
0.884
OS HR
95% CI 0.676–1.156
0.735
PFS HR
95% CI 0.585–0.923
FeatureCONTACT-01
PhasePhase 3
StatusCOMPLETED
ConditionCarcinoma, Non-Small-Cell Lung
PopulationPatients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy
DesignRandomized, parallel-group
MaskingNONE
AllocationRANDOMIZED
Primary purposeTREATMENT
Enrollment366
Primary endpointOverall Survival (OS)
Primary endpoint typeTime-to-event
Results postedYes
Statistical analyses posted15
ClinicalTrials.govNCT04471428

2. Clinical Question

The central statistical question was whether treatment with atezolizumab plus cabozantinib differed from docetaxel monotherapy with respect to overall survival in patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy.

Population

Patients with metastatic non-small cell lung cancer previously treated with an anti-PD-L1/PD-1 antibody and platinum-containing chemotherapy.

Intervention

Atezolizumab in combination with cabozantinib.

Comparator

Docetaxel monotherapy.

Primary question

Does the randomized treatment comparison produce a difference in overall survival?

3. Trial Design

01
Randomize366 participants
02
Arm AAtezolizumab + cabozantinib
03
Arm BDocetaxel monotherapy
04
Follow-upOS, PFS, response, patient-reported endpoints
05
AnalysisITT-based efficacy comparisons
ARM A · n = 185 at risk for reported serious-AE summary

Atezolizumab + Cabozantinib

  • Atezolizumab
  • Cabozantinib
  • Combination treatment arm
ARM B · n = 167 at risk for reported serious-AE summary

Docetaxel Monotherapy

  • Docetaxel
  • Monotherapy control arm

The ClinicalTrials.gov record describes a parallel, randomized, unmasked phase 3 treatment comparison. The primary efficacy analysis used the intention-to-treat population, which included all randomized participants whether or not they received the assigned treatment.

Allocation
Randomized allocation to one of two treatment arms.
Model
Parallel-group design rather than a crossover or factorial design.
Masking
NONE according to the ClinicalTrials.gov record.
Hypothesis
Superiority hypothesis for the reported statistical comparisons.

4. Endpoints

The registry lists one registered primary endpoint: overall survival. The statistical analyses posted on ClinicalTrials.gov also report secondary analyses for progression-free survival, objective response, patient-reported deterioration, PFS rates, and OS rates.

EndpointRegistry definition / time frameEndpoint type
Overall Survival (OS) OS was defined as the time from randomization to death from any cause. Participants alive at the time of the analysis were censored at the date when they were last known to be alive as documented by the investigator. Time frame: up to approximately 24 months. Time-to-event
Progression-Free Survival (PFS) as Determined by Investigator Time frame: up to approximately 24 months. Time-to-event
Confirmed Objective Response Rate (ORR) as Determined by Investigator Time frame: up to approximately 24 months. Binary
Time to Confirmed Deterioration (TTCD) in Patient-reported Physical Functioning (PF) Time frame: up to approximately 24 months. Participants without a confirmed deterioration at the time of analysis were censored according to the registry analysis description. Time-to-event
TTCD in Patient-reported Global Health Status (GHS) Time frame: up to approximately 24 months. Participants without a confirmed deterioration at the time of analysis were censored at the last time they were known to have not deteriorated. Time-to-event
PFS Rates Assessed by Investigator 6 months and 1 year. Binary
OS Rates 1 and 2 years. Binary
Endpoint hierarchy: The ClinicalTrials.gov record identifies OS as the sole registered primary endpoint. PFS, response, patient-reported deterioration, PFS rates, and OS rates are represented in the posted secondary analyses.

5. Analysis Populations and Stratification

The reported efficacy analyses use the intention-to-treat (ITT) population. The registry defines this population as all randomized participants, whether or not the participant received the assigned treatment. This is important because the treatment comparison remains anchored to randomization rather than being redefined after treatment exposure or discontinuation.

Analysis populationDefinition in the ClinicalTrials.gov recordRole
ITT All randomized participants, whether or not the participant received the assigned treatment. Primary efficacy and reported secondary efficacy analyses
Safety analysis at risk The registry-reported serious-adverse-event summary reports 58/167 for docetaxel monotherapy and 76/185 for atezolizumab + cabozantinib. Safety description

Several analyses were explicitly described as stratified. The ClinicalTrials.gov record identifies histology and prior NSCLC treatment regimens as stratification factors. The same factors appear in the descriptions of the stratified hazard-ratio and response analyses.

Why stratification matters

Stratified analysis allows the comparison to account for prespecified factors used to organize the randomized comparison. It can improve alignment between the analysis and the trial's randomization structure.

Stratified vs unstratified

The registry reports both stratified and unstratified analyses for several endpoints. These are different statistical analyses of the same randomized comparison, not separate treatment effects.

6. Statistical Methodology

Kaplan-Meier estimation

The registered OS endpoint is a time-to-event outcome. The registry states that the Kaplan-Meier method was used to estimate the median, with the 95% confidence interval for the median computed using the method of Brookmeyer and Crowley.

The Kaplan-Meier framework is appropriate when follow-up times differ among participants and some participants remain alive at the time of analysis. Rather than treating those participants as if they had experienced the event, their observations are censored at the last date they were known to be alive.

Kaplan-Meier survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni is the number at risk immediately before that time.

Log-rank test

The primary OS comparison was evaluated using a log-rank test. The same method was reported for PFS and the two patient-reported time-to-deterioration endpoints. The log-rank test compares the observed and expected event patterns between randomized treatment groups over follow-up.

Cox regression and hazard ratios

The registry states that the hazard ratio was estimated using a Cox regression model. For the stratified analyses, the registry-reported analysis notes identify histology and prior NSCLC treatment regimens as stratification factors.

Hazard-ratio interpretation
HR = instantaneous event rate in treatment group ÷ instantaneous event rate in control group

An HR below 1 indicates a lower estimated instantaneous event rate in the first-named treatment group relative to the comparator under the fitted model. It is not an absolute risk difference, a probability of survival, or a statement that every patient experiences the same proportional change.

Cochran-Mantel-Haenszel analysis

Confirmed objective response rate was analyzed using the Cochran-Mantel-Haenszel test in the stratified analysis. This method provides a way to compare categorical outcomes across treatment groups while accounting for prespecified strata.

Logistic regression

Odds ratios for confirmed objective response were estimated using logistic regression. The registry reports both stratified and unstratified logistic-regression analyses.

Wald / z-test methods

The reported PFS-rate and OS-rate comparisons used a z-test. The registry describes these as Wald / z-test analyses and states that confidence intervals for differences in rates were estimated using the normal approximation, with standard errors computed using the Greenwood method for the time-to-event-derived rates.

7. Primary Result: Overall Survival

Overall survival was the sole registered primary endpoint. The ClinicalTrials.gov record contains two formal analyses: a stratified analysis and an unstratified analysis. Both use the ITT population and compare docetaxel monotherapy with atezolizumab plus cabozantinib.

Stratified Overall Survival Analysis

Hazard ratio for overall survival

0.884

95% CI: 0.676–1.156   ·   P = 0.3668

Stratified log-rank analysis; hazard ratio estimated by Cox regression.

Primary OS analysisReported result
PopulationITT population
ComparisonDocetaxel Monotherapy vs Atezolizumab + Cabozantinib
MethodStratified log-rank test
Effect measureHazard ratio
Estimate0.884
95% CI0.676–1.156
P-value0.3668
Hypothesis typeSuperiority
Stratification factorsHistology and prior NSCLC treatment regimens
Clinical Biostats interpretation

The estimated hazard ratio of 0.884 indicates that the fitted Cox model estimated a lower instantaneous rate of death for the first-named group relative to the comparator, with the estimate corresponding to approximately a 11.6% lower estimated hazard because 1 − 0.884 = 0.116.

That statement is about the estimated relative hazard; it does not mean that 11.6% fewer participants died, that an individual patient's probability of death fell by 11.6%, or that survival time increased by 11.6%.

The 95% confidence interval, 0.676–1.156, expresses uncertainty around the estimated hazard ratio. Because the interval includes 1, the reported data are compatible with both a lower and a higher hazard under the model. The width of the interval also shows that the estimate is not precise enough to be interpreted as a narrowly defined treatment effect.

The P-value of 0.3668 addresses the statistical evidence against the relevant null hypothesis under the reported testing framework. It does not measure the size or clinical importance of the treatment effect. A P-value should therefore be interpreted alongside the hazard ratio and its confidence interval.

Finally, a Cox hazard ratio relies on a proportional-hazards framework for its usual interpretation as a common relative hazard over time. The ClinicalTrials.gov record does not provide a separate assessment of the proportional-hazards assumption, so the single HR should not be interpreted as proof that the relative hazard was constant throughout follow-up.

Unstratified Overall Survival Analysis

Unstratified hazard ratio for overall survival

0.907

95% CI: 0.696–1.182   ·   P = 0.4709

Unstratified log-rank analysis; hazard ratio estimated by Cox regression.

Primary OS analysisReported result
PopulationITT population
ComparisonDocetaxel Monotherapy vs Atezolizumab + Cabozantinib
MethodUnstratified log-rank test
Effect measureHazard ratio
Estimate0.907
95% CI0.696–1.182
P-value0.4709
Hypothesis typeSuperiority
Clinical Biostats interpretation

The unstratified hazard ratio of 0.907 corresponds to an estimated instantaneous hazard approximately 9.3% lower for the first-named group relative to the comparator, using the same simple interpretation of 1 − HR.

This does not mean that the probability of death was reduced by 9.3%, nor does it translate directly into a difference in median survival or a fixed difference in survival probability at a particular time point.

The 95% CI of 0.696–1.182 includes 1 and spans a range from a lower estimated hazard to a higher estimated hazard. The interval therefore communicates more information than the point estimate alone: considerable uncertainty remains about the magnitude and even the direction of the relative hazard compatible with the data.

The P-value of 0.4709 is not an effect-size measure. A larger or smaller P-value does not tell us how clinically large the hazard ratio is. Here, the appropriate interpretation is to consider the P-value together with the estimated HR, its confidence interval, the ITT population, censoring, and the prespecified superiority framework.

Two primary analyses, one endpoint: the two reported OS analyses should not be treated as two independent clinical outcomes. They are stratified and unstratified analyses of the same primary endpoint. The stratified analysis incorporates the registry-specified stratification factors, whereas the unstratified analysis does not.

8. Secondary Result: Progression-Free Survival

Progression-free survival as determined by investigator was a secondary time-to-event endpoint evaluated up to approximately 24 months. The registry reports both stratified and unstratified analyses using the ITT population.

Stratified PFS Analysis

Hazard ratio for progression-free survival

0.735

95% CI: 0.585–0.923   ·   P = 0.0079

Stratified log-rank analysis; hazard ratio estimated by Cox regression.

MeasureResult
EndpointProgression-Free Survival (PFS) as Determined by Investigator
Time frameUp to approximately 24 months
PopulationITT
MethodStratified log-rank test
Hazard ratio0.735
95% CI0.585–0.923
P-value0.0079
Clinical Biostats interpretation

An HR of 0.735 corresponds to approximately a 26.5% lower estimated instantaneous hazard for progression or death for the first-named group relative to the comparator, because 1 − 0.735 = 0.265.

The HR does not mean that 26.5% of participants avoided progression, nor does it specify how many additional months an individual participant remained progression-free. Those questions require absolute survival estimates or median time-to-event estimates, which are not included in the ClinicalTrials.gov record.

The 95% CI of 0.585–0.923 lies below 1. It provides a range of plausible relative hazard estimates under the statistical model and sampling framework, while still showing uncertainty about the exact magnitude of the effect.

The P-value of 0.0079 describes the evidence against the relevant superiority null hypothesis under the reported analysis. It does not say that there is a 0.79% probability that the null hypothesis is true, and it does not measure the size of the effect.

Because this is a Cox-model result, interpretation also depends on the proportional-hazards framework. In addition, PFS is subject to right censoring and depends on how progression and censoring are defined and observed. The ClinicalTrials.gov record does not provide a separate diagnostic of proportional hazards or a detailed missing-data assessment.

Unstratified PFS Analysis

Unstratified hazard ratio for progression-free survival

0.731

95% CI: 0.583–0.915   ·   P = 0.0061

Unstratified log-rank analysis; hazard ratio estimated by Cox regression.

Clinical Biostats interpretation

The unstratified PFS estimate of 0.731 corresponds to approximately a 26.9% lower estimated instantaneous hazard for progression or death for the first-named group relative to the comparator.

The confidence interval, 0.583–0.915, is entirely below 1 and is relatively narrow compared with the range of possible values represented by the point estimate alone. It therefore conveys both direction and uncertainty more completely than the HR by itself.

The P-value of 0.0061 provides a measure of statistical evidence under the specified test; it should not be read as the probability that the observed treatment effect is real or as a measure of how important the effect is clinically.

9. Secondary Result: Confirmed Objective Response Rate

Confirmed objective response rate as determined by investigator was a binary endpoint assessed up to approximately 24 months. The registry reports stratified and unstratified comparisons using both difference in response rates and odds-ratio measures.

Stratified Response-Rate Difference

Difference in response rates

−1.51

95% CI: −8.85 to 5.84   ·   P = 0.6846

Cochran-Mantel-Haenszel analysis with Wald confidence interval.

MeasureResult
EndpointConfirmed Objective Response Rate (ORR) as Determined by Investigator
Time frameUp to approximately 24 months
PopulationITT
MethodCochran-Mantel-Haenszel test
Effect measureDifference in Response Rates
Estimate−1.51
95% CI−8.85 to 5.84
P-value0.6846
Clinical Biostats interpretation

The reported difference in response rates is −1.51, with the direction defined by the registry's comparison of docetaxel monotherapy versus atezolizumab plus cabozantinib. The estimate is therefore a difference in percentage points, not a relative percentage change.

The 95% CI extends from −8.85 to 5.84. Because the interval includes 0, the data are compatible with either a lower or higher response rate for the first-named group under this analysis. The interval also illustrates why the point estimate alone should not be treated as a precise estimate of the underlying response-rate difference.

The P-value of 0.6846 measures statistical evidence under the reported test; it does not quantify the size of the difference. A response-rate comparison also answers a different question from a time-to-event analysis: ORR concerns whether a participant achieved a confirmed response, whereas PFS incorporates the timing of progression or death.

Stratified Odds Ratio for Response

Odds ratio for confirmed response

0.88

95% CI: 0.47–1.63

Odds ratio estimated by logistic regression; 95% CI computed using the Wald method.

An odds ratio of 0.88 means that the estimated odds of confirmed response for the first-named group were 0.88 times the odds for the comparator in the reported logistic-regression analysis. This is not the same as saying that the probability of response was 12% lower. Odds and probabilities are related but are not numerically interchangeable.

Unstratified Response Analyses

AnalysisEstimate95% CIP-value
Chi-squared test; difference in response rates−1.51−8.85 to 5.840.7216
Logistic regression; odds ratio0.870.47–1.62Not reported in the ClinicalTrials.gov record

The unstratified logistic-regression estimate of 0.87 has a 95% CI of 0.47–1.62. The ClinicalTrials.gov record does not provide a P-value for this particular logistic-regression result, so no additional formal significance statement is made here.

10. Patient-Reported Physical Functioning: Time to Confirmed Deterioration

Time to confirmed deterioration in patient-reported physical functioning was analyzed as a secondary time-to-event endpoint up to approximately 24 months.

Stratified TTCD in Physical Functioning

HR 0.82

95% CI: 0.59–1.16   ·   P = 0.2700

Stratified log-rank analysis; hazard ratio estimated by Cox regression.

The stratified estimate of 0.82 corresponds to an estimated instantaneous deterioration hazard approximately 18% lower for the first-named group relative to the comparator. The confidence interval of 0.59–1.16 includes 1, so the exact magnitude and direction of the underlying relative hazard remain uncertain within the reported interval.

The P-value of 0.2700 is a hypothesis-test quantity rather than an effect-size measure. It does not establish that the groups have identical patient-reported functioning, nor does it quantify the clinical importance of any difference.

Unstratified TTCD in Physical Functioning

Unstratified hazard ratio

0.84

95% CI: 0.60–1.17   ·   P = 0.3031

Unstratified log-rank analysis; hazard ratio estimated by Cox regression.

Clinical Biostats interpretation

The stratified and unstratified estimates, 0.82 and 0.84, are similar in direction and magnitude, but neither confidence interval excludes 1. These results should therefore be described as estimates with uncertainty rather than as evidence of a precisely quantified reduction in deterioration hazard.

Because deterioration is defined through patient-reported functioning, the endpoint also differs conceptually from OS and PFS. A time-to-deterioration endpoint asks when a prespecified deterioration occurs and requires appropriate handling of participants who have not deteriorated by the analysis time.

11. Patient-Reported Global Health Status: Time to Confirmed Deterioration

Time to confirmed deterioration in patient-reported global health status was another secondary time-to-event endpoint evaluated up to approximately 24 months.

Stratified TTCD in Global Health Status

HR 1.24

95% CI: 0.86–1.79   ·   P = 0.2408

Stratified log-rank analysis; hazard ratio estimated by Cox regression.

An HR of 1.24 corresponds to an estimated instantaneous deterioration hazard approximately 24% higher for the first-named group relative to the comparator under the reported model. That does not mean that 24% more participants deteriorated, because a hazard ratio is a relative time-to-event measure rather than an absolute event-rate difference.

The 95% CI of 0.86–1.79 includes 1 and spans both a possible lower and higher hazard. The P-value of 0.2408 does not measure the clinical importance of the estimate.

Unstratified TTCD in Global Health Status

Unstratified hazard ratio

1.26

95% CI: 0.88–1.81   ·   P = 0.1992

Unstratified log-rank analysis; hazard ratio estimated by Cox regression.

Clinical Biostats interpretation

The stratified and unstratified estimates of 1.24 and 1.26 are close to one another. Both confidence intervals include 1. This is useful descriptively because the two analysis approaches produce similar point estimates, but it does not turn the secondary endpoint into a confirmatory result.

The distinction between a hazard ratio above 1 and an odds ratio above 1 is also important. Here, the HR describes the relative instantaneous rate of confirmed deterioration over time; it is not an odds ratio for experiencing deterioration.

12. Landmark PFS Rates

The registry reports PFS rates assessed by investigator at 6 months and 1 year. These are binary, time-specific summaries derived from the time-to-event framework. The reported comparisons used a z-test, with confidence intervals based on the normal approximation and standard errors computed using the Greenwood method.

Time pointDifference in event-free rate95% CIP-value
6 months15.856.12–25.590.0014
1 year6.32−0.56 to 13.210.0719

At 6 Months

Difference in PFS rate

15.85

95% CI: 6.12–25.59   ·   P = 0.0014

The reported difference in event-free rate at 6 months was 15.85 percentage points under the registry's comparison. The confidence interval ranges from 6.12 to 25.59 percentage points. Unlike a hazard ratio, this measure is directly expressed as a difference between time-specific rates.

Clinical Biostats interpretation

This result illustrates why absolute landmark estimates complement hazard ratios. A hazard ratio summarizes relative event hazards over follow-up, while a 6-month event-free-rate difference describes the estimated separation at one specific time point.

The confidence interval quantifies uncertainty around that time-specific difference. The P-value of 0.0014 addresses the corresponding hypothesis test; it does not mean that the estimated difference is 0.14% likely to have arisen by chance, nor does it measure clinical magnitude.

At 1 Year

Difference in PFS rate

6.32

95% CI: −0.56 to 13.21   ·   P = 0.0719

At 1 year, the reported difference in event-free rate was 6.32 percentage points, with a 95% CI from −0.56 to 13.21 percentage points. The interval includes 0, in contrast with the 6-month confidence interval.

The difference between the two time points is a useful statistical teaching point: treatment-group separation can vary over time, and a statistically supported difference at one landmark does not automatically imply the same degree of separation at another landmark.

13. Landmark Overall Survival Rates

OS rates were reported at 1 and 2 years. The registry-reported statistical analysis reports a z-test comparison of event-free rates, with the 95% confidence interval for the difference estimated using the normal approximation and Greenwood standard errors.

Reported 1-year OS-rate difference

−0.85

95% CI: −11.63 to 9.92   ·   P = 0.8767

EndpointTime frameDifference95% CIP-value
OS Rates1 and 2 years−0.85−11.63 to 9.920.8767

The reported difference of −0.85 percentage points is accompanied by a broad 95% CI extending from −11.63 to 9.92 percentage points. The interval includes 0, and the P-value is 0.8767.

Clinical Biostats interpretation

The landmark OS result should not be confused with the primary OS hazard-ratio analysis. A landmark survival-rate difference is a time-specific absolute comparison, whereas the hazard ratio summarizes relative event hazards within a Cox-model framework.

The wide interval indicates substantial uncertainty around the reported time-specific difference. It is therefore inappropriate to infer a precise equivalence of survival probabilities from the point estimate alone.

14. Secondary Results Summary

EndpointAnalysisEffect95% CIP-value
PFSStratifiedHR 0.7350.585–0.9230.0079
PFSUnstratifiedHR 0.7310.583–0.9150.0061
Confirmed ORRStratified CMHDifference −1.51−8.85 to 5.840.6846
Confirmed ORRStratified logistic regressionOR 0.880.47–1.63Not reported
Confirmed ORRUnstratified chi-squaredDifference −1.51−8.85 to 5.840.7216
Confirmed ORRUnstratified logistic regressionOR 0.870.47–1.62Not reported
TTCD Physical FunctioningStratifiedHR 0.820.59–1.160.2700
TTCD Physical FunctioningUnstratifiedHR 0.840.60–1.170.3031
TTCD Global Health StatusStratifiedHR 1.240.86–1.790.2408
TTCD Global Health StatusUnstratifiedHR 1.260.88–1.810.1992
PFS rate6 monthsDifference 15.856.12–25.590.0014
PFS rate1 yearDifference 6.32−0.56 to 13.210.0719
OS rate1 yearDifference −0.85−11.63 to 9.920.8767
Reading the table: the direction of a reported difference follows the registry's stated comparison of Docetaxel Monotherapy vs Atezolizumab + Cabozantinib. Hazard ratios are similarly reported for that comparison. These conventions matter when interpreting whether a value above or below the null value represents a higher or lower outcome in a particular arm.

15. Safety Results

The ClinicalTrials.gov record includes serious adverse events by treatment arm. The reported affected/at-risk counts were 58/167 for docetaxel monotherapy and 76/185 for atezolizumab plus cabozantinib.

Treatment armParticipants with serious adverse eventsParticipants at risk
Docetaxel Monotherapy58167
Atezolizumab + Cabozantinib76185
Serious adverse events: affected participants
Docetaxel Monotherapy
58
Atezolizumab + Cabozantinib
76

The affected/at-risk figures should be kept distinct from efficacy event counts. A serious adverse event is a safety outcome, whereas OS and PFS are efficacy time-to-event endpoints. The ClinicalTrials.gov record does not provide a formal between-arm statistical test for serious adverse events, so no comparative P-value or relative-effect estimate is assigned to these counts here.

Safety denominator matters: the serious-adverse-event summary uses 167 and 185 as the reported at-risk denominators. Those values should not be silently replaced by the overall enrollment figure of 366 when describing the reported safety data.

16. Statistical Methods Explained

Why was the intention-to-treat population used?

The ITT population includes all randomized participants, whether or not they received the assigned treatment. This preserves the treatment comparison created by randomization. If participants were instead moved between analysis groups after randomization according to treatment received, prognostic differences could become entangled with treatment exposure.

Why use both stratified and unstratified analyses?

The stratified analysis accounts for the registry-specified stratification factors, while the unstratified analysis does not. Comparing the two can show how much the estimated treatment effect changes when those factors are incorporated. They should not be interpreted as two separate randomized experiments.

What does a hazard ratio of 0.735 mean?

For the reported PFS analysis, an HR of 0.735 means that the estimated instantaneous rate of progression or death for the first-named group was 0.735 times that of the comparator under the fitted Cox model. Equivalently, using the simple complement interpretation, the estimate corresponds to approximately a 26.5% lower hazard. It does not mean a 26.5% increase in median PFS or that 26.5% of participants were protected from progression.

Why is the confidence interval more informative than the point estimate alone?

A point estimate is only one estimate of an underlying treatment effect. The 95% confidence interval shows the uncertainty associated with that estimate under the model and sampling framework. For example, the stratified OS HR of 0.884 has a 95% CI of 0.676–1.156, demonstrating that the plausible range is substantially wider than the single point estimate.

Why is an odds ratio not the same as a risk ratio?

An odds ratio compares odds rather than probabilities. An OR of 0.88 means the estimated odds of response were 0.88 times the comparator's odds in the reported model. It does not mean that the response probability was exactly 12% lower. The difference becomes especially important when response probabilities are not small.

Why use a log-rank test for OS and PFS?

OS and PFS are time-to-event outcomes with potentially different follow-up times and right censoring. The log-rank test compares the event experience of the randomized groups over follow-up rather than reducing every participant to a simple event/no-event classification at one arbitrary time point.

Why report both hazard ratios and landmark rates?

A hazard ratio summarizes a relative time-to-event comparison, whereas a landmark rate provides an absolute estimate at a particular time. The registry-reported PFS analyses illustrate this distinction: the PFS HR is 0.735 in the stratified analysis, while the reported difference in PFS rate is 15.85 at 6 months and 6.32 at 1 year. These quantities describe different aspects of the same survival experience.

17. Multiplicity and Interpretation of Multiple Analyses

The ClinicalTrials.gov record contains 15 statistical analyses, including two analyses of the registered primary endpoint and multiple secondary endpoint analyses. The primary hypothesis type is identified as superiority.

Multiple reported analyses create an important interpretive distinction. A P-value from an individual secondary analysis should not automatically be interpreted as though it were the sole prespecified hypothesis test in the trial. Formal multiplicity control depends on the prespecified testing hierarchy, alpha allocation, and statistical analysis plan. Those details are not provided in the ClinicalTrials.gov record.

Analysis familyExamples in the ClinicalTrials.gov recordInterpretive issue
Primary endpointStratified and unstratified OSOne registered primary endpoint with two reported analysis approaches
Time-to-event secondary endpointsPFS; TTCD in PF; TTCD in GHSMultiple event-time comparisons
Binary response endpointORRMultiple methods and effect measures
Landmark ratesPFS at 6 months and 1 year; OS rates at 1 and 2 yearsMultiple time-specific comparisons

Accordingly, the numerical P-values should be read in the context of their individual analyses rather than assembled into a single informal list of "significant" and "nonsignificant" findings. The ClinicalTrials.gov record does not provide enough information to reconstruct a full multiplicity-adjustment hierarchy.

18. Censoring and Time-to-Event Interpretation

Time-to-event analysis depends on appropriate handling of participants whose event has not occurred by the analysis cutoff. For OS, the registry explicitly states that participants alive at the time of analysis were censored at the date when they were last known to be alive as documented by the investigator.

For the patient-reported deterioration endpoints, participants without confirmed deterioration were censored according to their last known non-deteriorated status. This preserves information up to the participant's last evaluable time rather than treating an unobserved future deterioration as if it had already occurred.

Censoring concept
Observed follow-up = time contributed before event or censoring

Censoring does not mean that a participant's outcome is known after the censoring time. It means that the analysis uses the information available through the last appropriate observation.

The Cox hazard ratio and Kaplan-Meier estimates therefore depend not only on the observed events but also on the timing and handling of censoring. A time-to-event estimate cannot be reconstructed accurately from an event count alone.

19. Non-Inferiority, Crossover, and Factorial Design

Superiority

The statistical analyses posted on ClinicalTrials.gov identify the hypothesis type as superiority. No non-inferiority margin is provided because the reported analyses are framed as superiority comparisons.

Crossover

The registry-reported trial design is described as a parallel-group randomized trial. The ClinicalTrials.gov record do not report a treatment crossover analysis.

Factorial design

The design model is PARALLEL rather than factorial. The ClinicalTrials.gov record therefore do not support a factorial treatment-interaction analysis.

Bayesian methods

No Bayesian method is listed among the normalized statistical methods in the ClinicalTrials.gov record.

These distinctions matter because the statistical logic of a superiority trial differs from that of a non-inferiority trial. In a non-inferiority analysis, the prespecified margin is central to the conclusion. Here, the ClinicalTrials.gov record identifies superiority as the hypothesis type and do not provide a non-inferiority margin.

20. Wald Confidence Intervals and z-Tests

The ORR analyses used Wald confidence intervals, and the landmark PFS and OS rate analyses used z-tests with confidence intervals based on normal approximation. These methods illustrate a common distinction between effect estimation and hypothesis testing.

Generic Wald form
Estimate ± critical value × standard error

The confidence interval combines an estimated effect with its standard error. The interval width therefore reflects both the variability of the estimate and the amount of information available for the analysis.

For the ORR difference, the reported estimate is −1.51 with a 95% CI of −8.85 to 5.84. For the 6-month PFS-rate difference, the estimate is 15.85 with a 95% CI of 6.12–25.59. These intervals are expressed on the natural scale of the reported difference rather than on the hazard-ratio scale.

The choice of effect scale matters. A confidence interval around an odds ratio is interpreted relative to the null value 1, while a confidence interval around a difference is interpreted relative to the null value 0.

21. Missing Data and Imputation

The ClinicalTrials.gov record provides censoring descriptions for the time-to-event endpoints, but they do not provide a detailed missing-data or imputation strategy for the full set of outcome measures.

For time-to-event endpoints, censoring is part of the primary analysis framework and should not automatically be equated with ordinary missing-data imputation. For patient-reported outcomes, missing questionnaires can raise additional questions about whether observations are missing independently of underlying health status.

Interpretive limit: the ClinicalTrials.gov record does not specify a complete imputation strategy, sensitivity-analysis framework, or missing-not-at-random analysis. Those features should therefore not be inferred from the reported hazard ratios or P-values.

22. What the Primary Hazard Ratio Does — and Does Not — Mean

Relative effect

The stratified OS HR of 0.884 means that the fitted model estimated the instantaneous rate of death for the first-named group at 0.884 times the rate for the comparator, under the reported Cox analysis.

Not an absolute effect

The HR does not state how many additional participants were alive at a particular time, how many deaths were prevented, or how much longer an individual participant lived. Those are absolute or individual-level quantities and require different summaries.

Confidence interval

The 95% CI of 0.676–1.156 indicates uncertainty around the estimated HR. Because it crosses the null value of 1, the reported interval includes both lower and higher hazards for the first-named group.

P-value

The P-value of 0.3668 measures statistical evidence under the reported superiority test. It is not the probability that the null hypothesis is true and is not a measure of the size or practical importance of the estimated HR.

Model assumptions

The Cox interpretation depends on the time-to-event model and its assumptions. The ClinicalTrials.gov record does not report a separate proportional-hazards diagnostic, so the HR should be treated as a model-based summary rather than as proof of a constant relative hazard at every point in time.

23. Why This Trial Matters Statistically

CONTACT-01 is a useful teaching example because a single randomized trial contains several distinct statistical estimands and analysis frameworks. The registered primary endpoint is time-to-event, but the posted secondary analyses move between survival analysis, categorical-data methods, logistic regression, and landmark rate comparisons.

ConceptHow it appears in CONTACT-01
RandomizationParallel-group randomized phase 3 design with 366 enrolled participants
ITT analysisEfficacy analyses include all randomized participants whether or not assigned treatment was received
Kaplan-Meier estimationUsed for OS estimation and median estimation
Log-rank testPrimary OS and secondary time-to-event comparisons
Hazard ratioPrimary OS and secondary PFS and patient-reported time-to-deterioration effects
Cox regressionUsed to estimate hazard ratios
Stratified analysisHistology and prior NSCLC treatment regimens appear as stratification factors
Cochran-Mantel-Haenszel testStratified analysis of confirmed objective response rate
Logistic regressionOdds-ratio estimation for confirmed objective response rate
Wald methodConfidence intervals for response-rate differences and odds ratios
z-testLandmark PFS-rate and OS-rate comparisons
Greenwood methodStandard errors for reported time-specific PFS and OS rate differences
Multiple analyses15 statistical analyses are posted across primary and secondary endpoints

The most important statistical lesson is that these methods are not interchangeable. A hazard ratio, odds ratio, response-rate difference, and landmark survival-rate difference each answer a different question. Correct interpretation requires preserving the effect scale, endpoint definition, analysis population, and time frame.

24. Important Limitations and Interpretation Issues

25. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The registered OS endpoint produced a stratified HR of 0.884 (95% CI 0.676–1.156; P = 0.3668). The secondary PFS analysis produced a stratified HR of 0.735 (95% CI 0.585–0.923; P = 0.0079). Other endpoints produced estimates on different scales with their own uncertainty.

Clinical interpretation

Clinical interpretation requires considering what each endpoint measures, the magnitude and precision of its effect estimate, the time frame, safety information, and the prespecified role of the endpoint. The ClinicalTrials.gov record alone do not justify reducing the entire trial to one numerical measure.

This distinction is especially important in a trial containing both OS and PFS results. A PFS hazard ratio and an OS hazard ratio do not represent interchangeable outcomes. OS incorporates survival through the entire follow-up period, while PFS captures the earlier occurrence of progression or death.

26. Overall Statistical Synthesis

CONTACT-01 provides a compact example of how clinical-trial evidence can change depending on the statistical lens used. The primary OS analysis estimates a hazard ratio of 0.884 in the stratified analysis, with a 95% CI of 0.676–1.156 and P = 0.3668. The corresponding unstratified estimate is 0.907, with a 95% CI of 0.696–1.182 and P = 0.4709.

For PFS, the stratified HR is 0.735 with a 95% CI of 0.585–0.923 and P = 0.0079, while the unstratified HR is 0.731 with a 95% CI of 0.583–0.915 and P = 0.0061. The landmark PFS analysis adds a time-specific perspective: the reported difference in event-free rate is 15.85 at 6 months and 6.32 at 1 year.

Other secondary endpoints illustrate why a single P-value cannot summarize the entire statistical record. Confirmed ORR has a reported stratified response-rate difference of −1.51 with a 95% CI of −8.85 to 5.84, while logistic regression gives an odds ratio of 0.88 with a 95% CI of 0.47–1.63. Patient-reported time-to-deterioration analyses similarly use hazard ratios but address different clinical constructs.

The statistically disciplined interpretation is therefore to retain the endpoint-specific effect measure, its confidence interval, its P-value, the analysis population, and the method used. The trial's results cannot be faithfully represented by treating hazard ratios, odds ratios, percentage-point differences, and landmark rates as though they were the same kind of quantity.

27. Related Tutorials

Learn more about the methods used in this trial:

28. Related Statistical Calculators

29. Sources

Continue through the Clinical Biostats statistical library

Explore the underlying statistical methods through tutorials and practical calculators for survival analysis, categorical data, regression, confidence intervals, and clinical-trial analysis.

30. Record Summary

CONTACT-01 is a randomized phase 3 parallel-group trial with 366 enrolled participants and a single registered primary endpoint, overall survival. The primary OS endpoint was analyzed in the ITT population using a log-rank test, with the hazard ratio estimated by Cox regression. The stratified OS analysis reported an HR of 0.884 (95% CI 0.676–1.156; P = 0.3668), while the unstratified analysis reported an HR of 0.907 (95% CI 0.696–1.182; P = 0.4709).

The secondary analyses demonstrate the breadth of modern clinical-trial statistics. PFS was evaluated with stratified and unstratified log-rank/Cox analyses; confirmed ORR used Cochran-Mantel-Haenszel, chi-squared, and logistic-regression approaches; patient-reported deterioration used time-to-event methods; and landmark PFS and OS rates used z-tests with Greenwood-based standard errors. The serious-adverse-event summary reports 58/167 for docetaxel monotherapy and 76/185 for atezolizumab plus cabozantinib.

The principal educational lesson is that a clinical trial does not have one universal "result." Each endpoint has an estimand, an analysis population, an effect measure, an uncertainty interval, and a statistical test. Reading CONTACT-01 rigorously therefore means keeping those elements together rather than collapsing the trial into a single hazard ratio or P-value.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation, preserve the original endpoint definitions and effect scales, and avoid manufacturing unreported quantities. Where the ClinicalTrials.gov record does not provide a result or methodological detail, this page does not infer one.