← Clinical Trials
Renal Cell Carcinoma Phase 3 Time-to-Event Analysis NCT03937219

COSMIC-313: Complete Statistical Analysis of Cabozantinib in Advanced or Metastatic Renal Cell Carcinoma

An independent statistical review of the randomized phase 3 COSMIC-313 trial evaluating cabozantinib in combination with nivolumab and ipilimumab versus cabozantinib-matched placebo with nivolumab and ipilimumab in previously untreated advanced or metastatic renal cell carcinoma.

Trial start: June 25, 2019  ·  Primary completion: January 31, 2022  ·  Status: Active, not recruiting
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical efficacy result presented here is the statistical analysis posted in the ClinicalTrials.gov record. The registry provides the official trial record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

COSMIC-313 was a randomized, triple-masked, parallel phase 3 trial with 855 enrolled participants. The primary endpoint was duration of progression-free survival (PFS) by Blinded Independent Radiology Committee (BIRC), analyzed as a time-to-event outcome using a log-rank test and summarized with a hazard ratio.

855
Enrolled
Phase 3
2
Arms
Parallel design
0.73
PFS HR
95% CI 0.57–0.94
0.0131
P-value
Superiority analysis
FeatureCOSMIC-313
Trial nameCOSMIC-313
PhasePhase 3
ConditionRenal Cell Carcinoma
PopulationPatients with previously untreated advanced or metastatic renal cell carcinoma
AllocationRandomized
Design modelParallel
MaskingTriple
Primary purposeTreatment
Enrollment855
Primary endpointDuration of Progression-Free Survival (PFS) by Blinded Independent Radiology Committee (BIRC)
Primary endpoint typeTime-to-event
Primary analysisLog-rank test
Effect measureHazard ratio
Hypothesis typeSuperiority
Lead sponsorExelixis
Sponsor typeIndustry
ClinicalTrials.govNCT03937219

2. Clinical Question

The trial asks whether adding cabozantinib to nivolumab and ipilimumab improves progression-free survival compared with cabozantinib-matched placebo combined with nivolumab and ipilimumab in patients with previously untreated advanced or metastatic renal cell carcinoma.

Population

Patients with previously untreated advanced or metastatic renal cell carcinoma.

Intervention

Cabozantinib in combination with nivolumab and ipilimumab.

Comparator

Cabozantinib-matched placebo in combination with nivolumab and ipilimumab.

Primary question

Does the cabozantinib-containing combination improve the duration of PFS relative to the placebo-containing combination?

3. Trial Design

01
Randomize 855 enrolled
02
Two arms Parallel design
03
Triple mask Masked trial design
04
Assess PFS BIRC evaluation
05
Compare Log-rank + HR
ARM A

Cabozantinib combination

  • Cabozantinib
  • Nivolumab
  • Ipilimumab
ARM B

Placebo combination

  • Cabozantinib-matched placebo
  • Nivolumab
  • Ipilimumab
Allocation
Randomized allocation was used to create the treatment comparison.
Masking
The trial is registered as triple-masked.
Design model
Parallel-group design with two arms.
Primary purpose
Treatment.

The use of a cabozantinib-matched placebo is important to the internal validity of the comparison because the randomized groups differ in the presence of cabozantinib while the background nivolumab and ipilimumab components are represented in both arms. The registry identifies the overall masking as triple.

4. Endpoints

EndpointRegistry definitionTime frameType
Duration of Progression-Free Survival (PFS) by Blinded Independent Radiology Committee (BIRC) Duration of PFS was defined as the time from randomization to the earlier of either the date of radiographic progression per BIRC or the date of death due to any cause. PFS (months) = (earliest date of progression, death, censoring - date of randomization + 1)/30.4375. PFS was determined as per Response Evaluation Criteria in Solid Tumors version (RECIST) v1.1. Up to 32 months Time-to-event

This endpoint combines two clinically important event types into a single time-to-event outcome: radiographic progression and death. The event is whichever occurs first. Participants who have not experienced either event at the relevant end of observation contribute follow-up through censoring.

Why PFS is a time-to-event endpoint

A simple proportion of patients who progress would discard information about when progression occurred. PFS retains the timing of progression or death and can also accommodate participants whose event status is not observed by the end of their available follow-up.

Registry definition
PFS = (earliest of progression, death, or censoring − randomization date + 1) / 30.4375

The registry expresses PFS in months using this specified calculation. The event definition is based on radiographic progression according to BIRC or death from any cause.

5. Statistical Methodology

Intention-to-treat analysis

The posted PFS analysis uses the PFS Intent-to-Treat (PITT) population. The registry defines this population as the first 550 randomized participants regardless of whether any study treatment or the correct study treatment was received. This is an important distinction from a safety population or an analysis based only on participants who actually received treatment.

Why this matters: analyzing randomized participants according to the trial's specified randomized population preserves the treatment comparison created by randomization more directly than restricting the analysis to participants who completed treatment or received a particular exposure.

Log-rank test

The registered statistical method for the primary PFS comparison is the log-rank test. This is a standard hypothesis test for comparing time-to-event distributions between randomized groups while accounting for the timing of events and censoring.

Hazard ratio

The primary effect measure is the hazard ratio (HR). Unlike a median or a fixed-time survival proportion, the HR summarizes the relative instantaneous event rate between the two groups over the analyzed time period under the model and analysis framework.

Conceptual interpretation
HR = hazard in cabozantinib + nivolumab + ipilimumab / hazard in placebo + nivolumab + ipilimumab

An HR below 1 indicates a lower estimated instantaneous event rate in the cabozantinib-containing group relative to the placebo-containing group. It is not itself an absolute difference in PFS time.

Blinded Independent Radiology Committee assessment

The primary endpoint is explicitly based on assessment by a Blinded Independent Radiology Committee (BIRC). This provides an endpoint-assessment framework in which radiographic progression is evaluated independently of the treatment assignment as part of the trial's blinded assessment process.

Superiority hypothesis

The registry classifies the hypothesis as superiority. Thus, the statistical question is whether the randomized treatment comparison provides evidence that the cabozantinib-containing regimen has a different, specifically improved, PFS experience rather than whether it merely meets a non-inferiority criterion.

6. Primary PFS Result

The posted statistical analysis compares cabozantinib + nivolumab + ipilimumab with placebo + nivolumab + ipilimumab in the PFS Intent-to-Treat population.

Hazard ratio for progression or death

0.73

95% CI: 0.57–0.94   ·   P = 0.0131

Analysis: two-sided 95% confidence interval; log-rank test; superiority hypothesis.

Relative hazard estimate
Cabozantinib combination
0.73
Reference
1.00
Clinical Biostats interpretation

An HR of 0.73 means that the estimated instantaneous rate of progression or death in the cabozantinib-containing group was about 73% of the corresponding rate in the placebo-containing group, under the time-to-event analysis. Equivalently, 1 − 0.73 = 0.27, so the estimated hazard was approximately 27% lower in relative terms.

The HR does not mean that 27% of participants avoided progression, that every participant experienced a 27% reduction in risk, or that PFS duration was 27% longer. It is a relative time-to-event measure, not an individual-level prediction.

The 95% CI of 0.57–0.94 describes the statistical uncertainty around the estimated HR under the analysis framework. It does not mean that individual patients have HRs somewhere between 0.57 and 0.94. The interval lies below 1, which is consistent with a lower event hazard in the cabozantinib-containing group under the reported analysis.

The P-value of 0.0131 addresses the compatibility of the observed data with the null hypothesis used for the superiority comparison. It does not measure the size or clinical importance of the treatment effect. Effect magnitude is better represented by the HR itself and its confidence interval, while absolute PFS quantities would provide additional clinical context when available.

Because this is a time-to-event analysis, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio representation. A single HR is most straightforward when the relative hazards are reasonably stable over time; the ClinicalTrials.gov record does not provide a time-varying hazard assessment or a proportional-hazards diagnostic.

7. Understanding the Primary Analysis Population

The posted analysis uses the PFS Intent-to-Treat (PITT) population, defined in the registry as the first 550 randomized participants regardless of whether any study treatment or the correct study treatment was received.

Population featureRegistry description
Analysis populationPFS Intent-to-Treat (PITT)
Population size specified in the analysis descriptionFirst 550 randomized participants
Treatment receipt requirementNone; analysis definition states that participants were included regardless of whether any study treatment or the correct study treatment was received
RolePrimary PFS efficacy analysis

This design choice is statistically important. Randomization creates the basis for a comparison between treatment assignments. If the primary analysis instead excluded participants because they did not receive treatment as planned, the groups could become systematically different after randomization. The PITT definition keeps the efficacy analysis tied to the randomized population specified by the registry.

8. How to Read the Hazard Ratio of 0.73

Relative effect

The reported HR of 0.73 corresponds to a 27% lower estimated instantaneous rate of progression or death for the cabozantinib-containing regimen relative to the placebo-containing regimen, using the reported hazard-ratio framework.

What it does not say

It does not say that the probability of progression or death was exactly 27% lower at every time point. It also does not give a median PFS, an absolute PFS difference, or the percentage of participants who benefited.

Why the confidence interval matters

The 95% CI of 0.57–0.94 communicates precision around the estimated HR. The width of the interval reflects uncertainty in the treatment-effect estimate; the interval is more informative about precision than the point estimate alone.

Why the P-value is different

The P-value of 0.0131 is evidence against the null hypothesis specified for the superiority analysis under the reported statistical test. It is not a probability that the treatment works, and it does not quantify how large the treatment effect is.

9. Statistical Methods Explained

Why was a log-rank test used?

PFS is measured as a time-to-event endpoint, with progression or death occurring at different times and with some participants potentially censored. The log-rank test is designed for comparing survival-type time-to-event distributions while using information from the timing of events rather than reducing the outcome to a single binary proportion.

What does an HR of 0.73 mean?

An HR of 0.73 means that the estimated instantaneous event rate for progression or death was 0.73 times the corresponding rate in the comparator group under the reported analysis. The simple relative interpretation is a 27% lower estimated hazard. It should not be translated into a 27% absolute reduction in the probability of an event.

Why is the confidence interval 0.57–0.94 important?

The point estimate is only one estimate of the treatment effect. The 95% CI shows the range of HR values compatible with the statistical uncertainty represented by the analysis. It provides information about precision that a P-value alone cannot provide.

Why does the P-value not measure effect size?

A P-value describes evidence against a null hypothesis within a specified statistical framework. It depends on both the observed effect and the amount of information available. Two studies can produce different P-values for effects of similar magnitude, and a small P-value does not automatically imply a large treatment effect.

Why is the analysis population important?

The primary analysis is based on the PFS Intent-to-Treat population, defined as the first 550 randomized participants regardless of whether study treatment or the correct study treatment was received. Anchoring the efficacy comparison to randomization helps preserve the comparability established by the randomized design.

Why does censoring matter?

Participants may reach the end of their evaluable follow-up without progression or death. Their observations are censored rather than treated as if they experienced the event at that time. Time-to-event methods such as Kaplan-Meier estimation and the log-rank test are designed to incorporate this partial information.

10. Kaplan-Meier Estimation and PFS

Although the posted primary analysis identifies the log-rank test and hazard ratio, the underlying PFS endpoint is naturally represented using a Kaplan-Meier framework. Kaplan-Meier estimation describes the estimated probability of remaining event-free over time while accounting for right-censored observations.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

where di is the number of events at time ti and ni is the number at risk immediately before that time.

The ClinicalTrials.gov record does not provide the event-by-event risk sets needed to reconstruct a Kaplan-Meier curve. It is therefore preferable to explain the method rather than manufacture a curve from the single reported HR and confidence interval.

Kaplan-Meier versus hazard ratio

The two approaches answer related but different questions. A Kaplan-Meier curve describes the estimated event-free probability over time. A hazard ratio summarizes the relative instantaneous event rate between groups under the fitted time-to-event framework. Reporting both can provide a more complete picture when the underlying survival estimates are available.

11. Blinding and Independent Radiology Assessment

COSMIC-313 is registered as triple-masked, and the primary PFS endpoint was determined by a Blinded Independent Radiology Committee. These features are particularly relevant for a radiographic endpoint because assessment of disease progression involves interpretation of imaging findings.

Blinding

Triple masking reduces the opportunity for knowledge of treatment assignment to influence aspects of trial conduct and assessment.

Independent review

BIRC assessment provides an independent framework for evaluating radiographic progression for the primary PFS endpoint.

RECIST v1.1

The registry states that PFS was determined according to Response Evaluation Criteria in Solid Tumors version (RECIST) v1.1.

Time-to-event structure

Radiographic progression and death are incorporated into a single prespecified PFS event definition.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm. These figures describe the number affected among the number at risk for each arm.

Safety measureCabozantinib + Nivolumab + IpilimumabPlacebo + Nivolumab + Ipilimumab
Serious adverse events 272/426 257/423
Serious adverse events: affected / at risk
Cabozantinib combination
272/426
Placebo combination
257/423

The ClinicalTrials.gov record reports these safety counts but do not provide a formal statistical comparison for serious adverse events. The figures should therefore be treated as descriptive safety information rather than as a separate hypothesis-tested efficacy-style result.

Safety interpretation: the serious-adverse-event counts use denominators of 426 and 423, respectively, rather than the overall enrollment of 855. These denominators are the at-risk populations reported with the safety data and should not be replaced with the overall enrollment when describing the reported safety measure.

13. Trial Timeline

June 25, 2019

Trial start

The registered study start date is June 25, 2019.

January 31, 2022

Primary completion

The registered primary completion date is January 31, 2022.

Current registry status

Active, not recruiting

The trial is listed as ACTIVE_NOT_RECRUITING in the ClinicalTrials.gov record.

14. What the Primary Result Establishes

The reported statistical analysis provides three closely related pieces of evidence: an estimated relative treatment effect, an uncertainty interval, and a hypothesis-test result.

QuantityReported valueStatistical role
Hazard ratio 0.73 Estimates the relative instantaneous rate of progression or death between the randomized groups.
95% confidence interval 0.57–0.94 Describes uncertainty around the estimated hazard ratio.
P-value 0.0131 Quantifies evidence against the null hypothesis within the reported superiority testing framework.
Analysis method Log-rank test Compares the time-to-event experience between the randomized groups.
Hypothesis type Superiority Frames the analysis as a test for improvement rather than non-inferiority.

Taken together, the posted analysis reports a hazard ratio below 1 with a 95% confidence interval that remains below 1 and a P-value of 0.0131. The statistical interpretation is therefore based on the reported superiority analysis, not on a comparison of unadjusted event percentages.

15. What the Primary Result Does Not Establish

Data boundary: The ClinicalTrials.gov record contains one posted statistical analysis for the primary PFS endpoint. No secondary-endpoint statistical analyses are included in the ClinicalTrials.gov record, so this page does not add secondary efficacy estimates, median PFS values, subgroup results, or other outcome measures that are not provided.

16. Important Limitations and Interpretation Issues

17. Why This Trial Matters Statistically

COSMIC-313 is a useful teaching example for understanding how randomized oncology trials translate a time-to-event clinical question into a formal statistical comparison. The endpoint is defined from randomization through progression, death, or censoring; radiographic progression is evaluated by BIRC; the primary analysis uses a PFS Intent-to-Treat population; and the treatment effect is summarized with a hazard ratio alongside a confidence interval and P-value.

ConceptHow it appears in COSMIC-313
RandomizationThe trial is registered as randomized with two parallel arms.
BlindingThe masking designation is triple.
Time-to-event endpointPrimary endpoint is duration of PFS by BIRC.
RECIST v1.1PFS is determined according to RECIST v1.1.
Intention-to-treat analysisThe posted efficacy analysis uses the PFS Intent-to-Treat population.
Kaplan-Meier frameworkPFS is a censored time-to-event outcome naturally represented through Kaplan-Meier estimation.
Log-rank testRegistry-reported method for the primary comparison.
Hazard ratioPrimary effect measure, reported as 0.73.
Confidence interval95% CI of 0.57–0.94 quantifies uncertainty around the HR.
P-valueReported as 0.0131 for the superiority analysis.
Independent assessmentBIRC is specified for the primary PFS endpoint.
Safety descriptionSerious adverse events are reported by treatment arm with affected/at-risk counts.

18. Statistical Methods Explained in More Depth

Randomization and causal comparison

Randomization is the structural foundation of the treatment comparison. By assigning participants to treatment groups randomly, the design aims to make the groups comparable in expectation with respect to measured and unmeasured baseline factors. The subsequent analysis can therefore attribute differences between randomized groups to treatment assignment more credibly than an observational comparison can.

Why the PITT population is different from an enrollment count

The trial enrolled 855 participants, but the registry-reported primary PFS analysis specifies the first 550 randomized participants as the PITT population. Enrollment describes the overall study population, whereas the PITT definition identifies the participants used for this particular efficacy analysis. Confusing these two quantities would produce an incorrect description of the primary analysis.

Why PFS includes death

If death were excluded from the endpoint, a participant who died before documented radiographic progression could potentially be treated differently from a participant who remained alive without progression. Including death in the definition ensures that death is itself a PFS event. In COSMIC-313, the registry explicitly defines the event as the earlier of radiographic progression per BIRC or death from any cause.

Why BIRC assessment matters

Radiographic progression depends on imaging assessments. A blinded independent committee provides an assessment process designed to reduce the influence of treatment assignment on the determination of progression. This is particularly relevant when progression is the event driving a primary efficacy endpoint.

Why a confidence interval is more informative than a P-value alone

The P-value indicates the degree of evidence against the null hypothesis under the specified test. The confidence interval adds information about the estimated magnitude and precision of the effect. In this trial, the HR of 0.73 is accompanied by a 95% CI of 0.57–0.94, giving a substantially richer description of the estimate than the P-value of 0.0131 alone.

Why a hazard ratio is not a median ratio

An HR compares instantaneous event rates within a time-to-event model. It is not calculated by dividing one group's median PFS by the other group's median PFS. Because the ClinicalTrials.gov record does not report median PFS, no median-based comparison is presented here.

19. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The reported PFS analysis produced an HR of 0.73 with a two-sided 95% CI of 0.57–0.94 and P = 0.0131 using a log-rank test under a superiority hypothesis.

Clinical interpretation

The statistical result describes a lower estimated relative hazard of progression or death in the cabozantinib-containing group. The ClinicalTrials.gov record does not provide median PFS or fixed-time PFS estimates, so absolute duration-based clinical measures cannot be stated here.

Keeping these two levels of interpretation separate is important. The statistical analysis tells us how the randomized groups compare under the prespecified endpoint and test. Clinical interpretation requires consideration of the magnitude of the effect, its precision, absolute outcome measures, safety, patient population, and the broader treatment context. Only some of those elements are contained in the ClinicalTrials.gov record.

20. A Worked Reading of the COSMIC-313 Result

Step 1 · Identify the endpoint

The primary endpoint is duration of PFS by BIRC, defined from randomization to the earlier of radiographic progression or death, with censoring handled according to the registered endpoint definition.

Step 2 · Identify the analysis population

The primary posted analysis uses the PFS Intent-to-Treat population, defined as the first 550 randomized participants regardless of whether any study treatment or the correct study treatment was received.

Step 3 · Identify the test

The registry reports a log-rank test for the primary comparison, with superiority as the hypothesis type.

Step 4 · Identify the effect estimate

The reported hazard ratio is 0.73, comparing cabozantinib + nivolumab + ipilimumab with placebo + nivolumab + ipilimumab.

Step 5 · Add uncertainty and hypothesis testing

The 95% CI is 0.57–0.94 and the P-value is 0.0131. The CI describes uncertainty around the HR; the P-value addresses the hypothesis test. Neither should be interpreted as an individual patient's probability of benefit.

21. What Is Available — and What Is Not — in the Supplied Results

InformationSupplied?How it is handled
Primary endpoint definitionYesReported using the registry wording and time frame.
Primary statistical methodYesLog-rank test explained in detail.
Primary effect estimateYesHR 0.73 reported exactly.
95% confidence intervalYes0.57–0.94 reported exactly.
P-valueYes0.0131 reported exactly.
Analysis populationYesFirst 550 randomized participants in the PITT population.
Secondary endpoint analysesNo formal analyses reportedNo secondary efficacy estimates are added.
Median PFSNot reportedNo median PFS is reported on this page.
Fixed-time PFS estimatesNot reportedNo fixed-time survival probabilities are reported.
Subgroup estimatesNot reportedNo subgroup results are added.
Interim-analysis detailsNot reportedNo interim schedule or alpha-spending procedure is inferred.
Stratification factorsNot reportedNo stratification variables are inferred.
Bayesian analysisNot reportedNo Bayesian method is attributed to the trial.

22. Limitations of the Statistical Evidence Presented Here

The most important limitation is the granularity of the ClinicalTrials.gov record. The registry provides a formal primary PFS analysis, but the statistical result is summarized at the hazard-ratio level. Without the underlying event and censoring data, a full reconstruction of the PFS distribution is not possible.

In particular, an HR of 0.73 does not reveal whether the treatment groups had similar or different PFS patterns at specific time points, whether hazards changed substantially over time, or what the median PFS was in either group. Those are separate descriptive questions that require additional survival information.

The analysis population also deserves attention. The trial enrolled 855 participants, while the posted primary PFS analysis is defined using the first 550 randomized participants. This distinction means that the enrollment count should not be substituted for the PITT analysis population when describing the primary efficacy result.

Finally, the ClinicalTrials.gov record does not include a multiplicity plan, interim-monitoring details, randomization strata, missing-data strategy beyond the endpoint's censoring definition, or formal assessment of proportional hazards. Those topics can materially affect detailed interpretation of a time-to-event trial, but they should not be reconstructed from assumptions when the ClinicalTrials.gov record does not report them.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through Clinical Biostats

Explore statistical tutorials, calculators, and additional clinical trial analyses covering the methods used in randomized clinical research.

26. Record Summary

COSMIC-313 provides a clear example of a randomized phase 3 time-to-event analysis. The trial enrolled 855 participants, used a randomized parallel design with triple masking, and evaluated cabozantinib + nivolumab + ipilimumab against placebo + nivolumab + ipilimumab. Its registered primary endpoint was BIRC-assessed duration of PFS up to 32 months, defined from randomization to the earlier of radiographic progression or death. The posted primary analysis used the PFS Intent-to-Treat population of the first 550 randomized participants, a log-rank test, and a hazard ratio as the effect measure.

The reported HR of 0.73, with a 95% CI of 0.57–0.94 and P = 0.0131, describes a lower estimated relative hazard of progression or death in the cabozantinib-containing group under the reported superiority analysis. The most important statistical discipline is to distinguish that relative estimate from absolute PFS measures, individual patient outcomes, and safety outcomes. The ClinicalTrials.gov record does not provide median PFS, fixed-time PFS estimates, subgroup estimates, or additional formal efficacy analyses, so none are inferred here.

Clinical Biostats methodology: A trial-results page should reconstruct the statistical story using the evidence actually reported, explain what each statistical quantity means, and make clear where the available data stop. For COSMIC-313, the central statistical lesson is the interpretation of a hazard ratio in a randomized, blinded, BIRC-assessed time-to-event analysis.