This page separates reported trial results from statistical interpretation. The numerical efficacy result presented here is the statistical analysis posted in the ClinicalTrials.gov record. The registry provides the official trial record.
1. Trial at a Glance
COSMIC-313 was a randomized, triple-masked, parallel phase 3 trial with 855 enrolled participants. The primary endpoint was duration of progression-free survival (PFS) by Blinded Independent Radiology Committee (BIRC), analyzed as a time-to-event outcome using a log-rank test and summarized with a hazard ratio.
| Feature | COSMIC-313 |
|---|---|
| Trial name | COSMIC-313 |
| Phase | Phase 3 |
| Condition | Renal Cell Carcinoma |
| Population | Patients with previously untreated advanced or metastatic renal cell carcinoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Triple |
| Primary purpose | Treatment |
| Enrollment | 855 |
| Primary endpoint | Duration of Progression-Free Survival (PFS) by Blinded Independent Radiology Committee (BIRC) |
| Primary endpoint type | Time-to-event |
| Primary analysis | Log-rank test |
| Effect measure | Hazard ratio |
| Hypothesis type | Superiority |
| Lead sponsor | Exelixis |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT03937219 |
2. Clinical Question
The trial asks whether adding cabozantinib to nivolumab and ipilimumab improves progression-free survival compared with cabozantinib-matched placebo combined with nivolumab and ipilimumab in patients with previously untreated advanced or metastatic renal cell carcinoma.
Population
Patients with previously untreated advanced or metastatic renal cell carcinoma.
Intervention
Cabozantinib in combination with nivolumab and ipilimumab.
Comparator
Cabozantinib-matched placebo in combination with nivolumab and ipilimumab.
Primary question
Does the cabozantinib-containing combination improve the duration of PFS relative to the placebo-containing combination?
3. Trial Design
Cabozantinib combination
- Cabozantinib
- Nivolumab
- Ipilimumab
Placebo combination
- Cabozantinib-matched placebo
- Nivolumab
- Ipilimumab
The use of a cabozantinib-matched placebo is important to the internal validity of the comparison because the randomized groups differ in the presence of cabozantinib while the background nivolumab and ipilimumab components are represented in both arms. The registry identifies the overall masking as triple.
4. Endpoints
| Endpoint | Registry definition | Time frame | Type |
|---|---|---|---|
| Duration of Progression-Free Survival (PFS) by Blinded Independent Radiology Committee (BIRC) | Duration of PFS was defined as the time from randomization to the earlier of either the date of radiographic progression per BIRC or the date of death due to any cause. PFS (months) = (earliest date of progression, death, censoring - date of randomization + 1)/30.4375. PFS was determined as per Response Evaluation Criteria in Solid Tumors version (RECIST) v1.1. | Up to 32 months | Time-to-event |
This endpoint combines two clinically important event types into a single time-to-event outcome: radiographic progression and death. The event is whichever occurs first. Participants who have not experienced either event at the relevant end of observation contribute follow-up through censoring.
Why PFS is a time-to-event endpoint
A simple proportion of patients who progress would discard information about when progression occurred. PFS retains the timing of progression or death and can also accommodate participants whose event status is not observed by the end of their available follow-up.
The registry expresses PFS in months using this specified calculation. The event definition is based on radiographic progression according to BIRC or death from any cause.
5. Statistical Methodology
Intention-to-treat analysis
The posted PFS analysis uses the PFS Intent-to-Treat (PITT) population. The registry defines this population as the first 550 randomized participants regardless of whether any study treatment or the correct study treatment was received. This is an important distinction from a safety population or an analysis based only on participants who actually received treatment.
Log-rank test
The registered statistical method for the primary PFS comparison is the log-rank test. This is a standard hypothesis test for comparing time-to-event distributions between randomized groups while accounting for the timing of events and censoring.
Hazard ratio
The primary effect measure is the hazard ratio (HR). Unlike a median or a fixed-time survival proportion, the HR summarizes the relative instantaneous event rate between the two groups over the analyzed time period under the model and analysis framework.
An HR below 1 indicates a lower estimated instantaneous event rate in the cabozantinib-containing group relative to the placebo-containing group. It is not itself an absolute difference in PFS time.
Blinded Independent Radiology Committee assessment
The primary endpoint is explicitly based on assessment by a Blinded Independent Radiology Committee (BIRC). This provides an endpoint-assessment framework in which radiographic progression is evaluated independently of the treatment assignment as part of the trial's blinded assessment process.
Superiority hypothesis
The registry classifies the hypothesis as superiority. Thus, the statistical question is whether the randomized treatment comparison provides evidence that the cabozantinib-containing regimen has a different, specifically improved, PFS experience rather than whether it merely meets a non-inferiority criterion.
6. Primary PFS Result
The posted statistical analysis compares cabozantinib + nivolumab + ipilimumab with placebo + nivolumab + ipilimumab in the PFS Intent-to-Treat population.
Hazard ratio for progression or death
95% CI: 0.57–0.94 · P = 0.0131
Analysis: two-sided 95% confidence interval; log-rank test; superiority hypothesis.
An HR of 0.73 means that the estimated instantaneous rate of progression or death in the cabozantinib-containing group was about 73% of the corresponding rate in the placebo-containing group, under the time-to-event analysis. Equivalently, 1 − 0.73 = 0.27, so the estimated hazard was approximately 27% lower in relative terms.
The HR does not mean that 27% of participants avoided progression, that every participant experienced a 27% reduction in risk, or that PFS duration was 27% longer. It is a relative time-to-event measure, not an individual-level prediction.
The 95% CI of 0.57–0.94 describes the statistical uncertainty around the estimated HR under the analysis framework. It does not mean that individual patients have HRs somewhere between 0.57 and 0.94. The interval lies below 1, which is consistent with a lower event hazard in the cabozantinib-containing group under the reported analysis.
The P-value of 0.0131 addresses the compatibility of the observed data with the null hypothesis used for the superiority comparison. It does not measure the size or clinical importance of the treatment effect. Effect magnitude is better represented by the HR itself and its confidence interval, while absolute PFS quantities would provide additional clinical context when available.
Because this is a time-to-event analysis, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio representation. A single HR is most straightforward when the relative hazards are reasonably stable over time; the ClinicalTrials.gov record does not provide a time-varying hazard assessment or a proportional-hazards diagnostic.
7. Understanding the Primary Analysis Population
The posted analysis uses the PFS Intent-to-Treat (PITT) population, defined in the registry as the first 550 randomized participants regardless of whether any study treatment or the correct study treatment was received.
| Population feature | Registry description |
|---|---|
| Analysis population | PFS Intent-to-Treat (PITT) |
| Population size specified in the analysis description | First 550 randomized participants |
| Treatment receipt requirement | None; analysis definition states that participants were included regardless of whether any study treatment or the correct study treatment was received |
| Role | Primary PFS efficacy analysis |
This design choice is statistically important. Randomization creates the basis for a comparison between treatment assignments. If the primary analysis instead excluded participants because they did not receive treatment as planned, the groups could become systematically different after randomization. The PITT definition keeps the efficacy analysis tied to the randomized population specified by the registry.
8. How to Read the Hazard Ratio of 0.73
The reported HR of 0.73 corresponds to a 27% lower estimated instantaneous rate of progression or death for the cabozantinib-containing regimen relative to the placebo-containing regimen, using the reported hazard-ratio framework.
It does not say that the probability of progression or death was exactly 27% lower at every time point. It also does not give a median PFS, an absolute PFS difference, or the percentage of participants who benefited.
The 95% CI of 0.57–0.94 communicates precision around the estimated HR. The width of the interval reflects uncertainty in the treatment-effect estimate; the interval is more informative about precision than the point estimate alone.
The P-value of 0.0131 is evidence against the null hypothesis specified for the superiority analysis under the reported statistical test. It is not a probability that the treatment works, and it does not quantify how large the treatment effect is.
9. Statistical Methods Explained
Why was a log-rank test used?
PFS is measured as a time-to-event endpoint, with progression or death occurring at different times and with some participants potentially censored. The log-rank test is designed for comparing survival-type time-to-event distributions while using information from the timing of events rather than reducing the outcome to a single binary proportion.
What does an HR of 0.73 mean?
An HR of 0.73 means that the estimated instantaneous event rate for progression or death was 0.73 times the corresponding rate in the comparator group under the reported analysis. The simple relative interpretation is a 27% lower estimated hazard. It should not be translated into a 27% absolute reduction in the probability of an event.
Why is the confidence interval 0.57–0.94 important?
The point estimate is only one estimate of the treatment effect. The 95% CI shows the range of HR values compatible with the statistical uncertainty represented by the analysis. It provides information about precision that a P-value alone cannot provide.
Why does the P-value not measure effect size?
A P-value describes evidence against a null hypothesis within a specified statistical framework. It depends on both the observed effect and the amount of information available. Two studies can produce different P-values for effects of similar magnitude, and a small P-value does not automatically imply a large treatment effect.
Why is the analysis population important?
The primary analysis is based on the PFS Intent-to-Treat population, defined as the first 550 randomized participants regardless of whether study treatment or the correct study treatment was received. Anchoring the efficacy comparison to randomization helps preserve the comparability established by the randomized design.
Why does censoring matter?
Participants may reach the end of their evaluable follow-up without progression or death. Their observations are censored rather than treated as if they experienced the event at that time. Time-to-event methods such as Kaplan-Meier estimation and the log-rank test are designed to incorporate this partial information.
10. Kaplan-Meier Estimation and PFS
Although the posted primary analysis identifies the log-rank test and hazard ratio, the underlying PFS endpoint is naturally represented using a Kaplan-Meier framework. Kaplan-Meier estimation describes the estimated probability of remaining event-free over time while accounting for right-censored observations.
where di is the number of events at time ti and ni is the number at risk immediately before that time.
The ClinicalTrials.gov record does not provide the event-by-event risk sets needed to reconstruct a Kaplan-Meier curve. It is therefore preferable to explain the method rather than manufacture a curve from the single reported HR and confidence interval.
Kaplan-Meier versus hazard ratio
The two approaches answer related but different questions. A Kaplan-Meier curve describes the estimated event-free probability over time. A hazard ratio summarizes the relative instantaneous event rate between groups under the fitted time-to-event framework. Reporting both can provide a more complete picture when the underlying survival estimates are available.
11. Blinding and Independent Radiology Assessment
COSMIC-313 is registered as triple-masked, and the primary PFS endpoint was determined by a Blinded Independent Radiology Committee. These features are particularly relevant for a radiographic endpoint because assessment of disease progression involves interpretation of imaging findings.
Blinding
Triple masking reduces the opportunity for knowledge of treatment assignment to influence aspects of trial conduct and assessment.
Independent review
BIRC assessment provides an independent framework for evaluating radiographic progression for the primary PFS endpoint.
RECIST v1.1
The registry states that PFS was determined according to Response Evaluation Criteria in Solid Tumors version (RECIST) v1.1.
Time-to-event structure
Radiographic progression and death are incorporated into a single prespecified PFS event definition.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized arm. These figures describe the number affected among the number at risk for each arm.
| Safety measure | Cabozantinib + Nivolumab + Ipilimumab | Placebo + Nivolumab + Ipilimumab |
|---|---|---|
| Serious adverse events | 272/426 | 257/423 |
The ClinicalTrials.gov record reports these safety counts but do not provide a formal statistical comparison for serious adverse events. The figures should therefore be treated as descriptive safety information rather than as a separate hypothesis-tested efficacy-style result.
13. Trial Timeline
Trial start
The registered study start date is June 25, 2019.
Primary completion
The registered primary completion date is January 31, 2022.
Active, not recruiting
The trial is listed as ACTIVE_NOT_RECRUITING in the ClinicalTrials.gov record.
14. What the Primary Result Establishes
The reported statistical analysis provides three closely related pieces of evidence: an estimated relative treatment effect, an uncertainty interval, and a hypothesis-test result.
| Quantity | Reported value | Statistical role |
|---|---|---|
| Hazard ratio | 0.73 | Estimates the relative instantaneous rate of progression or death between the randomized groups. |
| 95% confidence interval | 0.57–0.94 | Describes uncertainty around the estimated hazard ratio. |
| P-value | 0.0131 | Quantifies evidence against the null hypothesis within the reported superiority testing framework. |
| Analysis method | Log-rank test | Compares the time-to-event experience between the randomized groups. |
| Hypothesis type | Superiority | Frames the analysis as a test for improvement rather than non-inferiority. |
Taken together, the posted analysis reports a hazard ratio below 1 with a 95% confidence interval that remains below 1 and a P-value of 0.0131. The statistical interpretation is therefore based on the reported superiority analysis, not on a comparison of unadjusted event percentages.
15. What the Primary Result Does Not Establish
- It does not provide an absolute PFS difference. The registry-reported statistical analysis reports an HR and confidence interval but does not provide median PFS or fixed-time PFS estimates.
- It does not establish individual patient benefit. A population-level hazard ratio cannot predict the outcome of a particular participant.
- It does not imply a constant 27% reduction in risk at every time. The HR is a time-to-event summary rather than a fixed-time risk ratio.
- It does not provide a probability that the treatment is effective. The P-value is not the posterior probability that the alternative hypothesis is true.
- It does not describe all clinical outcomes. The registry-reported primary analysis concerns PFS; the ClinicalTrials.gov record do not include formal statistical analyses for additional efficacy endpoints.
- It does not replace safety assessment. Serious adverse events are reported separately and require their own clinical interpretation.
16. Important Limitations and Interpretation Issues
- Single primary analysis provided: the ClinicalTrials.gov record contains one formal statistical analysis, for BIRC-assessed PFS.
- No median PFS reported: the data provide an HR, 95% CI, and P-value but do not provide median PFS for either treatment group.
- No fixed-time PFS estimates reported: the ClinicalTrials.gov record does not include Kaplan-Meier estimates at particular time points.
- Hazard-ratio interpretation: the HR is a relative time-to-event measure and should not be interpreted as an absolute risk reduction or as a treatment effect experienced identically by every patient.
- Censoring: PFS analysis necessarily depends on how incomplete follow-up is handled. The registry definition explicitly includes censoring in the PFS calculation, but the ClinicalTrials.gov record does not describe the censoring distribution.
- Proportional-hazards consideration: a single HR is easiest to interpret as a summary of relative hazards when the proportional-hazards assumption is reasonable. The ClinicalTrials.gov record does not report a formal assessment of that assumption.
- Analysis population: the posted analysis uses the first 550 randomized participants in the PITT population, rather than simply all 855 enrolled participants.
- Safety denominators: serious adverse-event data use at-risk denominators of 426 and 423, so those figures should not be conflated with overall enrollment.
- Multiplicity: the ClinicalTrials.gov record identifies one primary endpoint analysis but do not provide an alpha-allocation or multiplicity-adjustment strategy. No such procedure is inferred here.
- Interim analysis: the ClinicalTrials.gov record does not describe an interim-analysis schedule or alpha-spending procedure. None is added to this analysis.
- Stratification: the ClinicalTrials.gov record does not specify randomization strata or stratified analysis factors. No stratification factors are inferred.
- Bayesian methods: the registry-reported statistical method is a log-rank test; no Bayesian analysis is reported in the ClinicalTrials.gov record.
17. Why This Trial Matters Statistically
COSMIC-313 is a useful teaching example for understanding how randomized oncology trials translate a time-to-event clinical question into a formal statistical comparison. The endpoint is defined from randomization through progression, death, or censoring; radiographic progression is evaluated by BIRC; the primary analysis uses a PFS Intent-to-Treat population; and the treatment effect is summarized with a hazard ratio alongside a confidence interval and P-value.
| Concept | How it appears in COSMIC-313 |
|---|---|
| Randomization | The trial is registered as randomized with two parallel arms. |
| Blinding | The masking designation is triple. |
| Time-to-event endpoint | Primary endpoint is duration of PFS by BIRC. |
| RECIST v1.1 | PFS is determined according to RECIST v1.1. |
| Intention-to-treat analysis | The posted efficacy analysis uses the PFS Intent-to-Treat population. |
| Kaplan-Meier framework | PFS is a censored time-to-event outcome naturally represented through Kaplan-Meier estimation. |
| Log-rank test | Registry-reported method for the primary comparison. |
| Hazard ratio | Primary effect measure, reported as 0.73. |
| Confidence interval | 95% CI of 0.57–0.94 quantifies uncertainty around the HR. |
| P-value | Reported as 0.0131 for the superiority analysis. |
| Independent assessment | BIRC is specified for the primary PFS endpoint. |
| Safety description | Serious adverse events are reported by treatment arm with affected/at-risk counts. |
18. Statistical Methods Explained in More Depth
Randomization and causal comparison
Randomization is the structural foundation of the treatment comparison. By assigning participants to treatment groups randomly, the design aims to make the groups comparable in expectation with respect to measured and unmeasured baseline factors. The subsequent analysis can therefore attribute differences between randomized groups to treatment assignment more credibly than an observational comparison can.
Why the PITT population is different from an enrollment count
The trial enrolled 855 participants, but the registry-reported primary PFS analysis specifies the first 550 randomized participants as the PITT population. Enrollment describes the overall study population, whereas the PITT definition identifies the participants used for this particular efficacy analysis. Confusing these two quantities would produce an incorrect description of the primary analysis.
Why PFS includes death
If death were excluded from the endpoint, a participant who died before documented radiographic progression could potentially be treated differently from a participant who remained alive without progression. Including death in the definition ensures that death is itself a PFS event. In COSMIC-313, the registry explicitly defines the event as the earlier of radiographic progression per BIRC or death from any cause.
Why BIRC assessment matters
Radiographic progression depends on imaging assessments. A blinded independent committee provides an assessment process designed to reduce the influence of treatment assignment on the determination of progression. This is particularly relevant when progression is the event driving a primary efficacy endpoint.
Why a confidence interval is more informative than a P-value alone
The P-value indicates the degree of evidence against the null hypothesis under the specified test. The confidence interval adds information about the estimated magnitude and precision of the effect. In this trial, the HR of 0.73 is accompanied by a 95% CI of 0.57–0.94, giving a substantially richer description of the estimate than the P-value of 0.0131 alone.
Why a hazard ratio is not a median ratio
An HR compares instantaneous event rates within a time-to-event model. It is not calculated by dividing one group's median PFS by the other group's median PFS. Because the ClinicalTrials.gov record does not report median PFS, no median-based comparison is presented here.
19. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The reported PFS analysis produced an HR of 0.73 with a two-sided 95% CI of 0.57–0.94 and P = 0.0131 using a log-rank test under a superiority hypothesis.
Clinical interpretation
The statistical result describes a lower estimated relative hazard of progression or death in the cabozantinib-containing group. The ClinicalTrials.gov record does not provide median PFS or fixed-time PFS estimates, so absolute duration-based clinical measures cannot be stated here.
Keeping these two levels of interpretation separate is important. The statistical analysis tells us how the randomized groups compare under the prespecified endpoint and test. Clinical interpretation requires consideration of the magnitude of the effect, its precision, absolute outcome measures, safety, patient population, and the broader treatment context. Only some of those elements are contained in the ClinicalTrials.gov record.
20. A Worked Reading of the COSMIC-313 Result
The primary endpoint is duration of PFS by BIRC, defined from randomization to the earlier of radiographic progression or death, with censoring handled according to the registered endpoint definition.
The primary posted analysis uses the PFS Intent-to-Treat population, defined as the first 550 randomized participants regardless of whether any study treatment or the correct study treatment was received.
The registry reports a log-rank test for the primary comparison, with superiority as the hypothesis type.
The reported hazard ratio is 0.73, comparing cabozantinib + nivolumab + ipilimumab with placebo + nivolumab + ipilimumab.
The 95% CI is 0.57–0.94 and the P-value is 0.0131. The CI describes uncertainty around the HR; the P-value addresses the hypothesis test. Neither should be interpreted as an individual patient's probability of benefit.
21. What Is Available — and What Is Not — in the Supplied Results
| Information | Supplied? | How it is handled |
|---|---|---|
| Primary endpoint definition | Yes | Reported using the registry wording and time frame. |
| Primary statistical method | Yes | Log-rank test explained in detail. |
| Primary effect estimate | Yes | HR 0.73 reported exactly. |
| 95% confidence interval | Yes | 0.57–0.94 reported exactly. |
| P-value | Yes | 0.0131 reported exactly. |
| Analysis population | Yes | First 550 randomized participants in the PITT population. |
| Secondary endpoint analyses | No formal analyses reported | No secondary efficacy estimates are added. |
| Median PFS | Not reported | No median PFS is reported on this page. |
| Fixed-time PFS estimates | Not reported | No fixed-time survival probabilities are reported. |
| Subgroup estimates | Not reported | No subgroup results are added. |
| Interim-analysis details | Not reported | No interim schedule or alpha-spending procedure is inferred. |
| Stratification factors | Not reported | No stratification variables are inferred. |
| Bayesian analysis | Not reported | No Bayesian method is attributed to the trial. |
22. Limitations of the Statistical Evidence Presented Here
The most important limitation is the granularity of the ClinicalTrials.gov record. The registry provides a formal primary PFS analysis, but the statistical result is summarized at the hazard-ratio level. Without the underlying event and censoring data, a full reconstruction of the PFS distribution is not possible.
In particular, an HR of 0.73 does not reveal whether the treatment groups had similar or different PFS patterns at specific time points, whether hazards changed substantially over time, or what the median PFS was in either group. Those are separate descriptive questions that require additional survival information.
The analysis population also deserves attention. The trial enrolled 855 participants, while the posted primary PFS analysis is defined using the first 550 randomized participants. This distinction means that the enrollment count should not be substituted for the PITT analysis population when describing the primary efficacy result.
Finally, the ClinicalTrials.gov record does not include a multiplicity plan, interim-monitoring details, randomization strata, missing-data strategy beyond the endpoint's censoring definition, or formal assessment of proportional hazards. Those topics can materially affect detailed interpretation of a time-to-event trial, but they should not be reconstructed from assumptions when the ClinicalTrials.gov record does not report them.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NCT03937219 — COSMIC-313.
- PubMed record: PMID 37163623.
- PubMed record: PMID 35383907.
- PubMed record: PMID 34550754.
- PubMed record: PMID 33058158.
Continue through Clinical Biostats
Explore statistical tutorials, calculators, and additional clinical trial analyses covering the methods used in randomized clinical research.
26. Record Summary
COSMIC-313 provides a clear example of a randomized phase 3 time-to-event analysis. The trial enrolled 855 participants, used a randomized parallel design with triple masking, and evaluated cabozantinib + nivolumab + ipilimumab against placebo + nivolumab + ipilimumab. Its registered primary endpoint was BIRC-assessed duration of PFS up to 32 months, defined from randomization to the earlier of radiographic progression or death. The posted primary analysis used the PFS Intent-to-Treat population of the first 550 randomized participants, a log-rank test, and a hazard ratio as the effect measure.
The reported HR of 0.73, with a 95% CI of 0.57–0.94 and P = 0.0131, describes a lower estimated relative hazard of progression or death in the cabozantinib-containing group under the reported superiority analysis. The most important statistical discipline is to distinguish that relative estimate from absolute PFS measures, individual patient outcomes, and safety outcomes. The ClinicalTrials.gov record does not provide median PFS, fixed-time PFS estimates, subgroup estimates, or additional formal efficacy analyses, so none are inferred here.