This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
METEOR was a randomized, parallel-group, open-label phase 3 trial comparing cabozantinib tablets with everolimus tablets in subjects with metastatic renal cell carcinoma. The registry reports 658 enrolled subjects, two treatment arms, one registered primary time-to-event endpoint, and formal statistical analyses for progression-free survival, overall survival, and objective response rate.
| Feature | METEOR |
|---|---|
| Trial name | METEOR |
| Brief title | A Study of Cabozantinib (XL184) vs Everolimus in Subjects With Metastatic Renal Cell Carcinoma |
| Phase | Phase 3 |
| Condition | Renal Cell Carcinoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 658 |
| Interventions | Cabozantinib tablets (drug); Everolimus (Afinitor) tablets (drug) |
| Lead sponsor | Exelixis |
| Sponsor type | Industry |
| Status | Completed |
| Start | 2013-06 |
| Primary completion | 2015-05-22 |
| ClinicalTrials.gov | NCT01865747 |
2. Clinical Question
The central statistical question was whether randomized assignment to cabozantinib, compared with everolimus, was associated with a difference in progression-free survival in subjects with metastatic renal cell carcinoma. The registry identifies progression-free survival as the single registered primary endpoint and classifies the formal hypothesis as one of superiority.
Population
Subjects with metastatic renal cell carcinoma enrolled in the randomized phase 3 METEOR trial.
Intervention
Cabozantinib (XL184) tablets.
Comparator
Everolimus (Afinitor) tablets.
Primary question
Does cabozantinib produce a superior progression-free survival outcome compared with everolimus?
3. Trial Design
Cabozantinib
- Cabozantinib (XL184) tablets
- Randomized treatment assignment
- Compared with everolimus for PFS, OS, and ORR analyses
Everolimus
- Everolimus (Afinitor) tablets
- Randomized treatment assignment
- Comparator arm for PFS, OS, and ORR analyses
Because this was an open-label randomized comparison, the statistical interpretation depends particularly strongly on how objective time-to-event endpoints were defined and analyzed. Randomization establishes the treatment comparison, while the survival-analysis methods translate follow-up information into an estimate of the difference between treatment groups.
4. Endpoints
| Endpoint | Registry definition / time frame | Type | Analysis |
|---|---|---|---|
| Progression-free Survival (PFS) | PFS is measured from the date of randomization until the date of first documented disease progression or date of death from any cause as determined by the Independent Radiology Committee (IRC) per RECIST 1.1, assessed for up to 17 months.. The primary analysis of PFS is the time from randomization to date of first documented tumor progression as determined by investigator (per RECIST 1.1 criteria) or death due to any cause, whichever occurred first. | Time-to-event | Stratified log-rank test; hazard ratio |
| Overall Survival (OS) | OS was measured from the time of randomization until 320 deaths, approximately 28 months. | Time-to-event | Stratified log-rank test; hazard ratio |
| Objective Response Rate (ORR) | ORR was assessed at 8 weeks post-randomization, every 8 weeks for 12 months, and every 12 weeks until date of disease progression or death, up to May 2015 (approximately 21 months). | Binary | Cochran-Mantel-Haenszel test |
The registry identifies one primary endpoint: progression-free survival. Overall survival and objective response rate are reported as secondary endpoints. The distinction matters because the inferential role of a result depends not only on its numerical value but also on where the endpoint sits in the prespecified trial structure.
5. Statistical Methodology
Stratified log-rank testing
The registry reports a log-rank test for the primary PFS analysis and for the secondary OS analysis. For PFS, the analysis text states that the log-rank test was stratified by the Memorial Sloan-Kettering Cancer Center (MSKCC) group and the number of prior VEGFR TKIs. For OS, the same two stratification concepts are reported: MSKCC risk group and number of prior VEGFR TKIs.
A log-rank test compares the observed pattern of event occurrence between randomized groups across follow-up. In a time-to-event trial, it uses the ordering of event times while accounting for the changing number of participants still at risk. Stratification allows the comparison to be performed within prespecified strata rather than treating all subjects as though the stratification factors were irrelevant.
The log-rank framework asks whether the event experience over follow-up differs systematically between randomized treatment groups, conditional on the analysis structure.
Hazard ratio as the effect measure
The registry reports the hazard ratio (HR) as the effect measure for both PFS and OS. A hazard ratio compares the estimated instantaneous event rates between treatment groups over the analyzed follow-up.
An HR below 1 indicates a lower estimated instantaneous event rate in the cabozantinib group relative to the everolimus group under the analysis model.
Cochran-Mantel-Haenszel analysis
The secondary ORR analysis used the Cochran-Mantel-Haenszel test. This is a categorical-data method that can compare treatment groups while accounting for stratification. The registry identifies the analysis as a superiority test in the ITT population.
Unlike PFS and OS, ORR is binary: a participant either meets the prespecified response definition or does not. That difference in endpoint structure explains why a categorical-data method was used for ORR while a survival-analysis method was used for PFS and OS.
Intention-to-treat analysis
The OS analysis explicitly used the Intent to Treat (ITT) population, consisting of all 658 randomized subjects: 330 assigned to cabozantinib and 328 assigned to everolimus. The ORR analysis also used the ITT population, with the same randomized group sizes. ITT analysis preserves the randomized treatment comparison rather than redefining groups according to treatment received after randomization.
Kaplan-Meier estimation
The registered PFS definition states that a Kaplan-Meier analysis was performed to estimate the median duration. Kaplan-Meier estimation is appropriate for time-to-event data because it can incorporate participants whose event has not occurred by the end of available follow-up.
Here, di represents events at an event time and ni represents the number at risk immediately before that time. The ClinicalTrials.gov record does not report the resulting median PFS, so this page does not introduce a median value.
6. Analysis Populations and Stratification
| Endpoint | Analysis population | Randomized group sizes | Important analysis features |
|---|---|---|---|
| Primary PFS | First 375 randomized subjects | 187 cabozantinib; 188 everolimus | Stratified log-rank test; stratified by MSKCC group and number of prior VEGFR TKIs |
| OS | ITT population | 330 cabozantinib; 328 everolimus | Second interim analysis; cutoff 31 December 2015; stratified log-rank test |
| ORR | ITT population | 330 cabozantinib; 328 everolimus | IRC response assessment per RECIST 1.1; Cochran-Mantel-Haenszel test |
The difference between the PFS and OS analysis populations is an important statistical feature. The prespecified primary PFS analysis was based on the first 375 randomized subjects, whereas the reported OS analysis used all 658 randomized subjects in the ITT population. These are therefore not interchangeable analyses, even though both compare cabozantinib with everolimus.
7. Results: Progression-Free Survival
PFS was the single registered primary endpoint. The primary analysis was based on the first 375 randomized subjects: 187 assigned to cabozantinib and 188 assigned to everolimus. The reported method was a log-rank test, stratified by MSKCC group and number of prior VEGFR TKIs.
Primary PFS treatment effect
95% CI: 0.45–0.74 · P < 0.0001
Cabozantinib vs everolimus · Superiority hypothesis
| Primary PFS element | Reported value |
|---|---|
| Analysis population | First 375 randomized subjects |
| Cabozantinib | 187 subjects |
| Everolimus | 188 subjects |
| Method | Log-rank test |
| Stratification | MSKCC group and number of prior VEGFR TKIs |
| Effect measure | Hazard ratio |
| Hazard ratio | 0.58 |
| 95% confidence interval | 0.45–0.74 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
An HR of 0.58 means that, under the time-to-event analysis, the estimated instantaneous rate of progression or death in the cabozantinib group was approximately 58% of that in the everolimus group. Equivalently, 0.58 corresponds to an estimated 42% lower instantaneous hazard relative to everolimus.
The HR does not mean that 42% of patients avoided progression, that every patient experienced exactly a 42% reduction in risk, or that the probability of progression was reduced by 42% at every individual time point. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.
The 95% CI of 0.45–0.74 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of effects that individual patients experienced.
The P-value of <0.0001 addresses the strength of evidence against the null hypothesis within the specified statistical test. It is not a measure of effect size, clinical importance, or the probability that the treatment is effective.
Because the analysis is based on time-to-event data, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio representation. The ClinicalTrials.gov record identifies the stratified log-rank analysis and HR but do not provide additional model-diagnostic information that would permit a separate assessment of the proportional-hazards assumption.
Why the PFS analysis is statistically distinctive
The primary PFS analysis combines several important design features: randomization, a prespecified subset of the randomized population, a time-to-event endpoint, stratification, a log-rank comparison, and a hazard-ratio effect measure. Each component answers a different statistical question.
Randomization
Creates the treatment groups whose outcomes are compared, reducing the role of measured and unmeasured baseline differences in the treatment comparison.
Time-to-event endpoint
Uses not only whether progression or death occurred, but also when the event occurred and how long participants remained event-free.
Stratification
Accounts for the prespecified MSKCC group and number of prior VEGFR TKIs in the log-rank comparison.
Hazard ratio
Provides a relative measure of the event rate between the two randomized treatment groups.
8. Results: Overall Survival
Overall survival was a secondary endpoint. The registry states that OS was measured from randomization until 320 deaths, approximately 28 months. The second interim analysis used the full ITT population of 658 randomized subjects, with a cutoff date of 31 December 2015.
Secondary OS treatment effect
95% CI: 0.53–0.83 · P = 0.0003
Cabozantinib vs everolimus · Superiority hypothesis
| OS analysis element | Reported value |
|---|---|
| Analysis population | ITT population |
| Cabozantinib | 330 subjects |
| Everolimus | 328 subjects |
| Interim analysis | Second interim analysis |
| Cutoff date | 31 December 2015 |
| Event target / time frame | 320 deaths, approximately 28 months |
| Method | Log-rank test |
| Stratification | MSKCC risk group and number of prior VEGFR TKIs |
| Effect measure | Hazard ratio |
| Hazard ratio | 0.66 |
| 95% confidence interval | 0.53–0.83 |
| P-value | 0.0003 |
An HR of 0.66 means that the estimated instantaneous rate of death in the cabozantinib group was approximately 66% of that in the everolimus group under the reported analysis. Equivalently, 0.66 corresponds to an estimated 34% lower instantaneous hazard relative to everolimus.
This is not the same as saying that 34% fewer patients died, that survival probability increased by 34%, or that each participant had exactly a 34% reduction in the probability of death. The hazard ratio summarizes a relative time-to-event comparison.
The 95% CI of 0.53–0.83 provides a measure of precision around the estimated HR. It indicates that the reported estimate is not a single exact population quantity known without uncertainty; the interval reflects statistical uncertainty under the analysis framework.
The P-value of 0.0003 quantifies evidence against the null hypothesis for the reported test. It does not quantify the magnitude of the treatment effect, the clinical value of the treatment, or the probability that the null hypothesis is true.
The OS analysis was explicitly an interim analysis. Interim analyses require careful interpretation because the timing of an analysis relative to accumulating events can affect statistical inference. The ClinicalTrials.gov record identifies the analysis as a second interim analysis but do not provide an alpha-spending boundary or numerical information-allocation schedule, so none is inferred here.
Why OS and PFS should not be treated as interchangeable
PFS and OS are both time-to-event outcomes, but they define different events. PFS ends at the first documented tumor progression or death, whichever occurs first. OS ends at death from any cause. A participant can therefore experience a PFS event without having experienced an OS event.
This distinction is important when interpreting the two HRs. The PFS HR of 0.58 and OS HR of 0.66 describe different event processes. One should not be interpreted as a direct surrogate for the other solely because both are reported as hazard ratios.
9. Results: Objective Response Rate
ORR was a secondary binary endpoint. The registry states that ORR was assessed at 8 weeks post-randomization, every 8 weeks for 12 months, and every 12 weeks thereafter until the date of disease progression as recorded in the registry wording.
The ORR analysis used the ITT population: all randomized participants, consisting of 330 cabozantinib and 328 everolimus. Response was determined by an Independent Radiology Committee (IRC) according to RECIST 1.1.
Secondary ORR comparison
Cochran-Mantel-Haenszel test · ITT population
Cabozantinib vs everolimus · Superiority hypothesis
| ORR analysis element | Reported value |
|---|---|
| Analysis population | ITT population |
| Cabozantinib | 330 randomized subjects |
| Everolimus | 328 randomized subjects |
| Response assessment | Independent Radiology Committee per RECIST 1.1 |
| Endpoint type | Binary |
| Method | Cochran-Mantel-Haenszel test |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The reported P < 0.0001 indicates strong evidence against the null hypothesis under the reported Cochran-Mantel-Haenszel comparison of ORR. It does not, by itself, tell us the magnitude of the difference in response rates.
That distinction is especially important here because the ClinicalTrials.gov record provides the formal ORR test and its P-value but do not provide the arm-specific response percentages or a confidence interval for the response-rate difference. Therefore, this page does not invent an ORR effect estimate that is absent from the ClinicalTrials.gov record.
ORR also answers a different question from PFS. A response endpoint classifies tumor response, whereas PFS incorporates the timing of progression or death. Consequently, the ORR P-value should not be interpreted as evidence for a particular magnitude of PFS benefit.
10. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm using affected participants over participants at risk. These data should be read separately from the efficacy analyses because safety and efficacy answer different questions and may use different analysis populations or definitions.
| Safety measure | Cabozantinib (XL184) | Everolimus (Afinitor) |
|---|---|---|
| Serious adverse events | 131 / 331 affected / at risk | 139 / 322 affected / at risk |
The registry-derived ClinicalTrials.gov record do not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events. Accordingly, the figures are presented descriptively rather than converted into an inferred hypothesis test.
11. Interim Analysis and Statistical Inference
The OS analysis was explicitly identified as the second interim analysis, with a cutoff date of 31 December 2015 and an endpoint time frame defined around 320 deaths, approximately 28 months.
Interim analysis is a design feature rather than merely a date label. When investigators examine accumulating outcome data before the trial is complete, the statistical design must account for the fact that multiple opportunities to evaluate the evidence can alter the probability of observing a nominally significant result by chance.
What the registry establishes
The ClinicalTrials.gov record establishes that the reported OS result came from a second interim analysis and that the analysis included the full ITT population.
What the registry data do not establish
The ClinicalTrials.gov record does not provide an alpha-spending function, interim boundary, or numerical allocation of type I error. These details are therefore not reconstructed here.
This distinction is important because a nominal P-value such as 0.0003 should be interpreted in the context of the prespecified interim-analysis framework. The existence of an interim analysis does not by itself invalidate the result; rather, it means that the timing and rules governing interim inference are part of the statistical design.
12. Stratified Analysis
Both time-to-event analyses used stratification. The PFS log-rank test was stratified by MSKCC group and number of prior VEGFR TKIs. The OS log-rank test was stratified by MSKCC risk group and number of prior VEGFR TKIs.
Stratification allows the survival comparison to account for specified baseline or disease-history categories rather than ignoring them completely in the primary test.
Stratification does not mean that the trial produced a separate primary conclusion for every MSKCC category or every level of prior VEGFR TKI exposure. It means that these factors were incorporated into the reported comparison. The ClinicalTrials.gov record does not provide subgroup-specific hazard ratios or interaction tests, so no subgroup treatment-effect conclusions are added.
13. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS is a time-to-event endpoint. Some participants experience progression or death during follow-up, while others may not experience the event during the period in which they are observed. The log-rank test is designed to compare the event-time distributions of two groups while incorporating the timing of observed events and censoring.
For METEOR, the reported PFS comparison was additionally stratified by MSKCC group and number of prior VEGFR TKIs. The resulting P-value therefore belongs to the specified stratified comparison, not to an unadjusted comparison invented after the fact.
What does an HR of 0.58 mean?
An HR of 0.58 means that the estimated instantaneous rate of progression or death was 0.58 times that of the comparator under the reported time-to-event analysis. Expressed as a relative difference, 1 − 0.58 = 0.42, so the estimate corresponds to a 42% lower instantaneous hazard relative to everolimus.
It does not mean a 42% reduction in every patient's probability of progression, and it does not provide an absolute percentage of patients who benefited.
Why does the confidence interval matter?
The PFS HR was estimated as 0.58 with a 95% CI of 0.45–0.74. The interval conveys statistical uncertainty around the point estimate. A narrower interval generally represents greater precision than a wider interval, although precision also depends on the underlying design and amount of information.
The confidence interval should not be interpreted as the range of individual patient responses. It concerns uncertainty in the estimated treatment effect at the population-analysis level.
Why doesn't the P-value measure effect size?
A P-value measures how compatible the observed data are with a specified null hypothesis under the statistical test. It depends on both the magnitude of the observed difference and the amount of statistical information available. Consequently, a small P-value does not by itself establish that an effect is large or clinically important.
For METEOR, the PFS result provides both pieces of information separately: the HR 0.58 describes the estimated relative treatment effect, while P < 0.0001 describes the strength of evidence against the tested null hypothesis.
Why use the ITT population for OS and ORR?
The ITT principle retains participants in the groups to which they were randomized. For METEOR, the OS and ORR analyses used all 658 randomized subjects, with 330 assigned to cabozantinib and 328 assigned to everolimus.
This preserves the treatment comparison generated by randomization. It also means that the analysis is not restricted only to participants who completed treatment or followed an ideal treatment course.
Why was a Cochran-Mantel-Haenszel test used for ORR?
ORR is a binary endpoint rather than a time-to-event endpoint. The Cochran-Mantel-Haenszel test is a categorical-data method that can compare groups while accounting for stratification. In METEOR, the registry reports this method for the ITT ORR analysis and reports a superiority hypothesis.
Why should PFS and OS results not be combined into one number?
PFS and OS measure different events. PFS counts the first documented progression or death, whereas OS counts death from any cause. Their hazard ratios therefore summarize different underlying event processes. Reporting both provides complementary information rather than two measurements of exactly the same outcome.
14. Reading the Hazard Ratios Together
| Endpoint | HR | 95% CI | P-value | Statistical role |
|---|---|---|---|---|
| PFS | 0.58 | 0.45–0.74 | <0.0001 | Primary endpoint |
| OS | 0.66 | 0.53–0.83 | 0.0003 | Secondary endpoint; second interim analysis |
The two HRs are both below 1, but they should be interpreted independently because their event definitions and analysis roles differ. The PFS estimate is based on the first 375 randomized subjects, while the OS estimate is based on all 658 randomized subjects. The OS analysis also has an explicitly identified second-interim-analysis context.
The most defensible summary is not simply that "both hazard ratios were below 1." A fuller reading identifies what event was analyzed, which population was analyzed, how the comparison was performed, what the HR estimates, how precise the estimates are, and what inferential framework generated the P-values.
For PFS, the HR of 0.58 corresponds to a 42% lower estimated instantaneous hazard relative to everolimus. For OS, the HR of 0.66 corresponds to a 34% lower estimated instantaneous hazard. These are relative hazard statements, not absolute survival differences.
15. Censoring and Time-to-Event Interpretation
Time-to-event analysis differs from a simple comparison of proportions because not every participant necessarily has the event observed during the relevant follow-up. Kaplan-Meier methods allow information from participants who remain event-free at the end of their observable follow-up to contribute until their censoring time.
This is one reason a hazard ratio cannot be translated directly into an absolute percentage reduction in events. The HR summarizes the relative event rate over time within a survival-analysis framework, whereas an absolute event probability would require a specified time point and corresponding survival estimates.
Censoring
A participant without an observed event can still contribute follow-up information until the censoring point.
Event timing
An event occurring earlier and an event occurring later are not treated as equivalent observations in a time-to-event analysis.
Hazard
The hazard describes an instantaneous event rate conditional on remaining event-free up to a particular time.
Survival probability
A survival probability describes the probability of remaining event-free through a specified time point and is conceptually different from the HR.
16. Limitations and Interpretation Issues
- Different analysis populations: the primary PFS analysis used the first 375 randomized subjects, whereas the OS and ORR analyses used the full 658-subject ITT population. These results should not be treated as though they came from one identical analysis set.
- Interim OS analysis: the reported OS result came from a second interim analysis with a cutoff date of 31 December 2015. Interim inference depends on the prespecified monitoring framework, and the ClinicalTrials.gov record does not provide the numerical alpha-spending details.
- Hazard-ratio interpretation: an HR is a relative time-to-event measure, not an absolute risk difference or the percentage of participants who benefit.
- Proportional-hazards considerations: a single HR is most straightforward to interpret when the relative hazard is reasonably represented by a common ratio over time. The ClinicalTrials.gov record does not provide diagnostics allowing a separate assessment of this assumption.
- Endpoint differences: PFS, OS, and ORR are different outcomes and require different statistical interpretations.
- ORR effect size not reported: the registry data provide the ORR analysis method and P-value but do not provide arm-specific ORR percentages or a confidence interval in the registry-reported statistical-analyses record. No response-rate estimate is inferred.
- Safety comparison: serious adverse events are reported descriptively as affected participants over participants at risk. The ClinicalTrials.gov record does not provide a formal statistical comparison.
- Open-label design: the registry classifies masking as none. This is a design characteristic that should be kept in mind when interpreting outcomes, particularly endpoints that can involve assessment decisions, although the ClinicalTrials.gov record does not quantify any resulting bias.
17. Why This Trial Matters Statistically
METEOR is a useful statistical teaching case because its registry record brings together several foundational clinical-trial concepts without requiring them to be treated as interchangeable. It has randomized treatment allocation, a time-to-event primary endpoint, stratified survival analysis, a hazard-ratio effect measure, an ITT population for secondary analyses, a binary response endpoint, a Cochran-Mantel-Haenszel comparison, and an interim OS analysis.
| Concept | How it appears in METEOR |
|---|---|
| Randomization | Randomized allocation to cabozantinib or everolimus. |
| Parallel-group design | Two treatment arms were evaluated in parallel. |
| Time-to-event endpoint | PFS was the single registered primary endpoint. |
| Kaplan-Meier estimation | The registered PFS definition states that Kaplan-Meier analysis was used to estimate median duration. |
| Hazard ratio | Reported for both PFS and OS. |
| Log-rank test | Used for the primary PFS and secondary OS comparisons. |
| Stratified analysis | PFS and OS log-rank tests were stratified by MSKCC group/risk group and number of prior VEGFR TKIs. |
| Intention-to-treat analysis | Used for OS and ORR in all 658 randomized subjects. |
| Binary endpoint analysis | ORR was analyzed with a Cochran-Mantel-Haenszel test. |
| Interim analysis | The reported OS analysis was the second interim analysis with a 31 December 2015 cutoff. |
| Confidence intervals | 95% two-sided confidence intervals were reported for the PFS and OS HRs. |
| Superiority testing | The reported PFS, OS, and ORR analyses used superiority hypotheses. |
The most important lesson is methodological: a clinical-trial result is not just a number. The meaning of a number depends on the endpoint definition, analysis population, statistical method, stratification, timing of the analysis, effect measure, and uncertainty interval surrounding the estimate.
18. Related Tutorials
Learn more about the methods used in this trial:
19. Related Calculators
20. Sources
- ClinicalTrials.gov: NCT01865747 — METEOR.
- PubMed: PMID 37656041.
- PubMed: PMID 34921022.
- PubMed: PMID 34364385.
- PubMed: PMID 31887537.
- PubMed: PMID 31371341.
Continue through the Clinical Biostats statistical learning pathway
Explore the underlying survival-analysis, categorical-data, confidence-interval, and clinical-trial methods used to interpret randomized evidence.
21. Record Summary
METEOR provides a compact example of how several clinical-trial statistical methods fit together. The trial randomized 658 subjects in a two-arm, parallel, unmasked phase 3 design comparing cabozantinib with everolimus. Its registered primary endpoint was PFS, analyzed in the first 375 randomized subjects using a stratified log-rank test, producing an HR of 0.58 with a 95% CI of 0.45–0.74 and P < 0.0001. The secondary OS analysis used the full ITT population at the second interim analysis and produced an HR of 0.66 with a 95% CI of 0.53–0.83 and P = 0.0003. ORR was analyzed as a binary endpoint in the ITT population using a Cochran-Mantel-Haenszel test, with P < 0.0001.
The statistical interpretation should remain tied to the specific analysis that generated each estimate. PFS and OS measure different events; the primary PFS and secondary OS analyses used different populations; the OS result came from an interim analysis; and the ORR record supplies a P-value without an arm-specific response estimate in the ClinicalTrials.gov record. Keeping those distinctions visible is essential for a precise reading of the evidence.