This page separates reported trial results from statistical interpretation. Numerical trial results and design facts are restricted to the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
COMET-1 was a randomized, quadruple-masked, parallel-group phase 3 treatment trial comparing cabozantinib with prednisone in men with metastatic castration-resistant prostate cancer previously treated with docetaxel and abiraterone or MDV3100. The registered primary endpoint was overall survival, a time-to-event outcome analyzed after 614 events.
| Feature | COMET-1 |
|---|---|
| Phase | Phase 3 |
| Population | Men with metastatic castration-resistant prostate cancer previously treated with docetaxel and abiraterone or MDV3100 |
| Design | Randomized, parallel-group, quadruple-masked |
| Allocation | Randomized |
| Primary endpoint | Overall survival (OS) |
| Primary endpoint type | Time-to-event |
| Primary hypothesis | Superiority |
| Enrollment | 1028 |
| Study start | 2012-07 |
| Primary completion | 2014-09 |
| Sponsor | Exelixis |
| ClinicalTrials.gov | NCT01605227 |
2. Clinical Question
The statistical question was whether cabozantinib produced a superior overall-survival outcome compared with prednisone in men with metastatic castration-resistant prostate cancer who had previously received docetaxel and abiraterone or MDV3100.
Population
Men with metastatic castration-resistant prostate cancer previously treated with docetaxel and abiraterone or MDV3100.
Intervention
Cabozantinib.
Comparator
Prednisone.
Primary question
Does cabozantinib demonstrate superiority over prednisone for overall survival?
3. Trial Design
The registry describes COMET-1 as a randomized, parallel-group, quadruple-masked phase 3 treatment trial with two intervention arms and 1028 enrolled participants. The primary efficacy analysis used the intention-to-treat population, preserving randomized treatment assignment for the primary comparison.
Cabozantinib
- Cabozantinib
- Included in the ITT efficacy population
- Serious adverse events: 420/681
Prednisone
- Prednisone
- Included in the ITT efficacy population
- Serious adverse events: 181/342
4. Randomization, Stratification, and Analysis Populations
The primary OS analysis was performed in the Intent to Treat (ITT) population, consisting of all 1028 randomized subjects: 682 assigned to cabozantinib and 346 assigned to prednisone.
The registry reports that the log-rank analyses were stratified by three factors:
- Prior cabazitaxel
- Baseline pain severity
- Baseline ECOG performance status
| Analysis feature | Registry-supported description |
|---|---|
| Primary efficacy population | ITT population |
| ITT size | 1028 randomized subjects |
| Cabozantinib ITT group | 682 |
| Prednisone ITT group | 346 |
| Survival stratification | Prior cabazitaxel, baseline pain severity, and baseline ECOG performance status |
| CMH stratification | Prior cabazitaxel, baseline pain severity, and baseline ECOG performance status |
Stratification is important because the analysis compares treatment groups while accounting for prespecified factors that may be associated with the outcome. It does not change the randomized treatment assignment; instead, it incorporates the design strata into the statistical comparison.
5. Endpoints
| Endpoint | Registry definition / time frame | Role |
|---|---|---|
| Overall Survival (OS) | The primary analysis of OS is defined as the time from randomization to death due to any cause. Participants that had not died or were permanently lost to follow-up were censored at the last known date alive. Median OS was calculated using Kaplan-Meier estimates. Analysis for OS was performed after 614 events had occurred. OS was measured from the time of randomization until 614 events, approximately 24 months after study start. | Primary |
| Bone Scan Response (BSR) | BSR was measured at the end of Week 12 as determined by the IRF. | Secondary |
| Progression-free Survival (PFS) | Duration of PFS was defined as time from the date of randomization to earlier of date of radiographic progression (bone/andor soft tissue) according to the investigator's assessment or death, assessed for up to approximately 24 months | Other pre-specified |
6. Statistical Methodology
Kaplan-Meier estimation
The OS registry definition states that median OS was calculated using Kaplan-Meier estimates. Kaplan-Meier estimation is appropriate for time-to-event data because participants may have different lengths of follow-up and some participants may not experience the event during observation.
Here, di is the number of events at event time ti, and ni is the number of participants at risk immediately before that time.
The important statistical feature is that censoring does not automatically mean that a participant is treated as having experienced the event. For OS, participants who had not died or were permanently lost to follow-up were censored at the last known date alive.
Stratified log-rank test
The primary OS analysis used a log-rank test, and the registry specifically states that it was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. The PFS analysis also used a stratified log-rank test with the same three stratification factors.
The log-rank framework compares the observed and expected numbers of events between treatment groups across follow-up. A stratified version performs this comparison within the prespecified strata and combines the resulting evidence rather than treating the strata as irrelevant to the analysis.
Cox proportional-hazards effect measure
The primary OS effect measure was reported as a Cox Proportional Hazard, normalized here as a hazard ratio. The estimated OS hazard ratio was 0.90 with a two-sided 95% confidence interval of 0.76 to 1.06.
A hazard ratio is a relative time-to-event measure. It is not an absolute survival probability, a median survival time, or the percentage of participants who benefit.
Stratified Cochran-Mantel-Haenszel analysis
Bone Scan Response at Week 12 was analyzed using the Cochran-Mantel-Haenszel (CMH) test. The registry states that this analysis was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status.
For a binary endpoint such as BSR, a CMH analysis provides a way to compare treatment groups while accounting for prespecified categorical strata. This is conceptually different from the log-rank test: the CMH framework is used here for a binary Week 12 outcome, whereas the log-rank framework is used for time-to-event outcomes.
7. Primary Result: Overall Survival
The primary endpoint was overall survival, measured from randomization until 614 events, approximately 24 months after study start. The analysis was performed in the ITT population of 1028 randomized subjects.
Overall survival hazard ratio
95% CI: 0.76–1.06 · P = 0.213
Analysis: stratified log-rank test; effect measure reported as a Cox proportional hazard.
| Primary endpoint | Cabozantinib | Prednisone | Effect estimate | P-value |
|---|---|---|---|---|
| Overall Survival | 682 ITT subjects | 346 ITT subjects | HR 0.90 95% CI 0.76–1.06 |
0.213 |
The estimated OS hazard ratio of 0.90 means that the fitted relative hazard of death for cabozantinib versus prednisone was estimated at 0.90. Expressed as a simple derived interpretation, this corresponds to an estimated hazard that was approximately 10% lower in the cabozantinib group relative to the prednisone group under the model.
The HR does not mean that 10% fewer participants died, that survival probability was increased by 10%, or that each individual participant experienced a 10% reduction in risk. A hazard ratio summarizes a relative event-rate comparison over the analyzed time-to-event framework.
The two-sided 95% confidence interval of 0.76 to 1.06 describes uncertainty around the estimated hazard ratio. Because the interval extends below and above 1, the estimate is compatible with both a lower and a higher hazard under the stated statistical framework. The interval should not be interpreted as a range containing the treatment effect for individual patients.
The P-value of 0.213 addresses evidence against the null hypothesis in the prespecified superiority comparison. It does not measure the size or clinical importance of the treatment effect. A P-value should therefore be read together with the hazard ratio and its confidence interval rather than used as a substitute for them.
The analysis was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. Interpretation also depends on the time-to-event framework, including censoring and the assumptions underlying a Cox proportional-hazards effect measure. The ClinicalTrials.gov record does not provide enough information to assess the proportional-hazards assumption directly.
8. Secondary Result: Bone Scan Response
Bone Scan Response was measured at the end of Week 12 as determined by the IRF. The analysis used the ITT population, consisting of 682 cabozantinib subjects and 346 prednisone subjects.
Cochran-Mantel-Haenszel comparison
Cabozantinib vs prednisone · Stratified CMH test
Strata: prior cabazitaxel, baseline pain severity, and baseline ECOG performance status.
The registry reports a P-value of <0.001 for the stratified CMH comparison of Bone Scan Response at Week 12. This provides statistical evidence of a difference between the randomized treatment groups under the specified CMH analysis.
The ClinicalTrials.gov record does not provide the BSR percentages, an odds ratio, a risk ratio, an absolute risk difference, or a confidence interval. Therefore, the result should not be converted into a treatment-effect magnitude that is not reported in the ClinicalTrials.gov record.
The P-value itself is not a measure of how large the difference in BSR was. A formal effect estimate with its confidence interval would be needed to quantify the magnitude and precision of the between-group difference.
The CMH analysis was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. This is important because the reported P-value is tied to that prespecified stratified comparison rather than to an unadjusted comparison that ignores the analysis strata.
9. Other Pre-Specified Result: Progression-Free Survival
Progression-free survival was listed as an other pre-specified outcome measure. The registry analysis used the ITT population of 1028 randomized subjects, with a data cutoff date of 07 July 2014.
Progression-free survival hazard ratio
95% CI: 0.40–0.57 · P < 0.001
Analysis: stratified log-rank test; effect measure: hazard ratio.
| Endpoint | Analysis population | Effect estimate | P-value |
|---|---|---|---|
| Progression-free Survival | ITT: 1028 randomized subjects 682 cabozantinib; 346 prednisone |
HR 0.48 95% CI 0.40–0.57 |
<0.001 |
The PFS hazard ratio of 0.48 means that the estimated instantaneous hazard of the PFS event for cabozantinib relative to prednisone was 0.48 under the fitted time-to-event analysis. As a simple derived statement, this corresponds to an estimated hazard approximately 52% lower with cabozantinib relative to prednisone.
This does not mean that 52% of participants avoided progression, that PFS probability increased by 52%, or that every participant experienced the same relative reduction. The hazard ratio is a model-based relative measure of the event process.
The two-sided 95% confidence interval of 0.40 to 0.57 describes the statistical precision of the estimated hazard ratio. The entire interval lies below 1, so the estimated treatment effect is consistently on the lower-hazard side within this confidence interval. The interval nevertheless does not describe individual-patient outcomes.
The P-value of <0.001 quantifies evidence against the relevant null hypothesis within the reported statistical framework. It does not tell us that the HR is “48% effective,” nor does it measure clinical magnitude. The HR and confidence interval are needed to describe the estimated effect itself.
The PFS analysis used the stratified log-rank test, with stratification by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. Because PFS is a time-to-event endpoint, censoring and the interpretation of the hazard ratio remain central to understanding the result.
10. Comparing the Three Reported Analyses
The registry provides three statistical analyses spanning two time-to-event endpoints and one binary endpoint. The methods match the structure of the outcomes: log-rank testing and hazard ratios for OS and PFS, and a CMH test for Week 12 BSR.
| Outcome | Role | Endpoint type | Method | Effect measure | P-value |
|---|---|---|---|---|---|
| Overall Survival | Primary | Time-to-event | Log-rank | Hazard ratio 0.90 95% CI 0.76–1.06 |
0.213 |
| Bone Scan Response | Secondary | Binary | Cochran-Mantel-Haenszel | Not reported in the ClinicalTrials.gov record | <0.001 |
| Progression-free Survival | Other pre-specified | Time-to-event | Log-rank | Hazard ratio 0.48 95% CI 0.40–0.57 |
<0.001 |
This distinction matters statistically. A P-value from a CMH analysis cannot be interpreted as though it were a hazard-ratio result, and a hazard ratio cannot be interpreted as though it were a simple binary response proportion. Each estimate belongs to a particular endpoint definition, analysis population, and statistical model.
11. What a Hazard Ratio of 0.90 Means
An HR of 0.90 is a ratio of estimated hazards. Because the value is below 1, the fitted analysis estimates a lower instantaneous event rate in the cabozantinib group than in the prednisone group.
It is not a 10-percentage-point difference in survival, not a 10% increase in the probability of being alive, and not a statement that every participant's personal risk was reduced by exactly 10%.
The 95% CI of 0.76–1.06 is more informative about precision than the point estimate alone. It shows that the estimate has uncertainty extending across the value 1, which is the conventional no-hazard-difference reference value.
The P-value of 0.213 is a statement about statistical evidence under the specified hypothesis test. It is not a probability that the treatment has no effect and is not a measure of how clinically important the HR is.
12. Why the PFS Result Looks Different From the OS Result
COMET-1 illustrates why different clinical endpoints can produce different statistical estimates even when they compare the same randomized treatment groups.
Because the event definitions differ, the hazard ratio for one endpoint should not be expected to equal the hazard ratio for another endpoint. The OS HR of 0.90 and the PFS HR of 0.48 are therefore separate statistical quantities answering different questions.
OS
Measures time from randomization to death due to any cause. The primary analysis occurred after 614 events.
PFS
Measures time from randomization according to the registered progression-free survival definition. The registry-reported analysis used a 07 July 2014 data cutoff.
Different event processes
Different endpoints can produce different hazard ratios because the events being modeled are not the same.
Different interpretation
A PFS hazard ratio should not be presented as though it were an OS hazard ratio or an absolute survival difference.
13. Why Stratification Was Used
Both the survival analyses and the CMH analysis were stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. The same factors therefore appear repeatedly in the statistical methodology of the trial.
Stratification is particularly useful in a randomized trial when the investigators have prespecified clinically relevant factors that may influence outcomes. Rather than ignoring these factors, the analysis incorporates the treatment comparison within the strata and then combines the evidence.
The purpose is not to create a new treatment assignment or to turn observational associations into randomized evidence. The randomized comparison remains the foundation; stratification specifies how the comparison is analyzed.
The registry does not provide enough information in the ClinicalTrials.gov record to quantify how much each individual stratification factor contributed to the final estimates. Therefore, the appropriate interpretation is methodological rather than attributing a particular numerical effect to any one factor.
14. Intention-to-Treat Analysis
The ITT population included 1028 randomized subjects: 682 in the cabozantinib group and 346 in the prednisone group. This population was used for the OS, BSR, and PFS analyses reported in the ClinicalTrials.gov record.
The central principle of ITT is that participants are analyzed according to the treatment group to which they were randomized. This preserves the treatment comparison generated by randomization rather than redefining groups according to treatment received after randomization.
Why it matters
Randomization creates the basis for comparing treatment groups. ITT maintains that randomized comparison in the efficacy analysis.
What ITT does not do
ITT does not remove censoring, eliminate missing information, or guarantee identical follow-up in every participant.
15. Safety
The ClinicalTrials.gov record reports serious adverse events by treatment arm using affected participants over participants at risk.
| Safety measure | Cabozantinib | Prednisone |
|---|---|---|
| Serious adverse events | 420/681 | 181/342 |
These figures should be interpreted as the reported counts over the corresponding at-risk denominators. They should not be converted into percentages unless that calculation is explicitly desired and permitted by the reporting framework; the ClinicalTrials.gov record instruct that numbers be reported exactly as given.
The safety data describe the number of participants with serious adverse events relative to the reported number at risk in each arm. The denominators differ from the ITT efficacy denominators of 682 and 346, which is an important distinction when comparing efficacy and safety populations.
A safety comparison should not be reduced to a single P-value or treated as another efficacy endpoint. Serious adverse events are an exposure-related safety outcome, and their interpretation depends on the definition, observation period, and population represented by the denominator.
16. Statistical Methods Explained
Why was a log-rank test used for OS?
OS is a time-to-event endpoint: participants can die at different times, and some participants may be censored rather than observed to death. The log-rank test is designed to compare survival experience between randomized groups while using the timing of events throughout follow-up.
What does the OS hazard ratio of 0.90 mean?
It means the estimated hazard of death for cabozantinib relative to prednisone was 0.90 under the reported Cox proportional-hazards framework. The simple derived interpretation is an approximately 10% lower estimated hazard, not a 10-percentage-point improvement in survival.
Why is the confidence interval important?
The 95% CI of 0.76–1.06 communicates uncertainty around the OS estimate. The point estimate alone cannot show how precisely the treatment effect was estimated. The interval also spans 1, so the estimate is not separated from the conventional no-difference value within this confidence interval.
Why doesn't the P-value measure effect size?
The P-value evaluates statistical evidence against a null hypothesis under the specified testing framework. It is influenced by the amount of information in the analysis and does not directly quantify the magnitude of an effect. Effect size is described by measures such as the hazard ratio, while precision is described by the confidence interval.
Why was a Cochran-Mantel-Haenszel test used for Bone Scan Response?
BSR is a binary endpoint measured at Week 12, unlike OS and PFS, which are time-to-event outcomes. The CMH framework allows the treatment comparison to account for the prespecified stratification factors rather than treating the analysis as an unstratified binary comparison.
Why are the OS and PFS hazard ratios not interchangeable?
OS and PFS define different events. OS concerns death due to any cause, whereas PFS uses a progression-related time-to-event definition. Consequently, the hazard ratio for PFS describes a different event process from the hazard ratio for OS.
What does censoring mean in the OS analysis?
The registry states that participants who had not died or were permanently lost to follow-up were censored at the last known date alive. Censoring means that the analysis uses the participant's observed follow-up up to that point rather than treating the participant as having experienced the event.
17. Censoring and the Time-to-Event Framework
The OS endpoint explicitly specifies censoring for participants who had not died or were permanently lost to follow-up. This is a fundamental feature of survival analysis because clinical trials rarely observe every participant from randomization through the event of interest.
| Component | COMET-1 registry definition | Statistical implication |
|---|---|---|
| Time origin | Randomization | Participants enter the OS clock at randomization. |
| Event | Death due to any cause | OS is an all-cause mortality endpoint. |
| Censoring | Last known date alive for participants who had not died or were permanently lost to follow-up | Observed follow-up contributes information up to the censoring time. |
| Primary event target | 614 events | The OS analysis was performed after the specified event count had occurred. |
The distinction between an event and censoring is important. A censored participant is not counted as having died at the censoring time. Instead, the available survival information is retained up to the last known time alive.
18. Primary Endpoint Timing and Event-Driven Analysis
The registered OS endpoint was measured until 614 events, approximately 24 months after study start. The registry states that the OS analysis was performed after 614 events had occurred.
Study start
COMET-1 began according to the registry.
Primary OS event target
The registered OS analysis was defined around 614 events.
Primary completion
The registry reports primary completion in 2014-09.
PFS data cutoff
The registry-reported PFS statistical analysis identifies this data cutoff date.
An event-driven endpoint is analyzed when the prespecified amount of event information has accumulated rather than simply when a fixed calendar date arrives. For survival outcomes, the number of observed events can determine the statistical information available for estimating and comparing hazards.
19. Interpreting the Confidence Intervals
Confidence intervals are especially useful for avoiding an overly narrow reading of a single point estimate.
OS
HR 0.90 with a two-sided 95% CI of 0.76–1.06. The interval crosses 1.
PFS
HR 0.48 with a two-sided 95% CI of 0.40–0.57. The interval lies below 1.
Precision is endpoint-specific
The width of each interval reflects uncertainty for that particular endpoint and analysis.
Not individual prediction intervals
A confidence interval around a hazard ratio does not describe the range of outcomes an individual participant will experience.
The difference between the OS and PFS intervals should not be interpreted simply as a difference in “certainty” without considering the endpoint, number of events, censoring, and statistical model. Confidence intervals describe uncertainty around the corresponding estimated parameter.
20. Multiplicity and the Reported P-values
The ClinicalTrials.gov record identifies the primary hypothesis type as superiority and identify one registered primary endpoint: OS. The registry also reports secondary BSR and other pre-specified PFS analyses.
| Outcome | Role | Hypothesis type | Reported statistical evidence |
|---|---|---|---|
| Overall Survival | Primary | Superiority | HR 0.90; 95% CI 0.76–1.06; P = 0.213 |
| Bone Scan Response | Secondary | Superiority | CMH P < 0.001 |
| Progression-free Survival | Other pre-specified | Superiority | HR 0.48; 95% CI 0.40–0.57; P < 0.001 |
The ClinicalTrials.gov record does not specify an alpha-spending scheme, a formal multiplicity-adjustment procedure, or an endpoint hierarchy beyond the registered roles shown above. Accordingly, this page does not infer a particular familywise-error procedure or assign a confirmatory status to analyses beyond the registry's classifications.
21. Proportional-Hazards Considerations
The OS effect measure was reported as a Cox proportional-hazards measure, and the PFS effect measure was reported as a hazard ratio. A Cox hazard ratio is most naturally interpreted as a relative hazard under a proportional-hazards framework.
The ClinicalTrials.gov record does not report a diagnostic assessment of the proportional-hazards assumption, nor do they provide time-varying hazard ratios or alternative estimands. Therefore, the appropriate educational interpretation is to recognize the assumption rather than claim that it was or was not satisfied.
If the relative hazard changes substantially over time, a single hazard ratio can become a less complete description of the treatment effect. The ClinicalTrials.gov record does not allow a direct assessment of this issue for COMET-1.
22. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The primary OS analysis estimated a hazard ratio of 0.90 with a two-sided 95% CI of 0.76–1.06 and P = 0.213. The PFS analysis estimated a hazard ratio of 0.48 with a two-sided 95% CI of 0.40–0.57 and P < 0.001. BSR was analyzed with a stratified CMH test with P < 0.001.
Clinical interpretation
The statistical results describe differences in specified trial endpoints; they do not by themselves summarize every aspect of clinical benefit, safety, patient experience, or treatment choice. The ClinicalTrials.gov record does not provide the additional outcome measures needed for a broader clinical assessment.
This distinction is important because a statistically strong result for one endpoint does not automatically determine the interpretation of another endpoint. COMET-1 provides a clear example: PFS and BSR show reported statistical evidence of between-group differences, while the primary OS estimate has a confidence interval that crosses 1 and a P-value of 0.213.
23. Limitations
- Limited registry detail: the ClinicalTrials.gov record does not provide baseline characteristics, median OS, median PFS, response percentages, or subgroup estimates.
- Hazard-ratio assumptions: the ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption.
- Censoring: OS interpretation depends on the registry's censoring rules and the validity of the assumptions underlying survival analysis.
- Analysis populations: efficacy analyses use the ITT population, whereas the serious-adverse-event denominators reported in the registry are 681 and 342, respectively. These populations should not be conflated.
- Multiplicity: the ClinicalTrials.gov record does not specify a complete multiplicity-adjustment strategy across the reported endpoints.
- Effect estimates for BSR: the registry reports a CMH P-value but does not provide a BSR effect estimate or confidence interval in the ClinicalTrials.gov record.
- Generalizability: the trial population is specifically defined by metastatic castration-resistant prostate cancer and prior treatment with docetaxel and abiraterone or MDV3100; applicability outside that population cannot be inferred from the registry data alone.
24. Why This Trial Matters Statistically
COMET-1 is a useful teaching case because the ClinicalTrials.gov record brings together several core principles of clinical-trial biostatistics without requiring the analyst to treat every endpoint as though it were the same type of outcome.
| Concept | How it appears in COMET-1 |
|---|---|
| Randomization | 1028 subjects were randomized to two parallel treatment arms. |
| Quadruple masking | The registry identifies the trial as quadruple-masked. |
| Intention-to-treat | The OS, BSR, and PFS analyses use the ITT population of 1028 randomized subjects. |
| Kaplan-Meier estimation | Median OS was calculated using Kaplan-Meier estimates. |
| Time-to-event analysis | OS and PFS are analyzed as time-to-event endpoints. |
| Log-rank test | OS and PFS were analyzed using log-rank tests. |
| Stratified analysis | Survival and CMH analyses were stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. |
| Hazard ratio | OS HR 0.90; PFS HR 0.48. |
| Confidence interval | Two-sided 95% CIs were reported for the OS and PFS hazard ratios. |
| Cochran-Mantel-Haenszel test | Used for Week 12 Bone Scan Response. |
| Event-driven analysis | The primary OS analysis was performed after 614 events. |
| Safety denominators | Serious adverse events were reported as 420/681 and 181/342. |
The most instructive feature is the alignment between endpoint type and statistical method. OS and PFS require methods that account for event timing and censoring, whereas BSR is a binary Week 12 outcome and was analyzed using a stratified categorical-data method.
25. Related Tutorials
Learn more about the methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: NCT01605227 — COMET-1.
- Linked publication: PubMed record — PMID 29272162.
- Linked publication: PubMed record — PMID 27400947.
Continue through Clinical Biostats
Use the trial's statistical methods as a starting point for deeper study of survival analysis, stratified testing, confidence intervals, and clinical-trial methodology.
28. Record Summary
COMET-1 provides a useful example of how a randomized phase 3 clinical trial can require several statistical methods because its endpoints have different structures. The primary endpoint, overall survival, was analyzed in the ITT population using a stratified log-rank test with a Cox proportional-hazards effect measure after 614 events. The reported OS hazard ratio was 0.90 with a two-sided 95% CI of 0.76–1.06 and P = 0.213. Bone Scan Response at Week 12 was analyzed using a stratified CMH test with P < 0.001, while the other pre-specified PFS analysis produced a hazard ratio of 0.48 with a two-sided 95% CI of 0.40–0.57 and P < 0.001.
The most important statistical lesson is that these results should be interpreted according to their endpoint definitions and analysis methods. A hazard ratio describes a relative time-to-event effect; a confidence interval describes uncertainty around that estimate; a P-value addresses statistical evidence rather than effect size; and a stratified CMH test answers a different type of question from a log-rank test. The reported serious-adverse-event counts also use denominators different from the ITT efficacy population, reinforcing the importance of identifying the analysis population before interpreting a numerical result.