← Clinical Trials
Metastatic Castration-Resistant Prostate Cancer Phase 3 Completed NCT01605227

COMET-1: Complete Statistical Analysis of Cabozantinib in Metastatic Castration-Resistant Prostate Cancer

An independent statistical review of the randomized phase 3 COMET-1 trial comparing cabozantinib with prednisone in men with metastatic castration-resistant prostate cancer previously treated with docetaxel and abiraterone or MDV3100.

Trial period: 2012-07 to 2014-09  ·  Sponsor: Exelixis  ·  Enrollment: 1028
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical trial results and design facts are restricted to the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.

Registry record: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

COMET-1 was a randomized, quadruple-masked, parallel-group phase 3 treatment trial comparing cabozantinib with prednisone in men with metastatic castration-resistant prostate cancer previously treated with docetaxel and abiraterone or MDV3100. The registered primary endpoint was overall survival, a time-to-event outcome analyzed after 614 events.

1028
Randomized
682 vs 346
2
Treatment arms
Cabozantinib vs prednisone
0.90
OS hazard ratio
95% CI 0.76–1.06
0.48
PFS hazard ratio
95% CI 0.40–0.57
FeatureCOMET-1
PhasePhase 3
PopulationMen with metastatic castration-resistant prostate cancer previously treated with docetaxel and abiraterone or MDV3100
DesignRandomized, parallel-group, quadruple-masked
AllocationRandomized
Primary endpointOverall survival (OS)
Primary endpoint typeTime-to-event
Primary hypothesisSuperiority
Enrollment1028
Study start2012-07
Primary completion2014-09
SponsorExelixis
ClinicalTrials.govNCT01605227

2. Clinical Question

The statistical question was whether cabozantinib produced a superior overall-survival outcome compared with prednisone in men with metastatic castration-resistant prostate cancer who had previously received docetaxel and abiraterone or MDV3100.

Population

Men with metastatic castration-resistant prostate cancer previously treated with docetaxel and abiraterone or MDV3100.

Intervention

Cabozantinib.

Comparator

Prednisone.

Primary question

Does cabozantinib demonstrate superiority over prednisone for overall survival?

3. Trial Design

The registry describes COMET-1 as a randomized, parallel-group, quadruple-masked phase 3 treatment trial with two intervention arms and 1028 enrolled participants. The primary efficacy analysis used the intention-to-treat population, preserving randomized treatment assignment for the primary comparison.

01
Randomization 1028 subjects
02
Two arms Cabozantinib or prednisone
03
Follow-up Time-to-event outcomes
04
Analysis Stratified comparisons
ARM A · n = 682

Cabozantinib

  • Cabozantinib
  • Included in the ITT efficacy population
  • Serious adverse events: 420/681
ARM B · n = 346

Prednisone

  • Prednisone
  • Included in the ITT efficacy population
  • Serious adverse events: 181/342
Quadruple masking: the registry identifies the study as quadruple-masked. The ClinicalTrials.gov record does not specify which four participant or study roles were masked, so this page does not infer those roles.

4. Randomization, Stratification, and Analysis Populations

The primary OS analysis was performed in the Intent to Treat (ITT) population, consisting of all 1028 randomized subjects: 682 assigned to cabozantinib and 346 assigned to prednisone.

The registry reports that the log-rank analyses were stratified by three factors:

Analysis featureRegistry-supported description
Primary efficacy populationITT population
ITT size1028 randomized subjects
Cabozantinib ITT group682
Prednisone ITT group346
Survival stratificationPrior cabazitaxel, baseline pain severity, and baseline ECOG performance status
CMH stratificationPrior cabazitaxel, baseline pain severity, and baseline ECOG performance status

Stratification is important because the analysis compares treatment groups while accounting for prespecified factors that may be associated with the outcome. It does not change the randomized treatment assignment; instead, it incorporates the design strata into the statistical comparison.

5. Endpoints

EndpointRegistry definition / time frameRole
Overall Survival (OS) The primary analysis of OS is defined as the time from randomization to death due to any cause. Participants that had not died or were permanently lost to follow-up were censored at the last known date alive. Median OS was calculated using Kaplan-Meier estimates. Analysis for OS was performed after 614 events had occurred. OS was measured from the time of randomization until 614 events, approximately 24 months after study start. Primary
Bone Scan Response (BSR) BSR was measured at the end of Week 12 as determined by the IRF. Secondary
Progression-free Survival (PFS) Duration of PFS was defined as time from the date of randomization to earlier of date of radiographic progression (bone/andor soft tissue) according to the investigator's assessment or death, assessed for up to approximately 24 months Other pre-specified

6. Statistical Methodology

Kaplan-Meier estimation

The OS registry definition states that median OS was calculated using Kaplan-Meier estimates. Kaplan-Meier estimation is appropriate for time-to-event data because participants may have different lengths of follow-up and some participants may not experience the event during observation.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di is the number of events at event time ti, and ni is the number of participants at risk immediately before that time.

The important statistical feature is that censoring does not automatically mean that a participant is treated as having experienced the event. For OS, participants who had not died or were permanently lost to follow-up were censored at the last known date alive.

Stratified log-rank test

The primary OS analysis used a log-rank test, and the registry specifically states that it was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. The PFS analysis also used a stratified log-rank test with the same three stratification factors.

The log-rank framework compares the observed and expected numbers of events between treatment groups across follow-up. A stratified version performs this comparison within the prespecified strata and combines the resulting evidence rather than treating the strata as irrelevant to the analysis.

Cox proportional-hazards effect measure

The primary OS effect measure was reported as a Cox Proportional Hazard, normalized here as a hazard ratio. The estimated OS hazard ratio was 0.90 with a two-sided 95% confidence interval of 0.76 to 1.06.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the treatment group

A hazard ratio is a relative time-to-event measure. It is not an absolute survival probability, a median survival time, or the percentage of participants who benefit.

Stratified Cochran-Mantel-Haenszel analysis

Bone Scan Response at Week 12 was analyzed using the Cochran-Mantel-Haenszel (CMH) test. The registry states that this analysis was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status.

For a binary endpoint such as BSR, a CMH analysis provides a way to compare treatment groups while accounting for prespecified categorical strata. This is conceptually different from the log-rank test: the CMH framework is used here for a binary Week 12 outcome, whereas the log-rank framework is used for time-to-event outcomes.

7. Primary Result: Overall Survival

The primary endpoint was overall survival, measured from randomization until 614 events, approximately 24 months after study start. The analysis was performed in the ITT population of 1028 randomized subjects.

Overall survival hazard ratio

0.90

95% CI: 0.76–1.06   ·   P = 0.213

Analysis: stratified log-rank test; effect measure reported as a Cox proportional hazard.

Primary endpointCabozantinibPrednisoneEffect estimateP-value
Overall Survival 682 ITT subjects 346 ITT subjects HR 0.90
95% CI 0.76–1.06
0.213
Clinical Biostats interpretation

The estimated OS hazard ratio of 0.90 means that the fitted relative hazard of death for cabozantinib versus prednisone was estimated at 0.90. Expressed as a simple derived interpretation, this corresponds to an estimated hazard that was approximately 10% lower in the cabozantinib group relative to the prednisone group under the model.

The HR does not mean that 10% fewer participants died, that survival probability was increased by 10%, or that each individual participant experienced a 10% reduction in risk. A hazard ratio summarizes a relative event-rate comparison over the analyzed time-to-event framework.

The two-sided 95% confidence interval of 0.76 to 1.06 describes uncertainty around the estimated hazard ratio. Because the interval extends below and above 1, the estimate is compatible with both a lower and a higher hazard under the stated statistical framework. The interval should not be interpreted as a range containing the treatment effect for individual patients.

The P-value of 0.213 addresses evidence against the null hypothesis in the prespecified superiority comparison. It does not measure the size or clinical importance of the treatment effect. A P-value should therefore be read together with the hazard ratio and its confidence interval rather than used as a substitute for them.

The analysis was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. Interpretation also depends on the time-to-event framework, including censoring and the assumptions underlying a Cox proportional-hazards effect measure. The ClinicalTrials.gov record does not provide enough information to assess the proportional-hazards assumption directly.

8. Secondary Result: Bone Scan Response

Bone Scan Response was measured at the end of Week 12 as determined by the IRF. The analysis used the ITT population, consisting of 682 cabozantinib subjects and 346 prednisone subjects.

Cochran-Mantel-Haenszel comparison

P < 0.001

Cabozantinib vs prednisone   ·   Stratified CMH test

Strata: prior cabazitaxel, baseline pain severity, and baseline ECOG performance status.

Clinical Biostats interpretation

The registry reports a P-value of <0.001 for the stratified CMH comparison of Bone Scan Response at Week 12. This provides statistical evidence of a difference between the randomized treatment groups under the specified CMH analysis.

The ClinicalTrials.gov record does not provide the BSR percentages, an odds ratio, a risk ratio, an absolute risk difference, or a confidence interval. Therefore, the result should not be converted into a treatment-effect magnitude that is not reported in the ClinicalTrials.gov record.

The P-value itself is not a measure of how large the difference in BSR was. A formal effect estimate with its confidence interval would be needed to quantify the magnitude and precision of the between-group difference.

The CMH analysis was stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. This is important because the reported P-value is tied to that prespecified stratified comparison rather than to an unadjusted comparison that ignores the analysis strata.

9. Other Pre-Specified Result: Progression-Free Survival

Progression-free survival was listed as an other pre-specified outcome measure. The registry analysis used the ITT population of 1028 randomized subjects, with a data cutoff date of 07 July 2014.

Progression-free survival hazard ratio

0.48

95% CI: 0.40–0.57   ·   P < 0.001

Analysis: stratified log-rank test; effect measure: hazard ratio.

EndpointAnalysis populationEffect estimateP-value
Progression-free Survival ITT: 1028 randomized subjects
682 cabozantinib; 346 prednisone
HR 0.48
95% CI 0.40–0.57
<0.001
Clinical Biostats interpretation

The PFS hazard ratio of 0.48 means that the estimated instantaneous hazard of the PFS event for cabozantinib relative to prednisone was 0.48 under the fitted time-to-event analysis. As a simple derived statement, this corresponds to an estimated hazard approximately 52% lower with cabozantinib relative to prednisone.

This does not mean that 52% of participants avoided progression, that PFS probability increased by 52%, or that every participant experienced the same relative reduction. The hazard ratio is a model-based relative measure of the event process.

The two-sided 95% confidence interval of 0.40 to 0.57 describes the statistical precision of the estimated hazard ratio. The entire interval lies below 1, so the estimated treatment effect is consistently on the lower-hazard side within this confidence interval. The interval nevertheless does not describe individual-patient outcomes.

The P-value of <0.001 quantifies evidence against the relevant null hypothesis within the reported statistical framework. It does not tell us that the HR is “48% effective,” nor does it measure clinical magnitude. The HR and confidence interval are needed to describe the estimated effect itself.

The PFS analysis used the stratified log-rank test, with stratification by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. Because PFS is a time-to-event endpoint, censoring and the interpretation of the hazard ratio remain central to understanding the result.

10. Comparing the Three Reported Analyses

The registry provides three statistical analyses spanning two time-to-event endpoints and one binary endpoint. The methods match the structure of the outcomes: log-rank testing and hazard ratios for OS and PFS, and a CMH test for Week 12 BSR.

OutcomeRoleEndpoint typeMethodEffect measureP-value
Overall Survival Primary Time-to-event Log-rank Hazard ratio 0.90
95% CI 0.76–1.06
0.213
Bone Scan Response Secondary Binary Cochran-Mantel-Haenszel Not reported in the ClinicalTrials.gov record <0.001
Progression-free Survival Other pre-specified Time-to-event Log-rank Hazard ratio 0.48
95% CI 0.40–0.57
<0.001

This distinction matters statistically. A P-value from a CMH analysis cannot be interpreted as though it were a hazard-ratio result, and a hazard ratio cannot be interpreted as though it were a simple binary response proportion. Each estimate belongs to a particular endpoint definition, analysis population, and statistical model.

11. What a Hazard Ratio of 0.90 Means

Relative effect

An HR of 0.90 is a ratio of estimated hazards. Because the value is below 1, the fitted analysis estimates a lower instantaneous event rate in the cabozantinib group than in the prednisone group.

What it does not mean

It is not a 10-percentage-point difference in survival, not a 10% increase in the probability of being alive, and not a statement that every participant's personal risk was reduced by exactly 10%.

Precision

The 95% CI of 0.76–1.06 is more informative about precision than the point estimate alone. It shows that the estimate has uncertainty extending across the value 1, which is the conventional no-hazard-difference reference value.

P-value

The P-value of 0.213 is a statement about statistical evidence under the specified hypothesis test. It is not a probability that the treatment has no effect and is not a measure of how clinically important the HR is.

12. Why the PFS Result Looks Different From the OS Result

COMET-1 illustrates why different clinical endpoints can produce different statistical estimates even when they compare the same randomized treatment groups.

Because the event definitions differ, the hazard ratio for one endpoint should not be expected to equal the hazard ratio for another endpoint. The OS HR of 0.90 and the PFS HR of 0.48 are therefore separate statistical quantities answering different questions.

OS

Measures time from randomization to death due to any cause. The primary analysis occurred after 614 events.

PFS

Measures time from randomization according to the registered progression-free survival definition. The registry-reported analysis used a 07 July 2014 data cutoff.

Different event processes

Different endpoints can produce different hazard ratios because the events being modeled are not the same.

Different interpretation

A PFS hazard ratio should not be presented as though it were an OS hazard ratio or an absolute survival difference.

13. Why Stratification Was Used

Both the survival analyses and the CMH analysis were stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status. The same factors therefore appear repeatedly in the statistical methodology of the trial.

Stratification is particularly useful in a randomized trial when the investigators have prespecified clinically relevant factors that may influence outcomes. Rather than ignoring these factors, the analysis incorporates the treatment comparison within the strata and then combines the evidence.

Conceptual structure
Treatment comparison = within-stratum comparisons combined across prespecified strata

The purpose is not to create a new treatment assignment or to turn observational associations into randomized evidence. The randomized comparison remains the foundation; stratification specifies how the comparison is analyzed.

The registry does not provide enough information in the ClinicalTrials.gov record to quantify how much each individual stratification factor contributed to the final estimates. Therefore, the appropriate interpretation is methodological rather than attributing a particular numerical effect to any one factor.

14. Intention-to-Treat Analysis

The ITT population included 1028 randomized subjects: 682 in the cabozantinib group and 346 in the prednisone group. This population was used for the OS, BSR, and PFS analyses reported in the ClinicalTrials.gov record.

The central principle of ITT is that participants are analyzed according to the treatment group to which they were randomized. This preserves the treatment comparison generated by randomization rather than redefining groups according to treatment received after randomization.

Why it matters

Randomization creates the basis for comparing treatment groups. ITT maintains that randomized comparison in the efficacy analysis.

What ITT does not do

ITT does not remove censoring, eliminate missing information, or guarantee identical follow-up in every participant.

15. Safety

The ClinicalTrials.gov record reports serious adverse events by treatment arm using affected participants over participants at risk.

Safety measureCabozantinibPrednisone
Serious adverse events 420/681 181/342

These figures should be interpreted as the reported counts over the corresponding at-risk denominators. They should not be converted into percentages unless that calculation is explicitly desired and permitted by the reporting framework; the ClinicalTrials.gov record instruct that numbers be reported exactly as given.

Clinical Biostats interpretation

The safety data describe the number of participants with serious adverse events relative to the reported number at risk in each arm. The denominators differ from the ITT efficacy denominators of 682 and 346, which is an important distinction when comparing efficacy and safety populations.

A safety comparison should not be reduced to a single P-value or treated as another efficacy endpoint. Serious adverse events are an exposure-related safety outcome, and their interpretation depends on the definition, observation period, and population represented by the denominator.

16. Statistical Methods Explained

Why was a log-rank test used for OS?

OS is a time-to-event endpoint: participants can die at different times, and some participants may be censored rather than observed to death. The log-rank test is designed to compare survival experience between randomized groups while using the timing of events throughout follow-up.

What does the OS hazard ratio of 0.90 mean?

It means the estimated hazard of death for cabozantinib relative to prednisone was 0.90 under the reported Cox proportional-hazards framework. The simple derived interpretation is an approximately 10% lower estimated hazard, not a 10-percentage-point improvement in survival.

Why is the confidence interval important?

The 95% CI of 0.76–1.06 communicates uncertainty around the OS estimate. The point estimate alone cannot show how precisely the treatment effect was estimated. The interval also spans 1, so the estimate is not separated from the conventional no-difference value within this confidence interval.

Why doesn't the P-value measure effect size?

The P-value evaluates statistical evidence against a null hypothesis under the specified testing framework. It is influenced by the amount of information in the analysis and does not directly quantify the magnitude of an effect. Effect size is described by measures such as the hazard ratio, while precision is described by the confidence interval.

Why was a Cochran-Mantel-Haenszel test used for Bone Scan Response?

BSR is a binary endpoint measured at Week 12, unlike OS and PFS, which are time-to-event outcomes. The CMH framework allows the treatment comparison to account for the prespecified stratification factors rather than treating the analysis as an unstratified binary comparison.

Why are the OS and PFS hazard ratios not interchangeable?

OS and PFS define different events. OS concerns death due to any cause, whereas PFS uses a progression-related time-to-event definition. Consequently, the hazard ratio for PFS describes a different event process from the hazard ratio for OS.

What does censoring mean in the OS analysis?

The registry states that participants who had not died or were permanently lost to follow-up were censored at the last known date alive. Censoring means that the analysis uses the participant's observed follow-up up to that point rather than treating the participant as having experienced the event.

17. Censoring and the Time-to-Event Framework

The OS endpoint explicitly specifies censoring for participants who had not died or were permanently lost to follow-up. This is a fundamental feature of survival analysis because clinical trials rarely observe every participant from randomization through the event of interest.

ComponentCOMET-1 registry definitionStatistical implication
Time originRandomizationParticipants enter the OS clock at randomization.
EventDeath due to any causeOS is an all-cause mortality endpoint.
CensoringLast known date alive for participants who had not died or were permanently lost to follow-upObserved follow-up contributes information up to the censoring time.
Primary event target614 eventsThe OS analysis was performed after the specified event count had occurred.

The distinction between an event and censoring is important. A censored participant is not counted as having died at the censoring time. Instead, the available survival information is retained up to the last known time alive.

18. Primary Endpoint Timing and Event-Driven Analysis

The registered OS endpoint was measured until 614 events, approximately 24 months after study start. The registry states that the OS analysis was performed after 614 events had occurred.

2012-07

Study start

COMET-1 began according to the registry.

Approximately 24 months after study start

Primary OS event target

The registered OS analysis was defined around 614 events.

2014-09

Primary completion

The registry reports primary completion in 2014-09.

07 July 2014

PFS data cutoff

The registry-reported PFS statistical analysis identifies this data cutoff date.

An event-driven endpoint is analyzed when the prespecified amount of event information has accumulated rather than simply when a fixed calendar date arrives. For survival outcomes, the number of observed events can determine the statistical information available for estimating and comparing hazards.

19. Interpreting the Confidence Intervals

Confidence intervals are especially useful for avoiding an overly narrow reading of a single point estimate.

OS

HR 0.90 with a two-sided 95% CI of 0.76–1.06. The interval crosses 1.

PFS

HR 0.48 with a two-sided 95% CI of 0.40–0.57. The interval lies below 1.

Precision is endpoint-specific

The width of each interval reflects uncertainty for that particular endpoint and analysis.

Not individual prediction intervals

A confidence interval around a hazard ratio does not describe the range of outcomes an individual participant will experience.

The difference between the OS and PFS intervals should not be interpreted simply as a difference in “certainty” without considering the endpoint, number of events, censoring, and statistical model. Confidence intervals describe uncertainty around the corresponding estimated parameter.

20. Multiplicity and the Reported P-values

The ClinicalTrials.gov record identifies the primary hypothesis type as superiority and identify one registered primary endpoint: OS. The registry also reports secondary BSR and other pre-specified PFS analyses.

OutcomeRoleHypothesis typeReported statistical evidence
Overall SurvivalPrimarySuperiorityHR 0.90; 95% CI 0.76–1.06; P = 0.213
Bone Scan ResponseSecondarySuperiorityCMH P < 0.001
Progression-free SurvivalOther pre-specifiedSuperiorityHR 0.48; 95% CI 0.40–0.57; P < 0.001

The ClinicalTrials.gov record does not specify an alpha-spending scheme, a formal multiplicity-adjustment procedure, or an endpoint hierarchy beyond the registered roles shown above. Accordingly, this page does not infer a particular familywise-error procedure or assign a confirmatory status to analyses beyond the registry's classifications.

Interpretation caution: the fact that multiple P-values are reported does not by itself establish that they were tested under a common multiplicity-adjusted error budget. The appropriate interpretation of each P-value depends on the prespecified statistical analysis plan, which is not reported here.

21. Proportional-Hazards Considerations

The OS effect measure was reported as a Cox proportional-hazards measure, and the PFS effect measure was reported as a hazard ratio. A Cox hazard ratio is most naturally interpreted as a relative hazard under a proportional-hazards framework.

The ClinicalTrials.gov record does not report a diagnostic assessment of the proportional-hazards assumption, nor do they provide time-varying hazard ratios or alternative estimands. Therefore, the appropriate educational interpretation is to recognize the assumption rather than claim that it was or was not satisfied.

Why the assumption matters
If HR(t) ≈ constant, a single HR summarizes the relative hazard over time

If the relative hazard changes substantially over time, a single hazard ratio can become a less complete description of the treatment effect. The ClinicalTrials.gov record does not allow a direct assessment of this issue for COMET-1.

22. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary OS analysis estimated a hazard ratio of 0.90 with a two-sided 95% CI of 0.76–1.06 and P = 0.213. The PFS analysis estimated a hazard ratio of 0.48 with a two-sided 95% CI of 0.40–0.57 and P < 0.001. BSR was analyzed with a stratified CMH test with P < 0.001.

Clinical interpretation

The statistical results describe differences in specified trial endpoints; they do not by themselves summarize every aspect of clinical benefit, safety, patient experience, or treatment choice. The ClinicalTrials.gov record does not provide the additional outcome measures needed for a broader clinical assessment.

This distinction is important because a statistically strong result for one endpoint does not automatically determine the interpretation of another endpoint. COMET-1 provides a clear example: PFS and BSR show reported statistical evidence of between-group differences, while the primary OS estimate has a confidence interval that crosses 1 and a P-value of 0.213.

23. Limitations

24. Why This Trial Matters Statistically

COMET-1 is a useful teaching case because the ClinicalTrials.gov record brings together several core principles of clinical-trial biostatistics without requiring the analyst to treat every endpoint as though it were the same type of outcome.

ConceptHow it appears in COMET-1
Randomization1028 subjects were randomized to two parallel treatment arms.
Quadruple maskingThe registry identifies the trial as quadruple-masked.
Intention-to-treatThe OS, BSR, and PFS analyses use the ITT population of 1028 randomized subjects.
Kaplan-Meier estimationMedian OS was calculated using Kaplan-Meier estimates.
Time-to-event analysisOS and PFS are analyzed as time-to-event endpoints.
Log-rank testOS and PFS were analyzed using log-rank tests.
Stratified analysisSurvival and CMH analyses were stratified by prior cabazitaxel, baseline pain severity, and baseline ECOG performance status.
Hazard ratioOS HR 0.90; PFS HR 0.48.
Confidence intervalTwo-sided 95% CIs were reported for the OS and PFS hazard ratios.
Cochran-Mantel-Haenszel testUsed for Week 12 Bone Scan Response.
Event-driven analysisThe primary OS analysis was performed after 614 events.
Safety denominatorsSerious adverse events were reported as 420/681 and 181/342.

The most instructive feature is the alignment between endpoint type and statistical method. OS and PFS require methods that account for event timing and censoring, whereas BSR is a binary Week 12 outcome and was analyzed using a stratified categorical-data method.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through Clinical Biostats

Use the trial's statistical methods as a starting point for deeper study of survival analysis, stratified testing, confidence intervals, and clinical-trial methodology.

28. Record Summary

COMET-1 provides a useful example of how a randomized phase 3 clinical trial can require several statistical methods because its endpoints have different structures. The primary endpoint, overall survival, was analyzed in the ITT population using a stratified log-rank test with a Cox proportional-hazards effect measure after 614 events. The reported OS hazard ratio was 0.90 with a two-sided 95% CI of 0.76–1.06 and P = 0.213. Bone Scan Response at Week 12 was analyzed using a stratified CMH test with P < 0.001, while the other pre-specified PFS analysis produced a hazard ratio of 0.48 with a two-sided 95% CI of 0.40–0.57 and P < 0.001.

The most important statistical lesson is that these results should be interpreted according to their endpoint definitions and analysis methods. A hazard ratio describes a relative time-to-event effect; a confidence interval describes uncertainty around that estimate; a P-value addresses statistical evidence rather than effect size; and a stratified CMH test answers a different type of question from a log-rank test. The reported serious-adverse-event counts also use denominators different from the ITT efficacy population, reinforcing the importance of identifying the analysis population before interpreting a numerical result.

Clinical Biostats methodology: A trial-results page should distinguish the registered endpoint, analysis population, statistical method, effect measure, and uncertainty rather than presenting every reported P-value as though it answered the same question. This page uses only the registry-reported COMET-1 registry data for trial-specific numerical and methodological claims.