This page separates reported trial results from statistical interpretation. Numerical results are taken from the ClinicalTrials.gov record. The registry reports hazard-ratio estimates and 95% confidence intervals for the posted primary-endpoint analyses, but it does not report p-values for those analyses in the ClinicalTrials.gov record.
1. Trial at a Glance
CheckMate-227 was a randomized, parallel-group, phase 3 treatment trial evaluating nivolumab, nivolumab plus ipilimumab, and nivolumab plus platinum-doublet chemotherapy against platinum-doublet chemotherapy in patients with non-small cell lung cancer.
| Feature | CheckMate-227 |
|---|---|
| Trial name | CheckMate-227 |
| ClinicalTrials.gov identifier | NCT02477826 |
| Phase | Phase 3 |
| Condition | Non-Small Cell Lung Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 2747 |
| Arms | 4 |
| Primary endpoints | Progression-Free Survival Per BICR; Overall Survival |
| Primary endpoint type | Time-to-event |
| Hypothesis type reported for analyses | Superiority |
| Results posted | Yes |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | Industry |
| Trial start | 2015-08-05 |
| Primary completion | 2024-10-25 |
2. Clinical Question
The registry describes a broad randomized comparison of nivolumab-based treatment strategies with platinum-doublet chemotherapy in patients with stage IV non-small cell lung cancer. The two registered primary endpoints are progression-free survival per blinded independent central review and overall survival.
Population
Patients with stage IV non-small cell lung cancer, as described in the trial's brief title.
Intervention strategies
Nivolumab; nivolumab plus ipilimumab; and nivolumab plus platinum-doublet chemotherapy.
Comparator
Platinum-doublet chemotherapy.
Primary question
How do the registered nivolumab-based comparisons affect the time from randomization to progression or death and the time from randomization to death?
3. Trial Design
Registered intervention set
The trial data list seven interventions: nivolumab, ipilimumab, carboplatin, cisplatin, gemcitabine, pemetrexed, and paclitaxel. The ClinicalTrials.gov record identifies specific randomized-arm comparisons using abbreviated arm labels such as Nivo + Ipi, Nivolumab, Nivo + Chemo, and Chemotherapy.
| Intervention | Registry classification |
|---|---|
| Nivolumab | Drug |
| Ipilimumab | Drug |
| Carboplatin | Drug |
| Cisplatin | Drug |
| Gemcitabine | Drug |
| Pemetrexed | Drug |
| Paclitaxel | Drug |
4. Endpoints
| Endpoint | Registered definition / time frame | Analysis type |
|---|---|---|
| Progression-Free Survival Per BICR | From randomization untill disease progression or death, whichever occurs first (up to approximately 481 weeks) | Time-to-event; Kaplan-Meier estimates are part of the registered definition |
| Overall Survival | From randomization untill death or last follow up whichever occurs first (up to approximately 481 weeks) | Time-to-event |
Progression-Free Survival Per BICR
The registry defines progression-free survival as the time between the date of randomization and the date of first documented disease progression, based on BICR assessments per RECIST v1.1, or death due to any cause, whichever occurs first. The registry definition explicitly refers to Kaplan-Meier estimates.
Overall Survival
The registry defines overall survival for all randomized participants as the time between the randomization date and the date of death from any cause. The registered time frame is from randomization until death or last follow-up, whichever occurs first, up to approximately 481 weeks.
5. Statistical Methodology
The ClinicalTrials.gov record identifies the effect measure as a hazard ratio and classify the hypothesis type as superiority. For every posted primary analysis, the analysis population is all randomized participants.
The normalized statistical-method field is reported as Not reported for the analyses posted on ClinicalTrials.gov, except that one PFS comparison identifies the reported method as "Cox Proportional Hazard." The ClinicalTrials.gov record does not provide a complete statistical analysis plan, test-statistic specification, stratification factors, multiplicity procedure, or interim-analysis procedure.
Kaplan-Meier estimation
Kaplan-Meier estimation is appropriate for the registered progression-free and overall-survival endpoints because it estimates the event-time distribution while accommodating right censoring. The PFS definition itself explicitly references Kaplan-Meier estimates.
Here, di represents events at an event time and ni represents participants at risk immediately before that time. This is an educational representation of the method; the ClinicalTrials.gov record does not provide the underlying individual event and censoring records.
Hazard ratios
The hazard ratio compares the estimated instantaneous event rates between two randomized groups under the fitted time-to-event model. An HR below 1 indicates a lower estimated event hazard for the first-named group relative to the second-named group.
For example, an HR of 0.61 corresponds mathematically to an estimated hazard approximately 39% lower than the comparator hazard. That is a relative hazard interpretation, not a statement that 39% of patients avoided an event.
Cox proportional-hazards modeling
One reported PFS analysis explicitly reports "Cox Proportional Hazard" for the comparison of Part 1-Arm D: Nivo + Ipi versus Part 1-Arm F: Chemotherapy. The other posted analyses do not state a statistical method, even though the effect measure is consistently a hazard ratio.
The proportional-hazards model has the conceptual form:
The hazard ratio associated with a treatment indicator is exp(β). The ClinicalTrials.gov record does not report whether proportional hazards were formally assessed or how any departures from the assumption were handled.
6. Primary Results: Progression-Free Survival Per BICR
ClinicalTrials.gov reports seven primary-endpoint PFS comparisons in the ClinicalTrials.gov record. All are based on all randomized participants, use a two-sided 95% confidence interval, and are classified as superiority analyses. The ClinicalTrials.gov record does not provide p-values for these comparisons.
Part 1-Arm B: Nivo + Ipi vs Part 1-Arm C: Chemotherapy
An HR of 0.79 means the estimated instantaneous rate of progression or death in the first-named group was about 79% of the estimated rate in the chemotherapy group under the time-to-event model. Equivalently, 0.79 corresponds to a 21% lower estimated hazard.
The HR does not mean that 21% of patients avoided progression or death, nor does it describe an absolute difference in the probability of being progression-free at a particular time. The 95% CI of 0.67–0.93 quantifies uncertainty around the estimated relative hazard; it is not a range in which individual patients' treatment effects must fall.
The ClinicalTrials.gov record does not report a p-value. A p-value, even when available, would address evidence against a null hypothesis under the specified testing procedure; it would not measure the magnitude or clinical importance of the HR. Interpretation of a Cox HR also depends on the proportional-hazards assumption and on how censoring is handled.
Part 1-Arm A: Nivolumab vs Part 1-Arm C: Chemotherapy
An HR of 0.95 indicates that the estimated instantaneous rate of progression or death in the nivolumab group was 95% of that in the chemotherapy comparator under the reported time-to-event analysis.
The 95% CI of 0.81–1.12 spans 1.00. Consequently, the interval is compatible with a range of relative hazards that includes no difference as represented by HR 1.00. It does not establish that the treatments are identical; it indicates uncertainty around the point estimate.
No p-value is reported in the ClinicalTrials.gov record. A p-value should not be reconstructed from the confidence interval when the task is to report registry results exactly. The analysis is also a time-to-event comparison, so the HR should not be interpreted as a simple ratio of cumulative event probabilities.
Part 1-Arm B: Nivo + Ipi vs Part 1-Arm A: Nivolumab
An HR of 0.85 corresponds to an estimated instantaneous progression-or-death hazard about 15% lower in the first-named group than in the nivolumab group.
The 95% CI of 0.72–0.99 is relatively close to 1.00 at its upper boundary. That illustrates why the point estimate should not be treated as a precise description of a fixed treatment effect. The interval communicates statistical uncertainty around the estimate.
The ClinicalTrials.gov record provides no p-value. The HR is not an absolute measure of treatment benefit, and it does not indicate how many additional months an individual participant will remain progression-free. Because the endpoint includes either progression or death, it is a composite time-to-event endpoint.
Part 1-Arm D: Nivo + Ipi vs Part 1-Arm F: Chemotherapy
An HR of 0.76 means the estimated instantaneous rate of progression or death was 76% of that in the comparator, corresponding to a 24% lower estimated hazard for the first-named group.
The 95% CI of 0.60–0.97 expresses uncertainty around the estimate and remains below 1.00 throughout the reported interval. It does not imply that the true effect for every patient lies between 0.60 and 0.97.
This comparison is the only reported PFS analysis for which the registry data explicitly identify a Cox proportional-hazard method. The ClinicalTrials.gov record does not report a p-value or describe whether the proportional-hazards assumption was formally evaluated.
Part 1-Arm G: Nivo + Chemo vs Part 1-Arm F: Chemotherapy
An HR of 0.74 corresponds to an estimated instantaneous progression-or-death hazard 26% lower for the first-named group than for the comparator.
The 95% CI of 0.59–0.93 gives the precision of the relative hazard estimate within the statistical framework used. It is not an interval for individual-level benefit and does not give the absolute probability of remaining progression-free.
No p-value is included in the registry analysis. The superiority designation describes the hypothesis type recorded for the analysis; it does not change the interpretation of the HR or confidence interval.
Part 2 - Arm H: Nivolumab + Chemotherapy vs Part 2 - Arm I: Chemotherapy
An HR of 0.61 means that the estimated instantaneous rate of progression or death was 61% of the comparator rate, corresponding to a 39% lower estimated hazard for the first-named group.
The 95% CI of 0.52–0.72 is relatively narrow compared with several other reported PFS intervals, indicating greater statistical precision around this particular HR estimate. Precision, however, is not the same as clinical importance.
The registry data do not report a p-value. The HR also should not be converted into a claim about the percentage of participants who benefited. Censoring and the time-varying nature of risk remain fundamental to interpretation of a time-to-event endpoint.
Part 3 -Arm J: Nivolumab + Ipilimumab vs Part 3 - Arm K: Chemotherapy
An HR of 0.66 corresponds to an estimated instantaneous progression-or-death hazard approximately 34% lower in the first-named group than in the chemotherapy comparator.
The 95% CI of 0.49–0.89 provides a range of statistical uncertainty around that estimate and remains below 1.00. It should not be interpreted as a range of possible outcomes for individual participants.
The ClinicalTrials.gov record does not report a p-value. Nor do they provide enough information to reconstruct a Kaplan-Meier curve, median PFS, or absolute event probabilities. Those quantities should not be inferred from the HR alone.
7. Primary Results: Overall Survival
The ClinicalTrials.gov record contains seven primary overall-survival comparisons. The endpoint is defined as time from randomization to death from any cause, with follow-up continuing until death or last follow-up, whichever occurs first, up to approximately 481 weeks.
Part 1-Arm B: Nivo + Ipi vs Part 1-Arm C: Chemotherapy
An OS HR of 0.78 means the estimated instantaneous rate of death in the first-named group was 78% of the estimated rate in the chemotherapy group, corresponding to a 22% lower estimated hazard.
The 95% CI of 0.67–0.91 communicates uncertainty around the estimated relative hazard. It does not mean that the mortality reduction for each patient is exactly 22%, nor does it specify an absolute survival difference.
The ClinicalTrials.gov record does not report a p-value. The hazard ratio also does not reveal whether hazards were proportional throughout follow-up. If hazards change materially over time, a single HR can summarize the comparison imperfectly.
Part 1-Arm A: Nivolumab vs Part 1-Arm C: Chemotherapy
An HR of 0.91 corresponds to an estimated instantaneous death rate 91% of the comparator rate, or a 9% lower estimated hazard for the first-named group.
The 95% CI of 0.78–1.05 crosses 1.00. Thus, the confidence interval includes a no-difference hazard ratio as well as values below and above 1.00. The point estimate alone should therefore not be treated as a definitive measure of the underlying effect.
No p-value is reported in the ClinicalTrials.gov record. A p-value would not substitute for the confidence interval because statistical evidence and effect-size precision answer different questions.
Part 1-Arm B: Nivo + Ipi vs Part 1-Arm A: Nivolumab
An HR of 0.86 means the estimated instantaneous death rate was 86% of the comparator rate, corresponding to a 14% lower estimated hazard for the first-named group.
The 95% CI of 0.74–1.01 is close to the null value of 1.00 and includes it. This indicates that the estimate has meaningful uncertainty and that the ClinicalTrials.gov record does not establish a precise non-null relative hazard.
The absence of a reported p-value is important: no p-value should be invented or reverse-engineered. The HR also does not describe median survival or absolute survival at any specific time point.
Part 1-Arm D: Nivo + Ipi vs Part 1-Arm F: Chemotherapy
An HR of 0.64 corresponds to an estimated instantaneous death hazard approximately 36% lower in the first-named group than in the chemotherapy comparator.
The 95% CI of 0.51–0.79 indicates the uncertainty around the estimated relative hazard within the reported analysis. It does not represent variability in treatment effects among individual patients.
The registry does not provide a p-value. Nor does the HR alone establish an absolute difference in survival. To interpret absolute benefit, time-specific survival probabilities or other absolute measures would be required.
Part 1-Arm G: Nivo + Chemo vs Part 1-Arm F: Chemotherapy
An HR of 0.77 means the estimated instantaneous death rate was 77% of the comparator rate, corresponding to a 23% lower estimated hazard for the first-named group.
The 95% CI of 0.63–0.96 provides the reported uncertainty around the HR. Because the interval remains below 1.00, the reported point estimate and interval describe a relative hazard lower than the comparator throughout the stated interval.
The ClinicalTrials.gov record does not provide a p-value, absolute survival probabilities, or median OS. The HR therefore should not be expanded into claims about the proportion of patients alive at a particular time.
Part 2 - Arm H: Nivolumab + Chemotherapy vs Part 2 - Arm I: Chemotherapy
An HR of 0.75 corresponds to an estimated instantaneous death hazard 25% lower in the first-named group than in the comparator.
The 95% CI of 0.64–0.87 gives the uncertainty surrounding the relative hazard estimate. It is not a confidence interval for the number of life-years gained, nor does it give an individual patient's probability of benefit.
No p-value is reported in the ClinicalTrials.gov record. The analysis is classified as a superiority analysis, but that classification does not by itself provide the magnitude, precision, or absolute clinical effect of treatment.
Part 3 -Arm J: Nivolumab + Ipilimumab vs Part 3 - Arm K: Chemotherapy
An HR of 0.85 means the estimated instantaneous death hazard was 85% of the comparator hazard, corresponding to a 15% lower estimated hazard for the first-named group.
The 95% CI of 0.64–1.14 crosses 1.00 and is therefore compatible with both lower and higher hazards relative to the comparator under the statistical model. The relatively broad interval also demonstrates uncertainty around the point estimate.
The ClinicalTrials.gov record does not provide a p-value. The HR should not be interpreted as an absolute mortality reduction, and no conclusion about median survival or time-specific survival can be derived without additional reported data.
8. Primary Results in One Table
Putting the 14 posted primary analyses together makes an important feature of CheckMate-227 visible: the registry contains several distinct randomized comparisons across the registered trial parts. These should not be collapsed into a single overall treatment effect.
| Endpoint | Comparison | HR | 95% CI |
|---|---|---|---|
| PFS per BICR | Part 1-B Nivo + Ipi vs Part 1-C Chemotherapy | 0.79 | 0.67–0.93 |
| PFS per BICR | Part 1-A Nivolumab vs Part 1-C Chemotherapy | 0.95 | 0.81–1.12 |
| PFS per BICR | Part 1-B Nivo + Ipi vs Part 1-A Nivolumab | 0.85 | 0.72–0.99 |
| PFS per BICR | Part 1-D Nivo + Ipi vs Part 1-F Chemotherapy | 0.76 | 0.60–0.97 |
| PFS per BICR | Part 1-G Nivo + Chemo vs Part 1-F Chemotherapy | 0.74 | 0.59–0.93 |
| PFS per BICR | Part 2-H Nivolumab + Chemotherapy vs Part 2-I Chemotherapy | 0.61 | 0.52–0.72 |
| PFS per BICR | Part 3-J Nivolumab + Ipilimumab vs Part 3-K Chemotherapy | 0.66 | 0.49–0.89 |
| Overall Survival | Part 1-B Nivo + Ipi vs Part 1-C Chemotherapy | 0.78 | 0.67–0.91 |
| Overall Survival | Part 1-A Nivolumab vs Part 1-C Chemotherapy | 0.91 | 0.78–1.05 |
| Overall Survival | Part 1-B Nivo + Ipi vs Part 1-A Nivolumab | 0.86 | 0.74–1.01 |
| Overall Survival | Part 1-D Nivo + Ipi vs Part 1-F Chemotherapy | 0.64 | 0.51–0.79 |
| Overall Survival | Part 1-G Nivo + Chemo vs Part 1-F Chemotherapy | 0.77 | 0.63–0.96 |
| Overall Survival | Part 2-H Nivolumab + Chemotherapy vs Part 2-I Chemotherapy | 0.75 | 0.64–0.87 |
| Overall Survival | Part 3-J Nivolumab + Ipilimumab vs Part 3-K Chemotherapy | 0.85 | 0.64–1.14 |
9. How to Read the Hazard Ratios Across the Trial
The 14 hazard ratios range from 0.61 to 0.95 for PFS and from 0.64 to 0.91 for OS. These values describe different comparisons rather than repeated estimates of one common treatment effect.
This display is descriptive rather than a formal forest plot. It does not show confidence intervals, and it should not be used to determine whether one comparison has a statistically different treatment effect from another. A formal comparison would require an appropriate interaction framework and the underlying analysis structure.
10. Statistical Methods Explained
Why is Kaplan-Meier estimation appropriate for PFS and OS?
Both endpoints measure time until an event. Some participants may reach the end of observation without the event, creating right-censored observations. Kaplan-Meier estimation uses the observed event times while retaining information from participants up to their censoring time.
What does an HR of 0.61 mean?
An HR of 0.61 means that the estimated instantaneous event hazard in the first-named group is 61% of the comparator hazard under the fitted time-to-event model. It corresponds to a 39% lower estimated hazard, but it does not mean a 39% absolute reduction in the probability of an event.
Why is HR 1.00 important?
For a treatment-versus-comparator hazard ratio, 1.00 represents equal estimated hazards. A confidence interval that includes 1.00 therefore includes the no-difference value within the interval. That does not prove equality; it indicates that the reported interval is compatible with a no-difference hazard ratio.
Why report the confidence interval instead of only the HR?
The point estimate gives one estimated relative effect. The confidence interval adds information about statistical precision. For example, an HR of 0.85 with a relatively narrow interval communicates a different degree of precision from an HR of 0.85 with a very wide interval.
Why doesn't the p-value measure treatment effect size?
A p-value addresses evidence against a specified null hypothesis under a specified statistical procedure. It is affected by both effect magnitude and information in the data. It does not directly quantify the size of a treatment effect. In the registry-reported CheckMate-227 registry data, p-values are not reported for the 14 posted primary analyses.
Why should an HR not automatically be interpreted as a risk ratio?
Risk ratios compare cumulative probabilities over a defined period. A hazard ratio compares instantaneous event rates within a time-to-event framework. These quantities are related but are not interchangeable. The HR cannot by itself determine a cumulative event probability at a particular time.
What is the proportional-hazards issue?
A conventional Cox proportional-hazards interpretation assumes that the relative hazard between groups is adequately represented by a constant hazard ratio over time. If the relative hazards change substantially over follow-up, one HR can compress a more complicated time-dependent pattern into a single summary. The ClinicalTrials.gov record does not report a formal assessment of this assumption.
11. Confidence Intervals and Statistical Precision
The analyses posted on ClinicalTrials.gov consistently report two-sided 95% confidence intervals. The width of these intervals provides useful information about the precision of the estimated hazard ratio.
| Endpoint | Comparison with narrowest reported CI | 95% CI |
|---|---|---|
| PFS | Part 2-H Nivolumab + Chemotherapy vs Part 2-I Chemotherapy | 0.52–0.72 |
| PFS | Part 1-B Nivo + Ipi vs Part 1-C Chemotherapy | 0.67–0.93 |
| OS | Part 2-H Nivolumab + Chemotherapy vs Part 2-I Chemotherapy | 0.64–0.87 |
| OS | Part 1-B Nivo + Ipi vs Part 1-C Chemotherapy | 0.67–0.91 |
| OS | Part 3-J Nivolumab + Ipilimumab vs Part 3-K Chemotherapy | 0.64–1.14 |
The intervals illustrate why a point estimate should never be read in isolation. The OS HR of 0.85 for Part 3-J versus Part 3-K has a 95% CI of 0.64–1.14, while the OS HR of 0.75 for Part 2-H versus Part 2-I has a 95% CI of 0.64–0.87. The difference is not simply that one point estimate is smaller than the other; the uncertainty around each estimate must also be considered.
A 95% confidence interval is a statement about the statistical estimation procedure, not a statement that there is a 95% probability that the true hazard ratio lies inside this particular interval. In repeated samples under the model and assumptions, the corresponding procedure would capture the true parameter at the stated nominal rate.
12. Analysis Population and Randomization
Every registry-reported statistical analysis identifies all randomized participants as its analysis population. This is an important feature because treatment assignment occurs before knowledge of subsequent outcomes.
Why randomization matters
Randomization creates the foundation for comparing outcomes between assigned groups without deliberately choosing participants for one treatment based on their expected outcome.
Why the analysis population matters
Analyzing all randomized participants preserves the connection between the comparison and the original treatment assignment. The registry-reported primary analyses explicitly use this population.
Why censoring matters
Time-to-event analyses depend on rules governing participants whose event status is not observed through the end of available follow-up.
Why follow-up matters
An HR is estimated from accumulated event and censoring information. Its interpretation therefore depends on the follow-up represented in the analysis.
13. Safety Results
The ClinicalTrials.gov record includes serious adverse events by arm. These figures are presented exactly as reported in the registry, including the arm labels and denominators. They should be interpreted as the number affected divided by the reported number at risk, rather than converted into unreported percentages.
| Trial arm | Serious adverse events affected / at risk |
|---|---|
| Part 1-Arm B: Nivo + Ipi | 275/391 |
| Part 1-Arm A: Nivolumab | 239/391 |
| Part 1-Arm C: Chemotherapy | 213/387 |
| Part 1-Arm D: Nivo + Ipi | 128/185 |
| Part 1-Arm G: Nivo + Chemo | 112/172 |
| Part 1-Arm F: Chemotherapy | 91/183 |
| Part 2 - Arm H: Nivolumab + Chemotherapy | 240/375 |
| Part 2 - Arm I: Chemotherapy | 180/3 |
Serious adverse-event counts should be kept conceptually separate from efficacy hazard ratios. A safety count describes affected participants within an exposure or risk set, whereas an HR describes a relative time-to-event measure. They answer different statistical questions and should not be combined into a single numerical measure of overall trial performance.
14. What the Results Do Not Tell Us
The ClinicalTrials.gov record provides 14 primary-endpoint hazard-ratio estimates and corresponding 95% confidence intervals, but they do not provide several quantities that would ordinarily be useful for a complete survival analysis.
| Information | Available in the ClinicalTrials.gov record? | Statistical implication |
|---|---|---|
| Hazard-ratio estimates | Yes | Relative time-to-event effects can be described. |
| 95% confidence intervals | Yes | Precision can be described. |
| P-values | No | No formal p-value should be reported or reconstructed. |
| Median PFS | No | Cannot be reported from the ClinicalTrials.gov record. |
| Median OS | No | Cannot be reported from the ClinicalTrials.gov record. |
| Time-specific survival probabilities | No | Cannot be reported from the ClinicalTrials.gov record. |
| Kaplan-Meier event tables | No | Cannot reconstruct a valid KM curve. |
| Baseline characteristics | No | No balance table should be inferred. |
| Subgroup estimates | No | No subgroup treatment-effect claims are made. |
| Stratification factors | No | No stratified-analysis variables are inferred. |
| Multiplicity procedure | No | No familywise-error interpretation is assigned beyond the registry-reported superiority designation. |
| Interim-analysis procedure | No | No alpha-spending or stopping-boundary claims are made. |
| Missing-data/imputation procedure | No | No imputation method is inferred. |
| Bayesian analysis | No | No Bayesian interpretation is applied. |
This distinction is important. A statistically responsible trial-results page should not fill gaps in the registry record with remembered values from publications or with quantities reverse-engineered from hazard ratios. The result is a narrower page, but a more reproducible one.
15. Multiplicity and Multiple Comparisons
CheckMate-227 contains multiple posted primary analyses: seven for progression-free survival and seven for overall survival in the ClinicalTrials.gov record. The registry classifies the hypothesis type for each analysis as superiority.
The ClinicalTrials.gov record does not specify how multiplicity across these comparisons was controlled. Therefore, the individual confidence intervals should be reported as given, but no additional claim about familywise type I error control or a global multiplicity hierarchy should be invented.
This is particularly important when several randomized comparisons appear on the same trial page. A collection of confidence intervals is not equivalent to a single hypothesis test. Whether a particular nominal comparison was confirmatory depends on the prespecified statistical design, including endpoint hierarchy, alpha allocation, and multiplicity procedures. Those details are not reported here.
16. Crossover, Interim Analysis, and Other Design Features
The ClinicalTrials.gov record supports some design conclusions directly and do not support others. This distinction is worth making explicit because these features can substantially alter interpretation of time-to-event results.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Randomization | Supported. Allocation is randomized. |
| Parallel design | Supported. Design model is parallel. |
| Masking | Supported. Masking is none. |
| Superiority | Supported for the posted statistical analyses. |
| Factorial design | Not reported in the ClinicalTrials.gov record. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Interim analysis | Not reported in the ClinicalTrials.gov record. |
| Multiplicity procedure | Not reported in the ClinicalTrials.gov record. |
| Stratification | Not reported in the ClinicalTrials.gov record. |
| Missing-data/imputation method | Not reported in the ClinicalTrials.gov record. |
| Bayesian methods | Not reported in the ClinicalTrials.gov record. |
| Non-inferiority margin | Not applicable to the registry-reported superiority analyses. |
The absence of a reported method in the ClinicalTrials.gov record should not be treated as proof that the underlying protocol or statistical analysis plan contained no such procedure. It means only that the ClinicalTrials.gov record does not establish it.
17. Understanding PFS Per BICR
The PFS endpoint is distinctive because the event is determined using disease-progression assessment based on BICR, or death from any cause if death occurs first. This means the endpoint combines radiographic disease assessment with mortality into a single time-to-event measure.
Starting point
The clock begins at randomization.
Progression event
The registry definition uses first documented disease progression based on BICR assessments per RECIST v1.1.
Death event
Death due to any cause is also an event and competes with progression as the first event defining PFS.
Follow-up limit
The registered time frame extends up to approximately 481 weeks.
The BICR component is important because endpoint classification can influence the reliability of progression assessments. The ClinicalTrials.gov record identifies BICR in the endpoint name and definition but do not provide the underlying radiologic review records or details sufficient to reproduce the assessments.
18. Understanding Overall Survival
Overall survival is defined more simply: the time from randomization to death from any cause. Participants who have not died by their last available follow-up are handled as censored observations in a conventional time-to-event framework.
OS is less dependent on the subjective classification of disease progression because death is the event of interest. At the same time, OS reflects the complete subsequent treatment pathway after randomization, so its interpretation can be influenced by treatments received after the initial assigned intervention. The ClinicalTrials.gov record does not report those subsequent-treatment details and therefore do not support a quantitative analysis of that issue here.
19. Results vs Statistical Interpretation
Reported result
For example, the Part 2 PFS comparison reports HR 0.61 with a 95% CI of 0.52–0.72.
Interpretation
The first-named group has an estimated instantaneous progression-or-death hazard 39% lower than the comparator under the reported HR framework.
What is not established
The HR does not establish a 39% absolute reduction in progression or death, a median PFS difference, or a particular patient's expected benefit.
Uncertainty
The 95% CI communicates uncertainty around the relative hazard estimate and should be considered alongside the point estimate.
20. Important Limitations
- Registry-level information: the analysis is restricted to the ClinicalTrials.gov record and does not import additional numerical results from publications.
- Incomplete statistical-method reporting: most posted analyses do not state a statistical method. One PFS analysis explicitly identifies a Cox proportional-hazard method.
- No p-values: the statistical analyses posted on ClinicalTrials.gov contain hazard ratios and 95% confidence intervals but no p-values. None are reconstructed here.
- No median event times: median PFS and median OS are not provided in the ClinicalTrials.gov record.
- No absolute survival estimates: the ClinicalTrials.gov record does not provide Kaplan-Meier survival probabilities at specific time points.
- No subgroup results: subgroup-specific hazard ratios and confidence intervals are not included in the ClinicalTrials.gov record.
- Multiple comparisons: 14 primary analyses are posted, but the ClinicalTrials.gov record does not specify a multiplicity-control strategy.
- Proportional-hazards assumption: the ClinicalTrials.gov record does not report whether this assumption was formally evaluated.
- Arm structure: the registry analysis records reference multiple Part 1, Part 2, and Part 3 arm labels. This page reports those comparisons exactly rather than reconstructing an unprovided arm-allocation schema.
- Safety denominator: the registry-reported Part 2 - Arm I serious-adverse-event entry is 180/3 and is reproduced without correction because the source data do not establish an alternative denominator.
21. Why This Trial Matters Statistically
CheckMate-227 is a useful teaching example because its the ClinicalTrials.gov record illustrates how complex randomized trials can generate multiple time-to-event comparisons while retaining a common statistical language: randomization, censoring, hazard ratios, and confidence intervals.
| Concept | How it appears in CheckMate-227 |
|---|---|
| Randomization | The registered allocation is randomized. |
| Parallel design | The design model is parallel. |
| Time-to-event endpoints | PFS per BICR and OS are the two registered primary endpoints. |
| Kaplan-Meier estimation | The PFS definition explicitly refers to Kaplan-Meier estimates. |
| Hazard ratio | All 14 registry-reported primary analyses use HR as the effect measure. |
| Confidence interval | Every registry-reported primary analysis reports a two-sided 95% CI. |
| Superiority testing | The hypothesis type for the analyses posted on ClinicalTrials.gov is superiority. |
| Multiple comparisons | Seven PFS and seven OS primary analyses are posted. |
| Analysis population | All registry-reported analyses use all randomized participants. |
| Safety analysis | Serious adverse-event counts are posted on ClinicalTrials.gov for multiple arm labels. |
The central statistical lesson is that a hazard ratio is only one component of the evidence. A careful interpretation also asks which groups were compared, which endpoint was measured, who entered the analysis, how censoring was handled, how precisely the effect was estimated, and what additional design information is available.
22. A Practical Framework for Reading Each HR
Identify the endpoint
Determine whether the HR refers to PFS per BICR or OS. The two endpoints have different event definitions.
Identify the comparison
Read both arm labels. CheckMate-227 contains several distinct randomized comparisons, so an HR cannot be interpreted without its comparator.
Read the HR
Values below 1 indicate a lower estimated instantaneous event hazard for the first-named group.
Read the confidence interval
Assess the uncertainty and whether the interval includes the reference value of 1.00.
Check what is missing
Ask whether p-values, absolute survival estimates, medians, subgroup results, and the full statistical-analysis plan are actually available in the source being used.
23. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: CheckMate-227, NCT02477826.
- Linked publication record: PubMed PMID 42579028.
- Linked publication record: PubMed PMID 39369790.
- Linked publication record: PubMed PMID 38383458.
- Linked publication record: PubMed PMID 36223558.
- Linked publication record: PubMed PMID 33485960.
Continue through the Clinical Biostats statistical pathway
Move from the trial's reported time-to-event results to deeper tutorials on survival analysis, hazard ratios, confidence intervals, and related statistical methods.
26. Record Summary
CheckMate-227 provides a useful statistical case study in interpreting a randomized phase 3 trial with multiple registered comparisons and two time-to-event primary endpoints. The ClinicalTrials.gov record reports 14 primary analyses, each using all randomized participants, hazard ratios, two-sided 95% confidence intervals, and superiority as the hypothesis type. The reported PFS hazard ratios range from 0.61 to 0.95, while the reported OS hazard ratios range from 0.64 to 0.91.
The most important statistical discipline is to preserve the distinction between these separate comparisons. An HR must be tied to its endpoint and comparator, and its interpretation should include the confidence interval rather than relying on the point estimate alone. The ClinicalTrials.gov record does not report p-values, median survival times, time-specific survival probabilities, subgroup effects, or the full multiplicity and interim-analysis framework. Those omissions matter because they limit how far the numerical results can be interpreted without introducing information from outside the ClinicalTrials.gov recordset.