This page separates reported registry information from statistical interpretation. The ClinicalTrials.gov record contains posted outcome measures but no formal statistical-analysis records.
1. Trial at a Glance
CheckMate-040 is a completed phase 1/2, non-randomized, parallel-design oncology study evaluating nivolumab or nivolumab in combination with other agents in patients with advanced hepatocellular carcinoma. The registry reports 657 participants, 9 arms, 8 registered primary endpoints, and 27 posted outcome measures.
| Feature | CheckMate-040 |
|---|---|
| Phase | Phase 1/2 |
| Condition | Hepatocellular Carcinoma |
| Allocation | NON_RANDOMIZED |
| Design model | PARALLEL |
| Masking | NONE |
| Primary purpose | TREATMENT |
| Enrollment | 657 |
| Arms | 9 |
| Interventions | Nivolumab; Sorafenib; Ipilimumab; Cabozantinib |
| Results posted | Yes |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | INDUSTRY |
| Status | COMPLETED |
2. Clinical Question
The registry describes CheckMate-040 as an immunotherapy study evaluating the effectiveness, safety, and tolerability of nivolumab or nivolumab in combination with other agents in patients with advanced liver cancer. The condition listed in the registry is hepatocellular carcinoma.
Population
Patients with hepatocellular carcinoma participating in a phase 1/2 treatment study. The ClinicalTrials.gov record describes the condition as hepatocellular carcinoma and the study population as patients with advanced liver cancer.
Interventions
The registry identifies nivolumab, sorafenib, ipilimumab, and cabozantinib among the study interventions.
Comparator structure
The study is non-randomized and has 9 arms. Therefore, it does not have the randomized treatment-versus-control structure of a conventional parallel-group confirmatory trial.
Primary question
The registry-defined questions include safety outcomes and objective response rate in specified cohorts, with response assessed by either blinded independent central review or investigator assessment.
3. Trial Design
The non-randomized structure is central to the statistical interpretation. Without randomization, comparisons among cohorts or interventions can reflect differences in the patients entering those cohorts in addition to differences associated with treatment. A descriptive response rate can therefore be informative without constituting an unbiased randomized estimate of comparative treatment effect.
4. Study Timeline and Scope
Study start
The registry lists 2012-10-30 as the study start date.
Primary completion
The registry lists 2024-11-12 as the primary completion date.
Registry status
The current registry-reported ClinicalTrials.gov record identifies the study status as COMPLETED.
The unusually broad study period and multi-cohort structure are important when interpreting the registry as a single statistical object. The 657 enrolled participants were studied across 9 arms, while the registered primary endpoints are tied to specific safety and response questions and, for the response endpoints, specific cohorts.
5. Registered Primary Endpoints
| Endpoint | Type | Registry time frame | Assessment |
|---|---|---|---|
| Number of Participants With Adverse Events (AEs) | Binary | From first dose of study medication through 100 days following last dose of study treatment | Posted results: Yes |
| Number of Participants With Serious Adverse Events (SAEs) | Binary | From first dose of study medication through 100 days following last dose of study treatment | Posted results: Yes |
| Number of Participants With Adverse Events Leading to Discontinuation | Binary | From first dose of study medication through 100 days following last dose of study treatment | Posted results: Yes |
| Number of Participants Who Died | Binary | From first dose of study medication until study closure, up to approximately 144 months | Posted results: Yes |
| Number of Participants With Laboratory Abnormalities in Specific Liver Tests | Binary | From first dose of study medication through 100 days following last dose of study treatment | Posted results: Yes |
| Number of Participants With Laboratory Abnormalities in Specific Thyroid Tests | Binary | From first dose of study medication through 100 days following last dose of study treatment | Posted results: Yes |
| Objective Response Rate (ORR) Assessed by Blinded Independent Central Review (BICR) for Cohort 2 | Binary | From start of study treatment until disease progression, or date of subsequent anti-cancer therapy, whichever occurs first | BICR; RECIST v1.1 |
| Objective Response Rate (ORR) by Investigator for Cohorts 3, 4, 5, and 6 | Binary | From date of randomization for Cohorts 3, 4, and 6, or date of first dose for Cohort 5, until disease progression, or date of subsequent anti-cancer therapy | Investigator; RECIST v1.1 |
The ClinicalTrials.gov record identifies all 8 registered primary endpoints as binary endpoints. That classification matters because the natural statistical unit for these outcomes is generally the participant-level occurrence or non-occurrence of the event or response, rather than a continuous measurement or a time-to-event quantity.
6. Endpoint Definitions
Adverse Events
The registry defines an adverse event as a new untoward medical occurrence or worsening of a preexisting medical condition in a participant administered study treatment, without requiring a causal relationship to treatment. The definition includes unfavorable or unintended signs, symptoms, or disease occurrences temporally associated with treatment.
Serious Adverse Events
A serious adverse event is defined in the registry as an untoward medical occurrence that, at any dose, results in death, is life-threatening, requires inpatient hospitalization, or causes prolongation of existing hospitalization.
Deaths
The death endpoint is the number of participants who died during the study. Its registered time frame extends from the first dose of study medication until study closure, up to approximately 144 months.
Liver laboratory abnormalities
The liver laboratory endpoint concerns abnormalities in liver function tests related to potential drug-induced liver injury. The registry identifies alanine aminotransferase (ALT) and aspartate aminotransferase (AST) among the liver function parameters assessed, with abnormalities evaluated relative to the upper limit of normal for each parameter.
Thyroid laboratory abnormalities
The thyroid laboratory endpoint concerns abnormalities in thyroid-related laboratory measurements. The registry identifies FT3 and FT4 and states that abnormalities were assessed against the lower and upper limits of normal. TSH is also identified in the endpoint definition.
Objective response rate
Objective response rate is defined as the percentage of treated participants whose best overall response is either complete response or partial response according to RECIST v1.1. For Cohort 2, the registry specifies blinded independent central review. For Cohorts 3, 4, 5, and 6, the registry specifies investigator assessment.
7. Statistical Methodology
The ClinicalTrials.gov record reports that results are posted and identifies 27 posted outcome measures, but it contains no formal statistical analyses. Consequently, the ClinicalTrials.gov record does not specify a formal hypothesis test, effect estimate, confidence interval, regression model, multiplicity procedure, or p-value for the primary endpoints.
This means the page can describe the endpoint structure and explain appropriate statistical approaches, but it cannot legitimately report a formal treatment effect or inferential comparison from the ClinicalTrials.gov record.
Binary endpoint analysis
For a binary endpoint such as ORR or the occurrence of an adverse event, a basic descriptive analysis would report the number of participants meeting the endpoint and the number evaluated, together with the resulting proportion. A confidence interval can be constructed around a proportion when the underlying participant-level data are available.
Comparative binary analysis
If two randomized groups were being compared, common approaches could include a risk difference, risk ratio, odds ratio, or an appropriate exact or asymptotic hypothesis test. CheckMate-040, however, is explicitly non-randomized, so a simple between-arm comparison would require particular caution because treatment cohorts may differ systematically.
Time-to-event analysis
The registered death endpoint is binary as specified in the registry, even though death accumulated over follow-up. If individual event times and censoring information were available, a time-to-event analysis could provide more information than a simple ever-died proportion. Such an analysis is not reported in the ClinicalTrials.gov record and is not attributed to the registry.
Response analysis
ORR is naturally summarized as a proportion of treated participants whose best overall response is CR or PR. Because the endpoint uses RECIST v1.1 and specified review methods, the assessment process is part of the endpoint definition rather than a separate statistical test.
Safety analysis
Safety endpoints such as AEs, SAEs, and discontinuations are naturally summarized by counts and percentages. Because the study is non-randomized and has multiple cohorts, interpretation should focus on the exposure-defined cohort and denominator rather than treating the entire study as a single homogeneous treatment group.
8. Planned Analysis
No formal comparative analysis is reported
The registry reports posted outcome measures, but the ClinicalTrials.gov record contains no formal statistical analyses. Therefore, this page does not state an outcome estimate, confidence interval, or p-value for any primary endpoint.
For this type of study, the appropriate analysis depends on the endpoint and the specific cohort being evaluated. The following describes the statistical framework that would ordinarily be used, without attributing these procedures to CheckMate-040 unless the ClinicalTrials.gov record explicitly identifies them.
| Endpoint type | Typical analysis | Key quantity |
|---|---|---|
| Adverse events | Descriptive binary safety analysis | Participants with at least one AE / participants evaluated |
| Serious adverse events | Descriptive binary safety analysis | Participants with at least one SAE / participants evaluated |
| Discontinuation due to AE | Descriptive binary analysis | Participants discontinuing because of an AE / participants evaluated |
| Death | Binary analysis or, if event times are available, time-to-event analysis | Proportion or survival distribution |
| Liver laboratory abnormalities | Descriptive binary laboratory-safety analysis | Participants meeting specified abnormality criteria |
| Thyroid laboratory abnormalities | Descriptive binary laboratory-safety analysis | Participants meeting specified abnormality criteria |
| ORR | Descriptive response analysis; comparative methods only if the design supports them | CR + PR / treated participants evaluated |
9. Statistical Methods Explained
Why is ORR treated as a binary endpoint?
The registered ORR endpoint classifies each treated participant according to whether the best overall response is complete response or partial response. That produces an outcome that can be coded as response versus no response for the purpose of calculating an overall response proportion.
What does BICR add to an ORR endpoint?
Blinded independent central review provides an assessment process separate from the treating investigator for the specified Cohort 2 endpoint. Statistically, the important point is that the definition of response depends not only on RECIST v1.1 but also on who performs the assessment.
Why does non-randomization matter?
Randomization is the mechanism that balances measured and unmeasured prognostic factors in expectation. In a non-randomized study, patients entering different cohorts can differ for reasons unrelated to treatment. Consequently, a difference in response proportions across cohorts cannot automatically be interpreted as a causal treatment effect.
How should a binary safety endpoint be summarized?
A safety endpoint such as an SAE is typically summarized using the number of participants affected and the number at risk or evaluated. The denominator is crucial: a rate without its analysis population can be misleading, particularly when cohorts differ in size or treatment exposure.
Why is the registered death endpoint different from a survival endpoint?
The registry classifies "Number of Participants Who Died" as binary. That means the registered measure is a participant-level occurrence of death during the specified follow-up. A conventional overall-survival analysis would instead use the timing of death and censoring information to estimate the distribution of time to death.
Why should ORR not automatically be compared across all 9 arms?
The study is non-randomized, and the registered ORR endpoints themselves refer to specific cohorts and assessment procedures. Combining or directly comparing all arms would require additional design and analysis assumptions that are not provided in the ClinicalTrials.gov record.
10. Cohort-Specific Response Assessment
| Cohort | Registered response endpoint | Assessment method | Starting point |
|---|---|---|---|
| Cohort 2 | Objective Response Rate | Blinded Independent Central Review | Start of study treatment |
| Cohorts 3, 4, and 6 | Objective Response Rate | Investigator | Date of randomization |
| Cohort 5 | Objective Response Rate | Investigator | Date of first dose |
The registry therefore does not define one universal response analysis beginning at the same time point for every cohort. The distinction between randomization and first-dose dates is particularly relevant when interpreting response follow-up, because time zero is part of the estimand being measured.
11. Safety Framework
Safety is a major component of the registered primary-endpoint structure. The study includes separate binary endpoints for adverse events, serious adverse events, adverse events leading to discontinuation, deaths, liver-test abnormalities, and thyroid-test abnormalities.
Adverse events
The registry measures participants with AEs from first dose through 100 days following the last dose of study treatment.
Serious adverse events
SAEs use the same registry-reported treatment-to-follow-up window and apply the registry's seriousness definition.
Discontinuation
The registry separately identifies participants with AEs leading to discontinuation, allowing this outcome to be distinguished from overall AE occurrence.
Laboratory monitoring
Specific liver and thyroid laboratory abnormalities are separately registered binary safety outcomes.
Serious adverse events by arm
12. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The registry provides binary endpoint definitions and reports that results are posted, but the ClinicalTrials.gov record contains no formal statistical-analysis records. Therefore, the available information supports endpoint-level interpretation rather than a reconstructed inferential treatment comparison.
Clinical interpretation
The study was designed to evaluate effectiveness, safety, and tolerability of nivolumab or nivolumab-based treatment approaches in advanced hepatocellular carcinoma. Clinical interpretation of any posted outcome must remain tied to the relevant cohort and endpoint definition.
This distinction is especially important in a multi-arm, non-randomized study. A reported response proportion can describe what happened among treated participants in a cohort. It does not, by itself, establish that treatment caused the observed response or establish superiority over another intervention.
13. Limitations
- Non-randomized design: the absence of randomized allocation limits causal interpretation of differences among treatment cohorts.
- Multiple arms: the study contains 9 arms, making the analysis structure more complex than a two-group randomized comparison.
- Cohort-specific endpoints: the registered ORR endpoints apply to specified cohorts rather than uniformly to the entire enrolled population.
- Different assessment methods: Cohort 2 uses blinded independent central review, whereas the registered ORR endpoint for Cohorts 3, 4, 5, and 6 uses investigator assessment.
- Different time origins: response follow-up begins at study treatment for Cohort 2 and at randomization or first dose for other specified cohorts.
- No registry-reported formal analyses: the ClinicalTrials.gov record contains no statistical-analysis records, so formal effect estimates, confidence intervals, and p-values cannot be responsibly reconstructed.
- Outcome heterogeneity: the primary endpoints span safety, mortality, laboratory abnormalities, and tumor response, each requiring different statistical interpretation.
- Follow-up windows: safety endpoints and mortality use different registered follow-up definitions, so their denominators and observation periods should not be conflated.
14. Why This Trial Matters Statistically
CheckMate-040 is a useful teaching example because its statistical structure differs substantially from the simple randomized two-arm trial that is often used to introduce clinical-trial analysis.
| Concept | How it appears in CheckMate-040 |
|---|---|
| Non-randomized design | Participants are not assigned through randomized allocation according to the registry. |
| Multi-arm structure | The study contains 9 arms. |
| Binary endpoints | All 8 registered primary endpoints are classified as binary. |
| Objective response rate | ORR is based on complete or partial response under RECIST v1.1. |
| Independent review | Cohort 2 ORR is assessed by blinded independent central review. |
| Investigator assessment | ORR for Cohorts 3, 4, 5, and 6 is assessed by investigators. |
| Safety analysis | AEs, SAEs, discontinuations, deaths, liver abnormalities, and thyroid abnormalities are separately registered. |
| Analysis populations | Endpoint-specific denominators are important because the study contains multiple cohorts and arms. |
| Missing inferential detail | The ClinicalTrials.gov record reports no formal statistical-analysis records. |
15. What a Complete Statistical Analysis Would Need
A full independent statistical reconstruction would require more than the endpoint names. For binary endpoints, the analyst would need the numerator and denominator for each prespecified cohort and the corresponding analysis population. For any comparative analysis, the analyst would also need the cohort definitions and the prespecified comparison framework.
The proportion alone is not sufficient to establish causality. In a non-randomized multi-arm study, interpretation also depends on how participants entered each cohort and whether the comparison was prespecified.
For the response endpoints, the analysis would additionally need the RECIST response assessments and the appropriate review population. For safety endpoints, the exposure period and at-risk denominator are essential. For a conventional survival analysis of death, participant-level event times and censoring information would be required.
16. Multiplicity, Interim Analysis, and Other Design Topics
The ClinicalTrials.gov record does not identify a formal multiplicity-control strategy, alpha-spending procedure, interim-analysis rule, non-inferiority margin, crossover procedure, factorial structure, missing-data imputation method, stratification scheme, or Bayesian analysis.
| Design topic | Supported by the ClinicalTrials.gov record? | Interpretation |
|---|---|---|
| Non-inferiority margin | No | No margin is reported in the ClinicalTrials.gov record. |
| Crossover | No | No crossover procedure is reported in the ClinicalTrials.gov record. |
| Factorial design | No | The registry identifies a parallel design, but does not describe a factorial structure. |
| Multiplicity procedure | No | No formal multiplicity procedure is reported. |
| Interim analysis | No | No formal interim-analysis specification is reported. |
| Missing-data/imputation method | No | No imputation method is reported. |
| Stratification | No | No statistical stratification factors are reported. |
| Bayesian methods | No | No Bayesian statistical method is identified. |
These omissions are not evidence that the corresponding procedures were absent from the underlying study documentation. They mean only that the ClinicalTrials.gov record does not establish them, so they are not attributed to CheckMate-040 on this page.
17. Interpreting Non-Randomized Response Rates
A response rate describes the proportion of treated participants who met the response definition. In a non-randomized study, that descriptive quantity should not automatically be interpreted as the causal effect of the intervention.
If patients entering two cohorts differ in disease characteristics, prior treatment, prognosis, or other factors, those differences can influence observed response independently of treatment. Without randomization or an appropriate causal-adjustment strategy, a raw difference in ORR can therefore mix treatment effects with selection effects.
ORR assessed by blinded independent central review and ORR assessed by investigators are related but distinct measurement processes. Differences between such estimates should not automatically be interpreted as treatment differences because the assessment mechanism itself differs.
18. Statistical Methods Explained: A Deeper Walkthrough
What is the estimand for ORR?
At its simplest, the ORR estimand is the proportion of treated participants whose best overall response is complete response or partial response under the specified RECIST v1.1 assessment. The estimand is therefore a proportion, not a hazard ratio or median time.
Why is the denominator just as important as the numerator?
Suppose a cohort has a certain number of responders. The clinical interpretation changes substantially depending on how many participants were actually evaluated. This is why arm-level safety notation such as affected/at-risk counts is more informative than a numerator alone.
What would a confidence interval for ORR tell us?
A confidence interval around a response proportion would quantify sampling uncertainty under the chosen statistical framework. It would not quantify heterogeneity among individual patients and would not transform a non-randomized cohort into a randomized comparison.
When would a risk ratio be useful?
If two comparable groups were available, a risk ratio could compare the probability of an endpoint in one group with the probability in another. In CheckMate-040, however, the non-randomized design means that a crude risk ratio across cohorts would require careful qualification.
When would a time-to-event model be preferable for death?
If the underlying data included dates of death and appropriate censoring information, a survival model could use not only whether a participant died but also when the event occurred. This can be substantially more informative than collapsing the entire follow-up into a binary death indicator.
Why should the primary endpoints not all be analyzed identically?
Although the registry classifies all 8 primary endpoints as binary, their scientific meanings differ. An adverse event, a laboratory abnormality, a death, and an objective response are not interchangeable outcomes. The analysis should respect the clinical definition and observation window of each endpoint.
19. Data Interpretation by Endpoint Family
| Endpoint family | What it answers | Primary statistical caution |
|---|---|---|
| Adverse events | How many participants experienced an AE during the specified window? | Do not confuse occurrence with causality. |
| Serious adverse events | How many participants experienced an SAE meeting the registry definition? | Denominators and exposure matter. |
| Discontinuations | How many participants stopped treatment because of an AE? | Requires clear identification of treatment exposure and reason for discontinuation. |
| Deaths | How many participants died during the registered follow-up? | A binary death count loses event timing. |
| Liver tests | How many participants developed specified liver laboratory abnormalities? | Threshold definitions and laboratory denominators matter. |
| Thyroid tests | How many participants developed specified thyroid laboratory abnormalities? | Normal-range thresholds and monitoring windows matter. |
| ORR | How many treated participants achieved CR or PR? | Assessment method, cohort definition, and non-randomization matter. |
20. Results Reporting: What Can Be Said From the Supplied Data
ClinicalTrials.gov reports that results are posted for all 8 registered primary endpoints. However, the ClinicalTrials.gov record does not include the actual statistical-analysis records or endpoint-level result estimates needed to reproduce those results here.
Registry reporting status
Formal statistical analyses in the ClinicalTrials.gov record: none.
Accordingly, no endpoint-specific treatment estimate, confidence interval, or p-value is reported in this analysis.
This distinction is important. "Results posted" indicates that the registry contains posted outcome information. It does not, by itself, establish that a particular inferential statistical test was used or that a specific treatment comparison was statistically significant.
21. Why No Conventional Results Section Is Presented
The usual results section on a clinical-trial statistical page would present the numerical estimate, its confidence interval, and its p-value for each primary comparison. The registry-reported CheckMate-040 data contains no formal statistical analyses, so inserting such values would require importing information outside the ClinicalTrials.gov record.
22. Why This Trial Matters Statistically
CheckMate-040 demonstrates why clinical-trial statistics cannot be reduced to a single p-value. The study combines a non-randomized multi-arm design with binary safety and response endpoints, different response-assessment procedures, cohort-specific time origins, and a long overall study period.
The most important statistical lesson is that the design determines what a reported number can legitimately mean. A response proportion can be an important descriptive measure of activity within a cohort. It is not automatically evidence of a causal treatment effect, and it is not interchangeable with a randomized hazard ratio or a survival probability.
The second major lesson is the importance of endpoint definitions. "Number of participants with an adverse event," "number of participants who died," and "objective response rate" all have binary classifications, but they answer very different clinical questions and use different observation frameworks.
23. Related Tutorials
Learn more about the methods used to understand this trial:
24. Related Calculators
25. Sources
- ClinicalTrials.gov: CheckMate-040, NCT01658878.
- Linked publication: PubMed record for PMID 36852452: PubMed 36852452.
- Linked publication: PubMed record for PMID 36512738: PubMed 36512738.
- Linked publication: PubMed record for PMID 36323435: PubMed 36323435.
- Linked publication: PubMed record for PMID 34051329: PubMed 34051329.
- Linked publication: PubMed record for PMID 33001135: PubMed 33001135.
The linked PubMed records are provided because they are identified in the ClinicalTrials.gov record. No publication-specific numerical results are imported into this page unless they are also present in the ClinicalTrials.gov record.
Continue through the Clinical Biostats statistical library
Use the trial as a starting point for deeper study of response endpoints, safety analysis, binary outcomes, and survival methods.
26. Record Summary
CheckMate-040 is a completed phase 1/2, non-randomized, parallel study enrolling 657 participants across 9 arms. Its 8 registered primary endpoints are binary and span adverse events, serious adverse events, treatment discontinuations due to adverse events, deaths, liver and thyroid laboratory abnormalities, and objective response rate in specified cohorts. The response endpoints use RECIST v1.1 and distinguish blinded independent central review from investigator assessment.
From a statistical perspective, the most important feature is the distinction between descriptive outcome reporting and causal comparative inference. The ClinicalTrials.gov record confirms that results were posted but provides no formal statistical-analysis records.