← Clinical Trials
Advanced Cancer Phase 1/2 Completed NCT02488759

CheckMate-358: Complete Statistical Analysis of Nivolumab in Advanced Cancer

An independent statistical review of the randomized phase 1/2 CheckMate-358 study of nivolumab and nivolumab combination therapy in virus-associated tumors and various advanced cancer, based strictly on the information reported in the ClinicalTrials.gov record.

CheckMate-358  ·  NCT02488759  ·  Randomized  ·  578 participants
Scope of this record

This page separates posted trial information from statistical interpretation. The ClinicalTrials.gov record reports eight outcome measures and results, but no formal statistical analyses are included in the posted results.

Registry record: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

View CheckMate-358 on ClinicalTrials.gov

1. Trial at a Glance

CheckMate-358 is a completed randomized, parallel-design phase 1/2 study investigating nivolumab and nivolumab combination therapy in virus-associated tumors and various advanced cancer. The ClinicalTrials.gov record describes five arms, four registered primary endpoints, eight posted outcome measures, and total enrollment of 578 participants.

578
Enrollment
Participants
5
Arms
Randomized parallel design
4
Primary endpoints
Registered
8
Outcomes posted
ClinicalTrials.gov
FeatureCheckMate-358
Trial nameCheckMate-358
Brief titleAn Investigational Immuno-therapy Study to Investigate the Safety and Effectiveness of Nivolumab, and Nivolumab Combination Therapy in Virus-associated Tumors
PhasePhase 1/2
StatusCOMPLETED
Therapeutic areaOncology
ConditionVarious Advanced Cancer
AllocationRANDOMIZED
Design modelPARALLEL
MaskingNONE
Primary purposeTREATMENT
Enrollment578.0
Arms5
InterventionsNivolumab; Ipilimumab; Relatlimab; Daratumumab
Lead sponsorBristol-Myers Squibb
Sponsor typeINDUSTRY
Start2015-10-13
Primary completion2021-03-19
ClinicalTrials.gov identifierNCT02488759

2. Clinical Question

The ClinicalTrials.gov record describes a study investigating the safety and effectiveness of nivolumab and nivolumab combination therapy in virus-associated tumors, within a broader condition description of various advanced cancer. Because the ClinicalTrials.gov record does not provide a single pooled comparator regimen or a single efficacy hypothesis spanning all five arms, the statistical question is best understood at the level of the registered cohorts and endpoints rather than as one conventional two-arm superiority comparison.

Population

Participants enrolled under the condition description Various Advanced Cancer, including a neoadjuvant cohort and metastatic cohorts represented in the posted safety data.

Interventions

The ClinicalTrials.gov record lists nivolumab, ipilimumab, relatlimab, and daratumumab as interventions.

Design question

The trial used randomized allocation with a parallel design and no masking.

Primary questions

The registered primary endpoints focus on drug-related adverse events, serious adverse events, surgery delay in the neoadjuvant cohort, and investigator-assessed objective response rate in the metastatic setting.

3. Trial Design

01
Enroll 578 participants
02
Randomize Randomized allocation
03
Parallel arms 5 arms
04
Assess Safety and effectiveness
05
Report 8 outcomes posted
Allocation
RANDOMIZED
Design model
PARALLEL
Masking
NONE
Primary purpose
TREATMENT

These design descriptors matter statistically because randomization creates the basis for comparing assigned groups, while the parallel design means participants are evaluated within their assigned study pathways rather than being described as a crossover or factorial experiment. The absence of masking also means that neither participants nor investigators were assigned a blinded treatment condition according to the registry's masking field.

Study timeline

2015-10-13

Study start

The trial began on October 13, 2015.

2021-03-19

Primary completion

The registry lists March 19, 2021 as the primary completion date.

Completed

Current trial status in the ClinicalTrials.gov record

The trial status is listed as COMPLETED.

4. Endpoints

The ClinicalTrials.gov record identifies four registered primary endpoints. All four are binary endpoints in the registry-reported endpoint classification. The registry also states that results were posted for each of these four primary endpoints.

Primary endpoint Time frame Endpoint type Definition reported in registry data
Neoadjuvant: Number of Participants With Drug-Related Select Adverse Events (AEs) From first dose to 30 days post last dose (Up to 2 months) Binary Number of participants with any grade of drug-related select adverse events including endocrine, gastrointestinal, hepatic, pulmonary, renal, skin, and hypersensitivity AEs in the neoadjuvant cohort.
Neoadjuvant: Number of Participants With Drug-Related Serious Adverse Events (SAEs) From first dose to 30 days post last dose (Up to 2 months) Binary Number of participants with any grade of drug-related serious adverse events in the neoadjuvant cohort.
Neoadjuvant: Rate of Surgery Delay Day 29 Binary Percentage of participants in the neoadjuvant cohort with surgery delayed > 4 weeks from the planned surgery date or planned start date for chemoradiation due to a drug-related adverse event. The registry-reported definition identifies HPV positive squamous cell carcinoma of the head and neck, HPV negative squamous cell carcinoma of the head and neck, and cervical cancer among the diseases assessed.
Metastatic: Investigator-Assessed Objective Response Rate (ORR) From the date of first dose to the date of the initial objectively documented tumor progression or the date of the last reported assessment in the registry wording Binary Percentage of participants with a best overall response of confirmed complete response or partial response using RECIST 1.1 criteria.
Registry endpoint thresholds for metastatic ORR: the registry-reported definition states that an ORR in excess of 10% will be considered of clinical interest, and an ORR of 25% or greater will be considered of strong clinical interest.

Why the endpoint structure matters

These endpoints are not all measuring the same clinical construct. The first two quantify adverse-event occurrence, the third captures an operational consequence of toxicity in the neoadjuvant setting, and the fourth measures tumor response in metastatic disease. Treating all four as if they were interchangeable efficacy endpoints would obscure the actual statistical structure of the trial.

5. Statistical Analysis Status

The ClinicalTrials.gov record states that Results Posted = Yes, but no formal statistical analyses were posted. That distinction is important. A registry can contain posted outcome data without supplying the formal treatment-comparison analysis objects needed to reproduce a complete inferential analysis.

No formal comparative estimates were reported. The ClinicalTrials.gov record contain no hazard ratios, odds ratios, risk ratios, confidence intervals, standard errors, or p-values. This page therefore does not create or infer any of those quantities. In particular, the presence of five arms does not justify calculating a pairwise comparison without knowing the prespecified comparison structure and the underlying participant-level or arm-level outcome data.

6. Posted Primary Endpoint Results and Planned Analysis

The registry indicates that results were posted for all four registered primary endpoints. However, because the ClinicalTrials.gov recordset does not include formal statistical analyses, the appropriate approach is to distinguish the reported outcome definitions from the statistical methods that would ordinarily be used for those binary outcomes.

Neoadjuvant drug-related select adverse events

Registered endpoint

Binary outcome

Number of participants with any grade of drug-related select adverse events

Time frame: From first dose to 30 days post last dose (Up to 2 months)

The endpoint is a participant-level binary outcome: a participant either meets the registry definition of having at least one qualifying drug-related select adverse event during the specified observation window or does not. The ClinicalTrials.gov record confirms that results were posted for this endpoint, but do not provide a numerical result or formal statistical comparison in the posted results.

Typical binary-endpoint analysis
Risk = participants with the event / participants evaluated

For a randomized comparison, an analysis might report event proportions by arm together with a risk difference, risk ratio, odds ratio, or an appropriate exact or asymptotic confidence interval. The correct choice depends on the prespecified analysis plan, sample size, number of arms, and comparison structure.

Neoadjuvant drug-related serious adverse events

Registered endpoint

Binary outcome

Number of participants with any grade of drug-related serious adverse events

Time frame: From first dose to 30 days post last dose (Up to 2 months)

The registry defines this endpoint as the number of participants with any grade of drug-related serious adverse events in the neoadjuvant cohort. The ClinicalTrials.gov record states that results were posted but does not provide a formal comparative analysis.

Neoadjuvant rate of surgery delay

Registered endpoint

Day 29

Surgery delay > 4 weeks because of a drug-related adverse event

Binary endpoint expressed as a percentage of participants in the neoadjuvant cohort

This endpoint has a particularly clear operational definition. It is not simply a measure of whether an adverse event occurred; it asks whether surgery was delayed by more than four weeks from the planned surgery date or planned start date for chemoradiation because of a drug-related adverse event. That distinction makes the endpoint clinically different from the broader adverse-event measures.

Metastatic investigator-assessed ORR

Registered endpoint

ORR

Confirmed complete response or partial response using RECIST 1.1 criteria

Assessment begins at first dose and continues to initial objectively documented tumor progression or the last reported assessment in the registry wording.

The registry-reported definition identifies objective response rate as the percentage of participants whose best overall response is confirmed complete response or partial response under RECIST 1.1. The record also specifies two interpretive thresholds: an ORR in excess of 10% is considered of clinical interest, while an ORR of 25% or greater is considered of strong clinical interest.

Because ORR is binary rather than time-to-event, Kaplan-Meier estimation is not the natural primary analysis for this endpoint. The core calculation is the proportion of evaluable participants achieving the defined response, followed by an appropriate confidence interval and, if prespecified, a comparison against a reference value or between randomized groups.

7. Clinical Biostats Interpretation: Why No Effect Estimate Is Reported

Statistical interpretation

A binary endpoint can support a clear effect estimate, but only when the underlying event counts, denominators, and comparison framework are available. The registry-reported CheckMate-358 data do not provide the statistical-analysis records needed to establish such a comparison for the four registered primary endpoints.

What this does not mean

The absence of a reported hazard ratio, odds ratio, confidence interval, or p-value in the ClinicalTrials.gov record does not mean that the trial had no statistical analysis. It means that those formal analyses are not contained in the trial-data object in the ClinicalTrials.gov record.

Why the distinction matters

Calculating a new p-value from incomplete registry information could produce a number that appears authoritative while answering a different question from the prespecified analysis. An educational statistical analysis should distinguish what is documented from what would normally be calculated.

8. Safety Results: Serious Adverse Events by Arm

the ClinicalTrials.gov record reports serious adverse events by named arm or cohort. These data are the clearest numerical safety results provided in the trial data.

Arm / cohort Affected At risk
Metastatic Monotherapy 59 113
Metastatic Combo A 130 195
Metastatic Combo B 99 133
Metastatic Combo C 5 8
Metastatic Combo D 5 6
Neoadjuvant 40 123

The denominators vary substantially across these reported groups. That matters when interpreting raw event counts: an affected count cannot be compared directly across groups without considering the corresponding number at risk. The ClinicalTrials.gov record also do not identify a formal statistical comparison among these safety groups.

Small denominators require caution. The reported groups with 8 and 6 participants at risk are particularly small. Even a single additional event would represent a large change in the observed proportion, so raw percentages from such groups would be unstable. The ClinicalTrials.gov record does not provide confidence intervals for these safety estimates.

9. Safety Analysis Explained

What does "affected / at risk" represent?

The affected count is the number of participants experiencing the reported serious adverse-event outcome, while the at-risk value is the denominator reported with that result. Together, they provide the basic ingredients for an observed event proportion.

Why should counts and denominators be shown together?

A count without its denominator is difficult to interpret. For example, 5 events could represent a very different safety frequency depending on whether the underlying group contained 6 participants or a substantially larger population. The registry-reported CheckMate-358 data explicitly provide both values, which allows the reader to see the scale of each cohort.

What statistical comparison would normally be used?

For a simple binary safety outcome, investigators might compare proportions using a risk difference, risk ratio, or odds ratio, with a confidence interval. With small expected cell counts, exact methods may be preferable. In a multi-arm trial, however, the correct comparison must follow the prespecified analysis plan rather than being selected after looking at the observed counts.

Why is "statistically different" not the same as "clinically different"?

A statistical test addresses evidence against a specified null hypothesis. Clinical interpretation additionally depends on the magnitude of the difference, the seriousness and reversibility of the events, the precision of the estimate, and the broader benefit-risk context. A p-value alone cannot supply that context.

10. Statistical Methodology

Binary endpoint analysis

All four registered primary endpoints are classified as binary in the ClinicalTrials.gov record. Binary endpoints reduce an outcome to two states according to a prespecified definition. For this trial, those definitions include presence or absence of selected drug-related adverse events, presence or absence of drug-related serious adverse events, whether surgery was delayed by more than four weeks for the specified reason, and whether a participant achieved a confirmed complete or partial response.

Core binary-endpoint quantity
p = x / n

where x is the number of participants meeting the endpoint definition and n is the relevant analysis denominator.

Risk difference

When two groups are compared, the risk difference is the difference between their observed event probabilities. It is directly interpretable on the probability scale and answers the question of how much the event frequency differs between groups.

Conceptual form
RD = p1 − p0

A risk difference of zero corresponds to equal observed event probabilities. The ClinicalTrials.gov record does not provide enough information to calculate a prespecified CheckMate-358 risk difference for the primary endpoints.

Risk ratio

A risk ratio compares event probabilities multiplicatively. A value below 1 indicates a lower observed event probability in the numerator group relative to the reference group, while a value above 1 indicates a higher probability.

Odds ratio

An odds ratio compares the odds of an event rather than the event probabilities themselves. Odds ratios can be useful in binary-outcome models, but they should not be casually described as risk ratios, particularly when the outcome is common.

Confidence intervals

A confidence interval describes statistical uncertainty around an estimated quantity under the chosen model and sampling framework. It is more informative than a point estimate alone because it shows how precisely the study estimates the underlying effect. No confidence intervals for the CheckMate-358 primary endpoints are reported in the ClinicalTrials.gov record.

Multiplicity in a multi-endpoint, multi-arm study

CheckMate-358 has four registered primary endpoints and five arms. Whenever several hypotheses or comparisons are evaluated, the distinction between an individual nominal p-value and the overall familywise error rate becomes important. A formal multiplicity strategy would need to be specified in the statistical analysis plan. The ClinicalTrials.gov record does not report such a strategy, so none is assumed here.

11. Statistical Methods Explained

Why is randomization important in CheckMate-358?

Randomization assigns participants to study groups using a chance-based allocation mechanism. Its statistical value is that, in expectation, measured and unmeasured baseline factors are balanced across randomized groups. This supports causal interpretation of between-group comparisons when the trial is analyzed according to the randomized assignment and the prespecified design.

Why are the primary endpoints analyzed as binary outcomes?

The registry definitions convert each primary endpoint into a yes-or-no participant outcome. A participant either has a qualifying drug-related select adverse event or does not; either has a qualifying serious adverse event or does not; either experiences the specified surgery delay or does not; and either achieves a confirmed complete or partial response or does not.

Why is ORR different from progression-free survival?

ORR asks whether a participant achieves a specified tumor response. It does not directly incorporate the amount of time until progression. A time-to-event endpoint such as progression-free survival would instead account for the timing of progression or another specified event and for censoring. The registry-reported CheckMate-358 primary metastatic endpoint is ORR, not a time-to-event endpoint.

Why does RECIST 1.1 matter for ORR?

The response endpoint is only reproducible if "response" has a defined operational rule. The registry definition specifies RECIST 1.1 and requires a best overall response of confirmed complete response or partial response. That standardization reduces ambiguity in deciding whether a participant meets the binary response endpoint.

Why should the surgery-delay endpoint not be treated as an ordinary adverse-event rate?

The surgery-delay endpoint adds a clinical consequence and a time threshold to the adverse-event concept. It requires the surgery to be delayed by more than four weeks from the planned date or planned start date for chemoradiation and attributes the delay to a drug-related adverse event. That makes it a distinct endpoint rather than a duplicate count of adverse events.

Why can the safety groups not simply be ranked by event count?

Because the denominators differ. The reported serious-adverse-event counts range from 5 to 130, but the corresponding groups range from 6 to 195 participants at risk. Comparing raw counts without denominators would therefore be statistically misleading. Even comparing proportions would require caution because several groups are small.

12. What the Registry Data Can and Cannot Establish

QuestionWhat the ClinicalTrials.gov record establishesWhat they do not establish
Was the study randomized? Yes. The allocation field is RANDOMIZED. The exact randomization ratio or allocation algorithm.
How many participants were enrolled? 578.0. Exact arm-specific enrollment for all five arms.
How many arms were specified? 5. A complete arm-to-intervention mapping.
Were results posted? Yes. A complete formal statistical-analysis record.
What primary endpoints were registered? Four binary endpoints are listed. Formal hypothesis-test results for those endpoints.
Are safety counts available? Yes, serious adverse-event affected/at-risk counts are posted on ClinicalTrials.gov for six named groups or cohorts. Confidence intervals, p-values, or prespecified pairwise safety comparisons.
Is a treatment-effect estimate available? Not in the posted results. Hazard ratios, risk ratios, odds ratios, risk differences, or corresponding confidence intervals.

13. Limitations

14. Why This Trial Matters Statistically

CheckMate-358 is a useful statistical teaching case because its registry structure illustrates a problem that appears frequently in complex clinical research: the presence of reported outcomes does not necessarily mean that the public registry record contains every component needed to reconstruct the formal inferential analysis.

ConceptHow it appears in CheckMate-358
RandomizationThe trial allocation is listed as RANDOMIZED.
Parallel designThe design model is PARALLEL.
Binary endpointsAll four registered primary endpoints are classified as binary in the ClinicalTrials.gov record.
Safety analysisSerious adverse-event counts and denominators are reported for named groups and cohorts.
Response analysisMetastatic ORR is defined using confirmed complete response or partial response under RECIST 1.1.
Denominator awarenessSafety counts must be interpreted alongside their corresponding at-risk denominators.
MultiplicityFour primary endpoints and five arms create a setting where formal multiplicity handling could be important.
Analysis transparencyPosted results and formal statistical analyses are distinct fields in the ClinicalTrials.gov record.
Small-sample inferenceSeveral reported safety cohorts have very small denominators, limiting the stability of percentage estimates.

15. Endpoint-Specific Statistical Interpretation

Select adverse events

This endpoint measures whether a participant experienced at least one qualifying drug-related select adverse event during the specified follow-up period. It is fundamentally a binary safety outcome.

Serious adverse events

This endpoint captures the occurrence of any drug-related serious adverse event in the neoadjuvant cohort during the defined observation window.

Surgery delay

This endpoint measures a clinically consequential delay exceeding four weeks when the delay is attributed to a drug-related adverse event.

Objective response

ORR measures the proportion achieving confirmed complete or partial response under RECIST 1.1 rather than measuring time until progression or death.

16. How a Formal Analysis Could Be Structured

The following describes the statistical logic that would ordinarily accompany the endpoints posted on ClinicalTrials.gov. It is not a reconstruction of the CheckMate-358 statistical analysis plan.

Step 1: Define the analysis population

The analysis would first identify which participants belong in the denominator for each endpoint. Safety endpoints may use an exposed or treated population, whereas efficacy analyses can follow the randomized population or another prespecified response-evaluable population. The ClinicalTrials.gov record does not specify the complete analysis-population definitions.

Step 2: Calculate the observed event proportion

For each binary endpoint, the number of participants meeting the endpoint definition is divided by the appropriate denominator. This provides the basic observed rate for each arm or cohort.

Step 3: Quantify uncertainty

A confidence interval can be constructed around each proportion. With small denominators, exact binomial intervals may be considered rather than relying on large-sample normal approximations.

Step 4: Define the comparison

In a multi-arm randomized study, the comparison must be prespecified. Possible approaches include comparison against a reference arm, pairwise comparisons, or a global multi-arm hypothesis. The ClinicalTrials.gov record does not specify which approach was prespecified for the four primary endpoints.

Step 5: Address multiplicity

If several hypotheses are tested, the statistical plan may control the familywise type I error rate using a hierarchical testing strategy, alpha allocation, adjustment procedures, or another prespecified method. No such strategy is reported here.

Step 6: Report effect size and precision

A complete results presentation would ideally include the observed proportions, an interpretable effect measure, its confidence interval, and the corresponding hypothesis-test result where applicable. The registry-reported CheckMate-358 statistical-analysis data do not contain those quantities.

17. Understanding the Posted Safety Counts

Serious adverse events: affected participants
Metastatic Monotherapy
59
Metastatic Combo A
130
Metastatic Combo B
99
Metastatic Combo C
5
Metastatic Combo D
5
Neoadjuvant
40

The chart above displays the reported number of affected participants, not a risk-adjusted or percentage comparison. It is intentionally not presented as a ranking of safety risk because the denominators differ substantially among the reported groups.

For example, the ClinicalTrials.gov record reports 5 affected participants among 8 at risk for Metastatic Combo C and 5 among 6 at risk for Metastatic Combo D. The equal event counts do not imply equal event frequencies because the denominators are different.

18. Randomization as the Central Design Concept

The trial's statistical concepts identifies Randomization as the key statistical concept associated with this trial. Randomization is especially important in a study with multiple treatment pathways because it provides the framework for comparing groups while reducing systematic allocation bias.

However, randomization does not automatically solve every statistical problem. The analysis still needs a clearly defined estimand, an appropriate analysis population, valid endpoint definitions, adequate handling of missing observations, and an explicit strategy for multiple comparisons when several hypotheses are considered.

Educational takeaway

Randomization tells us how treatment assignment was generated. It does not, by itself, tell us which statistical test should be used for every endpoint. The endpoint definition and prespecified analysis plan determine the appropriate inferential method.

19. Primary Endpoint Thresholds and Interpretation

The metastatic ORR endpoint includes explicit interpretive thresholds in the registry definition. These thresholds are particularly useful for teaching the distinction between an observed response proportion and a prespecified benchmark.

ORR thresholdRegistry wordingStatistical interpretation
In excess of 10% Considered of clinical interest An observed ORR above this threshold would meet the registry's stated descriptive criterion. The ClinicalTrials.gov record does not provide a formal hypothesis-test result against this threshold.
25% or greater Considered of strong clinical interest An observed ORR at or above this threshold would meet the registry's stated descriptive criterion. This does not by itself establish statistical significance.

This distinction is important. A clinical-interest threshold is not automatically equivalent to a null hypothesis, and crossing a descriptive threshold does not by itself establish a confidence interval or p-value. Formal inference requires the observed data, the specified null hypothesis, and the prespecified statistical method.

20. Limitations of Registry-Based Statistical Reconstruction

ClinicalTrials.gov is designed to provide structured information about clinical studies, including their design and outcome measures. It is not necessarily a complete substitute for the statistical analysis plan, protocol, participant-level dataset, or full clinical-study report.

For CheckMate-358, this limitation is particularly relevant because the ClinicalTrials.gov record contains a rich set of endpoint definitions and safety counts but an empty formal statistical-analysis array. A responsible educational analysis therefore has to resist the temptation to fill the gaps with assumptions.

What can be described

The design, population description, enrollment, number of arms, interventions, endpoint definitions, time frames, posted-result status, and registry-reported safety counts.

What should not be inferred

Formal effect estimates, confidence intervals, p-values, multiplicity adjustments, interim-analysis rules, missing-data methods, or exact arm-level statistical comparisons not contained in the ClinicalTrials.gov record.

21. Related Tutorials

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Sources

Continue through the Clinical Biostats statistical pathway

Explore the underlying trial-design and binary-outcome concepts used to interpret randomized clinical studies.

24. Record Summary

CheckMate-358 is a completed randomized, parallel-design phase 1/2 oncology study with 578.0 enrolled participants, five arms, four registered primary endpoints, and eight posted outcome measures. The primary endpoints are binary and span drug-related select adverse events, drug-related serious adverse events, surgery delay in the neoadjuvant cohort, and investigator-assessed objective response rate in the metastatic setting.

The most important statistical lesson from the ClinicalTrials.gov record is the distinction between reported outcomes and reported statistical analyses. The trial data confirm that results were posted, but the registry-reported statistical-analysis array is empty.

Clinical Biostats methodology: A rigorous trial-results page should make the statistical structure understandable while preserving the boundary between documented evidence and educational explanation.