← Clinical Trials
Plasma Cell Myeloma Phase 3 Randomized NCT02990338

ICARIA-MM: Complete Statistical Analysis of Isatuximab in Multiple Myeloma

An independent statistical analysis of the randomized phase 3 ICARIA-MM trial comparing isatuximab, pomalidomide, and dexamethasone with pomalidomide and dexamethasone in refractory or relapsed and refractory multiple myeloma.

Phase 3  ·  Completed  ·  Enrollment 307  ·  Primary endpoint: Progression-Free Survival
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information from that registry record.

1. Trial at a Glance

ICARIA-MM was a randomized, parallel, phase 3 study comparing isatuximab plus pomalidomide and dexamethasone (IPd) with pomalidomide and dexamethasone (Pd) in patients with refractory or relapsed and refractory multiple myeloma. The registered primary endpoint was progression-free survival (PFS), a time-to-event outcome analyzed in the intention-to-treat population.

307
Enrolled
Randomized trial
2
Arms
Parallel design
0.596
PFS HR
95% CI 0.436–0.814
0.0005
PFS P-value
Two-sided log-rank
FeatureICARIA-MM
Trial nameICARIA-MM
NCT identifierNCT02990338
Therapeutic areaHematology
ConditionPlasma Cell Myeloma
PhasePhase 3
StatusCompleted
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment307.0
Arms2
Lead sponsorSanofi
Sponsor typeIndustry
Start2016-12-22
Primary completion2018-11-22
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted24
Statistical analyses posted3

2. Clinical Question

The central statistical question was whether adding isatuximab to pomalidomide and dexamethasone changed progression-free survival compared with pomalidomide and dexamethasone alone in patients with refractory or relapsed and refractory multiple myeloma.

Population

Patients with refractory or relapsed and refractory multiple myeloma, as described by the trial's brief title and registry condition.

Intervention

Isatuximab, pomalidomide, and dexamethasone (IPd).

Comparator

Pomalidomide and dexamethasone (Pd).

Primary question

Does the addition of isatuximab improve progression-free survival relative to pomalidomide and dexamethasone?

This question is a classic randomized time-to-event comparison. The important statistical distinction is that PFS is not simply a proportion of participants who eventually progress. Each participant contributes information over time, with an event occurring at progression or death and censoring possible when neither event has been observed by the relevant assessment point.

3. Trial Design

01
Randomize 307.0 enrolled
02
Two arms IPd vs Pd
03
Follow Progression / death
04
Analyze ITT population
05
Compare Hazard ratio
ARM 1 · IPd

Isatuximab combination

  • Isatuximab
  • Pomalidomide
  • Dexamethasone
ARM 2 · Pd

Pomalidomide combination

  • Pomalidomide
  • Dexamethasone
Allocation
Randomized allocation was used, providing the design basis for an intention-to-treat comparison.
Structure
The registry describes a two-arm parallel design.
Masking
Masking was listed as none.
Primary purpose
Treatment.

The registry identifies the trial as a phase 3 randomized study with two parallel arms. Because the trial was unmasked, the statistical interpretation should distinguish the randomized treatment comparison from any potential effects that knowledge of assignment could have on assessments or subsequent care. The ClinicalTrials.gov record does not provide additional details that would allow those effects to be quantified here.

4. Randomization, Stratification, and Analysis Population

The formal efficacy analyses reported in the registry were conducted in the intention-to-treat (ITT) population, defined in the primary analysis as including all randomized participants. This is central to the interpretation of the randomized comparison because treatment assignment, rather than treatment actually received, defines the groups being compared.

Analysis featureRegistry-supported description
Primary efficacy populationIntent-to-treat population including all randomized participants.
PFS comparisonPd versus IPd.
OS comparisonPd versus IPd.
ORR comparisonPd versus IPd.
Stratification factorsAge (<75 years versus ≥75 years) and number of previous lines of therapy (2 or 3 versus >3), according to IRT.

The registry's analysis notes repeatedly identify age and number of previous lines of therapy as stratification factors. This matters because the reported time-to-event comparisons were not treated as if every participant came from a single homogeneous population. Stratified analysis preserves the planned structure of the randomized comparison while accounting for the factors used in the analysis.

ITT interpretation: An ITT analysis answers a treatment-assignment question: what was the outcome among participants according to the groups to which they were randomized? It is therefore different from a per-protocol or purely treatment-exposure comparison.

5. Endpoints

The registry lists one primary endpoint: Progression Free Survival (PFS). The posted analyses also include Overall Survival (OS) and Overall Response Rate (ORR) as secondary outcomes.

EndpointRegistry definition / time frameStatistical approach
Progression Free Survival (PFS) Time from date of randomization to date of first documentation of progressive disease (PD) determined by Independent Response Committee (IRC) or date of death from any cause, whichever comes first. If progression or death was not observed, participant was censored at date of last progression-free tumor assessment prior to study cut-off date. Kaplan-Meier method; log-rank comparison; hazard ratio.
Overall Survival (OS): Final Analysis From the date of randomization to date of death from any cause or data cut-off date, whichever was earlier. Stratified log-rank comparison; hazard ratio.
Overall Response Rate (ORR): Percentage of Participants With Disease Response as Per Independent Response Committee (IRC) From the date of randomization to the date of first documentation of progression or initiation of further anti-myeloma treatment. Stratified Cochran-Mantel-Haenszel test.

6. Primary Endpoint: Progression-Free Survival

PFS was the registered primary endpoint and was analyzed as a time-to-event outcome. The event was the first of two possibilities: progression determined by the Independent Response Committee or death from any cause. Participants without either event were censored at the date of their last progression-free tumor assessment before the study cut-off date.

Hazard ratio for progression or death

0.596

95% CI: 0.436–0.814   ·   P = 0.0005

Two-sided log-rank test   ·   Superiority hypothesis

Primary PFS analysisPdIPd
Analysis populationIntent-to-treat population, including all randomized participants
ComparisonPomalidomide + dexamethasone vs isatuximab + pomalidomide + dexamethasone
MethodLog-rank test
Effect measureHazard ratio
Hazard ratio0.596
95% confidence interval0.436–0.814
P-value0.0005
HypothesisSuperiority
Clinical Biostats interpretation

The reported PFS hazard ratio of 0.596 means that, under the time-to-event comparison represented by the reported hazard ratio, the estimated hazard of progression or death in the IPd group was about 40.4% lower than in the Pd group. This is a relative hazard statement, not a statement that 40.4% of participants avoided progression.

The estimate does not mean that every individual participant experienced the same reduction in risk, nor does it mean that the probability of progression or death was reduced by exactly 40.4% at every particular time point. The hazard ratio summarizes a relative comparison of event rates over the analyzed follow-up under the statistical framework used.

The 95% confidence interval of 0.436–0.814 describes uncertainty around the estimated hazard ratio. It is substantially below 1, but it is not an interval containing the individual treatment effects experienced by patients. It describes uncertainty in the estimated population-level relative effect under the analysis framework.

The P-value of 0.0005 addresses the statistical evidence against the null hypothesis under the specified testing procedure. It does not measure the magnitude of the treatment effect, does not describe clinical importance, and does not give the probability that the treatment works.

Because PFS is subject to censoring, the analysis depends on how follow-up and censoring are handled. The registry states that participants without progression or death were censored at their last progression-free tumor assessment before the study cut-off date. Interpretation of a hazard ratio also relies on the time-to-event modeling framework; a single hazard ratio can be less informative if relative hazards vary substantially over time.

Why Kaplan-Meier estimation was appropriate

The registry states that PFS was analyzed by the Kaplan-Meier method. This is appropriate for a time-to-event endpoint because participants may enter the analysis with different amounts of observable follow-up and some may be censored before experiencing the event.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

At each observed event time, the estimated survival probability is updated according to the number of events and the number of participants at risk immediately before that time. For PFS, the corresponding survival function represents remaining progression-free without death through time t.

The registry also states that confidence intervals for Kaplan-Meier estimates were calculated using a log-log transformation of the survival function and methods of Brookmeyer and Crowley. That is a detail about how uncertainty around the estimated survival function was constructed; it is separate from the hazard-ratio comparison itself.

7. Stratified Log-Rank Analysis

The primary PFS comparison used a log-rank test, with stratification based on age and number of previous lines of therapy according to the IRT. The same stratification structure is reported for the OS analysis.

Age stratum

<75 years versus ≥75 years.

Previous therapy stratum

2 or 3 previous lines versus >3 previous lines.

Comparison

The randomized Pd and IPd groups were compared within the time-to-event framework.

Effect measure

Hazard ratio, with a two-sided 95% confidence interval for the primary PFS analysis.

A stratified log-rank test is useful when the trial has prespecified factors that may influence the event-time distribution. Rather than ignoring those factors, the test incorporates the strata into the comparison. The resulting inference is still about the randomized treatment comparison, not about whether the stratification factors themselves are beneficial or harmful.

8. Secondary Endpoint: Overall Survival

Overall Survival (OS): Final Analysis was a secondary time-to-event outcome. The registry defines the endpoint as time from randomization to death from any cause or the data cut-off date, whichever was earlier. The analysis was performed in the ITT population.

Hazard ratio for overall survival

0.776

95% CI: 0.594–1.015   ·   P = 0.0319

Two-sided log-rank test   ·   Superiority hypothesis

OS final analysisRegistry result
Analysis populationITT population
ComparisonPd vs IPd
MethodLog-rank test
Effect measureHazard ratio
Estimate0.776
95% confidence interval0.594–1.015
P-value0.0319
Hypothesis typeSuperiority
StratificationAge (<75 years versus ≥75 years) and number of previous lines of therapy (2 or 3 versus >3) according to IRT
Clinical Biostats interpretation

The OS hazard ratio of 0.776 corresponds to an estimated hazard of death about 22.4% lower for IPd relative to Pd under the reported hazard-ratio framework. Again, this is a relative event-rate interpretation, not a statement that 22.4% of patients survived or that every participant experienced that reduction.

The 95% confidence interval, 0.594–1.015, is wider relative to the point estimate than the PFS interval and extends slightly above 1. This means the estimated OS effect is less precisely separated from the null value than the PFS estimate in this posted analysis.

The P-value of 0.0319 is evidence generated by the specified statistical test. It is not a measure of effect size. A P-value should therefore be interpreted together with the hazard ratio and confidence interval rather than used as a substitute for them.

Most importantly, the registry states that a closed test procedure was used to control the type I error rate: no further testing would be performed unless the significance level had been reached on PFS. Consequently, the OS result should be interpreted in the context of that prespecified testing hierarchy rather than as an isolated P-value detached from the primary endpoint strategy.

9. Secondary Endpoint: Overall Response Rate

Overall Response Rate (ORR) was evaluated as the percentage of participants with disease response according to the Independent Response Committee. Unlike PFS and OS, ORR is a binary endpoint: each participant is classified according to whether the specified response occurred.

Stratified comparison of response

P < 0.0001

One-sided Cochran-Mantel-Haenszel test

ITT population   ·   Superiority hypothesis

ORR analysisRegistry result
Analysis populationITT population
ComparisonPd vs IPd
MethodCochran-Mantel-Haenszel test
EndpointPercentage of participants with disease response as per Independent Response Committee
P-value<0.0001
Hypothesis typeSuperiority
Testing directionOne-sided
StratificationAge (<75 years versus ≥75 years) and number of previous lines (2 or 3 versus >3) according to IRT
Clinical Biostats interpretation

The registry reports a one-sided P-value <0.0001 for the stratified Cochran-Mantel-Haenszel comparison of ORR. The result supports a statistical comparison favoring the direction specified by the superiority hypothesis under the prespecified one-sided test.

The registry-reported statistical analysis does not provide the arm-specific ORR percentages, an odds ratio, a risk ratio, a risk difference, or a confidence interval. Those quantities therefore cannot be reconstructed from the ClinicalTrials.gov record.

This is an important distinction between a P-value and an effect estimate. A very small P-value indicates strong statistical evidence under the specified null model and testing procedure, but it does not tell the reader how large the absolute difference in response rates was.

The use of the Cochran-Mantel-Haenszel test also matters. The test was stratified by age and number of previous lines of therapy, so the comparison was not simply an unadjusted two-by-two test that ignored the prespecified strata.

10. Closed Testing and Multiplicity

The statistical analyses contain an important design feature: a closed test procedure was used to control the type I error rate. The registry states that no further testing would be performed unless the significance level had been reached on PFS.

FeatureRegistry-supported interpretation
Primary endpointProgression Free Survival.
Testing frameworkSuperiority.
Multiplicity controlClosed test procedure.
Gatekeeping conditionNo further testing unless the significance level had been reached on PFS.
Secondary analyses postedOverall Survival: Final Analysis; Overall Response Rate.

Multiplicity is important because multiple formal statistical tests can otherwise increase the chance of obtaining at least one apparently positive result by chance alone. A closed testing procedure creates a prespecified logical structure for the hypotheses so that later claims depend on the required earlier result.

For ICARIA-MM, this means the PFS result should be understood as the statistical gateway specified by the registry. The OS and ORR results are not simply three unrelated P-values. Their interpretation belongs to the hierarchy and error-control framework used for the trial.

Practical reading rule: A secondary endpoint P-value should never be interpreted in isolation when the trial uses a hierarchical or closed testing procedure. The question is not only whether the numerical P-value is small, but also whether the endpoint was eligible for formal testing under the prespecified sequence.

11. Statistical Methodology

Kaplan-Meier estimation

The registry states that PFS was analyzed using the Kaplan-Meier method. This is the standard nonparametric estimator for a time-to-event survival function in the presence of right censoring. It provides an estimate of the probability of remaining event-free through successive event times.

Log-rank test

The primary PFS and secondary OS analyses used the log-rank test. The test compares the observed and expected event patterns between treatment groups across follow-up time. In this trial, the analysis was stratified according to age and number of previous lines of therapy.

Conceptual time-to-event comparison
H0: no treatment difference in the event-time distributions

The log-rank framework evaluates whether the observed event experience differs between randomized groups over follow-up. The hazard ratio then provides a quantitative relative-effect estimate.

Hazard ratio

The hazard ratio is the principal effect measure reported for PFS and OS. For a treatment-versus-control comparison, an HR below 1 indicates a lower estimated instantaneous event rate in the treatment group within the model's framework.

Interpretation
HR < 1  →  lower estimated instantaneous event rate in IPd than Pd

The HR is not a probability, is not an absolute risk difference, and should not be read as the percentage of participants who benefit.

Cochran-Mantel-Haenszel test

ORR was compared using a stratified Cochran-Mantel-Haenszel test. This method is designed for categorical data and incorporates the prespecified strata rather than treating all observations as though they came from one unstratified table.

Confidence intervals for Kaplan-Meier estimates

For Kaplan-Meier estimates, the registry states that confidence intervals were calculated with a log-log transformation of the survival function and methods of Brookmeyer and Crowley. These details concern estimation of uncertainty around time-to-event quantities and are distinct from the PFS hazard-ratio confidence interval.

12. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint rather than a simple binary endpoint. Participants can experience progression or death at different times, while others may remain event-free at the time their observation ends. The log-rank test is designed to compare such event-time distributions while incorporating the available follow-up information.

What does a PFS hazard ratio of 0.596 mean?

It indicates a lower estimated hazard of progression or death in the IPd group relative to the Pd group under the reported time-to-event framework. Numerically, 1 − 0.596 = 0.404, so the estimated hazard is approximately 40.4% lower. That calculation does not convert the HR into an absolute probability or guarantee that individual patients experienced a 40.4% reduction.

Why is the confidence interval important?

The PFS estimate of 0.596 is accompanied by a 95% confidence interval of 0.436–0.814. The interval communicates the statistical uncertainty around the estimated hazard ratio. A point estimate alone cannot show how precisely the treatment effect was estimated.

Why doesn't the P-value measure treatment effect size?

A P-value is a measure of evidence against a specified null hypothesis under a specified statistical procedure. It depends on the effect, the variability and the amount of information in the data. It therefore cannot be used as a substitute for the hazard ratio, response difference, or another clinically interpretable effect measure.

Why was the Cochran-Mantel-Haenszel test used for ORR?

ORR is a categorical outcome, so a categorical-data method is appropriate. The Cochran-Mantel-Haenszel framework additionally incorporates the trial's prespecified strata, allowing the comparison to account for the age and previous-lines strata identified in the registry.

Why does the closed testing procedure matter?

The registry states that the closed test procedure controlled the type I error rate and required the significance level to be reached on PFS before further testing. This means the statistical interpretation of subsequent formal tests depends on the prespecified hierarchy rather than on each P-value being considered independently.

13. Confidence Intervals and Precision

The two posted hazard-ratio analyses illustrate why a point estimate should always be considered together with its confidence interval.

EndpointHR95% CIP-valueInterpretive feature
PFS0.5960.436–0.8140.0005Interval remains below 1.
OS0.7760.594–1.0150.0319Interval extends slightly above 1.

The PFS estimate is accompanied by a confidence interval entirely below 1, whereas the OS interval extends slightly above 1. These are different statistical statements about precision and the null value. They should not be collapsed into a single conclusion based only on the numerical size of the two point estimates.

There is also no basis for treating the difference between the two hazard ratios as proof that the treatment effect differs between PFS and OS. Different endpoints measure different clinical events, have different censoring structures, and may be influenced by subsequent treatment and the timing of events.

14. Censoring and Time-to-Event Interpretation

PFS illustrates an important feature of clinical-trial statistics: not every participant necessarily experiences the event during the period used for analysis. The registry specifies that participants without progression or death were censored at the date of their last progression-free tumor assessment prior to the study cut-off date.

Event

First documentation of progressive disease by the Independent Response Committee or death from any cause, whichever occurs first.

Censoring

If progression or death was not observed, the participant was censored at the last progression-free tumor assessment before the study cut-off.

Why censoring matters

The analysis uses information from participants who have not yet experienced an event without treating them as though they experienced the event at the end of observation.

Why HR is not a risk ratio

A hazard ratio compares event rates over time; it is not the same mathematical quantity as a ratio of cumulative event probabilities.

These distinctions are essential when interpreting PFS. A participant who remains progression-free at the end of observed follow-up has contributed information up to that point, but the analysis does not assume that participant remained progression-free indefinitely.

15. Overall Survival vs Progression-Free Survival

PFS and OS are both time-to-event outcomes, but their event definitions differ. PFS counts the first occurrence of progression or death, while OS counts death from any cause. Consequently, a treatment can have different hazard ratios for these endpoints without the results being statistically contradictory.

FeaturePFSOS
Endpoint rolePrimarySecondary
EventProgression by IRC or death, whichever comes firstDeath from any cause
Analysis populationITTITT
MethodLog-rank testLog-rank test
Effect measureHazard ratioHazard ratio
Estimate0.5960.776
95% CI0.436–0.8140.594–1.015
P-value0.00050.0319

The distinction is more than terminology. PFS captures an earlier disease-control event and therefore can provide information before enough deaths have accumulated for OS to be fully informative. OS, by contrast, is directly based on death from any cause and can reflect the entire subsequent treatment pathway.

16. Safety

The ClinicalTrials.gov record includes serious adverse event counts by treatment arm. These are reported as affected participants divided by participants at risk.

ArmSerious adverse eventsInterpretation of registry-reported count
Pd (Pomalidomide + Dexamethasone)91/14991 affected participants among 149 at risk.
IPd (Isatuximab + Pomalidomide + Dexamet)112/152112 affected participants among 152 at risk.

The ClinicalTrials.gov record does not provide a statistical comparison, confidence interval, or P-value for these serious adverse event counts. They therefore should not be converted into a formal between-arm hypothesis test on this page.

Safety interpretation: The denominators posted on ClinicalTrials.gov for serious adverse events are the reported at-risk populations, 149 for Pd and 152 for IPd. These figures should be kept distinct from the overall enrollment of 307.0 and from the ITT population used for efficacy analyses.

This distinction illustrates why efficacy and safety populations must be read separately. The primary PFS analysis is explicitly based on all randomized participants, whereas the serious-adverse-event ClinicalTrials.gov record is presented as affected participants over participants at risk. A safety rate should not be substituted for an ITT efficacy estimate.

17. What the Primary PFS Result Does — and Does Not — Establish

What the estimate says

The PFS hazard ratio of 0.596 is an estimated relative treatment effect comparing IPd with Pd. Under the reported framework, the estimated hazard of progression or death was approximately 40.4% lower with IPd.

What the estimate does not say

It does not say that 40.4% of patients benefited, that 40.4% of patients avoided progression, or that every participant experienced a 40.4% reduction in individual risk.

What the confidence interval adds

The 95% CI of 0.436–0.814 describes uncertainty around the estimated hazard ratio. It does not describe the range of outcomes that individual patients might experience.

What the P-value adds

The P-value of 0.0005 quantifies statistical evidence under the specified test. It does not quantify the magnitude of benefit or establish clinical importance by itself.

Why the analysis population matters

The analysis was conducted in the ITT population. That means the comparison remains anchored to randomized treatment assignment rather than being restricted to participants who completed a particular amount of treatment.

18. Secondary Results in Context

The three posted statistical analyses form a coherent statistical sequence: a primary time-to-event endpoint, a secondary time-to-event endpoint, and a secondary binary endpoint.

EndpointRoleMethodEffect / result reported
Progression Free Survival Primary Log-rank HR 0.596; 95% CI 0.436–0.814; P = 0.0005
Overall Survival: Final Analysis Secondary Log-rank HR 0.776; 95% CI 0.594–1.015; P = 0.0319
Overall Response Rate Secondary Cochran-Mantel-Haenszel One-sided P <0.0001

The statistical methods match the structures of the endpoints. PFS and OS are time-to-event outcomes and were analyzed with log-rank testing and hazard ratios. ORR is categorical and was analyzed with a stratified Cochran-Mantel-Haenszel test.

The different effect measures also mean that the three results should not be placed on one numerical scale. An HR of 0.596, an HR of 0.776, and a P-value <0.0001 answer different statistical questions. There is no valid operation that converts these three quantities into one overall effect estimate.

19. Stratification: Why Age and Previous Lines Matter

The registry reports two stratification variables for the time-to-event analyses: age and number of previous lines of therapy. The categories were <75 years versus ≥75 years and 2 or 3 versus >3 previous lines.

Stratification factorCategoriesRole in posted analyses
Age<75 years vs ≥75 yearsStratification of PFS and OS analyses
Number of previous lines of therapy2 or 3 vs >3Stratification of PFS and OS analyses

Stratification is not the same as adjusting for every baseline characteristic. The analysis specifically identifies these variables as the strata used according to IRT. The statistical interpretation should therefore remain tied to the prespecified analysis structure rather than implying that the model adjusted for characteristics that are not documented in the ClinicalTrials.gov record.

For the ORR analysis, the registry similarly states that the one-sided P-value was stratified by age and number of previous lines according to IRT. Thus, the same basic clinical factors were incorporated into the categorical-data comparison.

20. Statistical Testing and One-Sided vs Two-Sided Inference

The posted analyses use different testing directions. PFS and OS have two-sided 95% confidence intervals, while the ORR analysis reports a one-sided P-value.

EndpointTesting directionConfidence interval reported
PFSTwo-sided95% CI: 0.436–0.814
OSTwo-sided95% CI: 0.594–1.015
ORROne-sidedNo confidence interval reported in the ClinicalTrials.gov record

A two-sided test considers departures from the null in either direction, whereas a one-sided test evaluates evidence in a prespecified direction. These are different inferential procedures. The ORR P-value therefore should not be compared numerically with the two-sided P-values as though all three represented identical tests.

The distinction is particularly important when reading very small P-values. A P-value of <0.0001 from a one-sided test does not carry exactly the same inferential definition as a two-sided P-value of 0.0005. The correct interpretation requires attention to the hypothesis and testing direction specified for the endpoint.

21. Missing Data and Censoring

The ClinicalTrials.gov record provides a specific censoring rule for PFS: participants without progression or death were censored at their last progression-free tumor assessment before the study cut-off date.

The ClinicalTrials.gov record does not describe a separate imputation strategy for missing data, nor do they provide a detailed missing-assessment sensitivity analysis. It would therefore be inappropriate to claim that a particular missing-data imputation method was used.

Documented

PFS censoring at the last progression-free tumor assessment before the study cut-off when progression or death was not observed.

Not specified in the ClinicalTrials.gov record

A separate statistical imputation method for missing data is not reported in the registry analysis text.

This is a useful methodological boundary. Censoring and missing-data imputation are related but not interchangeable concepts. A censored time-to-event observation still contributes information up to its censoring time; that does not mean the missing future event time was imputed.

22. Bayesian Methods, Non-Inferiority, and Crossover

The registry-reported statistical analysis describes a superiority hypothesis. No non-inferiority margin is reported, and no Bayesian method is identified in the posted statistical analyses.

Design topicRegistry-supported status
Superiority testingYes.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
Bayesian methodsNot reported in the statistical analyses posted on ClinicalTrials.gov.
CrossoverNot reported in the ClinicalTrials.gov record.
Closed testingYes; used to control type I error.

This distinction matters because different trial objectives require different statistical logic. A superiority trial asks whether the evidence supports a difference in the prespecified direction. A non-inferiority trial instead requires a margin defining how much loss of effect could still be considered acceptable. The ICARIA-MM ClinicalTrials.gov record identify a superiority hypothesis and do not provide a non-inferiority margin.

23. Trial Timeline

2016-12-22

Trial start

The registry lists 2016-12-22 as the study start date.

Phase 3

Randomized parallel design

The study was a phase 3 randomized parallel trial with two arms and no masking.

2018-11-22

Primary completion

The registry lists 2018-11-22 as the primary completion date.

Completed

Results posted

The registry reports that results were posted, including 24 outcome measures and 3 statistical analyses.

24. Limitations

25. Why This Trial Matters Statistically

ICARIA-MM is a useful teaching example because its registry record brings together several core concepts in clinical-trial biostatistics: randomized treatment assignment, an ITT efficacy population, time-to-event endpoints, Kaplan-Meier estimation, stratified log-rank testing, hazard ratios, confidence intervals, a stratified categorical-data analysis, one-sided and two-sided inference, and closed testing for type I error control.

Statistical conceptHow it appears in ICARIA-MM
RandomizationRandomized phase 3 parallel-group comparison.
Intention-to-treat analysisPrimary PFS and secondary OS and ORR analyses were performed in the ITT population.
Kaplan-Meier estimationUsed for the primary PFS analysis.
Time-to-event endpointPFS was the registered primary endpoint; OS was a secondary time-to-event endpoint.
Hazard ratioReported for PFS and OS.
Confidence interval95% two-sided CIs were reported for the PFS and OS hazard ratios.
Log-rank testUsed for PFS and OS comparisons.
Stratified analysisAge and number of previous lines of therapy were used as strata.
Cochran-Mantel-Haenszel testUsed for the ORR comparison.
One-sided testingUsed for the posted ORR P-value.
Multiplicity controlClosed testing procedure with PFS as the required statistical gate.
Safety analysisSerious adverse events were reported as affected participants over those at risk in each arm.

The most instructive feature is the relationship between endpoint type and statistical method. PFS and OS require methods that account for the timing of events and censoring. ORR is categorical and therefore uses a different test. The trial consequently demonstrates why clinical-trial statistics cannot be reduced to a single generic hypothesis test.

26. Interpreting the Trial as a Statistical Story

The primary PFS analysis begins with randomization and ends with a time-to-event comparison. Every randomized participant belongs to the ITT population, and the PFS endpoint records the first qualifying progression or death. Participants without such an event are censored according to the registry-defined rule.

The Kaplan-Meier method then summarizes the event-free experience over time, while the stratified log-rank test evaluates the randomized treatment comparison. The hazard ratio supplies a relative effect estimate, and its confidence interval supplies information about precision. The resulting PFS estimate is 0.596 with a 95% CI of 0.436–0.814 and a P-value of 0.0005.

The secondary OS analysis uses the same broad time-to-event framework but a different event: death from any cause. Its hazard ratio is 0.776 with a 95% CI of 0.594–1.015 and a P-value of 0.0319. The difference between these numerical estimates should not be interpreted as an inconsistency. PFS and OS measure different events.

Finally, ORR changes the statistical problem. Instead of asking when an event occurs, it asks whether the specified response occurred. The Cochran-Mantel-Haenszel test therefore provides a stratified categorical comparison, with a one-sided P-value of <0.0001.

Across all three analyses, the closed testing procedure is a critical part of the inferential structure. Statistical significance is not just a property of an isolated P-value; it depends on the endpoint hierarchy, hypothesis, testing direction, analysis population, and multiplicity-control strategy.

27. Related Tutorials

Learn more about the methods used in this trial:

28. Related Statistical Calculators

29. Sources

Continue through the Clinical Biostats statistical library

Explore the underlying statistical methods through tutorials and practical calculators for clinical-trial analysis.

30. Record Summary

ICARIA-MM is a randomized phase 3 trial with 307.0 enrolled participants, two parallel treatment arms, and a registered primary endpoint of progression-free survival. The primary PFS analysis used the ITT population, Kaplan-Meier estimation, and a stratified log-rank comparison, with a hazard ratio of 0.596, a two-sided 95% CI of 0.436–0.814, and a P-value of 0.0005.

The posted secondary analyses extend the statistical story to overall survival and overall response. OS was analyzed with the same general time-to-event framework and produced an HR of 0.776 with a 95% CI of 0.594–1.015 and a P-value of 0.0319. ORR was analyzed with a stratified Cochran-Mantel-Haenszel test and produced a one-sided P-value of <0.0001. These analyses are connected by a closed testing procedure intended to control the type I error rate, with further testing contingent on reaching the significance level for PFS.

The key statistical lesson is that the trial cannot be understood from any single number. The hazard ratios quantify relative time-to-event effects, confidence intervals describe uncertainty, P-values address evidence under specified hypotheses, and the ITT population preserves the randomized comparison. Stratification, censoring, endpoint hierarchy, and multiplicity control are equally important to the interpretation.

Clinical Biostats methodology: A trial-results page should distinguish the reported numerical evidence from statistical interpretation. For ICARIA-MM, the analysis therefore preserves the registry's endpoint definitions, analysis populations, methods, effect estimates, confidence intervals, P-values, testing framework, and safety counts while avoiding unsupported reconstruction of unreported arm-level results.