This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information from that registry record.
1. Trial at a Glance
ICARIA-MM was a randomized, parallel, phase 3 study comparing isatuximab plus pomalidomide and dexamethasone (IPd) with pomalidomide and dexamethasone (Pd) in patients with refractory or relapsed and refractory multiple myeloma. The registered primary endpoint was progression-free survival (PFS), a time-to-event outcome analyzed in the intention-to-treat population.
| Feature | ICARIA-MM |
|---|---|
| Trial name | ICARIA-MM |
| NCT identifier | NCT02990338 |
| Therapeutic area | Hematology |
| Condition | Plasma Cell Myeloma |
| Phase | Phase 3 |
| Status | Completed |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 307.0 |
| Arms | 2 |
| Lead sponsor | Sanofi |
| Sponsor type | Industry |
| Start | 2016-12-22 |
| Primary completion | 2018-11-22 |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 24 |
| Statistical analyses posted | 3 |
2. Clinical Question
The central statistical question was whether adding isatuximab to pomalidomide and dexamethasone changed progression-free survival compared with pomalidomide and dexamethasone alone in patients with refractory or relapsed and refractory multiple myeloma.
Population
Patients with refractory or relapsed and refractory multiple myeloma, as described by the trial's brief title and registry condition.
Intervention
Isatuximab, pomalidomide, and dexamethasone (IPd).
Comparator
Pomalidomide and dexamethasone (Pd).
Primary question
Does the addition of isatuximab improve progression-free survival relative to pomalidomide and dexamethasone?
This question is a classic randomized time-to-event comparison. The important statistical distinction is that PFS is not simply a proportion of participants who eventually progress. Each participant contributes information over time, with an event occurring at progression or death and censoring possible when neither event has been observed by the relevant assessment point.
3. Trial Design
Isatuximab combination
- Isatuximab
- Pomalidomide
- Dexamethasone
Pomalidomide combination
- Pomalidomide
- Dexamethasone
The registry identifies the trial as a phase 3 randomized study with two parallel arms. Because the trial was unmasked, the statistical interpretation should distinguish the randomized treatment comparison from any potential effects that knowledge of assignment could have on assessments or subsequent care. The ClinicalTrials.gov record does not provide additional details that would allow those effects to be quantified here.
4. Randomization, Stratification, and Analysis Population
The formal efficacy analyses reported in the registry were conducted in the intention-to-treat (ITT) population, defined in the primary analysis as including all randomized participants. This is central to the interpretation of the randomized comparison because treatment assignment, rather than treatment actually received, defines the groups being compared.
| Analysis feature | Registry-supported description |
|---|---|
| Primary efficacy population | Intent-to-treat population including all randomized participants. |
| PFS comparison | Pd versus IPd. |
| OS comparison | Pd versus IPd. |
| ORR comparison | Pd versus IPd. |
| Stratification factors | Age (<75 years versus ≥75 years) and number of previous lines of therapy (2 or 3 versus >3), according to IRT. |
The registry's analysis notes repeatedly identify age and number of previous lines of therapy as stratification factors. This matters because the reported time-to-event comparisons were not treated as if every participant came from a single homogeneous population. Stratified analysis preserves the planned structure of the randomized comparison while accounting for the factors used in the analysis.
5. Endpoints
The registry lists one primary endpoint: Progression Free Survival (PFS). The posted analyses also include Overall Survival (OS) and Overall Response Rate (ORR) as secondary outcomes.
| Endpoint | Registry definition / time frame | Statistical approach |
|---|---|---|
| Progression Free Survival (PFS) | Time from date of randomization to date of first documentation of progressive disease (PD) determined by Independent Response Committee (IRC) or date of death from any cause, whichever comes first. If progression or death was not observed, participant was censored at date of last progression-free tumor assessment prior to study cut-off date. | Kaplan-Meier method; log-rank comparison; hazard ratio. |
| Overall Survival (OS): Final Analysis | From the date of randomization to date of death from any cause or data cut-off date, whichever was earlier. | Stratified log-rank comparison; hazard ratio. |
| Overall Response Rate (ORR): Percentage of Participants With Disease Response as Per Independent Response Committee (IRC) | From the date of randomization to the date of first documentation of progression or initiation of further anti-myeloma treatment. | Stratified Cochran-Mantel-Haenszel test. |
6. Primary Endpoint: Progression-Free Survival
PFS was the registered primary endpoint and was analyzed as a time-to-event outcome. The event was the first of two possibilities: progression determined by the Independent Response Committee or death from any cause. Participants without either event were censored at the date of their last progression-free tumor assessment before the study cut-off date.
Hazard ratio for progression or death
95% CI: 0.436–0.814 · P = 0.0005
Two-sided log-rank test · Superiority hypothesis
| Primary PFS analysis | Pd | IPd |
|---|---|---|
| Analysis population | Intent-to-treat population, including all randomized participants | |
| Comparison | Pomalidomide + dexamethasone vs isatuximab + pomalidomide + dexamethasone | |
| Method | Log-rank test | |
| Effect measure | Hazard ratio | |
| Hazard ratio | 0.596 | |
| 95% confidence interval | 0.436–0.814 | |
| P-value | 0.0005 | |
| Hypothesis | Superiority | |
The reported PFS hazard ratio of 0.596 means that, under the time-to-event comparison represented by the reported hazard ratio, the estimated hazard of progression or death in the IPd group was about 40.4% lower than in the Pd group. This is a relative hazard statement, not a statement that 40.4% of participants avoided progression.
The estimate does not mean that every individual participant experienced the same reduction in risk, nor does it mean that the probability of progression or death was reduced by exactly 40.4% at every particular time point. The hazard ratio summarizes a relative comparison of event rates over the analyzed follow-up under the statistical framework used.
The 95% confidence interval of 0.436–0.814 describes uncertainty around the estimated hazard ratio. It is substantially below 1, but it is not an interval containing the individual treatment effects experienced by patients. It describes uncertainty in the estimated population-level relative effect under the analysis framework.
The P-value of 0.0005 addresses the statistical evidence against the null hypothesis under the specified testing procedure. It does not measure the magnitude of the treatment effect, does not describe clinical importance, and does not give the probability that the treatment works.
Because PFS is subject to censoring, the analysis depends on how follow-up and censoring are handled. The registry states that participants without progression or death were censored at their last progression-free tumor assessment before the study cut-off date. Interpretation of a hazard ratio also relies on the time-to-event modeling framework; a single hazard ratio can be less informative if relative hazards vary substantially over time.
Why Kaplan-Meier estimation was appropriate
The registry states that PFS was analyzed by the Kaplan-Meier method. This is appropriate for a time-to-event endpoint because participants may enter the analysis with different amounts of observable follow-up and some may be censored before experiencing the event.
At each observed event time, the estimated survival probability is updated according to the number of events and the number of participants at risk immediately before that time. For PFS, the corresponding survival function represents remaining progression-free without death through time t.
The registry also states that confidence intervals for Kaplan-Meier estimates were calculated using a log-log transformation of the survival function and methods of Brookmeyer and Crowley. That is a detail about how uncertainty around the estimated survival function was constructed; it is separate from the hazard-ratio comparison itself.
7. Stratified Log-Rank Analysis
The primary PFS comparison used a log-rank test, with stratification based on age and number of previous lines of therapy according to the IRT. The same stratification structure is reported for the OS analysis.
Age stratum
<75 years versus ≥75 years.
Previous therapy stratum
2 or 3 previous lines versus >3 previous lines.
Comparison
The randomized Pd and IPd groups were compared within the time-to-event framework.
Effect measure
Hazard ratio, with a two-sided 95% confidence interval for the primary PFS analysis.
A stratified log-rank test is useful when the trial has prespecified factors that may influence the event-time distribution. Rather than ignoring those factors, the test incorporates the strata into the comparison. The resulting inference is still about the randomized treatment comparison, not about whether the stratification factors themselves are beneficial or harmful.
8. Secondary Endpoint: Overall Survival
Overall Survival (OS): Final Analysis was a secondary time-to-event outcome. The registry defines the endpoint as time from randomization to death from any cause or the data cut-off date, whichever was earlier. The analysis was performed in the ITT population.
Hazard ratio for overall survival
95% CI: 0.594–1.015 · P = 0.0319
Two-sided log-rank test · Superiority hypothesis
| OS final analysis | Registry result |
|---|---|
| Analysis population | ITT population |
| Comparison | Pd vs IPd |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.776 |
| 95% confidence interval | 0.594–1.015 |
| P-value | 0.0319 |
| Hypothesis type | Superiority |
| Stratification | Age (<75 years versus ≥75 years) and number of previous lines of therapy (2 or 3 versus >3) according to IRT |
The OS hazard ratio of 0.776 corresponds to an estimated hazard of death about 22.4% lower for IPd relative to Pd under the reported hazard-ratio framework. Again, this is a relative event-rate interpretation, not a statement that 22.4% of patients survived or that every participant experienced that reduction.
The 95% confidence interval, 0.594–1.015, is wider relative to the point estimate than the PFS interval and extends slightly above 1. This means the estimated OS effect is less precisely separated from the null value than the PFS estimate in this posted analysis.
The P-value of 0.0319 is evidence generated by the specified statistical test. It is not a measure of effect size. A P-value should therefore be interpreted together with the hazard ratio and confidence interval rather than used as a substitute for them.
Most importantly, the registry states that a closed test procedure was used to control the type I error rate: no further testing would be performed unless the significance level had been reached on PFS. Consequently, the OS result should be interpreted in the context of that prespecified testing hierarchy rather than as an isolated P-value detached from the primary endpoint strategy.
9. Secondary Endpoint: Overall Response Rate
Overall Response Rate (ORR) was evaluated as the percentage of participants with disease response according to the Independent Response Committee. Unlike PFS and OS, ORR is a binary endpoint: each participant is classified according to whether the specified response occurred.
Stratified comparison of response
One-sided Cochran-Mantel-Haenszel test
ITT population · Superiority hypothesis
| ORR analysis | Registry result |
|---|---|
| Analysis population | ITT population |
| Comparison | Pd vs IPd |
| Method | Cochran-Mantel-Haenszel test |
| Endpoint | Percentage of participants with disease response as per Independent Response Committee |
| P-value | <0.0001 |
| Hypothesis type | Superiority |
| Testing direction | One-sided |
| Stratification | Age (<75 years versus ≥75 years) and number of previous lines (2 or 3 versus >3) according to IRT |
The registry reports a one-sided P-value <0.0001 for the stratified Cochran-Mantel-Haenszel comparison of ORR. The result supports a statistical comparison favoring the direction specified by the superiority hypothesis under the prespecified one-sided test.
The registry-reported statistical analysis does not provide the arm-specific ORR percentages, an odds ratio, a risk ratio, a risk difference, or a confidence interval. Those quantities therefore cannot be reconstructed from the ClinicalTrials.gov record.
This is an important distinction between a P-value and an effect estimate. A very small P-value indicates strong statistical evidence under the specified null model and testing procedure, but it does not tell the reader how large the absolute difference in response rates was.
The use of the Cochran-Mantel-Haenszel test also matters. The test was stratified by age and number of previous lines of therapy, so the comparison was not simply an unadjusted two-by-two test that ignored the prespecified strata.
10. Closed Testing and Multiplicity
The statistical analyses contain an important design feature: a closed test procedure was used to control the type I error rate. The registry states that no further testing would be performed unless the significance level had been reached on PFS.
| Feature | Registry-supported interpretation |
|---|---|
| Primary endpoint | Progression Free Survival. |
| Testing framework | Superiority. |
| Multiplicity control | Closed test procedure. |
| Gatekeeping condition | No further testing unless the significance level had been reached on PFS. |
| Secondary analyses posted | Overall Survival: Final Analysis; Overall Response Rate. |
Multiplicity is important because multiple formal statistical tests can otherwise increase the chance of obtaining at least one apparently positive result by chance alone. A closed testing procedure creates a prespecified logical structure for the hypotheses so that later claims depend on the required earlier result.
For ICARIA-MM, this means the PFS result should be understood as the statistical gateway specified by the registry. The OS and ORR results are not simply three unrelated P-values. Their interpretation belongs to the hierarchy and error-control framework used for the trial.
11. Statistical Methodology
Kaplan-Meier estimation
The registry states that PFS was analyzed using the Kaplan-Meier method. This is the standard nonparametric estimator for a time-to-event survival function in the presence of right censoring. It provides an estimate of the probability of remaining event-free through successive event times.
Log-rank test
The primary PFS and secondary OS analyses used the log-rank test. The test compares the observed and expected event patterns between treatment groups across follow-up time. In this trial, the analysis was stratified according to age and number of previous lines of therapy.
The log-rank framework evaluates whether the observed event experience differs between randomized groups over follow-up. The hazard ratio then provides a quantitative relative-effect estimate.
Hazard ratio
The hazard ratio is the principal effect measure reported for PFS and OS. For a treatment-versus-control comparison, an HR below 1 indicates a lower estimated instantaneous event rate in the treatment group within the model's framework.
The HR is not a probability, is not an absolute risk difference, and should not be read as the percentage of participants who benefit.
Cochran-Mantel-Haenszel test
ORR was compared using a stratified Cochran-Mantel-Haenszel test. This method is designed for categorical data and incorporates the prespecified strata rather than treating all observations as though they came from one unstratified table.
Confidence intervals for Kaplan-Meier estimates
For Kaplan-Meier estimates, the registry states that confidence intervals were calculated with a log-log transformation of the survival function and methods of Brookmeyer and Crowley. These details concern estimation of uncertainty around time-to-event quantities and are distinct from the PFS hazard-ratio confidence interval.
12. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS is a time-to-event endpoint rather than a simple binary endpoint. Participants can experience progression or death at different times, while others may remain event-free at the time their observation ends. The log-rank test is designed to compare such event-time distributions while incorporating the available follow-up information.
What does a PFS hazard ratio of 0.596 mean?
It indicates a lower estimated hazard of progression or death in the IPd group relative to the Pd group under the reported time-to-event framework. Numerically, 1 − 0.596 = 0.404, so the estimated hazard is approximately 40.4% lower. That calculation does not convert the HR into an absolute probability or guarantee that individual patients experienced a 40.4% reduction.
Why is the confidence interval important?
The PFS estimate of 0.596 is accompanied by a 95% confidence interval of 0.436–0.814. The interval communicates the statistical uncertainty around the estimated hazard ratio. A point estimate alone cannot show how precisely the treatment effect was estimated.
Why doesn't the P-value measure treatment effect size?
A P-value is a measure of evidence against a specified null hypothesis under a specified statistical procedure. It depends on the effect, the variability and the amount of information in the data. It therefore cannot be used as a substitute for the hazard ratio, response difference, or another clinically interpretable effect measure.
Why was the Cochran-Mantel-Haenszel test used for ORR?
ORR is a categorical outcome, so a categorical-data method is appropriate. The Cochran-Mantel-Haenszel framework additionally incorporates the trial's prespecified strata, allowing the comparison to account for the age and previous-lines strata identified in the registry.
Why does the closed testing procedure matter?
The registry states that the closed test procedure controlled the type I error rate and required the significance level to be reached on PFS before further testing. This means the statistical interpretation of subsequent formal tests depends on the prespecified hierarchy rather than on each P-value being considered independently.
13. Confidence Intervals and Precision
The two posted hazard-ratio analyses illustrate why a point estimate should always be considered together with its confidence interval.
| Endpoint | HR | 95% CI | P-value | Interpretive feature |
|---|---|---|---|---|
| PFS | 0.596 | 0.436–0.814 | 0.0005 | Interval remains below 1. |
| OS | 0.776 | 0.594–1.015 | 0.0319 | Interval extends slightly above 1. |
The PFS estimate is accompanied by a confidence interval entirely below 1, whereas the OS interval extends slightly above 1. These are different statistical statements about precision and the null value. They should not be collapsed into a single conclusion based only on the numerical size of the two point estimates.
There is also no basis for treating the difference between the two hazard ratios as proof that the treatment effect differs between PFS and OS. Different endpoints measure different clinical events, have different censoring structures, and may be influenced by subsequent treatment and the timing of events.
14. Censoring and Time-to-Event Interpretation
PFS illustrates an important feature of clinical-trial statistics: not every participant necessarily experiences the event during the period used for analysis. The registry specifies that participants without progression or death were censored at the date of their last progression-free tumor assessment prior to the study cut-off date.
Event
First documentation of progressive disease by the Independent Response Committee or death from any cause, whichever occurs first.
Censoring
If progression or death was not observed, the participant was censored at the last progression-free tumor assessment before the study cut-off.
Why censoring matters
The analysis uses information from participants who have not yet experienced an event without treating them as though they experienced the event at the end of observation.
Why HR is not a risk ratio
A hazard ratio compares event rates over time; it is not the same mathematical quantity as a ratio of cumulative event probabilities.
These distinctions are essential when interpreting PFS. A participant who remains progression-free at the end of observed follow-up has contributed information up to that point, but the analysis does not assume that participant remained progression-free indefinitely.
15. Overall Survival vs Progression-Free Survival
PFS and OS are both time-to-event outcomes, but their event definitions differ. PFS counts the first occurrence of progression or death, while OS counts death from any cause. Consequently, a treatment can have different hazard ratios for these endpoints without the results being statistically contradictory.
| Feature | PFS | OS |
|---|---|---|
| Endpoint role | Primary | Secondary |
| Event | Progression by IRC or death, whichever comes first | Death from any cause |
| Analysis population | ITT | ITT |
| Method | Log-rank test | Log-rank test |
| Effect measure | Hazard ratio | Hazard ratio |
| Estimate | 0.596 | 0.776 |
| 95% CI | 0.436–0.814 | 0.594–1.015 |
| P-value | 0.0005 | 0.0319 |
The distinction is more than terminology. PFS captures an earlier disease-control event and therefore can provide information before enough deaths have accumulated for OS to be fully informative. OS, by contrast, is directly based on death from any cause and can reflect the entire subsequent treatment pathway.
16. Safety
The ClinicalTrials.gov record includes serious adverse event counts by treatment arm. These are reported as affected participants divided by participants at risk.
| Arm | Serious adverse events | Interpretation of registry-reported count |
|---|---|---|
| Pd (Pomalidomide + Dexamethasone) | 91/149 | 91 affected participants among 149 at risk. |
| IPd (Isatuximab + Pomalidomide + Dexamet) | 112/152 | 112 affected participants among 152 at risk. |
The ClinicalTrials.gov record does not provide a statistical comparison, confidence interval, or P-value for these serious adverse event counts. They therefore should not be converted into a formal between-arm hypothesis test on this page.
This distinction illustrates why efficacy and safety populations must be read separately. The primary PFS analysis is explicitly based on all randomized participants, whereas the serious-adverse-event ClinicalTrials.gov record is presented as affected participants over participants at risk. A safety rate should not be substituted for an ITT efficacy estimate.
17. What the Primary PFS Result Does — and Does Not — Establish
The PFS hazard ratio of 0.596 is an estimated relative treatment effect comparing IPd with Pd. Under the reported framework, the estimated hazard of progression or death was approximately 40.4% lower with IPd.
It does not say that 40.4% of patients benefited, that 40.4% of patients avoided progression, or that every participant experienced a 40.4% reduction in individual risk.
The 95% CI of 0.436–0.814 describes uncertainty around the estimated hazard ratio. It does not describe the range of outcomes that individual patients might experience.
The P-value of 0.0005 quantifies statistical evidence under the specified test. It does not quantify the magnitude of benefit or establish clinical importance by itself.
The analysis was conducted in the ITT population. That means the comparison remains anchored to randomized treatment assignment rather than being restricted to participants who completed a particular amount of treatment.
18. Secondary Results in Context
The three posted statistical analyses form a coherent statistical sequence: a primary time-to-event endpoint, a secondary time-to-event endpoint, and a secondary binary endpoint.
| Endpoint | Role | Method | Effect / result reported |
|---|---|---|---|
| Progression Free Survival | Primary | Log-rank | HR 0.596; 95% CI 0.436–0.814; P = 0.0005 |
| Overall Survival: Final Analysis | Secondary | Log-rank | HR 0.776; 95% CI 0.594–1.015; P = 0.0319 |
| Overall Response Rate | Secondary | Cochran-Mantel-Haenszel | One-sided P <0.0001 |
The statistical methods match the structures of the endpoints. PFS and OS are time-to-event outcomes and were analyzed with log-rank testing and hazard ratios. ORR is categorical and was analyzed with a stratified Cochran-Mantel-Haenszel test.
The different effect measures also mean that the three results should not be placed on one numerical scale. An HR of 0.596, an HR of 0.776, and a P-value <0.0001 answer different statistical questions. There is no valid operation that converts these three quantities into one overall effect estimate.
19. Stratification: Why Age and Previous Lines Matter
The registry reports two stratification variables for the time-to-event analyses: age and number of previous lines of therapy. The categories were <75 years versus ≥75 years and 2 or 3 versus >3 previous lines.
| Stratification factor | Categories | Role in posted analyses |
|---|---|---|
| Age | <75 years vs ≥75 years | Stratification of PFS and OS analyses |
| Number of previous lines of therapy | 2 or 3 vs >3 | Stratification of PFS and OS analyses |
Stratification is not the same as adjusting for every baseline characteristic. The analysis specifically identifies these variables as the strata used according to IRT. The statistical interpretation should therefore remain tied to the prespecified analysis structure rather than implying that the model adjusted for characteristics that are not documented in the ClinicalTrials.gov record.
For the ORR analysis, the registry similarly states that the one-sided P-value was stratified by age and number of previous lines according to IRT. Thus, the same basic clinical factors were incorporated into the categorical-data comparison.
20. Statistical Testing and One-Sided vs Two-Sided Inference
The posted analyses use different testing directions. PFS and OS have two-sided 95% confidence intervals, while the ORR analysis reports a one-sided P-value.
| Endpoint | Testing direction | Confidence interval reported |
|---|---|---|
| PFS | Two-sided | 95% CI: 0.436–0.814 |
| OS | Two-sided | 95% CI: 0.594–1.015 |
| ORR | One-sided | No confidence interval reported in the ClinicalTrials.gov record |
A two-sided test considers departures from the null in either direction, whereas a one-sided test evaluates evidence in a prespecified direction. These are different inferential procedures. The ORR P-value therefore should not be compared numerically with the two-sided P-values as though all three represented identical tests.
The distinction is particularly important when reading very small P-values. A P-value of <0.0001 from a one-sided test does not carry exactly the same inferential definition as a two-sided P-value of 0.0005. The correct interpretation requires attention to the hypothesis and testing direction specified for the endpoint.
21. Missing Data and Censoring
The ClinicalTrials.gov record provides a specific censoring rule for PFS: participants without progression or death were censored at their last progression-free tumor assessment before the study cut-off date.
The ClinicalTrials.gov record does not describe a separate imputation strategy for missing data, nor do they provide a detailed missing-assessment sensitivity analysis. It would therefore be inappropriate to claim that a particular missing-data imputation method was used.
Documented
PFS censoring at the last progression-free tumor assessment before the study cut-off when progression or death was not observed.
Not specified in the ClinicalTrials.gov record
A separate statistical imputation method for missing data is not reported in the registry analysis text.
This is a useful methodological boundary. Censoring and missing-data imputation are related but not interchangeable concepts. A censored time-to-event observation still contributes information up to its censoring time; that does not mean the missing future event time was imputed.
22. Bayesian Methods, Non-Inferiority, and Crossover
The registry-reported statistical analysis describes a superiority hypothesis. No non-inferiority margin is reported, and no Bayesian method is identified in the posted statistical analyses.
| Design topic | Registry-supported status |
|---|---|
| Superiority testing | Yes. |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record. |
| Bayesian methods | Not reported in the statistical analyses posted on ClinicalTrials.gov. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Closed testing | Yes; used to control type I error. |
This distinction matters because different trial objectives require different statistical logic. A superiority trial asks whether the evidence supports a difference in the prespecified direction. A non-inferiority trial instead requires a margin defining how much loss of effect could still be considered acceptable. The ICARIA-MM ClinicalTrials.gov record identify a superiority hypothesis and do not provide a non-inferiority margin.
23. Trial Timeline
Trial start
The registry lists 2016-12-22 as the study start date.
Randomized parallel design
The study was a phase 3 randomized parallel trial with two arms and no masking.
Primary completion
The registry lists 2018-11-22 as the primary completion date.
Results posted
The registry reports that results were posted, including 24 outcome measures and 3 statistical analyses.
24. Limitations
- Registry-level scope: This page is restricted to the ClinicalTrials.gov record and the listed PubMed records. It does not add numerical results from publications beyond the ClinicalTrials.gov record.
- Incomplete arm-level efficacy data: The registry-reported statistical analysis provides the PFS and OS hazard ratios and the ORR P-value, but does not provide arm-specific PFS or OS medians or arm-specific ORR percentages.
- Limited safety detail: The ClinicalTrials.gov record provides serious adverse event counts by arm but not a full adverse-event table or a formal statistical comparison.
- Censoring: PFS includes censoring when progression or death was not observed by the relevant assessment and study cut-off. Time-to-event estimates therefore depend on the censoring framework.
- Hazard-ratio interpretation: Hazard ratios are relative time-to-event measures and should not be interpreted as relative risks or absolute risk reductions.
- Proportional-hazards considerations: A single hazard ratio is most straightforward to interpret when the relative hazards are reasonably stable over time. The ClinicalTrials.gov record does not provide a formal assessment of that assumption.
- Multiplicity: The closed testing procedure means that secondary formal inference is connected to the prespecified PFS testing gate.
- Unreported methods: The ClinicalTrials.gov record does not report Bayesian methods, a non-inferiority margin, a crossover strategy, or a separate missing-data imputation method.
- Unmasked design: Masking was listed as none. The ClinicalTrials.gov record does not quantify how knowledge of treatment assignment may have affected assessments or subsequent care.
25. Why This Trial Matters Statistically
ICARIA-MM is a useful teaching example because its registry record brings together several core concepts in clinical-trial biostatistics: randomized treatment assignment, an ITT efficacy population, time-to-event endpoints, Kaplan-Meier estimation, stratified log-rank testing, hazard ratios, confidence intervals, a stratified categorical-data analysis, one-sided and two-sided inference, and closed testing for type I error control.
| Statistical concept | How it appears in ICARIA-MM |
|---|---|
| Randomization | Randomized phase 3 parallel-group comparison. |
| Intention-to-treat analysis | Primary PFS and secondary OS and ORR analyses were performed in the ITT population. |
| Kaplan-Meier estimation | Used for the primary PFS analysis. |
| Time-to-event endpoint | PFS was the registered primary endpoint; OS was a secondary time-to-event endpoint. |
| Hazard ratio | Reported for PFS and OS. |
| Confidence interval | 95% two-sided CIs were reported for the PFS and OS hazard ratios. |
| Log-rank test | Used for PFS and OS comparisons. |
| Stratified analysis | Age and number of previous lines of therapy were used as strata. |
| Cochran-Mantel-Haenszel test | Used for the ORR comparison. |
| One-sided testing | Used for the posted ORR P-value. |
| Multiplicity control | Closed testing procedure with PFS as the required statistical gate. |
| Safety analysis | Serious adverse events were reported as affected participants over those at risk in each arm. |
The most instructive feature is the relationship between endpoint type and statistical method. PFS and OS require methods that account for the timing of events and censoring. ORR is categorical and therefore uses a different test. The trial consequently demonstrates why clinical-trial statistics cannot be reduced to a single generic hypothesis test.
26. Interpreting the Trial as a Statistical Story
The primary PFS analysis begins with randomization and ends with a time-to-event comparison. Every randomized participant belongs to the ITT population, and the PFS endpoint records the first qualifying progression or death. Participants without such an event are censored according to the registry-defined rule.
The Kaplan-Meier method then summarizes the event-free experience over time, while the stratified log-rank test evaluates the randomized treatment comparison. The hazard ratio supplies a relative effect estimate, and its confidence interval supplies information about precision. The resulting PFS estimate is 0.596 with a 95% CI of 0.436–0.814 and a P-value of 0.0005.
The secondary OS analysis uses the same broad time-to-event framework but a different event: death from any cause. Its hazard ratio is 0.776 with a 95% CI of 0.594–1.015 and a P-value of 0.0319. The difference between these numerical estimates should not be interpreted as an inconsistency. PFS and OS measure different events.
Finally, ORR changes the statistical problem. Instead of asking when an event occurs, it asks whether the specified response occurred. The Cochran-Mantel-Haenszel test therefore provides a stratified categorical comparison, with a one-sided P-value of <0.0001.
Across all three analyses, the closed testing procedure is a critical part of the inferential structure. Statistical significance is not just a property of an isolated P-value; it depends on the endpoint hierarchy, hypothesis, testing direction, analysis population, and multiplicity-control strategy.
27. Related Tutorials
Learn more about the methods used in this trial:
28. Related Statistical Calculators
29. Sources
- ClinicalTrials.gov: ICARIA-MM, NCT02990338.
- PubMed record: PMID 38299578.
- PubMed record: PMID 36108425.
- PubMed record: PMID 35641409.
- PubMed record: PMID 34800109.
- PubMed record: PMID 33839618.
Continue through the Clinical Biostats statistical library
Explore the underlying statistical methods through tutorials and practical calculators for clinical-trial analysis.
30. Record Summary
ICARIA-MM is a randomized phase 3 trial with 307.0 enrolled participants, two parallel treatment arms, and a registered primary endpoint of progression-free survival. The primary PFS analysis used the ITT population, Kaplan-Meier estimation, and a stratified log-rank comparison, with a hazard ratio of 0.596, a two-sided 95% CI of 0.436–0.814, and a P-value of 0.0005.
The posted secondary analyses extend the statistical story to overall survival and overall response. OS was analyzed with the same general time-to-event framework and produced an HR of 0.776 with a 95% CI of 0.594–1.015 and a P-value of 0.0319. ORR was analyzed with a stratified Cochran-Mantel-Haenszel test and produced a one-sided P-value of <0.0001. These analyses are connected by a closed testing procedure intended to control the type I error rate, with further testing contingent on reaching the significance level for PFS.
The key statistical lesson is that the trial cannot be understood from any single number. The hazard ratios quantify relative time-to-event effects, confidence intervals describe uncertainty, P-values address evidence under specified hypotheses, and the ITT population preserves the randomized comparison. Stratification, censoring, endpoint hierarchy, and multiplicity control are equally important to the interpretation.