This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
QUAZAR AML-001 was a randomized, parallel, quadruple-masked, phase 3 trial comparing oral azacitidine plus best supportive care with placebo plus best supportive care as maintenance therapy in subjects with acute myeloid leukemia in complete remission. The registry reports overall survival as the primary time-to-event endpoint, analyzed in the intent-to-treat population using a log-rank test and a stratified Cox proportional-hazards model.
| Feature | QUAZAR AML-001 |
|---|---|
| Trial name | QUAZAR AML-001 |
| ClinicalTrials.gov identifier | NCT01757535 |
| Phase | Phase 3 |
| Condition | Leukemia, Myeloid, Acute |
| Brief title | Efficacy of Oral Azacitidine Plus Best Supportive Care as Maintenance Therapy in Subjects With Acute Myeloid Leukemia (AML) in Complete Remission |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 472.0 |
| Interventions | Oral Azacitidine; Placebo |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 10 |
| Statistical analyses posted | 3 |
| Lead sponsor | Celgene |
2. Clinical Question
The central statistical question was whether oral azacitidine plus best supportive care produces superior overall survival compared with placebo plus best supportive care in subjects with acute myeloid leukemia in complete remission.
Population
Subjects with acute myeloid leukemia in complete remission, as described by the trial's registered brief title.
Intervention
Oral azacitidine plus best supportive care.
Comparator
Placebo plus best supportive care.
Primary question
Does oral azacitidine plus best supportive care improve overall survival relative to placebo plus best supportive care?
The registered hypothesis type is superiority. This is important because the analysis is framed around whether the treatment group has a statistically detectable improvement in the time-to-event outcome rather than whether it stays within a prespecified non-inferiority margin.
3. Trial Design
Oral azacitidine + best supportive care
- Oral azacitidine
- Best supportive care
- Randomized parallel-group assignment
Placebo + best supportive care
- Placebo
- Best supportive care
- Randomized parallel-group assignment
The design features are statistically important because randomization establishes the treatment comparison, while masking is intended to reduce the influence of knowledge of treatment assignment on trial conduct and assessment. The parallel structure means participants remain part of the randomized comparison rather than moving through sequential treatment cohorts.
4. Endpoints
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Overall survival (primary) | Overall survival was defined as time from randomization to death from any cause; participants surviving at the end of the follow-up period, or who withdraw consent, or who were lost to follow up were censored at the date last known alive. Time frame: Day 1 (randomization) up to data cut off date of 15 July 2019; median follow-up for OS estimated by the reverse K-M method was 41.2 months for all participants.. | Log-rank test; hazard ratio from a stratified Cox proportional-hazards model |
| Relapse free survival (secondary) | Kaplan-Meier Estimate of Relapse Free Survival (RFS). Time frame: From day 1 (randomization) up to data cut off date of 06 August 2024; approximately 135.5 months. | Log-rank test; hazard ratio from a stratified Cox proportional-hazards model |
| Time to definitive clinically meaningful deterioration | Time to Definitive Clinically Meaningful Deterioration for ≥ 2 Consecutive Visits as Measured Using the EQ-5D HRQoL Scale. Time frame: From day 1 (randomization) up to data cut off date of 15 July 2019; approximately 74 months. | Cox proportional-hazards model |
The primary endpoint is therefore a classic time-to-event endpoint. It differs from a binary endpoint because both whether an event occurs and when it occurs contribute information. Participants who remain alive at the end of follow-up do not simply disappear from the analysis; they contribute survival information up to their censoring date.
5. Statistical Methodology
Kaplan-Meier estimation
The registered primary endpoint is a Kaplan-Meier estimate for overall survival. Kaplan-Meier estimation is designed for time-to-event data with right censoring. It estimates the probability of remaining event-free beyond successive event times while retaining information from participants whose observation ends before an event occurs.
Here, di is the number of events at time ti and ni is the number of participants at risk immediately before that time.
For QUAZAR AML-001, the registry specifically defines overall survival from randomization to death from any cause. Participants who were alive at the end of follow-up, withdrew consent, or were lost to follow-up were censored at the date last known alive.
Log-rank test
The formal primary comparison used a log-rank test. The log-rank test compares the observed pattern of events between treatment groups across follow-up rather than comparing only one fixed time point.
This is especially appropriate when the endpoint is time to death because participants can contribute different amounts of follow-up and can be censored. The test uses information accumulated over the observed event times.
Cox proportional-hazards model
The treatment effect was summarized using a hazard ratio from a Cox proportional-hazards model. The registry states that the hazard ratio was obtained from a model stratified by age, cytogenetic risk category, and whether participants received consolidation therapy.
A hazard ratio is a relative time-to-event measure. It is not a probability, not a percentage of patients who benefit, and not the difference between two survival percentages.
Intention-to-treat analysis
The primary overall-survival analysis used the intent-to-treat (ITT) population. The registry defines this population as participants who were randomized, regardless of whether they received treatment or not.
Analyzing randomized participants according to their assigned group preserves the treatment comparison created by randomization. This is particularly important in randomized trials because treatment discontinuation or other post-randomization events should not automatically redefine the original treatment groups for the primary efficacy comparison.
Stratified analysis
The primary hazard ratio was stratified by age, cytogenetic risk category, and received consolidation therapy or not. Stratification allows the time-to-event comparison to account for these prespecified factors when estimating the relative treatment effect.
Conceptually, the analysis does not require the treatment groups to be compared as though every participant had the same distribution of these factors. Instead, the Cox model estimates the treatment effect within the stratified structure and combines the information across strata.
Confidence interval construction
The registry states that the confidence interval for the difference was derived using Kosorok's method. The reported interval is two-sided and has a confidence level of 95%.
The confidence interval should be read together with the point estimate. A point estimate such as HR 0.69 gives one estimate of the relative treatment effect; the 95% confidence interval of 0.55 to 0.86 describes the statistical uncertainty surrounding that estimate under the analysis framework.
6. Statistical Methods Explained
Why was a Kaplan-Meier analysis appropriate?
Overall survival is inherently a time-to-event endpoint. Some participants may die during follow-up while others remain alive when their observation ends. Kaplan-Meier estimation handles this right-censoring structure without requiring every participant to experience the event.
The important distinction is that a censored participant is not treated as having survived forever. Instead, that participant contributes information up to the last date on which survival status is known.
What does the hazard ratio of 0.69 mean?
The reported hazard ratio of 0.69 means that the estimated hazard of death associated with oral azacitidine plus best supportive care was 0.69 times the corresponding hazard under placebo plus best supportive care, according to the stratified Cox model.
Expressed as a simple relative interpretation, 1 − 0.69 = 0.31, so the estimate corresponds to approximately a 31% lower estimated hazard. That statement concerns the estimated hazard, not a 31% absolute reduction in mortality and not a claim that every participant experienced the same reduction.
Why use both a log-rank test and a Cox model?
The two methods answer related but distinct questions. The log-rank test provides a formal comparison of the time-to-event distributions between the randomized groups. The Cox model provides a clinically interpretable effect measure, the hazard ratio, together with its confidence interval.
Using both makes the result easier to interpret: the hypothesis test addresses evidence against the null hypothesis, while the hazard ratio quantifies the estimated relative treatment effect.
Why does the confidence interval matter?
The 95% confidence interval of 0.55 to 0.86 communicates the precision of the estimated hazard ratio. The interval is substantially narrower than an interval spanning very broad values, indicating that the estimated treatment effect is not represented by a single point estimate alone.
The interval does not mean that 95% of individual patients experience hazard ratios between 0.55 and 0.86. It describes uncertainty in the estimated treatment effect under the statistical model and sampling framework.
Why does the p-value not measure effect size?
The primary p-value is 0.0009. A p-value quantifies the compatibility of the observed data with the null hypothesis under the specified testing framework; it does not measure how large or clinically important the treatment effect is.
The effect size is communicated by the hazard ratio and its confidence interval. A very small p-value can accompany either a modest or a large effect depending on the amount of information in the study.
What should be considered when interpreting a Cox hazard ratio?
A Cox hazard ratio is model-based and is most naturally interpreted as a relative comparison of instantaneous event rates. Its interpretation requires care when hazards are not reasonably proportional over time. The ClinicalTrials.gov record does not report a separate assessment of the proportional-hazards assumption, so the reported HR should be interpreted as the prespecified model-based summary rather than as proof that the hazards were proportional throughout follow-up.
7. Primary Result: Overall Survival
The primary endpoint was the Kaplan-Meier (K-M) Estimate for Overall Survival (OS). The registered time frame was Day 1 (randomization) up to the data cut off date of 15 July 2019, with median follow-up for OS estimated by the reverse K-M method as described in the registry record.
Hazard ratio for overall survival
95% CI: 0.55–0.86 · P = 0.0009
Two-sided 95% confidence interval · Superiority hypothesis
| Primary endpoint | Oral Azacitidine + Best Supportive Care | Placebo + Best Supportive Care | Effect estimate |
|---|---|---|---|
| Overall survival | ITT population | ITT population | HR 0.69 (95% CI 0.55–0.86); P = 0.0009 |
The registry reports the hazard ratio from a Cox proportional-hazards model stratified by age, cytogenetic risk category, and received consolidation therapy or not. The formal comparison was performed using the log-rank test.
What the estimate means: An HR of 0.69 indicates a lower estimated hazard of death in the oral azacitidine plus best supportive care group relative to the placebo plus best supportive care group under the reported stratified Cox model. Equivalently, 0.69 corresponds to an estimated 31% lower hazard relative to the comparator because 1 − 0.69 = 0.31.
What it does not mean: It does not mean that 31% fewer participants died, that 31% of participants benefited, or that an individual participant's probability of death was reduced by exactly 31%. Hazard is a time-dependent rate concept rather than a simple probability.
What the confidence interval says: The two-sided 95% CI of 0.55 to 0.86 quantifies uncertainty around the estimated hazard ratio. It provides a range of values compatible with the statistical estimation framework; it is not a range of individual patient outcomes.
What the p-value says: The p-value of 0.0009 provides evidence against the null hypothesis in the reported superiority analysis. It does not measure the magnitude of the treatment effect. The magnitude is conveyed by the HR and its confidence interval.
Model caution: Because the effect estimate comes from a Cox proportional-hazards model, interpretation should recognize the model's proportional-hazards framework. The ClinicalTrials.gov record does not report a separate diagnostic of that assumption.
Analysis population: The analysis was conducted in the ITT population, defined by randomization regardless of whether participants received treatment.
Why the primary result is a time-to-event result
Overall survival is not adequately represented by a simple comparison of the number of deaths because follow-up can differ among participants and because censoring occurs. The analysis instead uses the ordering and timing of deaths together with each participant's observed follow-up.
This means the reported HR should be read as the summary of a longitudinal comparison. It incorporates information from the complete observed follow-up rather than reducing the trial to a single binary endpoint.
8. Secondary Result: Relapse Free Survival
The registry also reports a secondary Kaplan-Meier Estimate of Relapse Free Survival (RFS). The analysis used the ITT population and compared oral azacitidine plus best supportive care with placebo plus best supportive care using a log-rank test and a stratified Cox proportional-hazards model.
Hazard ratio for relapse free survival
95% CI: 0.52–0.80 · P < 0.0001
Two-sided 95% confidence interval · Superiority hypothesis
| Secondary endpoint | Analysis population | Method | Effect estimate |
|---|---|---|---|
| Relapse Free Survival | Intent-to-treat population | Log-rank test; stratified Cox proportional-hazards model | HR 0.65 (95% CI 0.52–0.80); P < 0.0001 |
The registry states that the RFS hazard ratio was obtained from a Cox proportional-hazards model stratified by age, cytogenetic risk category, and received consolidation therapy or not. The registered time frame extends from day 1 (randomization) to the data cutoff date of 06 August 2024; the registry describes this as approximately 135.5 months.
What the estimate means: An HR of 0.65 corresponds to an estimated hazard of relapse or the relevant RFS event that is 0.65 times the comparator hazard under the reported Cox model. As a simple relative interpretation, this corresponds to approximately a 35% lower estimated hazard because 1 − 0.65 = 0.35.
What it does not mean: The HR is not a 35% absolute reduction in relapse risk and does not mean that 35% of patients were protected from relapse. It is a relative time-to-event measure.
Precision: The 95% CI of 0.52 to 0.80 describes uncertainty around the estimated HR. The interval should be interpreted as an estimate of statistical precision, not as the distribution of individual patient effects.
Evidence versus effect magnitude: The p-value of < 0.0001 indicates strong evidence against the null hypothesis under the reported superiority testing framework, but it does not itself quantify the size of the treatment effect.
9. Secondary Result: EQ-5D Health-Related Quality of Life Deterioration
A third posted statistical analysis evaluated Time to Definitive Clinically Meaningful Deterioration for ≥ 2 Consecutive Visits as Measured Using the EQ-5D HRQoL Scale.
Hazard ratio for time to deterioration
95% CI: 0.6136–1.4231 · P = 0.7522
Two-sided 95% confidence interval · Superiority hypothesis
| Secondary endpoint | Population | Method | Effect estimate |
|---|---|---|---|
| Time to definitive clinically meaningful deterioration for ≥ 2 consecutive visits as measured using EQ-5D HRQoL | HRQoL evaluable population | Cox proportional-hazards model | HR 0.9345 (95% CI 0.6136–1.4231); P = 0.7522 |
The registry defines the HRQoL evaluable population as the ITT participants with a non-missing EQ-5D-3L score at baseline (C1D1) and at least one post-baseline visit. The reported time frame was from day 1 (randomization) up to the data cutoff date of 15 July 2019, approximately 74 months.
What the estimate means: The HR of 0.9345 is close to 1. Under the reported Cox model, the estimated instantaneous hazard of definitive clinically meaningful deterioration was 0.9345 times the comparator hazard.
What the confidence interval means: The 95% CI of 0.6136 to 1.4231 is relatively broad and includes 1. This means the estimated treatment effect is compatible with both lower and higher hazards under the statistical uncertainty represented by the interval.
What the p-value means: The p-value of 0.7522 does not provide evidence against the null hypothesis under the reported superiority analysis. It should not be interpreted as proof that the treatment groups are identical.
Why this endpoint should not be merged with the OS result: Overall survival and HRQoL deterioration measure different clinical constructs. A treatment effect on survival does not automatically imply the same effect on a patient-reported quality-of-life time-to-event endpoint.
10. Safety
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. These counts should be interpreted as a safety description rather than as an efficacy endpoint.
| Safety measure | Oral Azacitidine Plus Best Supportive Care | Placebo Plus Best Supportive Care |
|---|---|---|
| Serious adverse events | 110/236 | 109/233 |
The denominators shown above are the reported numbers at risk for this serious-adverse-event measure. The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or p-value for serious adverse events, so this page does not construct one.
That distinction matters. A safety table can describe the observed number of affected participants without implying that a treatment-group difference has been statistically demonstrated. Formal inference would require the prespecified safety-analysis framework and appropriate handling of exposure, event definitions, and follow-up.
11. Understanding the Analysis Population
The primary OS and secondary RFS analyses were performed in the intent-to-treat population. The registry explicitly defines the ITT population as participants who were randomized, regardless of whether they received treatment or not.
Why ITT matters
Randomization creates the foundation for the treatment comparison. Retaining participants according to randomized assignment preserves that structure for the primary efficacy analysis.
Post-randomization events
The ITT definition means that failure to receive treatment does not automatically remove a randomized participant from the efficacy population.
HRQoL population
The EQ-5D deterioration analysis used a narrower HRQoL evaluable population requiring a non-missing baseline score and at least one post-baseline visit.
Interpretation consequence
Different analysis populations answer slightly different questions and should not be treated as though they represent exactly the same set of participants.
12. Why Stratification Matters
The primary OS hazard ratio was estimated using a Cox model stratified by age, cytogenetic risk category, and received consolidation therapy or not. The same stratification factors are identified in the RFS analysis.
| Stratification factor | Role in the reported analysis |
|---|---|
| Age | Stratification factor in the Cox model |
| Cytogenetic risk category | Stratification factor in the Cox model |
| Received consolidation therapy or not | Stratification factor in the Cox model |
Stratification is useful when prognostic factors are expected to influence the event process. Instead of estimating one unqualified comparison that ignores these factors, the model compares treatment groups within the stratified framework.
Importantly, stratification does not transform the hazard ratio into an absolute treatment effect. The resulting HR remains a relative time-to-event measure. The purpose is to obtain the treatment comparison while accounting for the specified stratification structure.
13. Confidence Intervals: Reading the Three Reported Results
| Endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Overall survival | 0.69 | 0.55–0.86 | 0.0009 |
| Relapse free survival | 0.65 | 0.52–0.80 | < 0.0001 |
| EQ-5D deterioration | 0.9345 | 0.6136–1.4231 | 0.7522 |
These three estimates illustrate why a clinical trial should not be reduced to its p-values. The OS and RFS estimates are below 1 with confidence intervals that do not include 1, while the EQ-5D deterioration estimate is close to 1 and its confidence interval includes 1.
The confidence intervals also demonstrate that the precision of an estimate matters. A point estimate alone can create a false impression of certainty. The interval supplies the surrounding statistical uncertainty and helps distinguish a relatively concentrated estimate from one that permits a much wider range of effects.
14. Hazard Ratios and Absolute Interpretation
A hazard ratio compares instantaneous event rates under the fitted survival model. It does not directly state an absolute survival probability at a specific time.
This distinction is especially important when communicating time-to-event results. A hazard ratio summarizes the relative event process, whereas a Kaplan-Meier survival probability at a specific time answers a different question: what proportion of the analyzed population remains event-free at that time under the estimated survival curve?
The ClinicalTrials.gov record provides the HR estimates and confidence intervals but do not provide median overall survival, median RFS, or specific Kaplan-Meier survival probabilities. Those quantities are therefore not introduced here.
15. Censoring and Follow-up
The registry definition of overall survival specifies that participants who survive to the end of follow-up, withdraw consent, or are lost to follow-up are censored at the date last known alive.
Censoring is a central feature of survival analysis. A participant who is censored still contributes information before the censoring date. The analysis then uses the information available from the remaining participants who are still at risk.
16. Statistical Meaning of the Primary P-value
The primary OS comparison produced P = 0.0009 under the reported log-rank analysis.
A p-value is calculated relative to a null hypothesis and the statistical testing framework. It asks how compatible the observed evidence is with the null hypothesis under that framework. It is not the probability that the null hypothesis is true, and it is not the probability that the observed treatment effect occurred by chance.
For this trial, the p-value should therefore be interpreted together with the HR of 0.69 and the 95% CI of 0.55 to 0.86. The three quantities serve different purposes:
Hazard ratio
Describes the estimated relative treatment effect on the hazard scale.
Confidence interval
Describes statistical uncertainty and precision around the estimated effect.
P-value
Quantifies evidence against the null hypothesis under the specified test.
Analysis population
Defines which randomized participants contribute to the primary efficacy analysis.
17. What the Primary Result Does — and Does Not — Establish
The registered primary analysis reports a hazard ratio of 0.69 for overall survival, with a two-sided 95% CI of 0.55–0.86 and P = 0.0009. The analysis used a log-rank test, and the HR came from a Cox proportional-hazards model stratified by age, cytogenetic risk category, and received consolidation therapy or not.
The result does not provide a median survival time, a specific absolute survival probability, an individual patient's probability of benefit, or a guarantee that the same relative effect applies uniformly across all participants and all points in time.
Clinical trial statistics are most informative when the estimand, analysis population, effect measure, uncertainty interval, and testing framework are considered together. The hazard ratio is one component of that statistical story rather than a complete description of patient-level outcomes.
18. Limitations
- Limited registry detail: the ClinicalTrials.gov record identifies the primary endpoint, analysis methods, effect estimates, confidence intervals, p-values, analysis populations, and selected safety data, but do not provide a full statistical analysis plan.
- No median survival values reported: median overall survival and median relapse free survival are not included in the ClinicalTrials.gov record and are therefore not reported on this page.
- No baseline table reported: detailed baseline demographic and disease characteristics are not included in the ClinicalTrials.gov record.
- No subgroup estimates reported: the ClinicalTrials.gov record does not provide treatment-effect estimates for individual subgroups beyond the stratification factors used in the Cox model.
- Proportional-hazards assumption: the Cox model produces a hazard ratio under its modeling framework. The ClinicalTrials.gov record does not report a separate diagnostic of proportional hazards.
- Safety inference: serious adverse-event counts are reported, but the ClinicalTrials.gov record does not include a formal statistical comparison of those safety outcomes.
- HRQoL population: the EQ-5D deterioration analysis uses a defined HRQoL evaluable population rather than the complete ITT population.
- Endpoint-specific interpretation: overall survival, relapse free survival, and EQ-5D deterioration are different endpoints and should not be treated as interchangeable measures of treatment effect.
- Multiplicity information: the ClinicalTrials.gov record identifies three posted statistical analyses but do not provide a complete multiplicity-control strategy for all registered and posted outcomes.
- Interim-analysis information: the ClinicalTrials.gov record does not provide an interim-analysis schedule or alpha-spending procedure, so no such design feature is inferred here.
- Missing-data and imputation details: except for the HRQoL population definition, the ClinicalTrials.gov record does not specify a broader missing-data or imputation strategy.
19. Why This Trial Matters Statistically
QUAZAR AML-001 is a useful teaching case because the ClinicalTrials.gov record connects a randomized parallel-group design with multiple time-to-event endpoints and several core survival-analysis concepts.
| Concept | How it appears in QUAZAR AML-001 |
|---|---|
| Randomization | The allocation is randomized across two parallel arms. |
| Blinding | The trial is described as quadruple masked. |
| Intention-to-treat analysis | The primary OS and secondary RFS analyses use randomized participants regardless of treatment receipt. |
| Kaplan-Meier estimation | The primary endpoint and RFS are registered as Kaplan-Meier estimates. |
| Log-rank test | The primary OS and secondary RFS comparisons use a log-rank test. |
| Cox model | Hazard ratios are derived from Cox proportional-hazards models. |
| Hazard ratio | OS, RFS, and EQ-5D deterioration are summarized using hazard ratios. |
| Stratified analysis | OS and RFS Cox models are stratified by age, cytogenetic risk category, and consolidation therapy status. |
| Confidence intervals | All three reported statistical analyses include two-sided 95% confidence intervals. |
| P-values | Formal p-values are reported for the OS, RFS, and EQ-5D deterioration analyses. |
| Time-to-event endpoints | Overall survival, relapse free survival, and time to clinically meaningful deterioration are analyzed as time-to-event outcomes. |
The trial is particularly useful for showing that a single clinical study can contain several related but non-identical statistical questions. The primary OS analysis asks about time to death. RFS asks a different time-to-event question. The EQ-5D analysis asks when a clinically meaningful deterioration occurs according to a patient-reported health-related quality-of-life measure. The same broad survival-analysis toolkit can be used across these endpoints, while the interpretation of each endpoint remains specific to what it measures.
20. A Statistical Reading of the Three Effect Estimates
| Endpoint | HR | Simple relative interpretation | 95% CI |
|---|---|---|---|
| Overall survival | 0.69 | Approximately 31% lower estimated hazard | 0.55–0.86 |
| Relapse free survival | 0.65 | Approximately 35% lower estimated hazard | 0.52–0.80 |
| EQ-5D deterioration | 0.9345 | Approximately 6.55% lower estimated hazard, based only on 1 − 0.9345 | 0.6136–1.4231 |
The simple percentage interpretations above are arithmetic descriptions of the reported hazard ratios; they are not separate clinical effect measures. For the first two endpoints, the point estimates are below 1 and their confidence intervals remain below 1. For EQ-5D deterioration, the confidence interval crosses 1, so the point estimate alone should not be treated as evidence of a treatment effect.
This comparison illustrates why effect estimates should always be accompanied by uncertainty. A point estimate near 1 can have a wide interval, while a point estimate farther from 1 can be estimated with greater precision. The confidence interval is therefore essential to understanding what the data can and cannot establish.
21. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: NCT01757535 — QUAZAR AML-001.
- PubMed: PMID 37952982.
- PubMed: PMID 36951156.
- PubMed: PMID 35960871.
- PubMed: PMID 35437111.
- PubMed: PMID 34995344.
Continue through Clinical Biostats
Explore the statistical methods behind clinical-trial endpoints and connect them with practical tutorials and analysis tools.
24. Record Summary
QUAZAR AML-001 provides a clear example of randomized clinical-trial survival analysis. The trial used randomized parallel-group allocation, quadruple masking, and an ITT primary efficacy population. Overall survival was the registered primary time-to-event endpoint, analyzed using Kaplan-Meier estimation and a log-rank test, with the treatment effect summarized by a Cox proportional-hazards model stratified by age, cytogenetic risk category, and consolidation therapy status.
The primary OS analysis reported HR 0.69 with a 95% CI of 0.55–0.86 and P = 0.0009. The secondary RFS analysis reported HR 0.65 with a 95% CI of 0.52–0.80 and P < 0.0001. The secondary EQ-5D deterioration analysis reported HR 0.9345 with a 95% CI of 0.6136–1.4231 and P = 0.7522.
The most important statistical lesson is that these results should be interpreted as time-to-event estimates with uncertainty, not as isolated p-values. The hazard ratio communicates relative event rates under the Cox model, the confidence interval communicates precision, the log-rank test supplies the formal survival comparison, and the ITT population preserves the randomized treatment comparison. The different secondary endpoints also demonstrate why the meaning of a treatment effect depends on the endpoint being analyzed.