This page separates reported trial results from statistical interpretation. Numerical trial results and design facts on this page are restricted to the ClinicalTrials.gov record for NCT01568866. The registry provides the official trial record.
1. Trial at a Glance
ENDEAVOR was a randomized, open-label, parallel-group phase 3 study evaluating carfilzomib and dexamethasone versus bortezomib and dexamethasone in patients with multiple myeloma.
| Feature | ENDEAVOR |
|---|---|
| Phase | Phase 3 |
| Condition | Multiple Myeloma |
| Design | Randomized, parallel-group, open-label |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 929 |
| Primary endpoint | Progression-free Survival |
| Primary endpoint type | Time-to-event |
| Primary hypothesis | Superiority |
| Lead sponsor | Amgen |
| Sponsor type | Industry |
| Trial status | Completed |
| ClinicalTrials.gov | NCT01568866 |
2. Clinical Question
The central question was whether treatment with carfilzomib plus dexamethasone produced a superior progression-free survival outcome compared with bortezomib plus dexamethasone in the randomized trial population with multiple myeloma.
Population
Participants with multiple myeloma enrolled in the phase 3 ENDEAVOR trial.
Intervention
Carfilzomib plus dexamethasone.
Comparator
Bortezomib plus dexamethasone.
Primary question
Does carfilzomib plus dexamethasone improve progression-free survival relative to bortezomib plus dexamethasone?
3. Trial Design
Carfilzomib + DEX
- Carfilzomib
- Dexamethasone
Bortezomib + DEX
- Bortezomib
- Dexamethasone
The ClinicalTrials.gov record identifies the study as randomized, parallel, and unmasked. The enrollment target recorded in the ClinicalTrials.gov record was 929 participants, with two arms.
Trial timeline
Study start
The registered study start date was 20 June 2012.
Primary completion
The registered primary completion date was 10 November 2014.
4. Endpoints
| Endpoint | Role | Definition / time frame | Analysis type |
|---|---|---|---|
| Progression-free Survival | Primary | From randomization until the data cut-off date of 10 November 2014; median follow-up time for PFS was 11.1.and 11.9 months in the bortezomib and carfilzomib arms respectively | Time-to-event |
| Overall Survival | Secondary | From randomization until the data cut-off date of 03 January 2017; median follow-up time for OS was 36.9 and 37.5 months for each treatment group respectively. | Time-to-event |
| Overall Response | Secondary | Disease response was assessed every 28 days until end of treatment or the data cut-off date of 10 November 2014; median duration of treatment was 27 weeks in the bortezomib group and 40 weeks in the carfilzomib treatment group. | Binary |
| Percentage of Participants With ≥ Grade 2 Peripheral Neuropathy | Secondary | From the first dose of study drug up to 30 days after the last dose of study drug as of the data cut-off date of 10 November 2014; median duration of treatment was 27 weeks in the bortezomib group and 40 weeks in the carfilzomib treatment group. | Binary |
Primary endpoint definition
Progression-free survival was defined as the time from randomization to the earlier of disease progression or death due to any cause. Participants were evaluated for disease response and progression according to the International Myeloma Working Group-Uniform Response Criteria (IMWG-URC) as assessed by an Independent Review Committee (IRC). The registry states that median PFS was estimated using the Kaplan-Meier method.
5. Analysis Populations and Comparison Framework
The primary progression-free survival and the reported overall survival and overall response analyses used the intent-to-treat population. The peripheral-neuropathy analysis used the safety population, defined as all participants who received at least one dose of study treatment.
| Analysis | Population | Comparison |
|---|---|---|
| Progression-free Survival | Intent-to-treat | Bortezomib + DEX vs Carfilzomib + DEX |
| Overall Survival | Intent-to-treat | Bortezomib + DEX vs Carfilzomib + DEX |
| Overall Response | Intent-to-treat | Bortezomib + DEX vs Carfilzomib + DEX |
| ≥ Grade 2 Peripheral Neuropathy | Safety population: all participants who received at least 1 dose of study treatment | Bortezomib + DEX vs Carfilzomib + DEX |
The use of an ITT population for efficacy preserves the randomized treatment assignment as the basis of the primary comparison. This is important because the estimand represented by an ITT analysis is tied to assignment at randomization rather than only to participants who remain on treatment.
6. Statistical Methodology
Stratified log-rank test
The primary PFS comparison was performed using a stratified log-rank test. The registry analysis notes state that the log-rank test was stratified by the randomization stratification factors.
A stratified log-rank analysis compares the observed and expected event patterns between treatment groups while accounting for the prespecified strata. Conceptually, this prevents the treatment comparison from ignoring the structure introduced by stratified randomization.
Kaplan-Meier estimation
Progression-free survival and overall survival are time-to-event endpoints. Kaplan-Meier estimation is appropriate because not every participant necessarily experiences the event during the observation period. Participants without an observed event contribute information until their censoring time.
Here, di is the number of events at time ti and ni is the number of participants at risk immediately before that time.
Cox proportional-hazards modeling
The registry analysis notes state that the hazard ratio for PFS was estimated using a Cox proportional-hazards approach. The overall-survival analysis notes likewise state that the hazard ratio was estimated using a Cox proportional-hazards model.
For the reported ENDEAVOR analyses, an HR below 1 indicates a lower estimated instantaneous event rate for carfilzomib relative to bortezomib under the fitted model.
Cochran-Mantel-Haenszel analysis
Overall response was analyzed using a stratified Cochran-Mantel-Haenszel test. The registry notes state that the odds ratio was calculated using the Cochran-Mantel-Haenszel method stratified by prior proteasome inhibitor treatment, lines of prior treatment, ISS stage, and choice of route of bortezomib administration.
The peripheral-neuropathy analysis used the unconditional Cochran-Mantel-Haenszel method to estimate the odds ratio.
Effect measures
| Endpoint | Effect measure | Statistical method |
|---|---|---|
| Progression-free Survival | Hazard ratio | Stratified log-rank test; Cox proportional-hazards estimation |
| Overall Survival | Hazard ratio | Stratified log-rank test; Cox proportional-hazards estimation |
| Overall Response | Odds ratio | Stratified Cochran-Mantel-Haenszel test |
| ≥ Grade 2 Peripheral Neuropathy | Odds ratio | Unconditional Cochran-Mantel-Haenszel method |
7. Results: Progression-Free Survival
Progression-free survival was the registered primary endpoint. The analysis was conducted in the intent-to-treat population using a stratified log-rank test, with the hazard ratio reported as the effect measure.
Primary endpoint: progression-free survival
95% CI: 0.437–0.651 · P < 0.0001
Analysis population: intent-to-treat · Hypothesis: superiority
| Endpoint | Comparison | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|---|
| Progression-free Survival | Carfilzomib + DEX vs Bortezomib + DEX | HR 0.533 | 0.437–0.651 | < 0.0001 | Stratified log-rank |
The PFS analysis used data from randomization until the data cut-off date of 10 November 2014. The registry time frame reports median follow-up time for PFS as 11.1.and 11.9 months in the bortezomib and carfilzomib arms respectively.
The hazard ratio of 0.533 means that the estimated instantaneous rate of progression or death in the carfilzomib group was about 53.3% of the corresponding rate in the bortezomib group under the reported time-to-event model. Equivalently, 1 − 0.533 = 0.467, so the estimate corresponds to approximately a 46.7% lower estimated hazard.
This does not mean that 46.7% of participants avoided progression, that 46.7% of participants were cured, or that every participant experienced exactly a 46.7% reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute probability.
The 95% confidence interval of 0.437–0.651 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual treatment effects across patients.
The p-value of < 0.0001 addresses evidence against the null hypothesis under the specified testing framework. It does not measure the magnitude of the treatment effect; the HR and its confidence interval provide that information.
The registry analysis notes specify a group-sequential monitoring plan using an O'Brien-Fleming-type efficacy stopping boundary constructed with the Lan-DeMets alpha-spending function to ensure a 1-sided Type I error rate ≤ 0.025. The log-rank test was stratified by the randomization stratification factors. Because the endpoint is time-to-event, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio model.
8. Results: Overall Survival
Overall survival was a secondary time-to-event endpoint. The reported analysis used the intent-to-treat population and a stratified log-rank test. The registry notes specify that the second interim analysis was to be conducted after 394 events had been reached.
Overall survival
95% CI: 0.648–0.964 · P = 0.0100
Analysis population: intent-to-treat · Hypothesis: superiority
| Endpoint | Comparison | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|---|
| Overall Survival | Carfilzomib + DEX vs Bortezomib + DEX | HR 0.791 | 0.648–0.964 | 0.0100 | Stratified log-rank |
The OS time frame was from randomization until the data cut-off date of 03 January 2017. The ClinicalTrials.gov record reports median follow-up time for OS as 36.9 and 37.5 months for each treatment group respectively..
The OS hazard ratio of 0.791 corresponds to an estimated instantaneous death rate in the carfilzomib group that was approximately 79.1% of the rate in the bortezomib group under the reported model. In relative terms, 1 − 0.791 = 0.209, corresponding to approximately a 20.9% lower estimated hazard of death.
The HR does not mean that 20.9% of patients lived longer, nor does it give an absolute difference in survival probability. It summarizes the relative event-rate comparison over the analyzed time-to-event data.
The 95% confidence interval of 0.648–0.964 gives the statistical uncertainty around the HR estimate. It does not indicate that individual patients' treatment effects fall within this interval.
The reported p-value of 0.0100 is a hypothesis-testing quantity, not an effect-size measure. The analysis notes specify a one-sided significance level of 0.0123, determined using an O'Brien-Fleming-type alpha-spending function based on the actual number of events. The reported two-sided 95% confidence interval and the specified one-sided interim-testing framework should therefore be kept conceptually distinct.
As with PFS, the Cox-model hazard ratio requires attention to the proportional-hazards assumption and to censoring. The registry does not provide enough information in the ClinicalTrials.gov record to reconstruct the full survival curve or assess that assumption directly.
9. Results: Overall Response
Overall response was a secondary binary endpoint. Disease response was assessed every 28 days until the end of treatment or the data cut-off date of 10 November 2014. The analysis used the intent-to-treat population and a stratified Cochran-Mantel-Haenszel method.
Overall response
95% CI: 1.519–2.718 · P < 0.0001
Odds ratio calculated as carfilzomib relative to bortezomib
| Endpoint | Comparison | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|---|
| Overall Response | Carfilzomib + DEX vs Bortezomib + DEX | OR 2.032 | 1.519–2.718 | < 0.0001 | Stratified Cochran-Mantel-Haenszel |
An odds ratio of 2.032 means that the estimated odds of overall response were about 2.032 times as high in the carfilzomib group as in the bortezomib group, using the stratified Cochran-Mantel-Haenszel analysis.
Odds are not the same as probabilities. An OR of 2.032 therefore cannot be read as "twice the response rate" or as a 103.2-percentage-point increase in response. Converting an odds ratio into an absolute probability difference requires the underlying response probability in a comparison group.
The 95% confidence interval of 1.519–2.718 describes uncertainty around the estimated odds ratio. The p-value of < 0.0001 addresses the hypothesis test and does not quantify how clinically large the response difference is.
The analysis was stratified by prior proteasome inhibitor treatment, lines of prior treatment, ISS stage, and choice of route of bortezomib administration. Stratification can improve the alignment between the analysis and the randomized comparison when these factors are part of the trial's allocation structure.
10. Results: ≥ Grade 2 Peripheral Neuropathy
The registry reports a secondary safety endpoint measuring the percentage of participants with ≥ Grade 2 peripheral neuropathy. The analysis used the safety population, defined as participants who received at least one dose of study treatment.
Peripheral neuropathy
95% CI: 0.089–0.210 · P < 0.0001
Odds ratio estimated for carfilzomib relative to bortezomib
| Endpoint | Comparison | Estimate | 95% CI | P-value | Population |
|---|---|---|---|---|---|
| ≥ Grade 2 Peripheral Neuropathy | Carfilzomib + DEX vs Bortezomib + DEX | OR 0.137 | 0.089–0.210 | < 0.0001 | Safety population |
The odds ratio of 0.137 indicates that the estimated odds of the reported ≥ Grade 2 peripheral-neuropathy endpoint were lower in the carfilzomib group than in the bortezomib group under the reported analysis. The corresponding estimate is about 13.7% of the comparator odds.
This is an odds ratio, not a risk ratio. It should not be interpreted as saying that exactly 13.7% as many participants experienced the event. The absolute event probabilities are required to translate the odds ratio into an absolute risk difference.
The 95% confidence interval of 0.089–0.210 describes uncertainty around the odds-ratio estimate. The p-value of < 0.0001 provides evidence against the null hypothesis under the reported testing method but does not itself quantify the magnitude of the safety difference.
Because this analysis uses the safety population rather than the ITT population, its interpretation is exposure-based rather than purely assignment-based. The registry states that the analysis used the unconditional Cochran-Mantel-Haenszel method.
11. Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.
| Treatment arm | Affected | At risk |
|---|---|---|
| Bortezomib + DEX | 182 | 456 |
| Carfilzomib + DEX | 272 | 463 |
The reported serious-adverse-event counts are descriptive. The ClinicalTrials.gov record does not provide a formal statistical analysis, confidence interval, or p-value for this endpoint, so this page does not construct one.
12. Interim Analysis, Alpha Spending, and Multiplicity
The registry-reported analysis notes explicitly support an interim-analysis framework for the time-to-event endpoints. The PFS interim analysis was to use a group sequential monitoring plan with an O'Brien-Fleming-type efficacy stopping boundary constructed using the Lan-DeMets alpha-spending function.
Lan-DeMets alpha spending
Alpha spending provides a way to control the overall Type I error while allowing interim looks at accumulating data. The registry-reported PFS analysis notes specify a 1-sided Type I error rate ≤ 0.025.
O'Brien-Fleming-type boundary
The stated boundary framework makes early efficacy decisions more stringent than later decisions, preserving the overall error-control objective while allowing the trial to stop early for compelling evidence.
OS interim analysis
The second interim OS analysis was to be conducted after 394 events had been reached.
One-sided OS threshold
The OS analysis notes specify a one-sided significance level of 0.0123 based on the O'Brien-Fleming-type alpha-spending function and the actual number of events.
These details matter because a trial that examines accumulating efficacy data cannot simply treat every interim p-value as though it came from a single, preplanned final analysis. Alpha spending adjusts the testing framework so that repeated opportunities to declare efficacy remain compatible with the prespecified Type I error objective.
The ClinicalTrials.gov record identifies the spending framework and its stated error-rate targets. They do not provide the complete statistical-analysis-plan implementation, so this page does not reconstruct additional boundaries or information fractions.
13. Stratification and the Cochran-Mantel-Haenszel Framework
Stratification appears repeatedly in the ENDEAVOR statistical analysis. The registry identifies the log-rank analyses as stratified by the randomization stratification factors. The overall-response analysis additionally specifies the factors used for its Cochran-Mantel-Haenszel calculation.
| Analysis | Stratification information reported in the registry |
|---|---|
| Progression-free Survival | Log-rank test stratified by the randomization stratification factors |
| Overall Survival | Log-rank test stratified by the randomization stratification factors |
| Overall Response | Prior proteasome inhibitor treatment; lines of prior treatment; ISS stage; choice of route of bortezomib administration |
| Peripheral Neuropathy | Unconditional Cochran-Mantel-Haenszel method |
For a stratified analysis, the treatment comparison is formed across strata rather than by ignoring them. This can be particularly useful when baseline factors are related to prognosis or were incorporated into the randomization scheme. The important point is that the analysis should be aligned with the trial's prespecified stratification structure.
14. Statistical Methods Explained
Why use a stratified log-rank test for PFS?
PFS is a time-to-event endpoint, so the analysis must account for both when events occur and for participants who are censored. The stratified log-rank test compares the event experience between treatment groups while incorporating the randomization strata. This is different from simply comparing the proportion of participants who had progressed by a fixed date.
What does an HR of 0.533 mean?
An HR of 0.533 means that the estimated instantaneous rate of progression or death in the carfilzomib group was approximately 53.3% of the estimated rate in the bortezomib group under the fitted model. The corresponding relative reduction in estimated hazard is approximately 46.7%. It is not a statement that 46.7% of patients avoided an event.
Why is the confidence interval important?
The point estimate alone does not communicate how precisely the treatment effect has been estimated. For PFS, the 95% confidence interval is 0.437–0.651. It describes statistical uncertainty around the hazard-ratio estimate under the analysis framework. It does not describe the range of possible outcomes for an individual patient.
Why does the p-value not measure effect size?
A p-value evaluates compatibility with a specified null hypothesis under the statistical model and testing framework. It depends on both the magnitude of the observed effect and the amount of information in the data. An effect estimate such as an HR or OR is therefore necessary to understand the size and direction of the observed treatment difference.
Why use a Cochran-Mantel-Haenszel test for overall response?
Overall response is binary rather than time-to-event. The stratified Cochran-Mantel-Haenszel approach provides a way to compare treatment groups across prespecified strata and produces an odds-ratio estimate that accounts for those strata rather than collapsing the data into one unstratified comparison.
Why does the interim analysis affect interpretation?
When efficacy is evaluated before the planned accumulation of all information, the nominal significance threshold must account for the repeated opportunities to stop early. ENDEAVOR's registry-reported analysis notes describe an O'Brien-Fleming-type boundary implemented through Lan-DeMets alpha spending. The reported p-values therefore need to be understood within that sequential-testing framework rather than as isolated numbers.
15. Understanding the Primary PFS Result
The primary PFS HR of 0.533 indicates a lower estimated hazard of progression or death for carfilzomib relative to bortezomib in the randomized comparison. The estimate is a relative measure of event rates over time.
The 95% CI of 0.437–0.651 places the statistical uncertainty around the point estimate. The interval is substantially narrower than the entire positive range of possible hazard ratios, reflecting a more informative estimate than the point estimate alone.
The reported p-value of < 0.0001 indicates strong evidence against the null hypothesis under the reported stratified log-rank testing framework. It should not be converted into a probability that the treatment works or a probability that the observed HR is correct.
The ClinicalTrials.gov record does not report median PFS values, survival probabilities at specific time points, or event counts by treatment arm for the primary endpoint. Consequently, this page does not derive an absolute PFS difference from the HR.
16. Hazard Ratios and Proportional Hazards
The PFS and OS analyses both use hazard ratios. A hazard is an instantaneous event rate conditional on remaining event-free up to a given time. A hazard ratio compares those rates between treatment groups.
For ENDEAVOR, the reported PFS and OS hazard ratios are below 1. The interpretation is relative and model-based; it is not equivalent to an absolute risk reduction or a ratio of median survival times.
The Cox proportional-hazards framework is most straightforward to interpret when the relative hazard between groups is reasonably stable over time. The ClinicalTrials.gov record does not provide enough information to assess the proportional-hazards assumption directly. Therefore, the HR should be treated as the reported summary measure rather than as proof that the hazard ratio was constant at every time point.
17. Odds Ratios and Binary Endpoints
ENDEAVOR provides two useful examples of odds-ratio interpretation: overall response and ≥ Grade 2 peripheral neuropathy.
| Endpoint | OR | Direction for carfilzomib | 95% CI |
|---|---|---|---|
| Overall Response | 2.032 | Higher estimated odds | 1.519–2.718 |
| ≥ Grade 2 Peripheral Neuropathy | 0.137 | Lower estimated odds | 0.089–0.210 |
The two estimates move in opposite directions because they represent different binary outcomes. For response, an OR above 1 means higher estimated odds of response in the carfilzomib group. For peripheral neuropathy, an OR below 1 means lower estimated odds of the reported safety event.
18. What the P-Values Do — and Do Not — Mean
Four formal statistical analyses are posted on ClinicalTrials.gov for ENDEAVOR: one primary PFS analysis and three secondary analyses. Their reported p-values are:
| Endpoint | P-value | Effect measure |
|---|---|---|
| Progression-free Survival | < 0.0001 | HR 0.533 |
| Overall Survival | 0.0100 | HR 0.791 |
| Overall Response | < 0.0001 | OR 2.032 |
| ≥ Grade 2 Peripheral Neuropathy | < 0.0001 | OR 0.137 |
A p-value is tied to a null hypothesis and a testing procedure. It does not tell the reader the probability that the null hypothesis is true, the probability that the treatment effect is clinically important, or the probability that a patient will benefit.
In ENDEAVOR, the interpretation is further shaped by the interim-monitoring framework. The registry-reported analysis notes specifically identify one-sided Type I error control for the interim PFS and OS analyses. This is one reason that a p-value should always be read alongside the design and analysis plan rather than treated as a standalone verdict.
19. Primary vs Secondary Endpoints
| Endpoint | Role | Primary statistical question | Reported effect |
|---|---|---|---|
| Progression-free Survival | Primary | Superiority of carfilzomib + DEX vs bortezomib + DEX | HR 0.533 (95% CI 0.437–0.651), P < 0.0001 |
| Overall Survival | Secondary | Superiority of carfilzomib + DEX vs bortezomib + DEX | HR 0.791 (95% CI 0.648–0.964), P = 0.0100 |
| Overall Response | Secondary | Comparison of response odds | OR 2.032 (95% CI 1.519–2.718), P < 0.0001 |
| ≥ Grade 2 Peripheral Neuropathy | Secondary | Comparison of event odds | OR 0.137 (95% CI 0.089–0.210), P < 0.0001 |
This hierarchy matters. The primary endpoint was PFS, while OS, overall response, and peripheral neuropathy were secondary endpoints. The statistical role of a result is therefore determined not only by its numerical p-value but also by where it sits in the prespecified trial design and how interim monitoring and multiplicity were handled.
20. Limitations and Interpretation Issues
- Limited numerical detail: the ClinicalTrials.gov record provides effect estimates and confidence intervals but do not provide median PFS, median OS, event counts for the time-to-event analyses, or time-specific survival probabilities.
- Hazard-ratio interpretation: a single HR summarizes a time-to-event comparison but does not provide an absolute difference in survival probability.
- Proportional-hazards assumption: the Cox model is model-based, and the ClinicalTrials.gov record does not permit an independent assessment of proportional hazards.
- Interim monitoring: PFS and OS analyses were subject to sequential monitoring. Their p-values should therefore be interpreted within the stated alpha-spending framework.
- Analysis populations differ: efficacy analyses use the ITT population, whereas peripheral neuropathy uses the safety population.
- Odds ratios require care: ORs are not risk ratios and cannot be converted into absolute percentage-point differences without the underlying event probabilities.
- Serious adverse events: the ClinicalTrials.gov record provides affected/at-risk counts by arm but no formal comparative analysis for this safety measure.
- Open-label design: the trial is recorded as having no masking. This is a relevant design feature when considering outcomes that can involve assessment or reporting.
- Generalizability: the ClinicalTrials.gov record identifies the trial population as participants with multiple myeloma, but does not provide a full baseline-characteristics table in the ClinicalTrials.gov record.
21. Why This Trial Matters Statistically
ENDEAVOR is a useful teaching case because a single randomized phase 3 trial connects several core statistical methods: time-to-event analysis, stratified testing, Cox modeling, binary-outcome analysis, odds ratios, ITT analysis, safety-population analysis, and interim alpha spending.
| Concept | How it appears in ENDEAVOR |
|---|---|
| Randomization | Randomized, two-arm, parallel-group phase 3 design |
| Intention-to-treat analysis | Primary PFS and reported OS and response analyses used the ITT population |
| Time-to-event endpoint | PFS was the registered primary endpoint; OS was secondary |
| Kaplan-Meier estimation | Registered methodology for estimating median PFS |
| Stratified log-rank test | Used for PFS and OS comparisons |
| Hazard ratio | Reported for PFS and OS |
| Cox model | Used to estimate the reported hazard ratios |
| Cochran-Mantel-Haenszel test | Used for overall response and peripheral-neuropathy analyses |
| Odds ratio | Reported for overall response and ≥ Grade 2 peripheral neuropathy |
| Interim analysis | PFS and OS analyses included sequential monitoring |
| Alpha spending | Lan-DeMets alpha spending with an O'Brien-Fleming-type efficacy boundary |
| One-sided testing | Specified in the interim PFS and OS analysis framework |
| Safety population | Used for the peripheral-neuropathy analysis |
22. Statistical Story of ENDEAVOR
The statistical structure can be viewed as a sequence of increasingly specific questions.
The primary PFS analysis asks a time-to-event question. The hazard ratio provides the relative effect, the confidence interval describes precision, and the stratified log-rank test provides the formal comparison. Secondary analyses then address overall survival and binary outcomes such as response and peripheral neuropathy using methods appropriate to their endpoint types.
This is a useful general principle in clinical-trial statistics: the endpoint determines the data structure, and the data structure determines the appropriate statistical framework. Time-to-event outcomes require methods that account for follow-up and censoring; binary outcomes can be summarized with odds ratios and analyzed with methods such as Cochran-Mantel-Haenszel procedures.
23. Interpreting the Trial as a Statistical Analyst
First question: What was randomized?
The treatment assignment was randomized between carfilzomib plus dexamethasone and bortezomib plus dexamethasone in a two-arm parallel design.
Second question: What was primary?
Progression-free survival was the single registered primary endpoint, making it the central confirmatory efficacy outcome represented in the ClinicalTrials.gov record.
Third question: How was PFS analyzed?
With a stratified log-rank test and a hazard-ratio estimate from a Cox proportional-hazards model, using the ITT population.
Fourth question: How should secondary results be separated?
OS, overall response, and peripheral neuropathy answer different questions and use different endpoint-specific methods and analysis populations.
This separation is important because it prevents a collection of different estimates from being treated as though they were interchangeable measures of one outcome. A hazard ratio for PFS, a hazard ratio for OS, an odds ratio for response, and an odds ratio for peripheral neuropathy each have a distinct statistical meaning.
24. Overall Statistical Interpretation
The reported primary PFS analysis estimated a hazard ratio of 0.533 with a 95% confidence interval of 0.437–0.651 and a p-value of < 0.0001, using a stratified log-rank test in the ITT population.
The reported OS analysis estimated an HR of 0.791 with a 95% confidence interval of 0.648–0.964 and a p-value of 0.0100. Overall response had an OR of 2.032 with a 95% confidence interval of 1.519–2.718 and a p-value of < 0.0001.
The reported ≥ Grade 2 peripheral-neuropathy analysis produced an OR of 0.137 with a 95% confidence interval of 0.089–0.210 and a p-value of < 0.0001. Serious adverse events were reported descriptively as 182/456 in the bortezomib arm and 272/463 in the carfilzomib arm.
These summary statistics do not provide individual-level treatment effects, absolute PFS or OS differences, or a complete description of the survival curves. They also do not eliminate the need to consider endpoint hierarchy, interim monitoring, censoring, model assumptions, analysis populations, and the distinction between efficacy and safety endpoints.
25. Related Tutorials
Learn more about the methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: ENDEAVOR, NCT01568866.
- PubMed: PMID 30610657.
- PubMed: PMID 26771810.
- PubMed: PMID 26671818.
Continue through the Clinical Biostats statistical pathway
Use the methods in this trial as a starting point for deeper study of survival analysis, categorical-data methods, clinical-trial design, and statistical inference.
28. Record Summary
ENDEAVOR provides a compact example of how several clinical-trial statistical methods work together. The trial used randomized parallel-group allocation and evaluated a registered primary time-to-event endpoint, progression-free survival, with a stratified log-rank test and Cox-model hazard ratio. The reported primary PFS estimate was HR 0.533 with a 95% CI of 0.437–0.651 and P < 0.0001.
The secondary analyses extended the statistical story to overall survival, overall response, and peripheral neuropathy. OS produced an HR of 0.791 (95% CI 0.648–0.964; P = 0.0100), overall response produced an OR of 2.032 (95% CI 1.519–2.718; P < 0.0001), and ≥ Grade 2 peripheral neuropathy produced an OR of 0.137 (95% CI 0.089–0.210; P < 0.0001).
The most important statistical lesson is that these estimates should not be collapsed into a single number. Hazard ratios summarize relative time-to-event rates, odds ratios summarize relative odds for binary outcomes, confidence intervals describe statistical precision, and p-values address hypothesis tests. Interim monitoring and alpha spending further determine how evidence should be interpreted within the trial's sequential testing framework.