← Clinical Trials
Multiple Myeloma Phase 3 Completed NCT01568866

ENDEAVOR: Complete Statistical Analysis of Carfilzomib in Relapsed Multiple Myeloma

An independent statistical review of the randomized phase 3 ENDEAVOR trial comparing carfilzomib and dexamethasone with bortezomib and dexamethasone in relapsed multiple myeloma, with emphasis on progression-free survival, overall survival, response, safety, and the methods used to analyze them.

Phase 3  ·  Randomized parallel design  ·  Enrollment 929  ·  Primary completion 10 November 2014
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical trial results and design facts on this page are restricted to the ClinicalTrials.gov record for NCT01568866. The registry provides the official trial record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ENDEAVOR was a randomized, open-label, parallel-group phase 3 study evaluating carfilzomib and dexamethasone versus bortezomib and dexamethasone in patients with multiple myeloma.

929
Enrolled
2 treatment arms
2
Arms
Randomized parallel design
0.533
PFS HR
95% CI 0.437–0.651
< 0.0001
PFS P-value
Stratified log-rank
FeatureENDEAVOR
PhasePhase 3
ConditionMultiple Myeloma
DesignRandomized, parallel-group, open-label
AllocationRandomized
Primary purposeTreatment
Enrollment929
Primary endpointProgression-free Survival
Primary endpoint typeTime-to-event
Primary hypothesisSuperiority
Lead sponsorAmgen
Sponsor typeIndustry
Trial statusCompleted
ClinicalTrials.govNCT01568866

2. Clinical Question

The central question was whether treatment with carfilzomib plus dexamethasone produced a superior progression-free survival outcome compared with bortezomib plus dexamethasone in the randomized trial population with multiple myeloma.

Population

Participants with multiple myeloma enrolled in the phase 3 ENDEAVOR trial.

Intervention

Carfilzomib plus dexamethasone.

Comparator

Bortezomib plus dexamethasone.

Primary question

Does carfilzomib plus dexamethasone improve progression-free survival relative to bortezomib plus dexamethasone?

3. Trial Design

01
Randomize929 participants
02
Parallel armsTwo treatment groups
03
TreatmentCarfilzomib + DEX or bortezomib + DEX
04
AssessmentPFS, OS, response, safety
05
AnalysisStratified methods
Allocation
Randomized allocation to two parallel treatment groups.
Masking
None; the registry classifies the study as open-label.
Primary purpose
Treatment.
Hypothesis framework
Superiority.
ARM A

Carfilzomib + DEX

  • Carfilzomib
  • Dexamethasone
ARM B

Bortezomib + DEX

  • Bortezomib
  • Dexamethasone

The ClinicalTrials.gov record identifies the study as randomized, parallel, and unmasked. The enrollment target recorded in the ClinicalTrials.gov record was 929 participants, with two arms.

Trial timeline

20 June 2012

Study start

The registered study start date was 20 June 2012.

10 November 2014

Primary completion

The registered primary completion date was 10 November 2014.

4. Endpoints

EndpointRoleDefinition / time frameAnalysis type
Progression-free Survival Primary From randomization until the data cut-off date of 10 November 2014; median follow-up time for PFS was 11.1.and 11.9 months in the bortezomib and carfilzomib arms respectively Time-to-event
Overall Survival Secondary From randomization until the data cut-off date of 03 January 2017; median follow-up time for OS was 36.9 and 37.5 months for each treatment group respectively. Time-to-event
Overall Response Secondary Disease response was assessed every 28 days until end of treatment or the data cut-off date of 10 November 2014; median duration of treatment was 27 weeks in the bortezomib group and 40 weeks in the carfilzomib treatment group. Binary
Percentage of Participants With ≥ Grade 2 Peripheral Neuropathy Secondary From the first dose of study drug up to 30 days after the last dose of study drug as of the data cut-off date of 10 November 2014; median duration of treatment was 27 weeks in the bortezomib group and 40 weeks in the carfilzomib treatment group. Binary

Primary endpoint definition

Progression-free survival was defined as the time from randomization to the earlier of disease progression or death due to any cause. Participants were evaluated for disease response and progression according to the International Myeloma Working Group-Uniform Response Criteria (IMWG-URC) as assessed by an Independent Review Committee (IRC). The registry states that median PFS was estimated using the Kaplan-Meier method.

5. Analysis Populations and Comparison Framework

The primary progression-free survival and the reported overall survival and overall response analyses used the intent-to-treat population. The peripheral-neuropathy analysis used the safety population, defined as all participants who received at least one dose of study treatment.

AnalysisPopulationComparison
Progression-free SurvivalIntent-to-treatBortezomib + DEX vs Carfilzomib + DEX
Overall SurvivalIntent-to-treatBortezomib + DEX vs Carfilzomib + DEX
Overall ResponseIntent-to-treatBortezomib + DEX vs Carfilzomib + DEX
≥ Grade 2 Peripheral NeuropathySafety population: all participants who received at least 1 dose of study treatmentBortezomib + DEX vs Carfilzomib + DEX

The use of an ITT population for efficacy preserves the randomized treatment assignment as the basis of the primary comparison. This is important because the estimand represented by an ITT analysis is tied to assignment at randomization rather than only to participants who remain on treatment.

6. Statistical Methodology

Stratified log-rank test

The primary PFS comparison was performed using a stratified log-rank test. The registry analysis notes state that the log-rank test was stratified by the randomization stratification factors.

A stratified log-rank analysis compares the observed and expected event patterns between treatment groups while accounting for the prespecified strata. Conceptually, this prevents the treatment comparison from ignoring the structure introduced by stratified randomization.

Kaplan-Meier estimation

Progression-free survival and overall survival are time-to-event endpoints. Kaplan-Meier estimation is appropriate because not every participant necessarily experiences the event during the observation period. Participants without an observed event contribute information until their censoring time.

Kaplan-Meier survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di is the number of events at time ti and ni is the number of participants at risk immediately before that time.

Cox proportional-hazards modeling

The registry analysis notes state that the hazard ratio for PFS was estimated using a Cox proportional-hazards approach. The overall-survival analysis notes likewise state that the hazard ratio was estimated using a Cox proportional-hazards model.

Hazard ratio
HR = estimated hazard in the carfilzomib group ÷ estimated hazard in the bortezomib group

For the reported ENDEAVOR analyses, an HR below 1 indicates a lower estimated instantaneous event rate for carfilzomib relative to bortezomib under the fitted model.

Cochran-Mantel-Haenszel analysis

Overall response was analyzed using a stratified Cochran-Mantel-Haenszel test. The registry notes state that the odds ratio was calculated using the Cochran-Mantel-Haenszel method stratified by prior proteasome inhibitor treatment, lines of prior treatment, ISS stage, and choice of route of bortezomib administration.

The peripheral-neuropathy analysis used the unconditional Cochran-Mantel-Haenszel method to estimate the odds ratio.

Effect measures

EndpointEffect measureStatistical method
Progression-free SurvivalHazard ratioStratified log-rank test; Cox proportional-hazards estimation
Overall SurvivalHazard ratioStratified log-rank test; Cox proportional-hazards estimation
Overall ResponseOdds ratioStratified Cochran-Mantel-Haenszel test
≥ Grade 2 Peripheral NeuropathyOdds ratioUnconditional Cochran-Mantel-Haenszel method

7. Results: Progression-Free Survival

Progression-free survival was the registered primary endpoint. The analysis was conducted in the intent-to-treat population using a stratified log-rank test, with the hazard ratio reported as the effect measure.

Primary endpoint: progression-free survival

HR 0.533

95% CI: 0.437–0.651   ·   P < 0.0001

Analysis population: intent-to-treat   ·   Hypothesis: superiority

EndpointComparisonEstimate95% CIP-valueMethod
Progression-free Survival Carfilzomib + DEX vs Bortezomib + DEX HR 0.533 0.437–0.651 < 0.0001 Stratified log-rank

The PFS analysis used data from randomization until the data cut-off date of 10 November 2014. The registry time frame reports median follow-up time for PFS as 11.1.and 11.9 months in the bortezomib and carfilzomib arms respectively.

Clinical Biostats interpretation

The hazard ratio of 0.533 means that the estimated instantaneous rate of progression or death in the carfilzomib group was about 53.3% of the corresponding rate in the bortezomib group under the reported time-to-event model. Equivalently, 1 − 0.533 = 0.467, so the estimate corresponds to approximately a 46.7% lower estimated hazard.

This does not mean that 46.7% of participants avoided progression, that 46.7% of participants were cured, or that every participant experienced exactly a 46.7% reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute probability.

The 95% confidence interval of 0.437–0.651 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual treatment effects across patients.

The p-value of < 0.0001 addresses evidence against the null hypothesis under the specified testing framework. It does not measure the magnitude of the treatment effect; the HR and its confidence interval provide that information.

The registry analysis notes specify a group-sequential monitoring plan using an O'Brien-Fleming-type efficacy stopping boundary constructed with the Lan-DeMets alpha-spending function to ensure a 1-sided Type I error rate ≤ 0.025. The log-rank test was stratified by the randomization stratification factors. Because the endpoint is time-to-event, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio model.

8. Results: Overall Survival

Overall survival was a secondary time-to-event endpoint. The reported analysis used the intent-to-treat population and a stratified log-rank test. The registry notes specify that the second interim analysis was to be conducted after 394 events had been reached.

Overall survival

HR 0.791

95% CI: 0.648–0.964   ·   P = 0.0100

Analysis population: intent-to-treat   ·   Hypothesis: superiority

EndpointComparisonEstimate95% CIP-valueMethod
Overall Survival Carfilzomib + DEX vs Bortezomib + DEX HR 0.791 0.648–0.964 0.0100 Stratified log-rank

The OS time frame was from randomization until the data cut-off date of 03 January 2017. The ClinicalTrials.gov record reports median follow-up time for OS as 36.9 and 37.5 months for each treatment group respectively..

Clinical Biostats interpretation

The OS hazard ratio of 0.791 corresponds to an estimated instantaneous death rate in the carfilzomib group that was approximately 79.1% of the rate in the bortezomib group under the reported model. In relative terms, 1 − 0.791 = 0.209, corresponding to approximately a 20.9% lower estimated hazard of death.

The HR does not mean that 20.9% of patients lived longer, nor does it give an absolute difference in survival probability. It summarizes the relative event-rate comparison over the analyzed time-to-event data.

The 95% confidence interval of 0.648–0.964 gives the statistical uncertainty around the HR estimate. It does not indicate that individual patients' treatment effects fall within this interval.

The reported p-value of 0.0100 is a hypothesis-testing quantity, not an effect-size measure. The analysis notes specify a one-sided significance level of 0.0123, determined using an O'Brien-Fleming-type alpha-spending function based on the actual number of events. The reported two-sided 95% confidence interval and the specified one-sided interim-testing framework should therefore be kept conceptually distinct.

As with PFS, the Cox-model hazard ratio requires attention to the proportional-hazards assumption and to censoring. The registry does not provide enough information in the ClinicalTrials.gov record to reconstruct the full survival curve or assess that assumption directly.

9. Results: Overall Response

Overall response was a secondary binary endpoint. Disease response was assessed every 28 days until the end of treatment or the data cut-off date of 10 November 2014. The analysis used the intent-to-treat population and a stratified Cochran-Mantel-Haenszel method.

Overall response

OR 2.032

95% CI: 1.519–2.718   ·   P < 0.0001

Odds ratio calculated as carfilzomib relative to bortezomib

EndpointComparisonEstimate95% CIP-valueMethod
Overall Response Carfilzomib + DEX vs Bortezomib + DEX OR 2.032 1.519–2.718 < 0.0001 Stratified Cochran-Mantel-Haenszel
Clinical Biostats interpretation

An odds ratio of 2.032 means that the estimated odds of overall response were about 2.032 times as high in the carfilzomib group as in the bortezomib group, using the stratified Cochran-Mantel-Haenszel analysis.

Odds are not the same as probabilities. An OR of 2.032 therefore cannot be read as "twice the response rate" or as a 103.2-percentage-point increase in response. Converting an odds ratio into an absolute probability difference requires the underlying response probability in a comparison group.

The 95% confidence interval of 1.519–2.718 describes uncertainty around the estimated odds ratio. The p-value of < 0.0001 addresses the hypothesis test and does not quantify how clinically large the response difference is.

The analysis was stratified by prior proteasome inhibitor treatment, lines of prior treatment, ISS stage, and choice of route of bortezomib administration. Stratification can improve the alignment between the analysis and the randomized comparison when these factors are part of the trial's allocation structure.

10. Results: ≥ Grade 2 Peripheral Neuropathy

The registry reports a secondary safety endpoint measuring the percentage of participants with ≥ Grade 2 peripheral neuropathy. The analysis used the safety population, defined as participants who received at least one dose of study treatment.

Peripheral neuropathy

OR 0.137

95% CI: 0.089–0.210   ·   P < 0.0001

Odds ratio estimated for carfilzomib relative to bortezomib

EndpointComparisonEstimate95% CIP-valuePopulation
≥ Grade 2 Peripheral Neuropathy Carfilzomib + DEX vs Bortezomib + DEX OR 0.137 0.089–0.210 < 0.0001 Safety population
Clinical Biostats interpretation

The odds ratio of 0.137 indicates that the estimated odds of the reported ≥ Grade 2 peripheral-neuropathy endpoint were lower in the carfilzomib group than in the bortezomib group under the reported analysis. The corresponding estimate is about 13.7% of the comparator odds.

This is an odds ratio, not a risk ratio. It should not be interpreted as saying that exactly 13.7% as many participants experienced the event. The absolute event probabilities are required to translate the odds ratio into an absolute risk difference.

The 95% confidence interval of 0.089–0.210 describes uncertainty around the odds-ratio estimate. The p-value of < 0.0001 provides evidence against the null hypothesis under the reported testing method but does not itself quantify the magnitude of the safety difference.

Because this analysis uses the safety population rather than the ITT population, its interpretation is exposure-based rather than purely assignment-based. The registry states that the analysis used the unconditional Cochran-Mantel-Haenszel method.

11. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.

Treatment armAffectedAt risk
Bortezomib + DEX182456
Carfilzomib + DEX272463
Serious adverse events: affected / at risk
Bortezomib + DEX
182 / 456
Carfilzomib + DEX
272 / 463

The reported serious-adverse-event counts are descriptive. The ClinicalTrials.gov record does not provide a formal statistical analysis, confidence interval, or p-value for this endpoint, so this page does not construct one.

Safety interpretation: Serious adverse events and the specific ≥ Grade 2 peripheral-neuropathy endpoint are different outcomes. The presence of a formal odds-ratio analysis for peripheral neuropathy should not be interpreted as a formal statistical comparison of the serious-adverse-event counts unless such an analysis is separately reported.

12. Interim Analysis, Alpha Spending, and Multiplicity

The registry-reported analysis notes explicitly support an interim-analysis framework for the time-to-event endpoints. The PFS interim analysis was to use a group sequential monitoring plan with an O'Brien-Fleming-type efficacy stopping boundary constructed using the Lan-DeMets alpha-spending function.

Lan-DeMets alpha spending

Alpha spending provides a way to control the overall Type I error while allowing interim looks at accumulating data. The registry-reported PFS analysis notes specify a 1-sided Type I error rate ≤ 0.025.

O'Brien-Fleming-type boundary

The stated boundary framework makes early efficacy decisions more stringent than later decisions, preserving the overall error-control objective while allowing the trial to stop early for compelling evidence.

OS interim analysis

The second interim OS analysis was to be conducted after 394 events had been reached.

One-sided OS threshold

The OS analysis notes specify a one-sided significance level of 0.0123 based on the O'Brien-Fleming-type alpha-spending function and the actual number of events.

These details matter because a trial that examines accumulating efficacy data cannot simply treat every interim p-value as though it came from a single, preplanned final analysis. Alpha spending adjusts the testing framework so that repeated opportunities to declare efficacy remain compatible with the prespecified Type I error objective.

Conceptual Type I error allocation
Total planned Type I error → allocated across information times by the alpha-spending function

The ClinicalTrials.gov record identifies the spending framework and its stated error-rate targets. They do not provide the complete statistical-analysis-plan implementation, so this page does not reconstruct additional boundaries or information fractions.

13. Stratification and the Cochran-Mantel-Haenszel Framework

Stratification appears repeatedly in the ENDEAVOR statistical analysis. The registry identifies the log-rank analyses as stratified by the randomization stratification factors. The overall-response analysis additionally specifies the factors used for its Cochran-Mantel-Haenszel calculation.

AnalysisStratification information reported in the registry
Progression-free Survival Log-rank test stratified by the randomization stratification factors
Overall Survival Log-rank test stratified by the randomization stratification factors
Overall Response Prior proteasome inhibitor treatment; lines of prior treatment; ISS stage; choice of route of bortezomib administration
Peripheral Neuropathy Unconditional Cochran-Mantel-Haenszel method

For a stratified analysis, the treatment comparison is formed across strata rather than by ignoring them. This can be particularly useful when baseline factors are related to prognosis or were incorporated into the randomization scheme. The important point is that the analysis should be aligned with the trial's prespecified stratification structure.

14. Statistical Methods Explained

Why use a stratified log-rank test for PFS?

PFS is a time-to-event endpoint, so the analysis must account for both when events occur and for participants who are censored. The stratified log-rank test compares the event experience between treatment groups while incorporating the randomization strata. This is different from simply comparing the proportion of participants who had progressed by a fixed date.

What does an HR of 0.533 mean?

An HR of 0.533 means that the estimated instantaneous rate of progression or death in the carfilzomib group was approximately 53.3% of the estimated rate in the bortezomib group under the fitted model. The corresponding relative reduction in estimated hazard is approximately 46.7%. It is not a statement that 46.7% of patients avoided an event.

Why is the confidence interval important?

The point estimate alone does not communicate how precisely the treatment effect has been estimated. For PFS, the 95% confidence interval is 0.437–0.651. It describes statistical uncertainty around the hazard-ratio estimate under the analysis framework. It does not describe the range of possible outcomes for an individual patient.

Why does the p-value not measure effect size?

A p-value evaluates compatibility with a specified null hypothesis under the statistical model and testing framework. It depends on both the magnitude of the observed effect and the amount of information in the data. An effect estimate such as an HR or OR is therefore necessary to understand the size and direction of the observed treatment difference.

Why use a Cochran-Mantel-Haenszel test for overall response?

Overall response is binary rather than time-to-event. The stratified Cochran-Mantel-Haenszel approach provides a way to compare treatment groups across prespecified strata and produces an odds-ratio estimate that accounts for those strata rather than collapsing the data into one unstratified comparison.

Why does the interim analysis affect interpretation?

When efficacy is evaluated before the planned accumulation of all information, the nominal significance threshold must account for the repeated opportunities to stop early. ENDEAVOR's registry-reported analysis notes describe an O'Brien-Fleming-type boundary implemented through Lan-DeMets alpha spending. The reported p-values therefore need to be understood within that sequential-testing framework rather than as isolated numbers.

15. Understanding the Primary PFS Result

Relative effect

The primary PFS HR of 0.533 indicates a lower estimated hazard of progression or death for carfilzomib relative to bortezomib in the randomized comparison. The estimate is a relative measure of event rates over time.

Precision

The 95% CI of 0.437–0.651 places the statistical uncertainty around the point estimate. The interval is substantially narrower than the entire positive range of possible hazard ratios, reflecting a more informative estimate than the point estimate alone.

Statistical evidence

The reported p-value of < 0.0001 indicates strong evidence against the null hypothesis under the reported stratified log-rank testing framework. It should not be converted into a probability that the treatment works or a probability that the observed HR is correct.

Absolute effects

The ClinicalTrials.gov record does not report median PFS values, survival probabilities at specific time points, or event counts by treatment arm for the primary endpoint. Consequently, this page does not derive an absolute PFS difference from the HR.

16. Hazard Ratios and Proportional Hazards

The PFS and OS analyses both use hazard ratios. A hazard is an instantaneous event rate conditional on remaining event-free up to a given time. A hazard ratio compares those rates between treatment groups.

Hazard-ratio interpretation
HR < 1 → lower estimated instantaneous event rate in the carfilzomib group

For ENDEAVOR, the reported PFS and OS hazard ratios are below 1. The interpretation is relative and model-based; it is not equivalent to an absolute risk reduction or a ratio of median survival times.

The Cox proportional-hazards framework is most straightforward to interpret when the relative hazard between groups is reasonably stable over time. The ClinicalTrials.gov record does not provide enough information to assess the proportional-hazards assumption directly. Therefore, the HR should be treated as the reported summary measure rather than as proof that the hazard ratio was constant at every time point.

17. Odds Ratios and Binary Endpoints

ENDEAVOR provides two useful examples of odds-ratio interpretation: overall response and ≥ Grade 2 peripheral neuropathy.

EndpointORDirection for carfilzomib95% CI
Overall Response2.032Higher estimated odds1.519–2.718
≥ Grade 2 Peripheral Neuropathy0.137Lower estimated odds0.089–0.210

The two estimates move in opposite directions because they represent different binary outcomes. For response, an OR above 1 means higher estimated odds of response in the carfilzomib group. For peripheral neuropathy, an OR below 1 means lower estimated odds of the reported safety event.

Odds are not probabilities: An odds ratio of 2.032 does not mean that the response percentage was 2.032 times as large. Similarly, an odds ratio of 0.137 does not mean that the event percentage was exactly 13.7% of the comparator percentage. Absolute percentages are needed for those interpretations.

18. What the P-Values Do — and Do Not — Mean

Four formal statistical analyses are posted on ClinicalTrials.gov for ENDEAVOR: one primary PFS analysis and three secondary analyses. Their reported p-values are:

EndpointP-valueEffect measure
Progression-free Survival< 0.0001HR 0.533
Overall Survival0.0100HR 0.791
Overall Response< 0.0001OR 2.032
≥ Grade 2 Peripheral Neuropathy< 0.0001OR 0.137

A p-value is tied to a null hypothesis and a testing procedure. It does not tell the reader the probability that the null hypothesis is true, the probability that the treatment effect is clinically important, or the probability that a patient will benefit.

In ENDEAVOR, the interpretation is further shaped by the interim-monitoring framework. The registry-reported analysis notes specifically identify one-sided Type I error control for the interim PFS and OS analyses. This is one reason that a p-value should always be read alongside the design and analysis plan rather than treated as a standalone verdict.

19. Primary vs Secondary Endpoints

EndpointRolePrimary statistical questionReported effect
Progression-free Survival Primary Superiority of carfilzomib + DEX vs bortezomib + DEX HR 0.533 (95% CI 0.437–0.651), P < 0.0001
Overall Survival Secondary Superiority of carfilzomib + DEX vs bortezomib + DEX HR 0.791 (95% CI 0.648–0.964), P = 0.0100
Overall Response Secondary Comparison of response odds OR 2.032 (95% CI 1.519–2.718), P < 0.0001
≥ Grade 2 Peripheral Neuropathy Secondary Comparison of event odds OR 0.137 (95% CI 0.089–0.210), P < 0.0001

This hierarchy matters. The primary endpoint was PFS, while OS, overall response, and peripheral neuropathy were secondary endpoints. The statistical role of a result is therefore determined not only by its numerical p-value but also by where it sits in the prespecified trial design and how interim monitoring and multiplicity were handled.

20. Limitations and Interpretation Issues

21. Why This Trial Matters Statistically

ENDEAVOR is a useful teaching case because a single randomized phase 3 trial connects several core statistical methods: time-to-event analysis, stratified testing, Cox modeling, binary-outcome analysis, odds ratios, ITT analysis, safety-population analysis, and interim alpha spending.

ConceptHow it appears in ENDEAVOR
RandomizationRandomized, two-arm, parallel-group phase 3 design
Intention-to-treat analysisPrimary PFS and reported OS and response analyses used the ITT population
Time-to-event endpointPFS was the registered primary endpoint; OS was secondary
Kaplan-Meier estimationRegistered methodology for estimating median PFS
Stratified log-rank testUsed for PFS and OS comparisons
Hazard ratioReported for PFS and OS
Cox modelUsed to estimate the reported hazard ratios
Cochran-Mantel-Haenszel testUsed for overall response and peripheral-neuropathy analyses
Odds ratioReported for overall response and ≥ Grade 2 peripheral neuropathy
Interim analysisPFS and OS analyses included sequential monitoring
Alpha spendingLan-DeMets alpha spending with an O'Brien-Fleming-type efficacy boundary
One-sided testingSpecified in the interim PFS and OS analysis framework
Safety populationUsed for the peripheral-neuropathy analysis

22. Statistical Story of ENDEAVOR

The statistical structure can be viewed as a sequence of increasingly specific questions.

01
RandomizeCreate comparable groups
02
Measure PFSTrack progression or death
03
Estimate HRQuantify relative event rate
04
Test responseCompare binary outcomes
05
Assess safetyExamine treatment exposure

The primary PFS analysis asks a time-to-event question. The hazard ratio provides the relative effect, the confidence interval describes precision, and the stratified log-rank test provides the formal comparison. Secondary analyses then address overall survival and binary outcomes such as response and peripheral neuropathy using methods appropriate to their endpoint types.

This is a useful general principle in clinical-trial statistics: the endpoint determines the data structure, and the data structure determines the appropriate statistical framework. Time-to-event outcomes require methods that account for follow-up and censoring; binary outcomes can be summarized with odds ratios and analyzed with methods such as Cochran-Mantel-Haenszel procedures.

23. Interpreting the Trial as a Statistical Analyst

First question: What was randomized?

The treatment assignment was randomized between carfilzomib plus dexamethasone and bortezomib plus dexamethasone in a two-arm parallel design.

Second question: What was primary?

Progression-free survival was the single registered primary endpoint, making it the central confirmatory efficacy outcome represented in the ClinicalTrials.gov record.

Third question: How was PFS analyzed?

With a stratified log-rank test and a hazard-ratio estimate from a Cox proportional-hazards model, using the ITT population.

Fourth question: How should secondary results be separated?

OS, overall response, and peripheral neuropathy answer different questions and use different endpoint-specific methods and analysis populations.

This separation is important because it prevents a collection of different estimates from being treated as though they were interchangeable measures of one outcome. A hazard ratio for PFS, a hazard ratio for OS, an odds ratio for response, and an odds ratio for peripheral neuropathy each have a distinct statistical meaning.

24. Overall Statistical Interpretation

Primary efficacy

The reported primary PFS analysis estimated a hazard ratio of 0.533 with a 95% confidence interval of 0.437–0.651 and a p-value of < 0.0001, using a stratified log-rank test in the ITT population.

Secondary efficacy

The reported OS analysis estimated an HR of 0.791 with a 95% confidence interval of 0.648–0.964 and a p-value of 0.0100. Overall response had an OR of 2.032 with a 95% confidence interval of 1.519–2.718 and a p-value of < 0.0001.

Safety endpoint

The reported ≥ Grade 2 peripheral-neuropathy analysis produced an OR of 0.137 with a 95% confidence interval of 0.089–0.210 and a p-value of < 0.0001. Serious adverse events were reported descriptively as 182/456 in the bortezomib arm and 272/463 in the carfilzomib arm.

What the statistics cannot establish alone

These summary statistics do not provide individual-level treatment effects, absolute PFS or OS differences, or a complete description of the survival curves. They also do not eliminate the need to consider endpoint hierarchy, interim monitoring, censoring, model assumptions, analysis populations, and the distinction between efficacy and safety endpoints.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through the Clinical Biostats statistical pathway

Use the methods in this trial as a starting point for deeper study of survival analysis, categorical-data methods, clinical-trial design, and statistical inference.

28. Record Summary

ENDEAVOR provides a compact example of how several clinical-trial statistical methods work together. The trial used randomized parallel-group allocation and evaluated a registered primary time-to-event endpoint, progression-free survival, with a stratified log-rank test and Cox-model hazard ratio. The reported primary PFS estimate was HR 0.533 with a 95% CI of 0.437–0.651 and P < 0.0001.

The secondary analyses extended the statistical story to overall survival, overall response, and peripheral neuropathy. OS produced an HR of 0.791 (95% CI 0.648–0.964; P = 0.0100), overall response produced an OR of 2.032 (95% CI 1.519–2.718; P < 0.0001), and ≥ Grade 2 peripheral neuropathy produced an OR of 0.137 (95% CI 0.089–0.210; P < 0.0001).

The most important statistical lesson is that these estimates should not be collapsed into a single number. Hazard ratios summarize relative time-to-event rates, odds ratios summarize relative odds for binary outcomes, confidence intervals describe statistical precision, and p-values address hypothesis tests. Interim monitoring and alpha spending further determine how evidence should be interpreted within the trial's sequential testing framework.

Clinical Biostats methodology: A trial-results page should not merely repeat a result. The goal is to reconstruct the statistical structure of the trial, explain why the reported methods fit the endpoint types, and clearly distinguish reported estimates from educational interpretation.