← Clinical Trials
Multiple Myeloma Phase 3 Completed NCT03110562

BOSTON: Complete Statistical Analysis of Selinexor, Bortezomib, and Dexamethasone in Multiple Myeloma

An independent statistical review of the randomized phase 3 BOSTON trial comparing selinexor plus bortezomib and dexamethasone with bortezomib and dexamethasone, with emphasis on progression-free survival, response, overall survival, and safety.

Trial start: May 24, 2017  ·  Primary completion: February 18, 2020  ·  Enrollment: 402
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information contained in the ClinicalTrials.gov record.

1. Trial at a Glance

BOSTON was a completed, randomized phase 3 parallel-group trial in multiple myeloma. The registry reports 402 enrolled participants, four arms, no masking, and a primary endpoint of progression-free survival assessed by an Independent Review Committee (IRC).

402
Enrollment
Participants
4
Arms
Parallel design
0.7020
Primary PFS HR
95% CI 0.5279–0.9335
0.0075
Primary PFS P-value
Two-sided
FeatureBOSTON
Trial nameBOSTON
ClinicalTrials.gov identifierNCT03110562
Therapeutic areaHematology
ConditionMultiple Myeloma
PhasePhase 3
StatusCompleted
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment402
Arms4
Lead sponsorKaryopharm Therapeutics Inc
Sponsor typeIndustry
Results postedYes
Outcome measures posted12
Statistical analyses posted5

2. Clinical Question

The primary statistical question registered for BOSTON was whether selinexor plus bortezomib and dexamethasone (SVd) produced a different progression-free survival experience from bortezomib plus dexamethasone (Vd) in participants with multiple myeloma. The registered hypothesis type was superiority.

Population

Participants in the phase 3 BOSTON trial with the condition of multiple myeloma.

Intervention

SVd: selinexor + bortezomib + dexamethasone.

Comparator

Vd: bortezomib + dexamethasone.

Primary question

Does SVd improve IRC-assessed progression-free survival relative to Vd?

3. Trial Design

01
Enroll402 participants
02
RandomizeRandomized allocation
03
TreatFour-arm parallel design
04
AssessPFS, response, OS, safety
05
AnalyzeStratified statistical methods
Allocation
Randomized.
Design model
Parallel.
Masking
None.
Primary purpose
Treatment.
COMPARISON ARM

SVd

  • Selinexor
  • Bortezomib
  • Dexamethasone
  • Primary comparison: SVd versus Vd for IRC-assessed PFS
COMPARATOR ARM

Vd

  • Bortezomib
  • Dexamethasone
  • Comparator for the primary PFS analysis
Four-arm design: the registry reports four arms overall. The primary statistical analysis in the ClinicalTrials.gov record is specifically the SVd-versus-Vd comparison. The registry analysis text states that data was not planned to be reported for the SVdX and SdX arms for the listed efficacy analyses.

4. Trial Timing and Analysis Framework

Trial featureRegistry value
Start date2017-05-24
Primary completion date2020-02-18
Enrollment402
Primary endpoint typeTime-to-event
Registered primary endpoints1
Posted outcome measures12
Posted statistical analyses5
Primary endpoint analyses posted1

The distinction between a registered endpoint and a posted statistical analysis is important. BOSTON has one registered primary endpoint, and the ClinicalTrials.gov record contains one formal statistical analysis for that endpoint. The remaining four posted analyses are secondary analyses.

5. Endpoints

EndpointRegistry definition / time frameType
SVd/Vd Arm: Progression-free Survival (PFS) as Assessed by Independent Review Committee (IRC) From date of randomization until IRC-confirmed documented PD or death, censored date, whichever occurred first. Time-to-event
SVd/Vd Arm: Overall Response Rate (ORR) as Assessed by IRC From date of randomization until disease progression or initiating a new MM treatment (up to 33 months). Binary
SVd/Vd Arm: Percentage of Participants With Response Rate of Very Good Partial Response (VGPR) or Better Based on IRC Assessment From date of randomization until confirmed PD or initiating a new MM treatment (up to 33 months). Binary
SVd/Vd Arm: Number of Participants With at Least One Grade Greater Than or Equal to [≥] 2 Peripheral Neuropathy Events From first dose of study treatment to 30 days after the last dose of study treatment inclusive, or the day before the study treatment start date for subsequent therapy. Binary
SVd/Vd Arm: Overall Survival (OS) From date of randomization to the date of death or censored date, whichever occurred first (up to 45 months). Time-to-event

The primary endpoint is therefore a classic time-to-event outcome. Unlike a simple binary endpoint, PFS uses both event information and follow-up time, while allowing participants who have not experienced progression or death by the relevant observation point to contribute censored information.

6. Primary Statistical Analysis

Progression-Free Survival

The registry-defined primary endpoint was IRC-assessed PFS comparing the SVd and Vd arms. The analysis used a stratified log-rank test, with the effect reported as a hazard ratio from a stratified Cox proportional-hazards model.

Primary PFS hazard ratio

0.7020

95% CI: 0.5279–0.9335   ·   P = 0.0075

Two-sided 95% confidence interval · Superiority hypothesis

Primary analysis featureReported result
EndpointIRC-assessed progression-free survival
ComparisonSVd versus Vd
Analysis populationITT population
MethodStratified log-rank test
Effect measureHazard ratio
Estimate0.7020
95% CI0.5279–0.9335
P-value0.0075
HypothesisSuperiority
Clinical Biostats interpretation

The reported HR of 0.7020 means that, under the stratified Cox model, the estimated instantaneous rate of the PFS event was approximately 70.20% of the corresponding rate in the Vd group. Expressed as a simple relative interpretation, this corresponds to an estimated 29.80% lower hazard for the SVd group relative to Vd.

The HR does not mean that 29.80% of participants avoided progression, that PFS was 29.80% longer for every individual, or that the absolute probability of progression or death was reduced by exactly 29.80%. Hazard is a time-dependent rate concept, not a simple proportion of patients.

The 95% CI of 0.5279–0.9335 describes uncertainty around the estimated relative hazard under the statistical model and sampling framework. It is not an interval containing the effects experienced by individual patients.

The P-value of 0.0075 addresses the compatibility of the observed data with the relevant null hypothesis under the prespecified testing framework. It does not measure the size of the treatment effect, the probability that the treatment works, or the probability that the null hypothesis is true.

The analysis is also model-dependent. The registry states that the estimate was based on a stratified Cox proportional-hazards model, so interpretation of a single HR should take the proportional-hazards framework into account. The analysis was performed in the ITT population, preserving randomized treatment assignment for the efficacy comparison.

7. Secondary Efficacy Results

Overall Response Rate

ORR was analyzed as a binary endpoint using the Cochran-Mantel-Haenszel test. The registry reports an odds ratio of 1.9626 for SVd versus Vd.

Odds ratio for overall response

1.9626

95% CI: 1.2641–3.0471   ·   P = 0.0012

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

An OR of 1.9626 means that the estimated odds of achieving the registry-defined response were approximately 1.96 times the corresponding odds in the Vd group, after the specified stratified analysis. An odds ratio is not the same as a risk ratio or a percentage-point difference in response rate.

The 95% CI of 1.2641–3.0471 describes uncertainty around the odds-ratio estimate. The interval does not state that individual participants' responses varied between those two values.

The P-value of 0.0012 is evidence against the relevant null hypothesis under the stated statistical framework, but it is not a measure of how large or clinically important the response difference is.

Very Good Partial Response or Better

The registry also reports the percentage of participants with a response rate of VGPR or better based on IRC assessment. This binary endpoint was analyzed with a stratified Cochran-Mantel-Haenszel test.

Odds ratio for VGPR or better

1.6594

95% CI: 1.0993–2.5049   ·   P = 0.0082

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The reported OR of 1.6594 indicates higher estimated odds of achieving VGPR or better in the SVd group relative to Vd within the specified stratified analysis. It should not be read as a 65.94% increase in the percentage of responders.

The 95% CI of 1.0993–2.5049 quantifies uncertainty around the estimated odds ratio. Because the measure is an odds ratio, its numerical interpretation depends on the underlying response probabilities and should not be substituted for absolute response rates when those are available.

The P-value of 0.0082 is a hypothesis-testing quantity. It does not quantify effect magnitude or establish the probability that the treatment is beneficial.

8. Overall Survival

Overall survival was a secondary time-to-event endpoint. It was analyzed using a stratified log-rank test, with the effect reported as a hazard ratio from a stratified Cox proportional-hazards model.

Overall survival hazard ratio

0.8764

95% CI: 0.6313–1.2168   ·   P = 0.2152

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The HR of 0.8764 represents an estimated instantaneous death rate in the SVd group relative to the Vd group under the stratified Cox model. Expressed mechanically, the estimate is below 1, corresponding to an estimated hazard approximately 12.36% lower than Vd.

The confidence interval of 0.6313–1.2168 is relatively wide compared with the point estimate and includes 1. This means the data are compatible with a range of relative hazard values under the model, including values below and above the null value of 1.

The P-value of 0.2152 does not measure the size of the observed HR. It indicates that the reported comparison did not provide a small P-value under the stated superiority-testing framework. It should not be converted into a probability that one treatment is effective or ineffective.

No median OS is reported in the ClinicalTrials.gov record used for this page, so the HR, confidence interval, and P-value are the appropriate reported effect measures here rather than an inferred survival-time difference.

9. Safety Results

The registry provides serious adverse-event counts by arm. These values are reported as affected participants over participants at risk.

ArmSerious adverse eventsReported format
SVd Arm: Selinexor + Bortezomib + Dexame109/195Affected / at risk
Vd Arm: Bortezomib + Dexamethasone79/204Affected / at risk
SVdX Arm: Selinexor + Bortezomib + Dexam28/66Affected / at risk
SdX Arm: Selinexor + Dexamethasone7/14Affected / at risk

These are descriptive safety counts rather than the formal efficacy comparison. The ClinicalTrials.gov record does not provide a confidence interval or P-value for the serious-adverse-event counts by arm, so this page does not create one from the reported numerators and denominators.

Safety interpretation: the four-arm safety table should not be collapsed into a single SVd-versus-Vd percentage without preserving the reported arm definitions and denominators. Safety populations and efficacy populations can differ, and the registry separately describes the safety population for the peripheral-neuropathy analysis as participants who received at least one dose of study treatment.

10. Peripheral Neuropathy Analysis

A secondary safety endpoint assessed the number of participants with at least one Grade ≥2 peripheral neuropathy event. The registry specifies a safety population consisting of participants who had received at least one dose of study treatment and states that data was not planned to be reported for the SVdX and SdX arms for this analysis.

Odds ratio for Grade ≥2 peripheral neuropathy events

0.5042

95% CI: 0.3216–0.7906   ·   P = 0.0013

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The OR of 0.5042 indicates that the estimated odds of experiencing at least one Grade ≥2 peripheral neuropathy event were approximately half as large in the SVd group as in the Vd group under the specified stratified analysis.

An odds ratio below 1 does not mean that exactly 50% fewer participants experienced neuropathy. The corresponding absolute event probabilities would be needed to describe an absolute risk difference.

The 95% CI of 0.3216–0.7906 represents uncertainty around the odds-ratio estimate. The P-value of 0.0013 addresses the hypothesis test rather than the magnitude or clinical importance of the observed association.

11. Statistical Methodology

Intention-to-treat analysis

The registry defines the ITT population as consisting of all participants who were randomized to the study treatment, regardless of whether or not they received the study treatment. The primary PFS, ORR, VGPR-or-better, and OS analyses in the ClinicalTrials.gov record use this population.

ITT analysis is particularly important in a randomized trial because treatment assignment is established before subsequent adherence, discontinuation, or other post-randomization events occur. Analyzing participants according to their randomized assignment helps preserve the treatment-group comparability created by randomization.

Stratified log-rank test

The primary PFS and secondary OS analyses used a stratified log-rank test. A log-rank test compares time-to-event experience between groups over follow-up rather than comparing only a single time point.

Conceptual interpretation
H0: survival experience is equivalent across the randomized comparison, subject to the stratified analysis framework

The test uses the observed timing of events while accounting for censoring. Stratification allows comparisons to be organized within the prespecified strata rather than treating all observations as though they came from one homogeneous stratum.

Stratified Cox proportional-hazards model

The registry states that the PFS and OS hazard ratios were based on a stratified Cox proportional-hazards model with Efron's Method of handling ties. The primary PFS analysis was stratified for prior PI therapies, number of prior anti-MM regimens, and R-ISS Stage at screening. The OS analysis used the same three stratification factors.

Hazard ratio
HR = estimated hazard in SVd ÷ estimated hazard in Vd

A value below 1 indicates a lower estimated instantaneous event rate for SVd under the fitted model; a value above 1 indicates a higher estimated instantaneous event rate.

Cochran-Mantel-Haenszel test

The ORR, VGPR-or-better, and peripheral-neuropathy analyses used the Cochran-Mantel-Haenszel method. For a binary outcome, this approach can estimate a treatment association while accounting for specified stratification factors.

For ORR and VGPR-or-better, the registry specifies stratification by region, prior PI therapies, number of prior anti-MM regimens, and R-ISS stage at screening. For peripheral neuropathy, the specified factors were prior PI therapies, number of prior anti-MM regimens, and R-ISS stage at study entry.

Odds ratio

Conceptual form
OR = odds of response or event in SVd ÷ odds in Vd

An odds ratio of 1 represents equal odds. Values above 1 indicate higher odds in the SVd group, while values below 1 indicate lower odds. Odds are not the same quantity as probabilities or risks.

Censoring in time-to-event analysis

PFS and OS include a censoring concept in their registry definitions. When a participant has not experienced the relevant event by the point at which follow-up ends for that participant, the available follow-up contributes information up to the censoring point rather than treating the participant as though the event had occurred.

12. Statistical Methods Explained

Why was a stratified log-rank test used for PFS?

PFS is a time-to-event endpoint, so the analysis needs to incorporate both event timing and censoring. A log-rank test is designed for comparing survival-type distributions, while stratification allows the comparison to account for specified clinical factors. In BOSTON, the primary PFS analysis used prior PI therapies, number of prior anti-MM regimens, and R-ISS Stage at screening as stratification factors in the associated Cox analysis.

What does an HR of 0.7020 mean?

It means that the fitted model estimates the instantaneous PFS event rate in the SVd group at 0.7020 times that of the Vd group. The simple complement, 1 − 0.7020, gives 0.2980, or 29.80%, as a relative hazard reduction interpretation. That is a model-based relative statement, not an absolute reduction in the proportion of participants progressing or dying.

Why is the confidence interval important?

A point estimate alone does not communicate how precisely the treatment effect has been estimated. The PFS HR of 0.7020 has a 95% CI of 0.5279–0.9335, showing the uncertainty around the estimated relative hazard. The confidence interval is not a prediction interval for individual patients and does not mean that every patient's treatment effect lies inside it.

Why was the Cochran-Mantel-Haenszel method used for ORR?

ORR is a binary outcome: a participant either meets the registry-defined response criterion or does not. The Cochran-Mantel-Haenszel framework provides a way to compare binary outcomes while accounting for the specified stratification factors. BOSTON reports an odds ratio rather than a hazard ratio for this endpoint because response is not itself the primary time-to-event quantity.

Why is the OR of 1.9626 not the same as a 96.26% higher response rate?

An odds ratio compares odds, where odds are defined as probability divided by one minus probability. Because odds and probabilities are different scales, an OR cannot be interpreted as a percentage-point increase or as a direct risk ratio. Absolute response percentages would be needed to describe the difference in response probability directly.

Why does the OS HR of 0.8764 require a different interpretation from the PFS HR?

Both are hazard ratios from time-to-event analyses, but they correspond to different endpoints. PFS counts the registry-defined progression-or-death event, whereas OS concerns death. Therefore, the HRs should not be combined into a single measure or treated as though they represent the same clinical event.

What does the P-value tell us?

A P-value measures the compatibility of the observed data with a specified null hypothesis under the statistical model and testing framework. It does not measure effect size, clinical importance, or the probability that the null hypothesis is true. For BOSTON, the primary PFS P-value is 0.0075, while the secondary OS P-value is 0.2152; those numbers answer hypothesis-testing questions and should not be treated as effect-size measures.

13. Stratification in BOSTON

Stratification is central to the statistical analysis because the registry specifies different stratification factors for the efficacy and safety analyses.

Endpoint / analysisStratification factors
Primary PFS Prior PI therapies; number of prior anti-MM regimens; R-ISS Stage at screening
ORR Region; prior PI therapies; number of prior anti-MM regimens; R-ISS stage at screening
VGPR or better Region; prior PI therapies; number of prior anti-MM regimens; R-ISS stage at screening
Grade ≥2 peripheral neuropathy Prior PI therapies; number of prior anti-MM regimens; R-ISS stage at study entry
OS Prior PI therapies; number of prior anti-MM regimens; R-ISS Stage at screening

The purpose of stratification is not to make the treatment effect automatically larger or smaller. Rather, it incorporates specified sources of clinical heterogeneity into the comparison. The precise interpretation remains tied to the analysis model and its assumptions.

Randomization

Randomization establishes the treatment assignment before subsequent outcomes are observed, supporting the causal comparison between randomized groups.

Stratification

Stratified analysis accounts for prespecified factors when estimating or testing the treatment comparison.

ITT population

The efficacy analyses follow randomized assignment rather than requiring participants to complete treatment.

Safety population

The peripheral-neuropathy analysis uses participants who received at least one dose of study treatment.

14. Results in Context

The five posted statistical analyses form a coherent set of endpoint-specific comparisons rather than five interchangeable measures of the same effect. The primary analysis addresses PFS; the secondary analyses address response, depth of response, peripheral neuropathy, and OS.

EndpointMeasureEstimate95% CIP-value
Primary PFSHazard ratio0.70200.5279–0.93350.0075
ORROdds ratio1.96261.2641–3.04710.0012
VGPR or betterOdds ratio1.65941.0993–2.50490.0082
Grade ≥2 peripheral neuropathyOdds ratio0.50420.3216–0.79060.0013
OSHazard ratio0.87640.6313–1.21680.2152

The table illustrates why endpoint-specific interpretation matters. PFS and OS use hazard ratios because they are time-to-event outcomes. ORR, VGPR-or-better, and peripheral neuropathy use odds ratios because the analyses posted on ClinicalTrials.gov treat them as binary outcomes. These effect measures have different mathematical meanings and should not be compared numerically as though they were on one common scale.

Important statistical distinction: the presence of several statistically reported secondary endpoints does not make them equivalent to the primary endpoint. The registry identifies PFS as the single registered primary endpoint and the remaining analyses as secondary.

15. Primary Endpoint: What the PFS Result Does and Does Not Establish

What the result establishes statistically

The registry analysis reports an SVd-versus-Vd PFS HR of 0.7020, with a two-sided 95% CI of 0.5279–0.9335 and a P-value of 0.0075, using a stratified log-rank test and a stratified Cox proportional-hazards model.

What the result does not establish

The PFS HR does not provide a median PFS, an absolute percentage-point difference, an individual patient's expected benefit, or a guarantee that hazards remain proportional throughout follow-up. None of those quantities should be inferred from the HR alone.

Why absolute measures matter

A relative hazard measure is useful for comparing event rates, but absolute time-to-event quantities can answer different questions. Because the ClinicalTrials.gov record does not report median PFS or specific PFS probabilities, this page does not manufacture those values.

16. Secondary Endpoint Interpretation

The response analyses complement the primary PFS analysis by addressing binary measures of tumor response. The ORR analysis has an odds ratio of 1.9626, while the VGPR-or-better analysis has an odds ratio of 1.6594. These estimates describe different binary outcomes and should not be interpreted as interchangeable measures of efficacy.

The OS analysis provides a different time-to-event perspective. Its HR of 0.8764 has a 95% CI of 0.6313–1.2168 and a P-value of 0.2152. The ClinicalTrials.gov record does not include a median OS, so the appropriate interpretation is restricted to the reported hazard ratio, confidence interval, and hypothesis-test result.

The peripheral-neuropathy analysis is also distinct because it concerns a safety event rather than a tumor-control endpoint. Its OR of 0.5042 describes the relative odds of the specified Grade ≥2 event under the safety analysis framework. A lower odds ratio for a safety endpoint is not conceptually equivalent to a higher odds ratio for response; the direction of desirability depends on what the endpoint represents.

17. Limitations and Interpretation Issues

18. Why This Trial Matters Statistically

BOSTON is a useful teaching example because it combines randomized clinical-trial design with several major branches of applied biostatistics. The same trial contains a time-to-event primary endpoint, binary secondary endpoints, stratified analyses, ITT efficacy analysis, a separate safety population, hazard ratios, odds ratios, confidence intervals, and hypothesis testing.

Statistical conceptHow it appears in BOSTON
RandomizationThe registry classifies allocation as randomized.
Parallel-group designThe design model is parallel with four arms.
ITT analysisThe efficacy analysis population includes all randomized participants regardless of treatment receipt.
Time-to-event analysisPFS is the registered primary endpoint and OS is a secondary endpoint.
Stratified log-rank testUsed for the primary PFS and secondary OS comparisons.
Cox proportional-hazards modelUsed for the reported PFS and OS hazard ratios.
Cochran-Mantel-Haenszel testUsed for ORR, VGPR-or-better, and peripheral-neuropathy analyses.
Hazard ratioPrimary PFS estimate 0.7020; OS estimate 0.8764.
Odds ratioReported for ORR, VGPR-or-better, and Grade ≥2 peripheral neuropathy.
Confidence intervals95% two-sided intervals accompany all five posted statistical analyses.
Stratified analysisClinical and disease-history factors are incorporated into the reported analyses.
Safety populationThe peripheral-neuropathy analysis uses participants receiving at least one dose.

19. A Practical Reading Sequence for the BOSTON Results

A statistically disciplined reading of the registry results can proceed in layers.

Step 1 · Identify the primary endpoint

PFS is the one registered primary endpoint, assessed by an Independent Review Committee.

Step 2 · Identify the analysis population

The primary efficacy comparison uses the ITT population of randomized participants.

Step 3 · Read the effect measure

The primary effect measure is an HR of 0.7020, not an odds ratio or an absolute risk difference.

Step 4 · Read the uncertainty

The 95% CI is 0.5279–0.9335 and should be interpreted alongside the point estimate.

Step 5 · Read the P-value separately

The P-value of 0.0075 addresses the hypothesis test and does not measure effect magnitude.

Step 6 · Examine secondary endpoints

ORR, VGPR-or-better, peripheral neuropathy, and OS answer different clinical and statistical questions.

This sequence prevents a common analytical error: beginning with a P-value and working backward toward a clinical conclusion. The more informative approach is to establish the endpoint, population, estimand or effect measure, uncertainty, and analysis framework first.

20. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore the time-to-event, categorical-data, stratification, and hypothesis-testing methods used in BOSTON.

21. Record Summary

BOSTON provides a compact example of how a randomized phase 3 trial can require several distinct statistical perspectives. Its registered primary endpoint was IRC-assessed PFS, analyzed with a stratified log-rank test and summarized using a stratified Cox hazard ratio of 0.7020 with a 95% CI of 0.5279–0.9335 and P = 0.0075. Secondary analyses used Cochran-Mantel-Haenszel methods for ORR, VGPR-or-better, and Grade ≥2 peripheral neuropathy, while OS used a stratified log-rank test and Cox model with an HR of 0.8764.

The most useful statistical reading therefore separates time-to-event effects from binary-outcome effects, distinguishes efficacy ITT analysis from the safety population, and interprets every estimate together with its confidence interval and P-value. Just as importantly, the absence of a reported absolute outcome measure should not be filled by calculation or assumption.

Clinical Biostats methodology: A trial-results page should reconstruct the statistical story of the trial without silently adding estimates from other sources. Reported evidence, statistical interpretation, and general methodological explanation should remain clearly distinguishable.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators