This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information contained in the ClinicalTrials.gov record.
1. Trial at a Glance
BOSTON was a completed, randomized phase 3 parallel-group trial in multiple myeloma. The registry reports 402 enrolled participants, four arms, no masking, and a primary endpoint of progression-free survival assessed by an Independent Review Committee (IRC).
| Feature | BOSTON |
|---|---|
| Trial name | BOSTON |
| ClinicalTrials.gov identifier | NCT03110562 |
| Therapeutic area | Hematology |
| Condition | Multiple Myeloma |
| Phase | Phase 3 |
| Status | Completed |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 402 |
| Arms | 4 |
| Lead sponsor | Karyopharm Therapeutics Inc |
| Sponsor type | Industry |
| Results posted | Yes |
| Outcome measures posted | 12 |
| Statistical analyses posted | 5 |
2. Clinical Question
The primary statistical question registered for BOSTON was whether selinexor plus bortezomib and dexamethasone (SVd) produced a different progression-free survival experience from bortezomib plus dexamethasone (Vd) in participants with multiple myeloma. The registered hypothesis type was superiority.
Population
Participants in the phase 3 BOSTON trial with the condition of multiple myeloma.
Intervention
SVd: selinexor + bortezomib + dexamethasone.
Comparator
Vd: bortezomib + dexamethasone.
Primary question
Does SVd improve IRC-assessed progression-free survival relative to Vd?
3. Trial Design
SVd
- Selinexor
- Bortezomib
- Dexamethasone
- Primary comparison: SVd versus Vd for IRC-assessed PFS
Vd
- Bortezomib
- Dexamethasone
- Comparator for the primary PFS analysis
4. Trial Timing and Analysis Framework
| Trial feature | Registry value |
|---|---|
| Start date | 2017-05-24 |
| Primary completion date | 2020-02-18 |
| Enrollment | 402 |
| Primary endpoint type | Time-to-event |
| Registered primary endpoints | 1 |
| Posted outcome measures | 12 |
| Posted statistical analyses | 5 |
| Primary endpoint analyses posted | 1 |
The distinction between a registered endpoint and a posted statistical analysis is important. BOSTON has one registered primary endpoint, and the ClinicalTrials.gov record contains one formal statistical analysis for that endpoint. The remaining four posted analyses are secondary analyses.
5. Endpoints
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| SVd/Vd Arm: Progression-free Survival (PFS) as Assessed by Independent Review Committee (IRC) | From date of randomization until IRC-confirmed documented PD or death, censored date, whichever occurred first. | Time-to-event |
| SVd/Vd Arm: Overall Response Rate (ORR) as Assessed by IRC | From date of randomization until disease progression or initiating a new MM treatment (up to 33 months). | Binary |
| SVd/Vd Arm: Percentage of Participants With Response Rate of Very Good Partial Response (VGPR) or Better Based on IRC Assessment | From date of randomization until confirmed PD or initiating a new MM treatment (up to 33 months). | Binary |
| SVd/Vd Arm: Number of Participants With at Least One Grade Greater Than or Equal to [≥] 2 Peripheral Neuropathy Events | From first dose of study treatment to 30 days after the last dose of study treatment inclusive, or the day before the study treatment start date for subsequent therapy. | Binary |
| SVd/Vd Arm: Overall Survival (OS) | From date of randomization to the date of death or censored date, whichever occurred first (up to 45 months). | Time-to-event |
The primary endpoint is therefore a classic time-to-event outcome. Unlike a simple binary endpoint, PFS uses both event information and follow-up time, while allowing participants who have not experienced progression or death by the relevant observation point to contribute censored information.
6. Primary Statistical Analysis
Progression-Free Survival
The registry-defined primary endpoint was IRC-assessed PFS comparing the SVd and Vd arms. The analysis used a stratified log-rank test, with the effect reported as a hazard ratio from a stratified Cox proportional-hazards model.
Primary PFS hazard ratio
95% CI: 0.5279–0.9335 · P = 0.0075
Two-sided 95% confidence interval · Superiority hypothesis
| Primary analysis feature | Reported result |
|---|---|
| Endpoint | IRC-assessed progression-free survival |
| Comparison | SVd versus Vd |
| Analysis population | ITT population |
| Method | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.7020 |
| 95% CI | 0.5279–0.9335 |
| P-value | 0.0075 |
| Hypothesis | Superiority |
The reported HR of 0.7020 means that, under the stratified Cox model, the estimated instantaneous rate of the PFS event was approximately 70.20% of the corresponding rate in the Vd group. Expressed as a simple relative interpretation, this corresponds to an estimated 29.80% lower hazard for the SVd group relative to Vd.
The HR does not mean that 29.80% of participants avoided progression, that PFS was 29.80% longer for every individual, or that the absolute probability of progression or death was reduced by exactly 29.80%. Hazard is a time-dependent rate concept, not a simple proportion of patients.
The 95% CI of 0.5279–0.9335 describes uncertainty around the estimated relative hazard under the statistical model and sampling framework. It is not an interval containing the effects experienced by individual patients.
The P-value of 0.0075 addresses the compatibility of the observed data with the relevant null hypothesis under the prespecified testing framework. It does not measure the size of the treatment effect, the probability that the treatment works, or the probability that the null hypothesis is true.
The analysis is also model-dependent. The registry states that the estimate was based on a stratified Cox proportional-hazards model, so interpretation of a single HR should take the proportional-hazards framework into account. The analysis was performed in the ITT population, preserving randomized treatment assignment for the efficacy comparison.
7. Secondary Efficacy Results
Overall Response Rate
ORR was analyzed as a binary endpoint using the Cochran-Mantel-Haenszel test. The registry reports an odds ratio of 1.9626 for SVd versus Vd.
Odds ratio for overall response
95% CI: 1.2641–3.0471 · P = 0.0012
Two-sided 95% confidence interval · Superiority hypothesis
An OR of 1.9626 means that the estimated odds of achieving the registry-defined response were approximately 1.96 times the corresponding odds in the Vd group, after the specified stratified analysis. An odds ratio is not the same as a risk ratio or a percentage-point difference in response rate.
The 95% CI of 1.2641–3.0471 describes uncertainty around the odds-ratio estimate. The interval does not state that individual participants' responses varied between those two values.
The P-value of 0.0012 is evidence against the relevant null hypothesis under the stated statistical framework, but it is not a measure of how large or clinically important the response difference is.
Very Good Partial Response or Better
The registry also reports the percentage of participants with a response rate of VGPR or better based on IRC assessment. This binary endpoint was analyzed with a stratified Cochran-Mantel-Haenszel test.
Odds ratio for VGPR or better
95% CI: 1.0993–2.5049 · P = 0.0082
Two-sided 95% confidence interval · Superiority hypothesis
The reported OR of 1.6594 indicates higher estimated odds of achieving VGPR or better in the SVd group relative to Vd within the specified stratified analysis. It should not be read as a 65.94% increase in the percentage of responders.
The 95% CI of 1.0993–2.5049 quantifies uncertainty around the estimated odds ratio. Because the measure is an odds ratio, its numerical interpretation depends on the underlying response probabilities and should not be substituted for absolute response rates when those are available.
The P-value of 0.0082 is a hypothesis-testing quantity. It does not quantify effect magnitude or establish the probability that the treatment is beneficial.
8. Overall Survival
Overall survival was a secondary time-to-event endpoint. It was analyzed using a stratified log-rank test, with the effect reported as a hazard ratio from a stratified Cox proportional-hazards model.
Overall survival hazard ratio
95% CI: 0.6313–1.2168 · P = 0.2152
Two-sided 95% confidence interval · Superiority hypothesis
The HR of 0.8764 represents an estimated instantaneous death rate in the SVd group relative to the Vd group under the stratified Cox model. Expressed mechanically, the estimate is below 1, corresponding to an estimated hazard approximately 12.36% lower than Vd.
The confidence interval of 0.6313–1.2168 is relatively wide compared with the point estimate and includes 1. This means the data are compatible with a range of relative hazard values under the model, including values below and above the null value of 1.
The P-value of 0.2152 does not measure the size of the observed HR. It indicates that the reported comparison did not provide a small P-value under the stated superiority-testing framework. It should not be converted into a probability that one treatment is effective or ineffective.
No median OS is reported in the ClinicalTrials.gov record used for this page, so the HR, confidence interval, and P-value are the appropriate reported effect measures here rather than an inferred survival-time difference.
9. Safety Results
The registry provides serious adverse-event counts by arm. These values are reported as affected participants over participants at risk.
| Arm | Serious adverse events | Reported format |
|---|---|---|
| SVd Arm: Selinexor + Bortezomib + Dexame | 109/195 | Affected / at risk |
| Vd Arm: Bortezomib + Dexamethasone | 79/204 | Affected / at risk |
| SVdX Arm: Selinexor + Bortezomib + Dexam | 28/66 | Affected / at risk |
| SdX Arm: Selinexor + Dexamethasone | 7/14 | Affected / at risk |
These are descriptive safety counts rather than the formal efficacy comparison. The ClinicalTrials.gov record does not provide a confidence interval or P-value for the serious-adverse-event counts by arm, so this page does not create one from the reported numerators and denominators.
10. Peripheral Neuropathy Analysis
A secondary safety endpoint assessed the number of participants with at least one Grade ≥2 peripheral neuropathy event. The registry specifies a safety population consisting of participants who had received at least one dose of study treatment and states that data was not planned to be reported for the SVdX and SdX arms for this analysis.
Odds ratio for Grade ≥2 peripheral neuropathy events
95% CI: 0.3216–0.7906 · P = 0.0013
Two-sided 95% confidence interval · Superiority hypothesis
The OR of 0.5042 indicates that the estimated odds of experiencing at least one Grade ≥2 peripheral neuropathy event were approximately half as large in the SVd group as in the Vd group under the specified stratified analysis.
An odds ratio below 1 does not mean that exactly 50% fewer participants experienced neuropathy. The corresponding absolute event probabilities would be needed to describe an absolute risk difference.
The 95% CI of 0.3216–0.7906 represents uncertainty around the odds-ratio estimate. The P-value of 0.0013 addresses the hypothesis test rather than the magnitude or clinical importance of the observed association.
11. Statistical Methodology
Intention-to-treat analysis
The registry defines the ITT population as consisting of all participants who were randomized to the study treatment, regardless of whether or not they received the study treatment. The primary PFS, ORR, VGPR-or-better, and OS analyses in the ClinicalTrials.gov record use this population.
ITT analysis is particularly important in a randomized trial because treatment assignment is established before subsequent adherence, discontinuation, or other post-randomization events occur. Analyzing participants according to their randomized assignment helps preserve the treatment-group comparability created by randomization.
Stratified log-rank test
The primary PFS and secondary OS analyses used a stratified log-rank test. A log-rank test compares time-to-event experience between groups over follow-up rather than comparing only a single time point.
The test uses the observed timing of events while accounting for censoring. Stratification allows comparisons to be organized within the prespecified strata rather than treating all observations as though they came from one homogeneous stratum.
Stratified Cox proportional-hazards model
The registry states that the PFS and OS hazard ratios were based on a stratified Cox proportional-hazards model with Efron's Method of handling ties. The primary PFS analysis was stratified for prior PI therapies, number of prior anti-MM regimens, and R-ISS Stage at screening. The OS analysis used the same three stratification factors.
A value below 1 indicates a lower estimated instantaneous event rate for SVd under the fitted model; a value above 1 indicates a higher estimated instantaneous event rate.
Cochran-Mantel-Haenszel test
The ORR, VGPR-or-better, and peripheral-neuropathy analyses used the Cochran-Mantel-Haenszel method. For a binary outcome, this approach can estimate a treatment association while accounting for specified stratification factors.
For ORR and VGPR-or-better, the registry specifies stratification by region, prior PI therapies, number of prior anti-MM regimens, and R-ISS stage at screening. For peripheral neuropathy, the specified factors were prior PI therapies, number of prior anti-MM regimens, and R-ISS stage at study entry.
Odds ratio
An odds ratio of 1 represents equal odds. Values above 1 indicate higher odds in the SVd group, while values below 1 indicate lower odds. Odds are not the same quantity as probabilities or risks.
Censoring in time-to-event analysis
PFS and OS include a censoring concept in their registry definitions. When a participant has not experienced the relevant event by the point at which follow-up ends for that participant, the available follow-up contributes information up to the censoring point rather than treating the participant as though the event had occurred.
12. Statistical Methods Explained
Why was a stratified log-rank test used for PFS?
PFS is a time-to-event endpoint, so the analysis needs to incorporate both event timing and censoring. A log-rank test is designed for comparing survival-type distributions, while stratification allows the comparison to account for specified clinical factors. In BOSTON, the primary PFS analysis used prior PI therapies, number of prior anti-MM regimens, and R-ISS Stage at screening as stratification factors in the associated Cox analysis.
What does an HR of 0.7020 mean?
It means that the fitted model estimates the instantaneous PFS event rate in the SVd group at 0.7020 times that of the Vd group. The simple complement, 1 − 0.7020, gives 0.2980, or 29.80%, as a relative hazard reduction interpretation. That is a model-based relative statement, not an absolute reduction in the proportion of participants progressing or dying.
Why is the confidence interval important?
A point estimate alone does not communicate how precisely the treatment effect has been estimated. The PFS HR of 0.7020 has a 95% CI of 0.5279–0.9335, showing the uncertainty around the estimated relative hazard. The confidence interval is not a prediction interval for individual patients and does not mean that every patient's treatment effect lies inside it.
Why was the Cochran-Mantel-Haenszel method used for ORR?
ORR is a binary outcome: a participant either meets the registry-defined response criterion or does not. The Cochran-Mantel-Haenszel framework provides a way to compare binary outcomes while accounting for the specified stratification factors. BOSTON reports an odds ratio rather than a hazard ratio for this endpoint because response is not itself the primary time-to-event quantity.
Why is the OR of 1.9626 not the same as a 96.26% higher response rate?
An odds ratio compares odds, where odds are defined as probability divided by one minus probability. Because odds and probabilities are different scales, an OR cannot be interpreted as a percentage-point increase or as a direct risk ratio. Absolute response percentages would be needed to describe the difference in response probability directly.
Why does the OS HR of 0.8764 require a different interpretation from the PFS HR?
Both are hazard ratios from time-to-event analyses, but they correspond to different endpoints. PFS counts the registry-defined progression-or-death event, whereas OS concerns death. Therefore, the HRs should not be combined into a single measure or treated as though they represent the same clinical event.
What does the P-value tell us?
A P-value measures the compatibility of the observed data with a specified null hypothesis under the statistical model and testing framework. It does not measure effect size, clinical importance, or the probability that the null hypothesis is true. For BOSTON, the primary PFS P-value is 0.0075, while the secondary OS P-value is 0.2152; those numbers answer hypothesis-testing questions and should not be treated as effect-size measures.
13. Stratification in BOSTON
Stratification is central to the statistical analysis because the registry specifies different stratification factors for the efficacy and safety analyses.
| Endpoint / analysis | Stratification factors |
|---|---|
| Primary PFS | Prior PI therapies; number of prior anti-MM regimens; R-ISS Stage at screening |
| ORR | Region; prior PI therapies; number of prior anti-MM regimens; R-ISS stage at screening |
| VGPR or better | Region; prior PI therapies; number of prior anti-MM regimens; R-ISS stage at screening |
| Grade ≥2 peripheral neuropathy | Prior PI therapies; number of prior anti-MM regimens; R-ISS stage at study entry |
| OS | Prior PI therapies; number of prior anti-MM regimens; R-ISS Stage at screening |
The purpose of stratification is not to make the treatment effect automatically larger or smaller. Rather, it incorporates specified sources of clinical heterogeneity into the comparison. The precise interpretation remains tied to the analysis model and its assumptions.
Randomization
Randomization establishes the treatment assignment before subsequent outcomes are observed, supporting the causal comparison between randomized groups.
Stratification
Stratified analysis accounts for prespecified factors when estimating or testing the treatment comparison.
ITT population
The efficacy analyses follow randomized assignment rather than requiring participants to complete treatment.
Safety population
The peripheral-neuropathy analysis uses participants who received at least one dose of study treatment.
14. Results in Context
The five posted statistical analyses form a coherent set of endpoint-specific comparisons rather than five interchangeable measures of the same effect. The primary analysis addresses PFS; the secondary analyses address response, depth of response, peripheral neuropathy, and OS.
| Endpoint | Measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Primary PFS | Hazard ratio | 0.7020 | 0.5279–0.9335 | 0.0075 |
| ORR | Odds ratio | 1.9626 | 1.2641–3.0471 | 0.0012 |
| VGPR or better | Odds ratio | 1.6594 | 1.0993–2.5049 | 0.0082 |
| Grade ≥2 peripheral neuropathy | Odds ratio | 0.5042 | 0.3216–0.7906 | 0.0013 |
| OS | Hazard ratio | 0.8764 | 0.6313–1.2168 | 0.2152 |
The table illustrates why endpoint-specific interpretation matters. PFS and OS use hazard ratios because they are time-to-event outcomes. ORR, VGPR-or-better, and peripheral neuropathy use odds ratios because the analyses posted on ClinicalTrials.gov treat them as binary outcomes. These effect measures have different mathematical meanings and should not be compared numerically as though they were on one common scale.
15. Primary Endpoint: What the PFS Result Does and Does Not Establish
The registry analysis reports an SVd-versus-Vd PFS HR of 0.7020, with a two-sided 95% CI of 0.5279–0.9335 and a P-value of 0.0075, using a stratified log-rank test and a stratified Cox proportional-hazards model.
The PFS HR does not provide a median PFS, an absolute percentage-point difference, an individual patient's expected benefit, or a guarantee that hazards remain proportional throughout follow-up. None of those quantities should be inferred from the HR alone.
A relative hazard measure is useful for comparing event rates, but absolute time-to-event quantities can answer different questions. Because the ClinicalTrials.gov record does not report median PFS or specific PFS probabilities, this page does not manufacture those values.
16. Secondary Endpoint Interpretation
The response analyses complement the primary PFS analysis by addressing binary measures of tumor response. The ORR analysis has an odds ratio of 1.9626, while the VGPR-or-better analysis has an odds ratio of 1.6594. These estimates describe different binary outcomes and should not be interpreted as interchangeable measures of efficacy.
The OS analysis provides a different time-to-event perspective. Its HR of 0.8764 has a 95% CI of 0.6313–1.2168 and a P-value of 0.2152. The ClinicalTrials.gov record does not include a median OS, so the appropriate interpretation is restricted to the reported hazard ratio, confidence interval, and hypothesis-test result.
The peripheral-neuropathy analysis is also distinct because it concerns a safety event rather than a tumor-control endpoint. Its OR of 0.5042 describes the relative odds of the specified Grade ≥2 event under the safety analysis framework. A lower odds ratio for a safety endpoint is not conceptually equivalent to a higher odds ratio for response; the direction of desirability depends on what the endpoint represents.
17. Limitations and Interpretation Issues
- Endpoint-specific effect measures: PFS and OS are reported with hazard ratios, while ORR, VGPR-or-better, and peripheral neuropathy use odds ratios. These measures should not be directly compared as if they represented the same quantity.
- Hazard-ratio interpretation: the HR is model-based and should be interpreted within the stratified Cox proportional-hazards framework reported by the registry.
- Censoring: PFS and OS incorporate censored observations, so the analysis depends on how follow-up and censoring are handled.
- Analysis populations: efficacy analyses use the ITT population, while the peripheral-neuropathy analysis uses a safety population defined by receipt of at least one dose.
- Four-arm structure: the trial has four arms, but the principal efficacy comparison in the ClinicalTrials.gov record is specifically SVd versus Vd. The registry states that data was not planned to be reported for SVdX and SdX for the listed efficacy analyses.
- Incomplete absolute outcome information: the ClinicalTrials.gov record contains effect estimates and confidence intervals but does not provide median PFS, median OS, or absolute response percentages for the reported secondary analyses.
- Multiple analyses: the registry contains five posted statistical analyses across primary and secondary endpoints. The primary endpoint is distinct from the secondary analyses and should not be interpreted as one undifferentiated family of estimates.
- Safety interpretation: serious adverse-event counts are reported by arm, but the ClinicalTrials.gov record does not provide a formal comparative statistical analysis for those counts.
18. Why This Trial Matters Statistically
BOSTON is a useful teaching example because it combines randomized clinical-trial design with several major branches of applied biostatistics. The same trial contains a time-to-event primary endpoint, binary secondary endpoints, stratified analyses, ITT efficacy analysis, a separate safety population, hazard ratios, odds ratios, confidence intervals, and hypothesis testing.
| Statistical concept | How it appears in BOSTON |
|---|---|
| Randomization | The registry classifies allocation as randomized. |
| Parallel-group design | The design model is parallel with four arms. |
| ITT analysis | The efficacy analysis population includes all randomized participants regardless of treatment receipt. |
| Time-to-event analysis | PFS is the registered primary endpoint and OS is a secondary endpoint. |
| Stratified log-rank test | Used for the primary PFS and secondary OS comparisons. |
| Cox proportional-hazards model | Used for the reported PFS and OS hazard ratios. |
| Cochran-Mantel-Haenszel test | Used for ORR, VGPR-or-better, and peripheral-neuropathy analyses. |
| Hazard ratio | Primary PFS estimate 0.7020; OS estimate 0.8764. |
| Odds ratio | Reported for ORR, VGPR-or-better, and Grade ≥2 peripheral neuropathy. |
| Confidence intervals | 95% two-sided intervals accompany all five posted statistical analyses. |
| Stratified analysis | Clinical and disease-history factors are incorporated into the reported analyses. |
| Safety population | The peripheral-neuropathy analysis uses participants receiving at least one dose. |
19. A Practical Reading Sequence for the BOSTON Results
A statistically disciplined reading of the registry results can proceed in layers.
Step 1 · Identify the primary endpoint
PFS is the one registered primary endpoint, assessed by an Independent Review Committee.
Step 2 · Identify the analysis population
The primary efficacy comparison uses the ITT population of randomized participants.
Step 3 · Read the effect measure
The primary effect measure is an HR of 0.7020, not an odds ratio or an absolute risk difference.
Step 4 · Read the uncertainty
The 95% CI is 0.5279–0.9335 and should be interpreted alongside the point estimate.
Step 5 · Read the P-value separately
The P-value of 0.0075 addresses the hypothesis test and does not measure effect magnitude.
Step 6 · Examine secondary endpoints
ORR, VGPR-or-better, peripheral neuropathy, and OS answer different clinical and statistical questions.
This sequence prevents a common analytical error: beginning with a P-value and working backward toward a clinical conclusion. The more informative approach is to establish the endpoint, population, estimand or effect measure, uncertainty, and analysis framework first.
20. Sources
- ClinicalTrials.gov: NCT03110562 — BOSTON.
- Linked publication: PubMed PMID 37743180.
- Linked publication: PubMed PMID 37393120.
- Linked publication: PubMed PMID 34882831.
- Linked publication: PubMed PMID 34062004.
- Linked publication: PubMed PMID 33849608.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to explore the time-to-event, categorical-data, stratification, and hypothesis-testing methods used in BOSTON.
21. Record Summary
BOSTON provides a compact example of how a randomized phase 3 trial can require several distinct statistical perspectives. Its registered primary endpoint was IRC-assessed PFS, analyzed with a stratified log-rank test and summarized using a stratified Cox hazard ratio of 0.7020 with a 95% CI of 0.5279–0.9335 and P = 0.0075. Secondary analyses used Cochran-Mantel-Haenszel methods for ORR, VGPR-or-better, and Grade ≥2 peripheral neuropathy, while OS used a stratified log-rank test and Cox model with an HR of 0.8764.
The most useful statistical reading therefore separates time-to-event effects from binary-outcome effects, distinguishes efficacy ITT analysis from the safety population, and interprets every estimate together with its confidence interval and P-value. Just as importantly, the absence of a reported absolute outcome measure should not be filled by calculation or assumption.
22. Related Tutorials
Learn more about the methods used in this trial: