← Clinical Trials
Multiple Myeloma Phase 3 Randomized NCT00644228

SWOG S0777: Complete Statistical Analysis of Lenalidomide and Dexamethasone With or Without Bortezomib in Multiple Myeloma

An independent statistical review of the randomized phase 3 SWOG S0777 trial evaluating lenalidomide and dexamethasone with or without bortezomib in patients with previously untreated multiple myeloma.

SWOG S0777  ·  Phase 3  ·  NCT00644228  ·  Enrollment 525
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the statistical analyses and safety information contained in the ClinicalTrials.gov record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SWOG S0777 is a randomized, parallel, open-label phase 3 treatment trial in patients with previously untreated multiple myeloma. The registered primary endpoint is progression-free survival, with overall survival and response rates reported as secondary outcomes.

525
Enrollment
Registered trial enrollment
2
Arms
Randomized parallel design
0.712
PFS HR
96% CI 0.560–0.906
0.0018
PFS P-value
Primary endpoint
FeatureSWOG S0777
Trial nameSWOG S0777
PhasePhase 3
Therapeutic areaHematology
PopulationPatients with DS Stage I, DS Stage II, or DS Stage III multiple myeloma; the trial title specifies previously untreated multiple myeloma.
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment525
Primary endpointProgression-free Survival
Primary endpoint typeTime-to-event
Results postedYes
Lead sponsorNational Cancer Institute (NCI)
Sponsor typeNIH
ClinicalTrials.govNCT00644228

2. Clinical Question

The trial asks whether adding bortezomib to lenalidomide and dexamethasone improves outcomes compared with lenalidomide and dexamethasone alone in patients with previously untreated multiple myeloma.

Population

Patients with DS Stage I, DS Stage II, or DS Stage III multiple myeloma. The trial title identifies the population as previously untreated.

Intervention

Bortezomib, dexamethasone, and lenalidomide.

Comparator

Dexamethasone and lenalidomide without bortezomib.

Primary question

Does the bortezomib-containing regimen improve progression-free survival relative to lenalidomide and dexamethasone alone?

3. Trial Design

01
Randomize525 enrolled
02
Arm ILenalidomide + dexamethasone
03
Arm IILenalidomide + dexamethasone + bortezomib
04
FollowTime-to-event and response outcomes
05
ComparePrespecified statistical analyses
Allocation
Randomized allocation to two treatment arms.
Design model
Parallel-group design.
Masking
None; the registered trial is open-label.
Primary purpose
Treatment.
ARM I · DEXAMETHASONE + LENALIDOMIDE

Control regimen

  • Dexamethasone
  • Lenalidomide
ARM II · DEXAMETHASONE + LENALIDOMIDE + BORTEZOMIB

Bortezomib-containing regimen

  • Dexamethasone
  • Lenalidomide
  • Bortezomib

The intervention list also includes Laboratory Biomarker Analysis as an other intervention. The ClinicalTrials.gov record does not provide a statistical analysis of a biomarker-defined subgroup, so no biomarker efficacy result is presented here.

4. Trial Timing and Status

2008-07-28

Trial start

The registered trial start date was July 28, 2008.

2015-11-05

Primary completion

The registered primary completion date was November 5, 2015.

Current registry status

Active, not recruiting

The ClinicalTrials.gov record lists the study status as ACTIVE_NOT_RECRUITING.

5. Endpoints

EndpointRoleRegistry time frameEndpoint type
Progression-free Survival Primary From date of registration to date of first documentation of progression or symptomatic deterioration, or death due to any cause, assessed up to 6 years Time-to-event
Overall Survival Secondary Up to 6 years Time-to-event
Response Rates Secondary Up to 6 years Binary

The registry defines the primary endpoint as unstratified median progression-free survival in months. The registry-reported endpoint text ends with the phrase “death due to any cause, assessed up to 6 years”; no additional words are reported in the ClinicalTrials.gov record, so the definition is not expanded here.

6. Statistical Methodology

Log-rank test

The primary progression-free survival comparison was performed using a log-rank test. This is a standard method for comparing time-to-event distributions between randomized groups while accounting for the timing of events and right-censored observations.

The secondary overall survival analysis was also performed using a log-rank test. The response-rate analysis used a Cochran-Mantel-Haenszel test, which is designed for comparing categorical outcomes while accounting for stratification variables.

Methods reported in the registry
Primary PFS → Log-rank test
Overall survival → Log-rank test
Response rates → Cochran-Mantel-Haenszel test

The normalized statistical methods reported for the trial are the log-rank test and Cochran-Mantel-Haenszel test. The primary and overall-survival analyses report hazard ratios as their effect measure.

Hazard ratio

The primary progression-free survival analysis reports a hazard ratio of 0.712 comparing Arm I with Arm II as specified in the analysis record. The registry's secondary overall-survival analysis reports a hazard ratio of 0.709 and explicitly describes the comparison as bortezomib/lenalidomide/dexamethasone against lenalidomide/dexamethasone.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate for the numerator treatment group

A hazard ratio is a relative time-to-event measure. It is not a probability, an absolute risk difference, a median survival difference, or the percentage of patients who benefit.

Analysis population

For the primary progression-free survival analysis, all eligible and analyzable patients are included. An analyzable patient is defined in the registry as one who provided valid consent and who did not withdraw consent prior to initiating treatment.

The same analysis-population definition is reported for overall survival. For response rates, the registry states that all eligible, analyzable patients are included.

Stratified analysis

The primary analysis notes identify stratified analysis as an associated statistical concept. The power calculation is specifically described as being based on a one-sided stratified log-rank test at level 0.025. The ClinicalTrials.gov record does not identify the individual stratification factors, so no specific factors are listed.

7. Primary Result: Progression-free Survival

Progression-free survival was the single registered primary endpoint. The statistical analysis compares Arm I, dexamethasone and lenalidomide, with Arm II, dexamethasone, lenalidomide, and bortezomib.

Hazard ratio for progression-free survival

0.712

96% two-sided CI: 0.560–0.906   ·   P = 0.0018

Effect measure: Hazard Ratio (HR)  ·  Hypothesis type: Superiority

Primary endpointEstimateConfidence intervalP-valueMethod
Progression-free Survival HR 0.712 96% two-sided CI 0.560–0.906 0.0018 Log-rank test
Clinical Biostats interpretation

An HR of 0.712 means that, under the time-to-event comparison represented by the analysis, the estimated instantaneous rate of the progression-free survival event was about 28% lower for the numerator group than for the comparator group. This is a relative hazard interpretation, not a statement that 28% of patients avoided progression or that every individual patient experienced the same reduction.

The 96% two-sided confidence interval of 0.560–0.906 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. Because the entire interval is below 1, the interval is consistent with a lower estimated event hazard for the numerator group across the range represented by the interval.

The P-value of 0.0018 addresses the evidence against the null hypothesis in the specified statistical test. It does not measure the magnitude of the treatment effect. Effect size is conveyed by the hazard ratio and its confidence interval.

The registry analysis notes are especially important here: the design calculation was based on a one-sided stratified log-rank test at level 0.025 with two interim analyses. Because interim analyses can affect type I error control, the nominal P-value should be interpreted in the context of that prespecified sequential design rather than as though the data had been inspected only once at the end.

The ClinicalTrials.gov record does not provide median PFS, Kaplan-Meier estimates, event counts, or the individual stratification factors. Those quantities therefore cannot be reconstructed from the reported hazard ratio alone.

What the primary estimate tells us

Relative effect

HR 0.712 indicates a lower estimated instantaneous event rate in the numerator group under the reported time-to-event analysis.

Precision

The 96% CI of 0.560–0.906 gives the reported range of uncertainty around the hazard-ratio estimate.

Statistical evidence

P = 0.0018 indicates strong evidence against the null hypothesis under the specified test framework.

What is not reported

The ClinicalTrials.gov record does not report median PFS, event counts, or absolute PFS probabilities.

8. Secondary Results

Overall Survival

Hazard ratio for overall survival

0.709

95% two-sided CI: 0.524–0.959   ·   P = 0.0250

Time frame: Up to 6 years  ·  Method: Log-rank test

Secondary endpointEstimateConfidence intervalP-valueMethod
Overall Survival HR 0.709 95% two-sided CI 0.524–0.959 0.0250 Log-rank test
Response Rates Not reported Not reported 0.20 Cochran-Mantel-Haenszel test

The overall-survival analysis uses the same eligible-and-analyzable population definition described for the primary analysis. The registry explicitly states that the hazard ratio compares bortezomib/lenalidomide/dexamethasone against lenalidomide/dexamethasone.

Response Rates

Response Rates were analyzed as a binary endpoint using the Cochran-Mantel-Haenszel test. The posted analysis reports P = 0.20, with a superiority hypothesis. No response-rate effect estimate or confidence interval is provided in the registry-reported statistical analysis.

Interpretation: Because the registry supplies a P-value but no response-rate estimates or confidence interval, the statistical result can be described as a Cochran-Mantel-Haenszel comparison with P = 0.20. The P-value alone does not quantify the size of any difference in response rates.

9. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.

ArmSerious adverse eventsAffected / at risk
Arm I: Dexamethasone and Lenalidomide Serious adverse events 105/222
Arm II: Dexamethasone, Lenalidomide, Bortezomib Serious adverse events 124/237
Safety interpretation: These figures are reported as affected participants over participants at risk. The ClinicalTrials.gov record does not provide a formal between-arm statistical test, confidence interval, or relative effect measure for serious adverse events, so none is calculated here.

10. Statistical Design: Power and Interim Analysis

The registry's primary analysis notes provide unusually useful information about the planned statistical design. With four years of patient accrual and two and a half years of follow-up, the design calculation states that 220 patients per arm yields 87% power to detect an increase of PFS of 50%, from a median of 3 years to 4.5 years, corresponding to a hazard ratio of 1.5.

Planned power

The registry analysis note specifies 87% power for the stated design scenario.

Planned effect

The design scenario describes an increase in PFS from 3 years to 4.5 years, corresponding to HR 1.5.

Interim analyses

The power calculation explicitly incorporates two interim analyses.

Alpha level

The calculation uses a one-sided stratified log-rank test at level 0.025.

Why interim analysis changes the statistical problem

If a trial is examined repeatedly while data accumulate, repeatedly applying an ordinary final-analysis significance threshold can increase the probability of rejecting a true null hypothesis by chance. A prespecified group-sequential design addresses this by defining how evidence is evaluated at interim and final information times.

For SWOG S0777, the registry-reported analysis note explicitly incorporates two interim analyses into the power calculation and specifies a one-sided stratified log-rank test at level 0.025. This means that the PFS result should be understood as part of a sequential testing framework rather than as an isolated calculation performed after a single final look.

Design quantities reported in the registry
Accrual: 4 years  ·  Follow-up: 2.5 years
Patients per arm: 220  ·  Power: 87%
PFS scenario: 3 years → 4.5 years
Corresponding HR: 1.5
One-sided stratified log-rank level: 0.025
Interim analyses: 2

11. Multiplicity and Hypothesis Testing

The registry-reported statistical analysis identifies superiority as the hypothesis type for the primary and secondary analyses. The primary PFS design uses a one-sided test at level 0.025 and incorporates two interim analyses.

FeatureReported information
Primary hypothesisSuperiority
Primary statistical testOne-sided stratified log-rank test in the design calculation
Nominal design level0.025
Interim analyses2
Primary effect measureHazard ratio
Secondary OS hypothesisSuperiority
Response-rate hypothesisSuperiority

There are two distinct multiplicity issues worth separating. First, the interim analyses require sequential type I error control because the data are examined more than once. Second, multiple clinical endpoints are reported. The ClinicalTrials.gov record does not specify a hierarchical testing procedure or an alpha-allocation rule across the primary and secondary endpoints, so no such procedure is attributed to the trial here.

12. Stratification and the Cochran-Mantel-Haenszel Framework

Stratification appears explicitly in the trial's primary analysis notes and in the normalized statistical-method fields. The primary power calculation is based on a stratified log-rank test, while the response-rate analysis uses a Cochran-Mantel-Haenszel test.

Why stratification can matter
Stratified analysis → compare treatment groups while accounting for prespecified strata

When prognostic or design-related factors are represented by strata, a stratified analysis can combine information across those strata while respecting the structure of the randomized comparison. The ClinicalTrials.gov record identifies stratified analysis but do not identify the individual strata.

The Cochran-Mantel-Haenszel approach serves a related purpose for a binary endpoint: it can produce a treatment comparison while accounting for categorical stratification. In SWOG S0777, it was used for response rates. The registry does not provide the response-rate effect estimate or confidence interval, so the analysis is reported through its method and P-value rather than an inferred effect size.

13. Statistical Methods Explained

Why was a log-rank test used for progression-free survival?

Progression-free survival is a time-to-event endpoint. Patients can experience the event at different times, while others may be censored because their event status is not observed through the relevant follow-up. The log-rank test is designed to compare survival-type event-time distributions while incorporating the timing of events and censoring.

What does an HR of 0.712 mean?

An HR of 0.712 is a relative comparison of the event hazard between the groups under the reported time-to-event analysis. Numerically, 0.712 corresponds to a 28.8% lower estimated hazard for the numerator group relative to the comparator because \(1 - 0.712 = 0.288\). This is a derived interpretation of the reported HR; it is not a probability of avoiding progression.

Why does the confidence interval matter?

The point estimate 0.712 is only one estimate of the treatment effect. The 96% two-sided confidence interval, 0.560–0.906, describes the uncertainty around that estimate under the specified statistical framework. A confidence interval should not be interpreted as the range of effects experienced by individual patients.

Why doesn't the P-value measure effect size?

The P-value measures the evidence against a specified null hypothesis under the statistical model and test procedure. It depends on both the observed data and the amount of information available. The hazard ratio describes relative effect size, while its confidence interval communicates precision. These are different statistical quantities.

Why does the one-sided design matter?

The power calculation specifies a one-sided stratified log-rank test at level 0.025. A one-sided test places the rejection region in a prespecified direction. That direction must be defined before examining the results; it is not a justification for ignoring an effect in the opposite direction after the data are seen.

Why do the two interim analyses matter?

Two interim analyses mean that the accumulating trial data were considered at more than one point in time. Without appropriate sequential error control, repeated opportunities to declare significance can alter the false-positive rate. The registry-reported analysis note incorporates those interim analyses into the design calculation.

14. Reading the Primary Hazard Ratio Carefully

Relative effect

The reported primary HR of 0.712 corresponds to a 28.8% lower estimated hazard for the numerator group relative to the comparator. This describes the relative event rate represented by the time-to-event analysis; it does not describe an absolute reduction in the probability of progression.

Confidence interval

The 96% two-sided CI of 0.560–0.906 indicates uncertainty around the HR estimate. The interval remains below 1, which is consistent with a lower estimated event hazard for the numerator group across the reported interval.

P-value

The P-value of 0.0018 indicates evidence against the null hypothesis under the specified log-rank testing framework. It should not be converted into a statement such as “there is a 0.18% chance that the null hypothesis is true.” A P-value is not the probability that the treatment effect is real.

What the HR cannot provide

The ClinicalTrials.gov record does not allow the HR alone to determine median PFS, absolute PFS at a particular time point, the proportion of patients progressing, or the number of patients who benefited. Those quantities require additional event-time information.

15. Primary Analysis Population

The registry specifies that the primary PFS analysis includes all eligible and analyzable patients. An analyzable patient is defined as someone who provided valid consent and did not withdraw consent prior to initiating treatment.

EndpointAnalysis population
Progression-free Survival All eligible and analyzable patients; valid consent and no withdrawal of consent prior to initiating treatment.
Overall Survival All eligible and analyzable patients; valid consent and no withdrawal of consent prior to initiating treatment.
Response Rates All eligible, analyzable patients.

This definition is important because a randomized clinical trial's treatment comparison is meaningful only when the analysis population is clearly specified. The ClinicalTrials.gov record does not provide a separate per-protocol population, an as-treated population, or a missing-data/imputation strategy, so none is described here.

16. Results Summary

EndpointRoleEffect / resultCIP-valueMethod
Progression-free Survival Primary HR 0.712 96% two-sided CI 0.560–0.906 0.0018 Log-rank
Overall Survival Secondary HR 0.709 95% two-sided CI 0.524–0.959 0.0250 Log-rank
Response Rates Secondary Effect estimate not reported Not reported 0.20 Cochran-Mantel-Haenszel

This table should be read as a statistical summary rather than a substitute for the endpoint definitions. PFS and OS are time-to-event outcomes, while response rates are binary. Consequently, the hazard ratios for PFS and OS cannot be directly compared with the response-rate P-value.

17. Important Limitations and Interpretation Issues

18. Why This Trial Matters Statistically

SWOG S0777 is a useful teaching example because the registry record contains several core elements of randomized time-to-event analysis in one trial: randomized treatment allocation, a time-to-event primary endpoint, a log-rank comparison, a hazard-ratio effect measure, confidence intervals, a superiority hypothesis, stratified analysis, interim analyses, and a binary secondary endpoint analyzed with the Cochran-Mantel-Haenszel test.

ConceptHow it appears in SWOG S0777
RandomizationThe trial uses randomized allocation.
Parallel designTwo treatment arms are compared in a parallel-group design.
Time-to-event endpointProgression-free survival is the registered primary endpoint.
Log-rank testUsed for the primary PFS analysis and secondary OS analysis.
Hazard ratioUsed as the reported effect measure for PFS and OS.
Confidence intervalReported around both PFS and OS hazard ratios.
Stratified analysisExplicitly incorporated into the PFS design and analysis framework.
One-sided testingThe primary power calculation uses a one-sided stratified log-rank test at level 0.025.
Interim analysisThe design calculation incorporates two interim analyses.
PowerThe registry design note reports 87% power for its specified PFS scenario.
Cochran-Mantel-Haenszel testUsed for the binary response-rate endpoint.
Safety comparisonSerious adverse events are reported as affected participants over participants at risk by arm.

19. Statistical Methods Explained in Context

Time-to-event analysis

PFS and OS are not ordinary continuous outcomes because the event may occur at different times and some observations may be censored.

Log-rank testing

The log-rank test compares the event-time distributions between randomized groups using information accumulated over follow-up.

Hazard ratio

The HR provides a relative measure of the event hazard. Values below 1 indicate a lower estimated hazard for the numerator group.

Confidence interval

The interval communicates uncertainty around the estimated treatment effect and should be considered alongside the point estimate.

Stratified analysis

Stratification incorporates the trial's specified strata into the statistical comparison rather than treating the entire sample as an undifferentiated group.

Binary endpoint analysis

Response rates are analyzed differently from PFS and OS because response is represented as a binary outcome rather than an event time.

20. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary PFS analysis reports HR 0.712 with a 96% two-sided CI of 0.560–0.906 and P = 0.0018 under a log-rank testing framework. The secondary OS analysis reports HR 0.709 with a 95% two-sided CI of 0.524–0.959 and P = 0.0250.

Clinical interpretation

The statistical results describe relative differences in time-to-event outcomes between the randomized treatment groups. The ClinicalTrials.gov record does not provide enough absolute outcome information to characterize median survival or time-specific survival probabilities.

The distinction matters because statistical significance and clinical magnitude are related but different concepts. A P-value can indicate evidence against a null hypothesis, but it does not establish whether an effect is large or small in absolute clinical terms. Conversely, a hazard ratio requires its confidence interval and the endpoint definition to be interpreted correctly.

21. A Practical Reading of the SWOG S0777 Results

Step 1 · Identify the endpoint

The primary endpoint is progression-free survival, a time-to-event outcome measured from registration according to the registered definition.

Step 2 · Identify the effect measure

The primary analysis reports a hazard ratio of 0.712. Because the HR is below 1, the estimated event hazard is lower for the numerator treatment group under the reported comparison.

Step 3 · Examine uncertainty

The 96% two-sided confidence interval is 0.560–0.906. This interval provides the reported uncertainty around the HR rather than a range of individual patient outcomes.

Step 4 · Examine the test

The P-value is 0.0018. This provides evidence against the null hypothesis under the reported log-rank framework, but it does not quantify the size of the treatment effect.

Step 5 · Examine the design

The design calculation used a one-sided stratified log-rank test at level 0.025 and incorporated two interim analyses. Those design features are part of the correct interpretation of the statistical evidence.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Statistical Calculators

24. Sources

The four PubMed records above are listed because they are identified as linked publications in the ClinicalTrials.gov record. This page does not add numerical results from those publications beyond the statistical analyses and safety data in the ClinicalTrials.gov record.

25. Record Summary

SWOG S0777 provides a compact example of a randomized phase 3 time-to-event analysis. The registered primary endpoint is progression-free survival, analyzed with a log-rank test and expressed through a hazard ratio. The reported primary estimate is HR 0.712, with a 96% two-sided confidence interval of 0.560–0.906 and P = 0.0018. Overall survival was analyzed with a log-rank test and reported as HR 0.709, with a 95% two-sided confidence interval of 0.524–0.959 and P = 0.0250. Response rates were analyzed using the Cochran-Mantel-Haenszel test with P = 0.20.

The design information is equally important to the numerical results. The primary power calculation was based on four years of accrual, two and a half years of follow-up, 220 patients per arm, 87% power, an assumed PFS increase from 3 years to 4.5 years corresponding to HR 1.5, a one-sided stratified log-rank test at level 0.025, and two interim analyses. These features show why a clinical-trial result should be read as a combination of endpoint definition, analysis population, effect measure, uncertainty, hypothesis test, and prespecified design.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical evidence from the educational interpretation of that evidence. For SWOG S0777, that means preserving the registry's endpoint definitions and reported estimates while avoiding unsupported median outcomes, subgroup results, or additional numerical calculations that are not contained in the ClinicalTrials.gov record.