← Clinical Trials
Malignant Melanoma Phase 3 Time-to-Event Analysis NCT01006980

BRIM-3: Complete Statistical Analysis of Vemurafenib in Metastatic Melanoma

An independent statistical review of the randomized phase 3 BRIM-3 trial comparing vemurafenib with dacarbazine in previously untreated patients with metastatic melanoma, focusing on overall survival, progression-free survival, hazard ratios, confidence intervals, and log-rank testing.

Trial start: January 2010  ·  Primary completion: December 2010  ·  Enrollment: 675
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

BRIM-3 was a randomized, parallel-group, open-label phase 3 trial evaluating vemurafenib versus dacarbazine in previously untreated patients with metastatic melanoma. The registry reports two primary time-to-event endpoints: overall survival and progression-free survival.

675
Enrolled
Randomized trial
2
Arms
Vemurafenib vs dacarbazine
0.37
OS HR
95% CI 0.26–0.55
0.26
PFS HR
95% CI 0.20–0.33
FeatureBRIM-3
Trial nameBRIM-3
PhasePhase 3
ConditionMalignant melanoma
PopulationPreviously untreated patients with metastatic melanoma
DesignRandomized, parallel-group, open-label
AllocationRandomized
Primary purposeTreatment
Primary endpointsOverall survival and progression-free survival
Primary endpoint typeTime-to-event
Enrollment675
StatusCompleted
ClinicalTrials.govNCT01006980
Lead sponsorHoffmann-La Roche
Sponsor typeIndustry

2. Clinical Question

The central question was whether treatment with vemurafenib improved time-to-event outcomes compared with dacarbazine in previously untreated patients with metastatic melanoma.

Population

Previously untreated patients with metastatic melanoma.

Intervention

Vemurafenib.

Comparator

Dacarbazine.

Primary question

Does vemurafenib improve overall survival and progression-free survival relative to dacarbazine?

3. Trial Design

01
Randomize675 participants
02
VemurafenibActive treatment arm
03
DacarbazineComparator arm
04
AssessOS and PFS
05
AnalyzeLog-rank / HR
ARM A · VEMURAFENIB

Vemurafenib

  • Drug intervention
  • Randomized treatment assignment
  • Compared with dacarbazine for the primary time-to-event endpoints
ARM B · DACARBAZINE

Dacarbazine

  • Drug comparator
  • Randomized treatment assignment
  • Reference group for the primary time-to-event comparisons
Open-label design: the registry identifies masking as none. The randomized comparison therefore was not described as blinded in the ClinicalTrials.gov record.

4. Trial Timing and Follow-Up

January 2010

Trial initiated

The registry specifies that randomization was initiated in January 2010.

December 2010

Primary completion

The registry lists December 2010 as the primary completion date.

December 30, 2010

Clinical cutoff

The primary endpoint time frames for overall survival and progression-free survival extend from randomization to December 30, 2010.

The registry reports a median follow-up time in the vemurafenib group of 3.75 for the overall-survival endpoint. The registry wording does not specify a unit after this value, so this page does not add one.

5. Primary Endpoints

EndpointRegistry definitionTime framePrimary analysis
Overall Survival An overall survival event was defined as death due to any cause. The number of participants with overall survival events is reported. From randomization (initiated January 2010) to December 30, 2010. Median follow-up time in the vemurafenib group was 3.75. Log-rank test; hazard ratio
Progression-free Survival A progression-free survival event was defined as disease progression or death due to any cause. Tumor response (progression) was assessed according to RECIST version 1.1 criteria using CT scans or MRI. From randomization (initiated January 2010) to December 30, 2010. Log-rank test; hazard ratio

Both primary endpoints are time-to-event endpoints. This is important statistically because participants can have different follow-up times, and some may be censored before experiencing the event. The registry identifies the hypothesis type for both primary analyses as superiority.

6. Statistical Methodology

Log-rank test

The registry reports the log-rank test for both primary endpoints. The log-rank framework compares the observed and expected pattern of events between randomized treatment groups over follow-up, making it appropriate for comparing time-to-event distributions.

Conceptual comparison
H0: survival distributions are equivalent across treatment groups

For a superiority analysis, evidence against the null hypothesis is evaluated using the prespecified statistical framework. The reported BRIM-3 analyses provide both a hazard ratio and a P-value, allowing the statistical evidence to be considered alongside the magnitude and precision of the estimated treatment effect.

Hazard ratio

The effect measure reported for both primary endpoints was the hazard ratio (HR). The HR compares the estimated instantaneous event rate between treatment groups within the time-to-event modeling framework.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the vemurafenib group

An HR of 0.37 for overall survival, for example, corresponds to an estimated hazard approximately 37% as large in the vemurafenib group relative to the dacarbazine group. Expressed as a relative reduction in estimated hazard, 1 − 0.37 = 0.63, or 63%. This is a statement about the estimated hazard ratio, not a statement that 63% of participants survived or that every participant experienced a 63% reduction in risk.

Intention-to-treat analysis

The registry defines the ITT population for the overall-survival analysis as all randomized participants, whether or not study treatment was received, with participants analyzed according to the treatment assigned at randomization. This preserves the treatment comparison created by randomization.

Stratified analysis

The statistical-analysis record identifies stratified analysis as an analysis concept for overall survival and for the broader statistical methodology. The ClinicalTrials.gov record does not provide the specific stratification variables in the complete analysis text, so this page does not infer or add them.

PFS analysis population

Cox regression for PFS

For progression-free survival, the registry specifically states that hazard ratios for treatment with vemurafenib compared with dacarbazine were estimated using unstratified Cox regression. This is distinct from the registry's separate identification of stratified analysis as an analysis concept.

7. Overall Survival Results

The registry reports a formal primary-endpoint analysis comparing vemurafenib with dacarbazine using the log-rank test and a hazard ratio. The analysis population was the ITT population as defined in the registry.

Hazard ratio for death

0.37

95% CI: 0.26–0.55   ·   P < 0.0001

Two-sided 95% confidence interval · Superiority hypothesis

Estimated hazard ratio
Vemurafenib
0.37
Reference
1.00
Clinical Biostats interpretation

The reported HR of 0.37 means that the estimated instantaneous rate of death under the fitted time-to-event analysis was 37% of the corresponding rate in the dacarbazine group. Equivalently, the estimated hazard was 63% lower for vemurafenib relative to dacarbazine.

The HR does not mean that 63% of participants avoided death, that 63% of participants benefited, or that an individual patient's probability of death was reduced by exactly 63%.

The 95% CI of 0.26–0.55 describes statistical uncertainty around the estimated hazard ratio. It is an interval for the treatment-effect estimate under the analysis framework, not a range containing the effects experienced by individual participants.

The P < 0.0001 value addresses evidence against the relevant null hypothesis under the specified testing framework. It does not measure the size of the treatment effect. Effect size is conveyed by the HR, while precision is conveyed by the confidence interval.

Because this is a time-to-event analysis, interpretation also depends on censoring and on the assumptions underlying the hazard-based model. A single HR is a relative summary of event rates over follow-up rather than a complete description of the survival experience at every time point.

Prespecified design assumptions

The registry states that the trial had 80% power to detect a hazard ratio of 0.65 for overall survival with an alpha level of 0.045. The design specified an increase in median survival from 8 months for dacarbazine to 12.3 months for vemurafenib. The registry also reports one interim analysis for overall survival at 50% information.

Design target

The planned detectable overall-survival hazard ratio was 0.65, with 80% power and an alpha level of 0.045.

Interim look

One interim analysis for overall survival was planned at 50% information.

8. Progression-Free Survival Results

The second primary endpoint was progression-free survival. A PFS event was defined as disease progression or death due to any cause, with progression assessed using RECIST version 1.1 criteria based on CT scans or MRI.

Hazard ratio for progression or death

0.26

95% CI: 0.20–0.33   ·   P < .0001

Two-sided 95% confidence interval · Superiority hypothesis

Estimated hazard ratio
Vemurafenib
0.26
Reference
1.00
Clinical Biostats interpretation

The reported HR of 0.26 means that the estimated instantaneous rate of progression or death under the fitted analysis was 26% of the corresponding rate in the dacarbazine group. Expressed as a relative reduction in estimated hazard, 1 − 0.26 = 0.74, or 74%.

That interpretation should not be translated into a claim that 74% of patients were protected from progression or death. A hazard ratio is a relative time-to-event measure, not a percentage of patients who benefit.

The 95% CI of 0.20–0.33 quantifies uncertainty around the estimated HR. Its relatively narrow span compared with the point estimate indicates that the reported estimate is accompanied by a defined range of statistical uncertainty under the stated model and sampling framework.

The P < .0001 value measures evidence against the null hypothesis within the specified testing framework; it is not an estimate of effect magnitude or clinical importance.

For PFS specifically, the registry states that treatment hazard ratios were estimated using unstratified Cox regression. Interpretation therefore depends on the Cox model and on the handling of censoring and event times.

Prespecified PFS design

The registry states that the trial had 90% power to detect a hazard ratio of 0.55 for progression-free survival with an alpha level of 0.005. The design description gives an increase in median survival from 2.5 months for dacarbazine to 4.5 months for vemurafenib.

Analysis population: The PFS analysis population consisted of ITT participants randomized by October 27, 2010, at least 9 weeks before the December 30, 2010 clinical cutoff. This prespecified eligibility for the PFS analysis should not be silently treated as identical to the overall-survival ITT population.

9. Primary Results Side by Side

Primary endpointEffect measureEstimate95% CIP-valueHypothesis
Overall Survival Hazard ratio 0.37 0.26–0.55 <0.0001 Superiority
Progression-free Survival Hazard ratio 0.26 0.20–0.33 <.0001 Superiority

Both reported primary analyses produced hazard ratios below 1, with two-sided 95% confidence intervals entirely below 1 and very small reported P-values. Statistically, the two endpoints therefore tell a consistent story within the reported analysis framework: the estimated event hazard was lower in the vemurafenib group for both death and progression or death.

The two endpoints nevertheless represent different clinical events. Overall survival counts death from any cause, whereas progression-free survival counts either disease progression or death. A treatment can affect these endpoints differently because progression is an earlier event and because some participants can experience progression without immediately experiencing death.

10. Safety: Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These figures should be interpreted as safety counts using the affected/at-risk denominators provided in the registry record.

Safety groupAffectedAt riskReported ratio
Vemurafenib165336165/336
Dacarbazine5229352/293
Vemurafenib after crossover448444/84

The crossover category is reported separately in the ClinicalTrials.gov record. It should not be combined with the randomized vemurafenib and dacarbazine groups when describing the primary randomized safety comparison.

Safety interpretation: the ClinicalTrials.gov record reports serious adverse events as affected/at-risk counts. They do not provide a complete adverse-event table or definitions for every safety category, so this page does not add other safety measures.

11. Interim Analysis and Alpha Spending

Interim monitoring is an important component of the BRIM-3 statistical design. The registry states that the trial included one interim analysis for overall survival at 50% information.

Why conduct an interim analysis?

An interim analysis permits evaluation of accumulating trial information before all planned information has been collected. This can be useful in a time-to-event trial where events accumulate over calendar time.

Why alpha matters

Repeatedly examining accumulating efficacy data can alter the probability of a false-positive conclusion. A prespecified alpha framework is therefore part of the design of a group-sequential trial.

The registry reports different alpha levels for the two primary endpoint designs: 0.045 for the overall-survival power calculation and 0.005 for the progression-free-survival power calculation. These values should be understood as elements of the prespecified design rather than as generic P-value thresholds to be applied retrospectively.

Important distinction: an interim analysis does not mean that the final reported P-value should be interpreted without reference to the prespecified monitoring design. In a group-sequential setting, the timing of the interim look and the allocation of type I error are part of the inferential framework.

12. Intention-to-Treat Analysis

The registry explicitly defines the overall-survival ITT population as all randomized participants, whether or not study treatment was received. Participants were analyzed according to the treatment assigned at randomization.

Why ITT matters
Randomized assignment → analyze according to assigned treatment

The main statistical advantage is preservation of the treatment comparison generated by randomization. Treatment discontinuation, deviations from assigned therapy, or subsequent treatment can occur after randomization, but the ITT framework retains participants in their originally assigned groups for the primary efficacy comparison.

For PFS, the registry defines a more specific analysis population consisting of ITT participants randomized by October 27, 2010, at least 9 weeks before the December 30, 2010 cutoff. That distinction matters because an analysis population is part of the definition of the reported estimate.

13. Statistical Methods Explained

Why was a log-rank test used?

Overall survival and progression-free survival are time-to-event outcomes. Participants can be followed for different lengths of time, and some may be censored before experiencing the event. The log-rank test is designed to compare survival distributions while incorporating the timing of events and censoring rather than reducing every participant to a simple binary outcome.

What does an overall-survival HR of 0.37 mean?

Within the reported time-to-event analysis, the estimated instantaneous rate of death in the vemurafenib group was 0.37 times that in the dacarbazine group. This corresponds to a 63% lower estimated hazard. It does not mean that 63% of patients survived or that every patient had the same reduction in individual risk.

What does the 95% CI of 0.26–0.55 tell us?

The confidence interval describes uncertainty around the estimated HR. It gives a statistical interval for the underlying treatment-effect parameter under the specified model and sampling framework. It is not a range of individual patient outcomes and does not mean that future patients' hazards will necessarily fall inside the interval.

Why does the P-value not measure effect size?

A P-value describes the degree of statistical evidence against a null hypothesis under the specified testing framework. It is affected by both the magnitude of an observed effect and the amount of information available. The HR communicates relative effect magnitude, while the confidence interval communicates precision.

Why are OS and PFS analyzed separately?

OS and PFS are different endpoints. An OS event is death from any cause. A PFS event is disease progression or death from any cause, with progression assessed according to RECIST version 1.1 using CT or MRI. Their event definitions, timing, and censoring patterns can therefore differ.

Why does the ITT definition matter?

The ITT approach keeps randomized participants in their assigned treatment groups for efficacy analysis. This maintains the treatment comparison created by randomization and avoids redefining the comparison based on treatment received after randomization.

Why does the interim analysis affect interpretation?

Because the trial included an interim overall-survival analysis at 50% information, the inferential framework has to account for that planned look at the data. The design's alpha level and interim-analysis structure are therefore relevant when interpreting the statistical evidence rather than treating the reported P-value as an isolated number.

14. Understanding the Two Primary Hazard Ratios

FeatureOverall SurvivalProgression-free Survival
EventDeath due to any causeDisease progression or death due to any cause
Effect measureHazard ratioHazard ratio
Estimate0.370.26
95% CI0.26–0.550.20–0.33
P-value<0.0001<.0001
Analysis methodLog-rank testLog-rank test
Additional model informationRegistry identifies stratified analysis as an analysis conceptUnstratified Cox regression used to estimate HR

The HR estimates are not interchangeable. An OS HR summarizes the relative death hazard, while a PFS HR summarizes the relative hazard of either progression or death. The smaller PFS HR therefore should not be described as proof that the treatment effect on death itself was larger than the OS effect.

15. Censoring and Time-to-Event Interpretation

Time-to-event analyses are designed for settings in which not every participant experiences the event during the observation period. A participant who has not experienced the specified event by the time their usable follow-up ends can contribute information up to the censoring time.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

where di is the number of events at time ti and ni is the number at risk immediately before that time.

The registry-reported BRIM-3 registry data do not provide the underlying participant-level event and censoring times. Consequently, this page does not construct a Kaplan-Meier curve or attempt to reconstruct medians from the reported hazard ratios.

Educational note: a valid Kaplan-Meier reconstruction requires appropriate event and censoring information or sufficiently detailed source data. A hazard ratio and confidence interval alone are not enough to recreate the underlying survival curve.

16. Limitations

17. Why This Trial Matters Statistically

BRIM-3 is a useful teaching case because it combines randomized treatment assignment with two clinically distinct time-to-event endpoints and a prespecified interim analysis. The trial also illustrates why a statistical result should be read as a package rather than as a single P-value.

ConceptHow it appears in BRIM-3
RandomizationParticipants were randomized to vemurafenib or dacarbazine.
Parallel designThe registry identifies a parallel-group design with two arms.
Open-label treatmentMasking is listed as none.
ITT analysisOverall survival was analyzed using the registry-defined ITT population.
Time-to-event endpointsOverall survival and progression-free survival were both primary endpoints.
Log-rank testReported as the primary comparison method for both endpoints.
Hazard ratioReported as the effect measure for both primary endpoints.
Confidence intervalTwo-sided 95% CIs were reported for both primary HR estimates.
Interim analysisOne overall-survival interim analysis was planned at 50% information.
Stratified analysisIdentified as an analysis concept in the registry's statistical-analysis record.
Cox regressionUnstratified Cox regression was used to estimate the PFS hazard ratio.
Safety by armSerious adverse events were reported using affected/at-risk counts.

18. What the Hazard Ratio Does — and Does Not — Mean

Overall survival

An HR of 0.37 means that, under the reported time-to-event analysis, the estimated instantaneous rate of death for vemurafenib was 37% of that for dacarbazine. The corresponding relative reduction in estimated hazard is 63%.

It does not mean that 63% of participants were cured, that 63% of participants avoided death, or that every individual patient had exactly a 63% reduction in risk.

Progression-free survival

An HR of 0.26 means that the estimated instantaneous rate of progression or death for vemurafenib was 26% of that for dacarbazine. The corresponding relative reduction in estimated hazard is 74%.

It does not mean that 74% of participants remained progression-free or that the probability of progression or death for every patient was reduced by exactly 74%.

Why confidence intervals matter

The OS 95% CI of 0.26–0.55 and PFS 95% CI of 0.20–0.33 provide information about the statistical precision of the corresponding HR estimates. Neither interval describes the range of individual patient experiences.

Why P-values are not effect sizes

The reported P-values, <0.0001 for OS and <.0001 for PFS, quantify statistical evidence under the respective testing frameworks. They do not tell us how large the treatment effect is. The HR gives the relative effect estimate; the confidence interval gives its statistical uncertainty.

19. Design Features That Affect Interpretation

Superiority framework

Both primary endpoints were analyzed under a superiority hypothesis rather than a non-inferiority framework. A non-inferiority margin therefore does not form part of the reported primary interpretation.

No factorial design

The registry identifies a parallel two-arm design. No factorial design is reported in the ClinicalTrials.gov record.

Interim monitoring

One interim analysis for overall survival was planned at 50% information, so interim monitoring is an explicit component of the trial's statistical design.

Bayesian methods

No Bayesian statistical method is identified in the ClinicalTrials.gov record. The reported framework is based on frequentist survival-analysis methods.

20. What Is Not Reported in the Supplied Registry Data

The ClinicalTrials.gov record deliberately constrain the analysis to the information reported in the registry extract. Several commonly presented clinical-trial elements are not included in that extract.

TopicStatus in the ClinicalTrials.gov record
Baseline characteristicsNot provided.
Median overall survivalNot provided as a reported result.
Median progression-free survivalNot provided as a reported result.
Subgroup resultsNot provided.
Kaplan-Meier event counts or curvesNot provided.
Detailed adverse-event categoriesNot provided beyond serious adverse-event counts by arm.
Complete stratification factorsNot provided in the registry-reported analysis text.
Complete interim-analysis boundaryNot provided.
Missing-data or imputation proceduresNot provided.

Leaving these items out is intentional. A statistical analysis page is more useful when it distinguishes an unreported quantity from an inferred one.

21. Learning Pathway: Statistical Concepts in BRIM-3

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Primary Sources

Continue through the Clinical Biostats statistical learning pathway

Explore the statistical concepts behind randomized time-to-event trials, then connect those concepts to calculators and clinical-trial analysis workflows.

24. Record Summary

BRIM-3 provides a clear example of a randomized phase 3 superiority trial in which both primary endpoints are time-to-event outcomes. The registry reports a 0.37 hazard ratio for overall survival with a two-sided 95% CI of 0.26–0.55 and P < 0.0001, and a 0.26 hazard ratio for progression-free survival with a two-sided 95% CI of 0.20–0.33 and P < .0001. Both analyses used the log-rank test as the reported comparison method and hazard ratio as the effect measure.

The statistical interpretation depends on more than those four numbers. The overall-survival analysis used the registry-defined ITT population, the PFS analysis used a specified ITT subset based on the October 27, 2010 randomization cutoff, and the trial included an interim overall-survival analysis at 50% information. The ClinicalTrials.gov record also identify stratified analysis as an analysis concept and unstratified Cox regression as the method used to estimate the PFS hazard ratio.

Most importantly, the hazard ratios should be read as relative time-to-event measures, not as percentages of patients benefiting. Confidence intervals describe uncertainty around the estimates, while P-values describe statistical evidence under the specified testing framework. The combination of randomized allocation, ITT analysis, log-rank testing, hazard-ratio estimation, confidence intervals, and interim monitoring makes BRIM-3 a useful case study in the statistical analysis of clinical time-to-event data.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. Where the ClinicalTrials.gov record does not provide a result, subgroup estimate, baseline characteristic, or methodological detail, the page does not reconstruct it from outside information.