← Clinical Trials
Glioblastoma Phase 3 Time-to-Event NCT00884741

RTOG 0825: Complete Statistical Analysis of Temozolomide and Radiation Therapy With or Without Bevacizumab in Glioblastoma

An independent statistical analysis of the randomized phase 3 RTOG 0825 trial evaluating temozolomide and radiation therapy with or without bevacizumab in patients with newly diagnosed glioblastoma, with particular attention to overall survival, progression-free survival, time-to-event methodology, and safety.

Trial start: 2009-04-15  ·  Primary completion: 2013-03-17  ·  Enrollment: 637  ·  National Cancer Institute
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

RTOG 0825 was a randomized, double-blind, phase 3 trial evaluating temozolomide and radiation therapy with or without bevacizumab in patients with newly diagnosed glioblastoma, gliosarcoma, or supratentorial glioblastoma. The registry reports 637 enrolled participants and two randomized treatment arms.

637
Enrollment
Phase 3 trial
2
Treatment arms
Parallel randomized design
1.13
OS HR
95% CI 0.93–1.37
0.79
PFS HR
95% CI 0.66–0.94
FeatureRTOG 0825
Trial nameRTOG 0825
ClinicalTrials.gov identifierNCT00884741
PhasePhase 3
StatusCompleted
ConditionsGlioblastoma; Gliosarcoma; Supratentorial Glioblastoma
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment637
Primary endpointsOverall Survival (OS) and Progression-free Survival (PFS)
Primary endpoint typeTime-to-event
Primary hypothesis typeSuperiority
Lead sponsorNational Cancer Institute (NCI)
Sponsor typeNIH

2. Clinical Question

The trial addressed whether adding bevacizumab to temozolomide and radiation therapy could improve overall survival and progression-free survival compared with temozolomide and radiation therapy with placebo in the randomized comparison.

Population

Patients with newly diagnosed glioblastoma, gliosarcoma, or supratentorial glioblastoma represented in the RTOG 0825 trial.

Intervention

Temozolomide and radiation therapy with bevacizumab.

Comparator

Temozolomide and radiation therapy with placebo.

Primary question

Does adding bevacizumab improve the two primary time-to-event outcomes, OS and PFS, relative to temozolomide and radiation therapy with placebo?

3. Trial Design

01
Enroll 637 participants
02
Randomize Parallel allocation
03
Double-blind Bevacizumab or placebo
04
Follow OS / PFS / safety
05
Analyze Time-to-event endpoints
RANDOMIZED ARM 1 · n = 300 at risk for serious AEs

TMZ + RT + Placebo

  • Temozolomide
  • Radiation therapy
  • Placebo
  • Randomized comparator arm for the primary OS and PFS analyses
RANDOMIZED ARM 2 · n = 303 at risk for serious AEs

TMZ + RT + Bevacizumab

  • Temozolomide
  • Radiation therapy
  • Bevacizumab
  • Randomized intervention arm for the primary OS and PFS analyses

The intervention list in the registry also includes 3-Dimensional Conformal Radiation Therapy, Intensity-Modulated Radiation Therapy, Laboratory Biomarker Analysis, and Quality-of-Life Assessment. These are listed as trial interventions or activities in the registry data but are not assigned separate treatment-effect estimates in the statistical analyses posted on ClinicalTrials.gov.

Why the double-blind design matters: masking can reduce the influence of knowledge of treatment assignment on aspects of trial conduct and assessment. The registry identifies RTOG 0825 as double-blind; the ClinicalTrials.gov record does not specify the operational details of who was masked.

4. Endpoints

Both registered primary endpoints were time-to-event outcomes. The registry defines each endpoint from randomization and specifies Kaplan-Meier estimation, censoring of patients last known to be alive at their last contact, and analysis after 390 deaths had been reported.

EndpointRegistered definition / time framePrimary statistical method
Overall Survival (OS) From randomization to date of death or last follow-up. Survival time was defined as time from randomization to date of death from any cause. Patients last known to be alive were censored at the date of last contact. Analysis was planned after all 390 deaths had been reported. Kaplan-Meier estimation and log-rank testing; Cox proportional-hazard effect measure
Progression-free Survival (PFS) From randomization to date of progression, death, or last follow-up for progression-free survival. PFS was estimated by the Kaplan-Meier method, with patients last known to be alive censored at the date of last contact. Kaplan-Meier estimation and log-rank testing; Cox proportional-hazard effect measure
Grade 3 and Higher Treatment-related Toxicity Incidence of Grade 3 and Higher Treatment-related Toxicity as Assessed by the National Cancer Institute Common Terminology Criteria for Adverse Events (AEs) Version 3.0, up to 30 days. Chi-squared test

5. Statistical Methodology

Primary efficacy analysis

The two primary endpoints were analyzed as time-to-event outcomes. For both OS and PFS, the registry reports the log-rank test as the primary comparison method and the Cox proportional-hazard measure as the effect estimate. The analysis population was all eligible randomized patients.

Analysis framework
Randomized comparison → Kaplan-Meier estimation → log-rank test → Cox proportional-hazard estimate

This sequence separates estimation of the time-to-event distributions from formal comparison and from quantification of the relative event hazard.

Analysis population

The registry specifies all eligible randomized patients for both primary endpoint analyses. This anchors the efficacy comparison to randomized treatment assignment rather than restricting the primary analysis to participants who completed a particular amount of treatment.

One-sided testing

The registry-reported analysis text identifies one-sided testing as a concept in the primary analyses. The trial's hypothesis type was superiority. The analysis note states that the study was designed to provide 80% power for detecting a 25% relative reduction in mortality hazard, corresponding to a hazard ratio of .75, and a 30% reduction in progression hazard, corresponding to a hazard ratio of .70.

Secondary toxicity analysis

The registry reports a chi-squared test for the binary secondary endpoint measuring the incidence of Grade 3 and Higher Treatment-related Toxicity up to 30 days. The analysis population was eligible randomized patients with adverse event data who started study treatment.

6. Results: Overall Survival

Overall survival was one of the two primary endpoints. The analysis compared Randomized Arm 1, TMZ+RT + Placebo, with Randomized Arm 2, TMZ+RT + Bevacizumab, among all eligible randomized patients.

Hazard ratio for death

1.13

95% CI: 0.93–1.37   ·   P = 0.11

Log-rank analysis; Cox proportional-hazard effect measure.

Primary endpointComparisonEstimate95% CIP-value
Overall Survival (OS) TMZ+RT + Placebo vs TMZ+RT + Bevacizumab HR 1.13 0.93–1.37 0.11
Clinical Biostats interpretation

The estimated hazard ratio of 1.13 means that, under the Cox proportional-hazard model used for the analysis, the estimated instantaneous rate of death was 1.13 times as high in the TMZ+RT + Placebo group as in the TMZ+RT + Bevacizumab group. Equivalently, because the comparison is reported in the placebo-versus-bevacizumab direction, the point estimate does not indicate a lower estimated death hazard for the placebo group.

The hazard ratio does not mean that 13% more patients died, nor does it describe an absolute difference in survival probability. It is a relative time-to-event measure based on the fitted model.

The 95% confidence interval of 0.93–1.37 describes statistical uncertainty around the estimated hazard ratio. It includes 1, so the interval is compatible with both a modestly lower and a higher estimated hazard for the placebo group under this model.

The P-value of 0.11 measures evidence against the specified null hypothesis under the trial's statistical testing framework; it is not a measure of the magnitude or clinical importance of the treatment effect. It should not be interpreted as the probability that the treatment works or does not work.

Because this is a Cox-model hazard ratio, interpretation also depends on the proportional-hazards framework. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption, so no additional conclusion about that assumption is made here.

7. Results: Progression-free Survival

Progression-free survival was the second co-primary endpoint. As with OS, the analysis population was all eligible randomized patients, and the registry reports a log-rank test with a Cox proportional-hazard effect measure.

Hazard ratio for progression or death

0.79

95% CI: 0.66–0.94   ·   P = 0.004

Log-rank analysis; Cox proportional-hazard effect measure.

Primary endpointComparisonEstimate95% CIP-value
Progression-free Survival (PFS) TMZ+RT + Placebo vs TMZ+RT + Bevacizumab HR 0.79 0.66–0.94 0.004
Clinical Biostats interpretation

The estimated hazard ratio of 0.79 means that, under the fitted Cox model and using the reported comparison direction, the estimated instantaneous rate of progression or death in the TMZ+RT + Placebo group was approximately 79% of that in the TMZ+RT + Bevacizumab group. In relative terms, HR 0.79 corresponds to a 21% lower estimated hazard in the placebo group relative to the bevacizumab group under that comparison direction.

The hazard ratio does not mean that 21% of patients avoided progression, nor does it represent a 21-percentage-point difference in progression-free survival. It is a model-based relative time-to-event measure.

The 95% confidence interval of 0.66–0.94 quantifies uncertainty around the estimated hazard ratio and lies below 1. The interval therefore indicates that the estimated relative hazard favored the placebo group under the reported comparison direction throughout the interval.

The P-value of 0.004 describes the strength of statistical evidence under the specified testing framework. It does not measure the size of the effect and should not be treated as a probability that the observed treatment difference is clinically meaningful.

Educational note: a Kaplan-Meier curve is not reconstructed here from the reported hazard ratio, confidence interval, and P-value. A valid Kaplan-Meier reconstruction requires event and censoring information or sufficiently detailed underlying data.

8. Secondary Endpoint: Grade 3 and Higher Treatment-related Toxicity

The registry reports a secondary binary endpoint measuring the incidence of Grade 3 and Higher Treatment-related Toxicity as assessed by the National Cancer Institute Common Terminology Criteria for Adverse Events, Version 3.0, up to 30 days. The comparison used a chi-squared test.

Chi-squared comparison

P = 0.42

Eligible randomized patients with adverse event data who started study treatment.

Secondary endpointAnalysis methodP-valueTime frame
Incidence of Grade 3 and Higher Treatment-related Toxicity Chi-squared test 0.42 Up to 30 days
Clinical Biostats interpretation

The chi-squared test evaluates whether the observed categorical distribution of Grade 3 and higher treatment-related toxicity differs between the randomized groups under the specified analysis. The reported P-value of 0.42 does not provide strong statistical evidence of a difference under that test.

The ClinicalTrials.gov record does not provide the treatment-specific incidence estimates for this endpoint, so the P-value should not be converted into a treatment effect or an absolute risk difference. A P-value alone also does not establish that the two groups have identical toxicity rates.

The analysis population is narrower than the primary efficacy population because it requires adverse-event data and initiation of study treatment. That distinction is important when comparing efficacy and safety analyses.

9. Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events as affected participants over participants at risk. These figures are descriptive safety data and are distinct from the Grade 3 and Higher Treatment-related Toxicity endpoint analyzed with the chi-squared test.

GroupAffected / at risk
Pre-Randomization TMZ+RT5/10
Randomized Arm 1: TMZ+RT + Placebo115/300
Randomized Arm 2: TMZ+RT + Bevacizumab141/303

The randomized-arm figures show that serious adverse events were reported in 115/300 participants in the placebo arm and 141/303 participants in the bevacizumab arm among those at risk. The pre-randomization TMZ+RT group is presented separately because it is not one of the two randomized comparison arms.

Do not combine safety definitions: serious adverse events, Grade 3 and higher treatment-related toxicity, and all adverse events are different concepts. The ClinicalTrials.gov record reports serious adverse events by arm and separately report a chi-squared P-value for Grade 3 and higher treatment-related toxicity. The two measures should not be treated as interchangeable.

10. Why Kaplan-Meier Analysis Fits These Endpoints

OS and PFS are time-to-event endpoints because both the event and the time until the event are relevant. In OS, the event is death from any cause. In PFS, the event is progression or death. Not every participant necessarily experiences the event during the observation period, so a statistical method must accommodate right censoring.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di is the number of events at time ti, while ni is the number of participants at risk immediately before that time. The registry specifically states that OS and PFS were estimated by the Kaplan-Meier method.

The registry also specifies censoring for patients last known to be alive at their last contact. Censoring allows participants to contribute information for the period during which their event status is known without treating the last observed contact as if it were an event.

11. How the Log-rank Test and Hazard Ratio Work Together

The log-rank test and the Cox proportional-hazard estimate answer related but different questions.

Log-rank test

Provides a formal statistical comparison of the time-to-event experience between the randomized groups while accounting for event timing and censoring.

Hazard ratio

Quantifies the relative event hazard between the groups through the Cox proportional-hazard model.

Kaplan-Meier

Describes the estimated survival or progression-free survival function over time rather than reducing the entire follow-up to one number.

Confidence interval

Shows the statistical uncertainty around the reported hazard-ratio estimate rather than the variability of individual patient outcomes.

This distinction is particularly important for RTOG 0825 because the primary endpoints are both time-to-event outcomes. A P-value without the hazard ratio would not describe the direction or magnitude of the treatment comparison, while a hazard ratio without uncertainty or a formal comparison would provide an incomplete statistical picture.

12. Statistical Methods Explained

Why was a log-rank test used?

OS and PFS are time-to-event outcomes with censoring. The log-rank test is designed to compare survival-type distributions between groups while incorporating the timing of events and the participants who remain event-free at different points in follow-up. The registry explicitly reports the log-rank test for both primary endpoints.

What does an OS hazard ratio of 1.13 mean?

With the reported comparison of TMZ+RT + Placebo versus TMZ+RT + Bevacizumab, an HR of 1.13 means the estimated instantaneous death hazard in the placebo group was 1.13 times that of the bevacizumab group under the Cox model. It does not mean that 13% more participants died or that survival differed by 13 percentage points.

What does a PFS hazard ratio of 0.79 mean?

With the same reported comparison direction, HR 0.79 means the estimated instantaneous hazard of progression or death in the placebo group was 0.79 times that in the bevacizumab group. The corresponding relative difference is 21% lower estimated hazard for the placebo group under this comparison direction. That statement is about the estimated hazard, not an absolute probability of remaining progression-free.

Why does the confidence interval matter?

A point estimate is only one estimate from the observed data. The 95% confidence interval communicates statistical precision around that estimate. For OS, the interval is 0.93–1.37; for PFS, it is 0.66–0.94. These intervals provide substantially more information about uncertainty than the point estimates alone.

Why is the P-value not an effect size?

A P-value quantifies evidence against a specified null hypothesis under the statistical model and testing framework. It does not tell us how large the treatment effect is. Effect size is conveyed here by the hazard ratio, while its precision is conveyed by the confidence interval.

Why does one-sided testing matter?

The ClinicalTrials.gov record identifies one-sided testing for the primary analyses and specify a superiority hypothesis. A one-sided test places the rejection region in the prespecified direction of interest. The interpretation therefore depends on the direction defined in the statistical analysis plan and should not be treated as interchangeable with a two-sided test.

Why are OS and PFS treated as co-primary endpoints?

The registry identifies two primary endpoints and the analysis note describes type I error control for the co-primary endpoints. When multiple primary endpoints are tested, the statistical design must account for the multiplicity so that the overall false-positive error rate remains controlled according to the prespecified strategy.

13. Power and Prespecified Treatment Effects

The registry-reported statistical analysis note states that the trial was designed to concurrently provide 80% power for detection of a 25% relative reduction in mortality hazard, corresponding to a hazard ratio of .75, and a 30% reduction in progression hazard, corresponding to a hazard ratio of .70 for the addition of bevacizumab to temozolomide and radiation.

Design targetPrespecified value
Power80%
Mortality hazard target25% relative reduction
Corresponding OS hazard ratio.75
Progression hazard target30% relative reduction
Corresponding PFS hazard ratio.70

These are design assumptions, not observed treatment effects. The distinction is fundamental: a trial can be powered to detect a particular effect size while the eventual estimate is larger, smaller, or in the opposite direction.

14. The Role of 390 Deaths in the Analysis Plan

Why event counts matter

For time-to-event trials, statistical information depends strongly on the number of observed events, not simply on the number of enrolled participants.

Why follow-up matters

Participants contribute information over time. Event accumulation therefore determines when a planned event-driven analysis becomes available.

The 390-death target is therefore a feature of the planned information structure. It should not be confused with the enrollment count of 637 or with the number of participants experiencing a particular adverse event.

15. Randomization and the Analysis Population

Randomization is central to the interpretation of the primary efficacy comparison. The registry identifies RTOG 0825 as randomized, with a parallel design, and specifies all eligible randomized patients as the analysis population for both primary endpoints.

Analysis conceptRTOG 0825 application
RandomizationParticipants were randomly allocated in a parallel design.
Primary efficacy populationAll eligible randomized patients.
Primary endpoint typeTime-to-event.
Primary comparisonTMZ+RT + Placebo vs TMZ+RT + Bevacizumab.
Safety toxicity populationEligible randomized patients with adverse event data who started study treatment.

The difference between these populations illustrates an important general principle. The primary efficacy analysis is tied to randomized assignment, whereas the safety analysis requires actual treatment exposure and available adverse-event information.

16. Interpreting the Two Primary Results Together

The two primary endpoints provide different pieces of information about the same randomized comparison.

EndpointHazard ratio95% CIP-valueStatistical method
Overall Survival 1.13 0.93–1.37 0.11 Log-rank test; Cox proportional hazard
Progression-free Survival 0.79 0.66–0.94 0.004 Log-rank test; Cox proportional hazard

The estimates point in different directions because the comparison is expressed as placebo versus bevacizumab. For OS, the point estimate is above 1; for PFS, it is below 1. The confidence intervals and P-values provide the corresponding measures of uncertainty and statistical evidence.

This is precisely why a trial should not be summarized using a single P-value. The endpoint definition, comparison direction, effect estimate, confidence interval, analysis population, censoring rules, and testing framework all contribute to the statistical interpretation.

17. Limitations

18. Why This Trial Matters Statistically

RTOG 0825 is a useful statistical teaching case because the ClinicalTrials.gov record bring together randomized phase 3 design, double blinding, co-primary time-to-event endpoints, event-driven analysis, one-sided superiority testing, Kaplan-Meier estimation, log-rank testing, Cox proportional-hazard estimation, confidence intervals, and categorical safety analysis.

ConceptHow it appears in RTOG 0825
RandomizationRandomized parallel-group phase 3 design.
BlindingDouble-blind trial.
Time-to-event endpointsOS and PFS are both primary endpoints.
Kaplan-Meier estimationUsed for both OS and PFS.
Log-rank testReported primary comparison method for OS and PFS.
Hazard ratioCox proportional-hazard measure for both primary endpoints.
Confidence intervals95% two-sided intervals accompany both primary hazard-ratio estimates.
One-sided testingIdentified in the primary analysis information.
MultiplicityType I error control was part of the co-primary endpoint design.
Event-driven analysisPrimary analyses were planned after 390 deaths had been reported.
Categorical analysisChi-squared test used for Grade 3 and higher treatment-related toxicity.
Safety populationsSecondary toxicity analysis used eligible randomized patients with adverse-event data who started study treatment.

19. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The OS analysis reported HR 1.13 with a 95% CI of 0.93–1.37 and P = 0.11. The PFS analysis reported HR 0.79 with a 95% CI of 0.66–0.94 and P = 0.004. Both were analyzed using log-rank testing with Cox proportional-hazard effect estimates.

What the statistics do not establish

These summary measures do not by themselves provide median survival, absolute survival differences, individual patient benefit, subgroup effects, or a complete assessment of clinical benefit and risk.

The distinction is important because a statistically estimated hazard ratio is not itself a complete clinical description. Clinical interpretation requires the endpoint definition, absolute event probabilities where available, duration of follow-up, safety findings, and the context of the prespecified trial design.

20. Important Questions When Reading the Results

Does HR 0.79 mean bevacizumab reduced progression risk by 21%?

Not in the usual absolute-risk sense. The registry-reported analysis reports the comparison as TMZ+RT + Placebo versus TMZ+RT + Bevacizumab. HR 0.79 therefore describes a 21% lower estimated hazard for the placebo group relative to the bevacizumab group under the reported comparison direction. It does not mean that 21% of patients were prevented from progressing.

Does P = 0.004 tell us how large the PFS effect was?

No. The P-value addresses statistical evidence under the specified testing framework. The hazard ratio of 0.79 describes the relative effect estimate, while the 95% CI of 0.66–0.94 describes its statistical uncertainty.

Does OS HR 1.13 prove that bevacizumab was harmful?

No. The point estimate is above 1 for the placebo-versus-bevacizumab comparison, but the 95% CI is 0.93–1.37 and includes 1. A single hazard-ratio estimate should not be transformed into a categorical causal conclusion without considering the confidence interval, testing framework, and full trial evidence.

Why is the OS P-value compared with a special threshold?

The registry-reported analysis note states that the co-primary endpoint design used a significance criterion of 0.023 (one-sided) for OS to control type I error. That is a design feature of the hypothesis-testing framework rather than a property of the observed OS hazard ratio itself.

Why can safety use a chi-squared test while efficacy uses survival analysis?

The endpoints have different data structures. Grade 3 and higher treatment-related toxicity is a binary incidence endpoint over a specified time frame, making a categorical comparison appropriate. OS and PFS incorporate the timing of events and censoring, requiring time-to-event methods.

21. Sources

Source restriction: This page intentionally uses only the ClinicalTrials.gov record for numerical and trial-specific claims. No additional numerical results have been imported from publications or other external sources.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

Continue through Clinical Biostats

Explore additional statistical tutorials, calculators, and clinical trial analyses covering study design, survival analysis, hypothesis testing, and other methods used in clinical research.

24. Record Summary

RTOG 0825 provides a compact example of how a randomized phase 3 oncology trial can be analyzed using multiple statistical frameworks. The two primary endpoints were time-to-event outcomes analyzed with Kaplan-Meier estimation and log-rank testing, with Cox proportional-hazard estimates used to quantify relative effects. The OS analysis reported HR 1.13 (95% CI 0.93–1.37; P = 0.11), while the PFS analysis reported HR 0.79 (95% CI 0.66–0.94; P = 0.004). The trial's design incorporated one-sided testing, a superiority hypothesis, an event-driven analysis planned after 390 deaths, and type I error control for co-primary endpoints. A secondary toxicity endpoint was analyzed using a chi-squared test and had a reported P-value of 0.42.

The most useful way to read these results is to keep several statistical layers separate: endpoint definition, analysis population, effect estimate, confidence interval, P-value, and prespecified testing framework. The hazard ratio describes a relative time-to-event effect; it does not substitute for absolute survival probabilities or other measures that are not included in the ClinicalTrials.gov record.

Clinical Biostats methodology: A trial-results page should not merely repeat an abstract. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and avoiding numerical claims that are not supported by the underlying trial record.