This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
RTOG 0825 was a randomized, double-blind, phase 3 trial evaluating temozolomide and radiation therapy with or without bevacizumab in patients with newly diagnosed glioblastoma, gliosarcoma, or supratentorial glioblastoma. The registry reports 637 enrolled participants and two randomized treatment arms.
| Feature | RTOG 0825 |
|---|---|
| Trial name | RTOG 0825 |
| ClinicalTrials.gov identifier | NCT00884741 |
| Phase | Phase 3 |
| Status | Completed |
| Conditions | Glioblastoma; Gliosarcoma; Supratentorial Glioblastoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 637 |
| Primary endpoints | Overall Survival (OS) and Progression-free Survival (PFS) |
| Primary endpoint type | Time-to-event |
| Primary hypothesis type | Superiority |
| Lead sponsor | National Cancer Institute (NCI) |
| Sponsor type | NIH |
2. Clinical Question
The trial addressed whether adding bevacizumab to temozolomide and radiation therapy could improve overall survival and progression-free survival compared with temozolomide and radiation therapy with placebo in the randomized comparison.
Population
Patients with newly diagnosed glioblastoma, gliosarcoma, or supratentorial glioblastoma represented in the RTOG 0825 trial.
Intervention
Temozolomide and radiation therapy with bevacizumab.
Comparator
Temozolomide and radiation therapy with placebo.
Primary question
Does adding bevacizumab improve the two primary time-to-event outcomes, OS and PFS, relative to temozolomide and radiation therapy with placebo?
3. Trial Design
TMZ + RT + Placebo
- Temozolomide
- Radiation therapy
- Placebo
- Randomized comparator arm for the primary OS and PFS analyses
TMZ + RT + Bevacizumab
- Temozolomide
- Radiation therapy
- Bevacizumab
- Randomized intervention arm for the primary OS and PFS analyses
The intervention list in the registry also includes 3-Dimensional Conformal Radiation Therapy, Intensity-Modulated Radiation Therapy, Laboratory Biomarker Analysis, and Quality-of-Life Assessment. These are listed as trial interventions or activities in the registry data but are not assigned separate treatment-effect estimates in the statistical analyses posted on ClinicalTrials.gov.
4. Endpoints
Both registered primary endpoints were time-to-event outcomes. The registry defines each endpoint from randomization and specifies Kaplan-Meier estimation, censoring of patients last known to be alive at their last contact, and analysis after 390 deaths had been reported.
| Endpoint | Registered definition / time frame | Primary statistical method |
|---|---|---|
| Overall Survival (OS) | From randomization to date of death or last follow-up. Survival time was defined as time from randomization to date of death from any cause. Patients last known to be alive were censored at the date of last contact. Analysis was planned after all 390 deaths had been reported. | Kaplan-Meier estimation and log-rank testing; Cox proportional-hazard effect measure |
| Progression-free Survival (PFS) | From randomization to date of progression, death, or last follow-up for progression-free survival. PFS was estimated by the Kaplan-Meier method, with patients last known to be alive censored at the date of last contact. | Kaplan-Meier estimation and log-rank testing; Cox proportional-hazard effect measure |
| Grade 3 and Higher Treatment-related Toxicity | Incidence of Grade 3 and Higher Treatment-related Toxicity as Assessed by the National Cancer Institute Common Terminology Criteria for Adverse Events (AEs) Version 3.0, up to 30 days. | Chi-squared test |
5. Statistical Methodology
Primary efficacy analysis
The two primary endpoints were analyzed as time-to-event outcomes. For both OS and PFS, the registry reports the log-rank test as the primary comparison method and the Cox proportional-hazard measure as the effect estimate. The analysis population was all eligible randomized patients.
This sequence separates estimation of the time-to-event distributions from formal comparison and from quantification of the relative event hazard.
Analysis population
The registry specifies all eligible randomized patients for both primary endpoint analyses. This anchors the efficacy comparison to randomized treatment assignment rather than restricting the primary analysis to participants who completed a particular amount of treatment.
One-sided testing
The registry-reported analysis text identifies one-sided testing as a concept in the primary analyses. The trial's hypothesis type was superiority. The analysis note states that the study was designed to provide 80% power for detecting a 25% relative reduction in mortality hazard, corresponding to a hazard ratio of .75, and a 30% reduction in progression hazard, corresponding to a hazard ratio of .70.
Secondary toxicity analysis
The registry reports a chi-squared test for the binary secondary endpoint measuring the incidence of Grade 3 and Higher Treatment-related Toxicity up to 30 days. The analysis population was eligible randomized patients with adverse event data who started study treatment.
6. Results: Overall Survival
Overall survival was one of the two primary endpoints. The analysis compared Randomized Arm 1, TMZ+RT + Placebo, with Randomized Arm 2, TMZ+RT + Bevacizumab, among all eligible randomized patients.
Hazard ratio for death
95% CI: 0.93–1.37 · P = 0.11
Log-rank analysis; Cox proportional-hazard effect measure.
| Primary endpoint | Comparison | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Overall Survival (OS) | TMZ+RT + Placebo vs TMZ+RT + Bevacizumab | HR 1.13 | 0.93–1.37 | 0.11 |
The estimated hazard ratio of 1.13 means that, under the Cox proportional-hazard model used for the analysis, the estimated instantaneous rate of death was 1.13 times as high in the TMZ+RT + Placebo group as in the TMZ+RT + Bevacizumab group. Equivalently, because the comparison is reported in the placebo-versus-bevacizumab direction, the point estimate does not indicate a lower estimated death hazard for the placebo group.
The hazard ratio does not mean that 13% more patients died, nor does it describe an absolute difference in survival probability. It is a relative time-to-event measure based on the fitted model.
The 95% confidence interval of 0.93–1.37 describes statistical uncertainty around the estimated hazard ratio. It includes 1, so the interval is compatible with both a modestly lower and a higher estimated hazard for the placebo group under this model.
The P-value of 0.11 measures evidence against the specified null hypothesis under the trial's statistical testing framework; it is not a measure of the magnitude or clinical importance of the treatment effect. It should not be interpreted as the probability that the treatment works or does not work.
Because this is a Cox-model hazard ratio, interpretation also depends on the proportional-hazards framework. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption, so no additional conclusion about that assumption is made here.
7. Results: Progression-free Survival
Progression-free survival was the second co-primary endpoint. As with OS, the analysis population was all eligible randomized patients, and the registry reports a log-rank test with a Cox proportional-hazard effect measure.
Hazard ratio for progression or death
95% CI: 0.66–0.94 · P = 0.004
Log-rank analysis; Cox proportional-hazard effect measure.
| Primary endpoint | Comparison | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Progression-free Survival (PFS) | TMZ+RT + Placebo vs TMZ+RT + Bevacizumab | HR 0.79 | 0.66–0.94 | 0.004 |
The estimated hazard ratio of 0.79 means that, under the fitted Cox model and using the reported comparison direction, the estimated instantaneous rate of progression or death in the TMZ+RT + Placebo group was approximately 79% of that in the TMZ+RT + Bevacizumab group. In relative terms, HR 0.79 corresponds to a 21% lower estimated hazard in the placebo group relative to the bevacizumab group under that comparison direction.
The hazard ratio does not mean that 21% of patients avoided progression, nor does it represent a 21-percentage-point difference in progression-free survival. It is a model-based relative time-to-event measure.
The 95% confidence interval of 0.66–0.94 quantifies uncertainty around the estimated hazard ratio and lies below 1. The interval therefore indicates that the estimated relative hazard favored the placebo group under the reported comparison direction throughout the interval.
The P-value of 0.004 describes the strength of statistical evidence under the specified testing framework. It does not measure the size of the effect and should not be treated as a probability that the observed treatment difference is clinically meaningful.
8. Secondary Endpoint: Grade 3 and Higher Treatment-related Toxicity
The registry reports a secondary binary endpoint measuring the incidence of Grade 3 and Higher Treatment-related Toxicity as assessed by the National Cancer Institute Common Terminology Criteria for Adverse Events, Version 3.0, up to 30 days. The comparison used a chi-squared test.
Chi-squared comparison
Eligible randomized patients with adverse event data who started study treatment.
| Secondary endpoint | Analysis method | P-value | Time frame |
|---|---|---|---|
| Incidence of Grade 3 and Higher Treatment-related Toxicity | Chi-squared test | 0.42 | Up to 30 days |
The chi-squared test evaluates whether the observed categorical distribution of Grade 3 and higher treatment-related toxicity differs between the randomized groups under the specified analysis. The reported P-value of 0.42 does not provide strong statistical evidence of a difference under that test.
The ClinicalTrials.gov record does not provide the treatment-specific incidence estimates for this endpoint, so the P-value should not be converted into a treatment effect or an absolute risk difference. A P-value alone also does not establish that the two groups have identical toxicity rates.
The analysis population is narrower than the primary efficacy population because it requires adverse-event data and initiation of study treatment. That distinction is important when comparing efficacy and safety analyses.
9. Serious Adverse Events by Arm
The ClinicalTrials.gov record reports serious adverse events as affected participants over participants at risk. These figures are descriptive safety data and are distinct from the Grade 3 and Higher Treatment-related Toxicity endpoint analyzed with the chi-squared test.
| Group | Affected / at risk |
|---|---|
| Pre-Randomization TMZ+RT | 5/10 |
| Randomized Arm 1: TMZ+RT + Placebo | 115/300 |
| Randomized Arm 2: TMZ+RT + Bevacizumab | 141/303 |
The randomized-arm figures show that serious adverse events were reported in 115/300 participants in the placebo arm and 141/303 participants in the bevacizumab arm among those at risk. The pre-randomization TMZ+RT group is presented separately because it is not one of the two randomized comparison arms.
10. Why Kaplan-Meier Analysis Fits These Endpoints
OS and PFS are time-to-event endpoints because both the event and the time until the event are relevant. In OS, the event is death from any cause. In PFS, the event is progression or death. Not every participant necessarily experiences the event during the observation period, so a statistical method must accommodate right censoring.
Here, di is the number of events at time ti, while ni is the number of participants at risk immediately before that time. The registry specifically states that OS and PFS were estimated by the Kaplan-Meier method.
The registry also specifies censoring for patients last known to be alive at their last contact. Censoring allows participants to contribute information for the period during which their event status is known without treating the last observed contact as if it were an event.
11. How the Log-rank Test and Hazard Ratio Work Together
The log-rank test and the Cox proportional-hazard estimate answer related but different questions.
Log-rank test
Provides a formal statistical comparison of the time-to-event experience between the randomized groups while accounting for event timing and censoring.
Hazard ratio
Quantifies the relative event hazard between the groups through the Cox proportional-hazard model.
Kaplan-Meier
Describes the estimated survival or progression-free survival function over time rather than reducing the entire follow-up to one number.
Confidence interval
Shows the statistical uncertainty around the reported hazard-ratio estimate rather than the variability of individual patient outcomes.
This distinction is particularly important for RTOG 0825 because the primary endpoints are both time-to-event outcomes. A P-value without the hazard ratio would not describe the direction or magnitude of the treatment comparison, while a hazard ratio without uncertainty or a formal comparison would provide an incomplete statistical picture.
12. Statistical Methods Explained
Why was a log-rank test used?
OS and PFS are time-to-event outcomes with censoring. The log-rank test is designed to compare survival-type distributions between groups while incorporating the timing of events and the participants who remain event-free at different points in follow-up. The registry explicitly reports the log-rank test for both primary endpoints.
What does an OS hazard ratio of 1.13 mean?
With the reported comparison of TMZ+RT + Placebo versus TMZ+RT + Bevacizumab, an HR of 1.13 means the estimated instantaneous death hazard in the placebo group was 1.13 times that of the bevacizumab group under the Cox model. It does not mean that 13% more participants died or that survival differed by 13 percentage points.
What does a PFS hazard ratio of 0.79 mean?
With the same reported comparison direction, HR 0.79 means the estimated instantaneous hazard of progression or death in the placebo group was 0.79 times that in the bevacizumab group. The corresponding relative difference is 21% lower estimated hazard for the placebo group under this comparison direction. That statement is about the estimated hazard, not an absolute probability of remaining progression-free.
Why does the confidence interval matter?
A point estimate is only one estimate from the observed data. The 95% confidence interval communicates statistical precision around that estimate. For OS, the interval is 0.93–1.37; for PFS, it is 0.66–0.94. These intervals provide substantially more information about uncertainty than the point estimates alone.
Why is the P-value not an effect size?
A P-value quantifies evidence against a specified null hypothesis under the statistical model and testing framework. It does not tell us how large the treatment effect is. Effect size is conveyed here by the hazard ratio, while its precision is conveyed by the confidence interval.
Why does one-sided testing matter?
The ClinicalTrials.gov record identifies one-sided testing for the primary analyses and specify a superiority hypothesis. A one-sided test places the rejection region in the prespecified direction of interest. The interpretation therefore depends on the direction defined in the statistical analysis plan and should not be treated as interchangeable with a two-sided test.
Why are OS and PFS treated as co-primary endpoints?
The registry identifies two primary endpoints and the analysis note describes type I error control for the co-primary endpoints. When multiple primary endpoints are tested, the statistical design must account for the multiplicity so that the overall false-positive error rate remains controlled according to the prespecified strategy.
13. Power and Prespecified Treatment Effects
The registry-reported statistical analysis note states that the trial was designed to concurrently provide 80% power for detection of a 25% relative reduction in mortality hazard, corresponding to a hazard ratio of .75, and a 30% reduction in progression hazard, corresponding to a hazard ratio of .70 for the addition of bevacizumab to temozolomide and radiation.
| Design target | Prespecified value |
|---|---|
| Power | 80% |
| Mortality hazard target | 25% relative reduction |
| Corresponding OS hazard ratio | .75 |
| Progression hazard target | 30% relative reduction |
| Corresponding PFS hazard ratio | .70 |
These are design assumptions, not observed treatment effects. The distinction is fundamental: a trial can be powered to detect a particular effect size while the eventual estimate is larger, smaller, or in the opposite direction.
14. The Role of 390 Deaths in the Analysis Plan
Why event counts matter
For time-to-event trials, statistical information depends strongly on the number of observed events, not simply on the number of enrolled participants.
Why follow-up matters
Participants contribute information over time. Event accumulation therefore determines when a planned event-driven analysis becomes available.
The 390-death target is therefore a feature of the planned information structure. It should not be confused with the enrollment count of 637 or with the number of participants experiencing a particular adverse event.
15. Randomization and the Analysis Population
Randomization is central to the interpretation of the primary efficacy comparison. The registry identifies RTOG 0825 as randomized, with a parallel design, and specifies all eligible randomized patients as the analysis population for both primary endpoints.
| Analysis concept | RTOG 0825 application |
|---|---|
| Randomization | Participants were randomly allocated in a parallel design. |
| Primary efficacy population | All eligible randomized patients. |
| Primary endpoint type | Time-to-event. |
| Primary comparison | TMZ+RT + Placebo vs TMZ+RT + Bevacizumab. |
| Safety toxicity population | Eligible randomized patients with adverse event data who started study treatment. |
The difference between these populations illustrates an important general principle. The primary efficacy analysis is tied to randomized assignment, whereas the safety analysis requires actual treatment exposure and available adverse-event information.
16. Interpreting the Two Primary Results Together
The two primary endpoints provide different pieces of information about the same randomized comparison.
| Endpoint | Hazard ratio | 95% CI | P-value | Statistical method |
|---|---|---|---|---|
| Overall Survival | 1.13 | 0.93–1.37 | 0.11 | Log-rank test; Cox proportional hazard |
| Progression-free Survival | 0.79 | 0.66–0.94 | 0.004 | Log-rank test; Cox proportional hazard |
The estimates point in different directions because the comparison is expressed as placebo versus bevacizumab. For OS, the point estimate is above 1; for PFS, it is below 1. The confidence intervals and P-values provide the corresponding measures of uncertainty and statistical evidence.
This is precisely why a trial should not be summarized using a single P-value. The endpoint definition, comparison direction, effect estimate, confidence interval, analysis population, censoring rules, and testing framework all contribute to the statistical interpretation.
17. Limitations
- Limited registry result detail: the ClinicalTrials.gov record reports hazard ratios, confidence intervals, and P-values for the two primary endpoints but do not provide median OS, median PFS, time-specific survival estimates, event counts by randomized arm, or subgroup estimates.
- Hazard-ratio interpretation: Cox proportional-hazard estimates depend on the proportional-hazards framework. The ClinicalTrials.gov record does not report a formal proportional-hazards assessment.
- No subgroup estimates reported: subgroup-specific treatment effects cannot be inferred from the overall hazard ratios.
- No baseline table reported: the ClinicalTrials.gov record does not provide numerical baseline characteristics, so balance between randomized groups cannot be evaluated from those characteristics here.
- Safety endpoint distinction: the Grade 3 and higher treatment-related toxicity endpoint and the reported serious-adverse-event counts are different safety measures and should not be merged.
- Safety analysis population: the secondary toxicity analysis excludes eligible randomized patients who did not start study treatment or lacked adverse-event data, according to the registry definition.
- Absolute treatment effects: the ClinicalTrials.gov record does not provide absolute OS or PFS probabilities, so the hazard ratios cannot be translated into an absolute survival benefit from the provided information alone.
18. Why This Trial Matters Statistically
RTOG 0825 is a useful statistical teaching case because the ClinicalTrials.gov record bring together randomized phase 3 design, double blinding, co-primary time-to-event endpoints, event-driven analysis, one-sided superiority testing, Kaplan-Meier estimation, log-rank testing, Cox proportional-hazard estimation, confidence intervals, and categorical safety analysis.
| Concept | How it appears in RTOG 0825 |
|---|---|
| Randomization | Randomized parallel-group phase 3 design. |
| Blinding | Double-blind trial. |
| Time-to-event endpoints | OS and PFS are both primary endpoints. |
| Kaplan-Meier estimation | Used for both OS and PFS. |
| Log-rank test | Reported primary comparison method for OS and PFS. |
| Hazard ratio | Cox proportional-hazard measure for both primary endpoints. |
| Confidence intervals | 95% two-sided intervals accompany both primary hazard-ratio estimates. |
| One-sided testing | Identified in the primary analysis information. |
| Multiplicity | Type I error control was part of the co-primary endpoint design. |
| Event-driven analysis | Primary analyses were planned after 390 deaths had been reported. |
| Categorical analysis | Chi-squared test used for Grade 3 and higher treatment-related toxicity. |
| Safety populations | Secondary toxicity analysis used eligible randomized patients with adverse-event data who started study treatment. |
19. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The OS analysis reported HR 1.13 with a 95% CI of 0.93–1.37 and P = 0.11. The PFS analysis reported HR 0.79 with a 95% CI of 0.66–0.94 and P = 0.004. Both were analyzed using log-rank testing with Cox proportional-hazard effect estimates.
What the statistics do not establish
These summary measures do not by themselves provide median survival, absolute survival differences, individual patient benefit, subgroup effects, or a complete assessment of clinical benefit and risk.
The distinction is important because a statistically estimated hazard ratio is not itself a complete clinical description. Clinical interpretation requires the endpoint definition, absolute event probabilities where available, duration of follow-up, safety findings, and the context of the prespecified trial design.
20. Important Questions When Reading the Results
Does HR 0.79 mean bevacizumab reduced progression risk by 21%?
Not in the usual absolute-risk sense. The registry-reported analysis reports the comparison as TMZ+RT + Placebo versus TMZ+RT + Bevacizumab. HR 0.79 therefore describes a 21% lower estimated hazard for the placebo group relative to the bevacizumab group under the reported comparison direction. It does not mean that 21% of patients were prevented from progressing.
Does P = 0.004 tell us how large the PFS effect was?
No. The P-value addresses statistical evidence under the specified testing framework. The hazard ratio of 0.79 describes the relative effect estimate, while the 95% CI of 0.66–0.94 describes its statistical uncertainty.
Does OS HR 1.13 prove that bevacizumab was harmful?
No. The point estimate is above 1 for the placebo-versus-bevacizumab comparison, but the 95% CI is 0.93–1.37 and includes 1. A single hazard-ratio estimate should not be transformed into a categorical causal conclusion without considering the confidence interval, testing framework, and full trial evidence.
Why is the OS P-value compared with a special threshold?
The registry-reported analysis note states that the co-primary endpoint design used a significance criterion of 0.023 (one-sided) for OS to control type I error. That is a design feature of the hypothesis-testing framework rather than a property of the observed OS hazard ratio itself.
Why can safety use a chi-squared test while efficacy uses survival analysis?
The endpoints have different data structures. Grade 3 and higher treatment-related toxicity is a binary incidence endpoint over a specified time frame, making a categorical comparison appropriate. OS and PFS incorporate the timing of events and censoring, requiring time-to-event methods.
21. Sources
- ClinicalTrials.gov: RTOG 0825, NCT00884741.
- Linked publication: PubMed record for PMID 24552317.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
Continue through Clinical Biostats
Explore additional statistical tutorials, calculators, and clinical trial analyses covering study design, survival analysis, hypothesis testing, and other methods used in clinical research.
24. Record Summary
RTOG 0825 provides a compact example of how a randomized phase 3 oncology trial can be analyzed using multiple statistical frameworks. The two primary endpoints were time-to-event outcomes analyzed with Kaplan-Meier estimation and log-rank testing, with Cox proportional-hazard estimates used to quantify relative effects. The OS analysis reported HR 1.13 (95% CI 0.93–1.37; P = 0.11), while the PFS analysis reported HR 0.79 (95% CI 0.66–0.94; P = 0.004). The trial's design incorporated one-sided testing, a superiority hypothesis, an event-driven analysis planned after 390 deaths, and type I error control for co-primary endpoints. A secondary toxicity endpoint was analyzed using a chi-squared test and had a reported P-value of 0.42.
The most useful way to read these results is to keep several statistical layers separate: endpoint definition, analysis population, effect estimate, confidence interval, P-value, and prespecified testing framework. The hazard ratio describes a relative time-to-event effect; it does not substitute for absolute survival probabilities or other measures that are not included in the ClinicalTrials.gov record.