← Clinical Trials
Ovarian Cancer Phase 3 Time-to-Event NCT00951496

GOG-0252: Complete Statistical Analysis of Bevacizumab and Intraperitoneal Chemotherapy in Ovarian Cancer

An independent statistical review of the randomized phase 3 GOG-0252 trial evaluating bevacizumab with intravenous or intraperitoneal chemotherapy in patients with stage II-III ovarian epithelial cancer, fallopian tube cancer, or primary peritoneal cancer.

GOG-0252  ·  Phase 3  ·  Randomized  ·  Completed
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

GOG-0252 was a randomized, parallel, open-label phase 3 trial with 1,560 enrolled patients and three treatment arms. The registered primary endpoint was median progression-free survival, measured from randomization until the first indication of progression based on RECIST criteria.

1,560
Enrollment
All enrolled patients
3
Arms
Parallel design
0.94
PFS HR
Arm II vs Arm I
0.99
PFS HR
Arm III vs Arm I
FeatureGOG-0252
TrialGOG-0252
ClinicalTrials.gov identifierNCT00951496
PhasePhase 3
StatusCompleted
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment1,560
Number of arms3
Lead sponsorNational Cancer Institute (NCI)
Sponsor typeNIH
Start2009-08-11
Primary completion2016-01-11
Primary endpoint typeTime-to-event
Registered primary endpoints1
Statistical analyses posted2

2. Clinical Question

The clinical question represented by the registered trial is whether different ways of administering chemotherapy with bevacizumab produce different progression-free survival outcomes in patients with stage II-III ovarian epithelial cancer, fallopian tube cancer, or primary peritoneal cancer.

Population

Patients with stage II-III ovarian epithelial cancer, fallopian tube cancer, or primary peritoneal cancer, as represented by the trial's registered condition list.

Reference regimen

Arm I: paclitaxel, carboplatin, and bevacizumab administered with intravenous chemotherapy.

Alternative regimen

Arm II: paclitaxel, carboplatin administered intraperitoneally, and bevacizumab.

Third regimen

Arm III: paclitaxel administered intraperitoneally, cisplatin, and bevacizumab.

The posted primary analyses make two direct comparisons against Arm I: Arm II versus Arm I and Arm III versus Arm I. The ClinicalTrials.gov record does not provide a separate formal statistical comparison between Arm II and Arm III.

3. Trial Design

01
Randomize 1,560 enrolled
02
Three arms Parallel treatment groups
03
Treatment Bevacizumab + chemotherapy
04
Assess Progression or censoring
05
Compare Stratified log-rank analysis
ARM I

Intravenous reference regimen

  • Paclitaxel
  • Carboplatin
  • Bevacizumab
ARM II

Intraperitoneal carboplatin regimen

  • Paclitaxel
  • Carboplatin administered intraperitoneally
  • Bevacizumab administered intraperitoneally
ARM III

Intraperitoneal cisplatin regimen

  • Paclitaxel administered intraperitoneally
  • Cisplatin
  • Bevacizumab
Three-arm interpretation: The existence of three randomized arms does not mean that every possible pairwise comparison is represented in the posted primary analyses. The statistical analyses posted on ClinicalTrials.gov specify Arm II versus Arm I and Arm III versus Arm I. Those are therefore the comparisons evaluated on this page.

4. Primary Endpoint

EndpointDefinition / time frameStatistical analysis
Median Progression-free Survival Progression-free survival is measured from date of randomization until first indication of progression based on RECIST criteria. Stratified log-rank test; hazard ratio

The registry defines progression using Response Evaluation Criteria in Solid Tumors criteria (RECIST v1.0): a 20% increase in the sum of the longest diameter of target lesions, a measurable increase in a non-target lesion, or the appearance of new lesions.

This is a time-to-event endpoint. The analysis therefore uses information about both whether progression occurred and when it occurred. Patients without a recorded progression by the relevant observation point contribute censored follow-up rather than being treated as if progression had occurred.

5. Analysis Population, Stratification, and Covariate Adjustment

Both posted primary analyses specify an intention-to-treat analysis population consisting of all enrolled patients. The analysis text also states that the progression-free survival comparison was stratified by stage of disease and size of residual disease.

Statistical featureGOG-0252 specification
Primary efficacy populationIntention-to-treat: all enrolled patients
Comparison 1Arm II vs Arm I
Comparison 2Arm III vs Arm I
Primary methodLog-rank test
Stratification factorsStage of disease; size of residual disease
Effect measureHazard ratio
AdjustmentAdjusted for stage of disease and residual disease size

Why stratification matters

Stratification allows the time-to-event comparison to account for prespecified disease characteristics while preserving the randomized comparison. Instead of treating all patients as though they came from one homogeneous risk set, the analysis compares treatment groups within the specified strata and combines the information across them.

For this trial, stage of disease and residual disease size are particularly relevant because both describe aspects of disease status at baseline. The posted analysis therefore does not simply report an unstratified comparison of all progression times.

6. Statistical Methodology

Kaplan-Meier estimation

A time-to-event endpoint such as progression-free survival is commonly described using the Kaplan-Meier estimator. The estimator accounts for patients who remain free of progression at their last assessment by censoring their follow-up rather than assigning an artificial progression time.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di is the number of events at time ti, and ni is the number of patients at risk immediately before that time.

Stratified log-rank test

The posted primary analyses use a stratified log-rank test to assess equality of progression-free survival hazards between the specified treatment groups. The analysis is stratified by stage of disease and size of residual disease.

The important statistical point is that the log-rank test evaluates evidence against a null hypothesis concerning the survival experience of the groups. It does not itself produce the hazard ratio. The hazard ratio is the effect measure reported alongside the test result.

Hazard ratio

Interpretation
HR < 1  →  lower estimated instantaneous event rate in the first-named treatment group relative to the reference group

For GOG-0252, the reported HRs compare progression hazards for Arm II or Arm III with Arm I. A hazard ratio close to 1 indicates that the estimated relative hazard is close to that of the reference arm.

Intention-to-treat analysis

The posted primary analyses use all enrolled patients as the intention-to-treat population. The central principle is to retain patients according to their randomized treatment assignment for the efficacy comparison. This protects the treatment contrast created by randomization and avoids redefining the comparison according to later treatment exposure.

Covariate adjustment and stratification

The analysis notes explicitly describe adjustment for stage of disease and residual disease size. In a stratified survival analysis, these variables help account for differences in underlying event risk across clinically defined strata while estimating the treatment contrast across those strata.

7. Primary Results: Arm II vs Arm I

The first posted primary analysis compares Arm II with Arm I for median progression-free survival. The analysis population was all enrolled patients, and the registry reports a stratified log-rank test with the comparison adjusted for stage of disease and residual disease size.

Progression-free survival hazard ratio

0.94

95% CI: 0.81–1.09   ·   P = 0.341

Arm II relative to Arm I; two-sided 95% confidence interval.

FeatureArm II vs Arm I
EndpointMedian Progression-free Survival
Analysis populationIntention-to-treat: all enrolled patients
MethodStratified log-rank test
StratificationStage of disease and size of residual disease
Effect measureHazard ratio
Hazard ratio0.94
95% CI0.81–1.09
P-value0.341
Hypothesis typeNon-inferiority or equivalence
Clinical Biostats interpretation

The reported HR of 0.94 means that the estimated progression hazard for Arm II relative to Arm I was 0.94 under the reported stratified survival analysis. Expressed descriptively, the point estimate is below 1, but it is close to 1.

The HR does not mean that 6% of patients benefited, that progression was reduced by exactly 6% for every patient, or that the two treatment strategies have identical clinical effects. A hazard ratio is a relative time-to-event measure, not an individual-level probability.

The 95% CI of 0.81–1.09 describes statistical uncertainty around the estimated hazard ratio. It includes 1, so the interval is compatible with a lower hazard, a hazard close to equality, or a higher hazard for Arm II relative to Arm I.

The P-value of 0.341 is evidence used in the hypothesis test; it is not a measure of effect size. A p-value does not tell us that the treatment effect is 34.1%, nor does it quantify the clinical importance of the difference.

Most importantly, the registry identifies the hypothesis type as non-inferiority or equivalence. Non-inferiority is not established merely because a conventional p-value is greater than 0.05. It depends on a prespecified non-inferiority margin and the corresponding confidence-interval decision rule. The ClinicalTrials.gov record does not provide a numeric margin, so the posted HR and CI should be reported without assigning a non-inferiority conclusion that is not explicitly contained in the ClinicalTrials.gov record.

What the design comment adds

The registry states that the study was designed to provide 80% power when Arm II reduces the progression-free survival event rate 20%. It also states that the critical p-value accounts for correlation between the two primary hypotheses.

That statement is important because the trial was not simply a two-group superiority comparison. The statistical design contemplated two primary hypotheses, and the registry description explicitly indicates that the relationship between those hypotheses was incorporated into the critical-value framework.

8. Primary Results: Arm III vs Arm I

The second posted primary analysis compares Arm III with Arm I using the same primary endpoint and the same general stratified survival-analysis framework.

Progression-free survival hazard ratio

0.99

95% CI: 0.86–1.15   ·   P = 0.587

Arm III relative to Arm I; two-sided 95% confidence interval.

FeatureArm III vs Arm I
EndpointMedian Progression-free Survival
Analysis populationIntention-to-treat: all enrolled patients
MethodLog-rank test
StratificationStage of disease and size of residual disease
Effect measureHazard ratio
Hazard ratio0.99
95% CI0.86–1.15
P-value0.587
Hypothesis typeNon-inferiority or equivalence
Clinical Biostats interpretation

The reported HR of 0.99 places the Arm III estimate very close to 1 relative to Arm I. Under the reported model, the estimated progression hazard for Arm III was 0.99 times that of Arm I.

This does not establish that the two regimens are clinically identical. A point estimate near 1 is compatible with several underlying patterns, and the confidence interval provides the more informative description of statistical precision.

The 95% CI of 0.86–1.15 includes 1. The interval therefore allows for a lower, approximately equal, or higher progression hazard for Arm III relative to Arm I within the uncertainty represented by this analysis.

The P-value of 0.587 describes the evidence from the reported hypothesis test; it does not describe the magnitude of any treatment difference. It should not be converted into a percentage treatment effect.

As with the Arm II comparison, the registry labels the hypothesis type as non-inferiority or equivalence. The ClinicalTrials.gov record does not provide a numeric non-inferiority margin or the exact confidence-interval criterion needed to make a formal non-inferiority determination. Therefore, the statistical interpretation should remain tied to the reported HR, CI, and p-value rather than extending them into an unsupported non-inferiority conclusion.

Power and the second primary hypothesis

The registry states that the study was designed to provide 80% power when Arm III reduced the true progression-free survival event rate 20% compared with Arm I. It also states that the critical p-value accounts for correlation between the two primary hypotheses.

This is a useful reminder that the observed HR of 0.99 should not be substituted into the original power calculation. Power is a property of a prespecified design under assumed alternatives, whereas the hazard ratio reported above is an estimate obtained from the observed trial data.

9. Comparing the Two Primary Estimates

Primary comparisonHR95% CIP-value
Arm II vs Arm I0.940.81–1.090.341
Arm III vs Arm I0.990.86–1.150.587

The two point estimates are both close to 1, with Arm II having an estimated HR of 0.94 and Arm III an estimated HR of 0.99 relative to Arm I. The confidence intervals for both comparisons include 1.

How not to over-interpret the pair of results

It would be inappropriate to treat the smaller HR for Arm II as proof that Arm II is statistically different from Arm III. The posted analyses are each comparisons with Arm I; they do not provide a direct Arm II-versus-Arm III hypothesis test.

Similarly, the fact that the two p-values differ does not itself demonstrate that the treatment effects differ. Comparing two estimates requires an appropriate statistical test of their difference or interaction, not a comparison of their individual p-values.

10. Understanding the Non-Inferiority / Equivalence Framework

The registry identifies the hypothesis type for both primary analyses as non-inferiority or equivalence. This changes how the statistical evidence should be interpreted compared with a conventional superiority trial.

Superiority question

A superiority analysis asks whether the data provide evidence that one treatment differs from another in the specified direction, usually relative to a null effect of 1 for a hazard ratio.

Non-inferiority question

A non-inferiority analysis asks whether the alternative treatment is not worse than the reference by more than a prespecified clinically acceptable margin.

Why the margin matters

The numerical margin defines how much loss of efficacy can be tolerated while still meeting the non-inferiority objective. The ClinicalTrials.gov record does not state a numeric margin.

Why the CI matters

For a non-inferiority claim, the confidence interval is interpreted against the prespecified margin. A conventional p-value alone is not the decision rule.

Important: The ClinicalTrials.gov record identifies the trial as having a non-inferiority or equivalence hypothesis and describe the power assumptions, but they do not provide a numeric non-inferiority margin. The HRs and two-sided 95% CIs can therefore be explained, but a formal non-inferiority or equivalence conclusion should not be inferred beyond the posted information.

11. Multiplicity and the Two Primary Hypotheses

GOG-0252 has one registered primary endpoint but two posted primary analyses: Arm II versus Arm I and Arm III versus Arm I. The analysis comments explicitly state that the critical p-value accounts for correlation between the two primary hypotheses.

FeatureStatistical implication
One registered primary endpointMedian progression-free survival
Two primary comparisonsArm II vs Arm I; Arm III vs Arm I
Common referenceArm I
Multiplicity issueTwo related primary hypotheses require control of the overall false-positive framework
Registry statementCritical p-value accounts for correlation between the two primary hypotheses

The word correlation matters here. Because both hypotheses use Arm I as the reference group, the statistical tests are not independent in the simple sense that two unrelated experiments would be. The registry analysis explicitly says that this correlation was incorporated into the critical p-value.

The exact critical p-value is not included in the ClinicalTrials.gov record, so this page does not substitute the conventional 0.05 threshold for the trial's prespecified critical value.

12. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected patients divided by patients at risk. These are presented below exactly as reported in the registry rather than converting the counts into percentages.

ArmSerious adverse eventsAffected / at risk
Arm IPaclitaxel, carboplatin, bevacizumab156/511
Arm IIPaclitaxel, carboplatin IP, bevacizumab179/510
Arm IIIPaclitaxel IP, cisplatin, bevacizumab215/508
Safety interpretation: These serious-adverse-event figures are descriptive arm-level safety data. They should not be combined with the progression-free survival hazard ratios into a single statistical measure of overall benefit or harm. Efficacy and safety answer different questions and are analyzed using different statistical frameworks.

The ClinicalTrials.gov record does not provide formal statistical analyses of serious adverse events, confidence intervals, p-values, or adjusted comparisons by arm. Accordingly, this page does not construct such analyses.

13. Trial Timeline

2009-08-11

Trial start

GOG-0252 began enrollment under the randomized phase 3 study design.

2016-01-11

Primary completion

The registry lists January 11, 2016 as the primary completion date.

Completed

Results posted

The registry record is marked completed and contains seven posted outcome measures and two posted statistical analyses.

14. Statistical Methods Explained

Why was a stratified log-rank test used?

Progression-free survival is a time-to-event endpoint, so a log-rank framework is appropriate for comparing the survival experience of randomized groups over follow-up. GOG-0252 used a stratified version because the posted analysis specifies stage of disease and residual disease size as stratification factors.

What does an HR of 0.94 mean?

An HR of 0.94 for Arm II versus Arm I means that the estimated progression hazard for Arm II was 0.94 times the corresponding hazard for Arm I under the reported analysis. It is a relative time-to-event measure, not a statement that 6% of patients avoided progression.

What does an HR of 0.99 mean?

An HR of 0.99 for Arm III versus Arm I places the estimated progression hazard very close to the reference value of 1. The associated 95% CI of 0.86–1.15 shows that the estimate has statistical uncertainty extending on both sides of 1.

Why is the confidence interval more informative than the p-value alone?

The confidence interval provides both an effect estimate and a range representing statistical uncertainty around that estimate. The p-value summarizes evidence against the specified null hypothesis; it does not tell the reader how large or clinically important the treatment difference is.

Why doesn't a p-value of 0.341 prove equivalence?

Failure to reject a conventional equality null is not the same as demonstrating that two treatments are sufficiently close. Equivalence and non-inferiority require a prespecified margin and a corresponding confidence-interval decision rule. The ClinicalTrials.gov record identifies the hypothesis type but do not give the numeric margin.

Why were the analyses based on intention-to-treat?

The posted primary analyses use all enrolled patients. An intention-to-treat approach retains patients in the comparison associated with their randomized assignment, preserving the treatment contrast created by randomization rather than redefining treatment groups according to subsequent exposure.

Why is the two-hypothesis structure important?

There are two primary treatment comparisons against the same reference arm. Testing multiple primary hypotheses can increase the chance of a false-positive finding if treated as unrelated tests. The registry specifically states that the critical p-value accounts for correlation between the two primary hypotheses.

15. What the Hazard Ratio Does — and Does Not — Mean

Arm II vs Arm I

The HR of 0.94 is a model-based relative measure of the progression hazard for Arm II compared with Arm I. The point estimate is below 1, but the 95% CI of 0.81–1.09 includes 1.

It does not mean that 94% of patients remained progression-free, that progression probability was 94%, or that every patient experienced the same relative change in progression risk.

Arm III vs Arm I

The HR of 0.99 is the corresponding relative progression-hazard estimate for Arm III compared with Arm I. Its 95% CI of 0.86–1.15 includes 1.

Again, the HR is not an absolute risk, a median, a probability of benefit, or a patient-level prediction.

Why censoring matters

Progression-free survival incorporates follow-up time. A patient who has not experienced progression by the end of observed follow-up cannot automatically be treated as having had an infinitely long progression-free interval. Survival methods account for this through censoring.

Why the proportional-hazards issue matters

The hazard ratio is a relative hazard measure from a time-to-event model. Its interpretation is strongest when the underlying hazard relationship is reasonably represented by the model. The ClinicalTrials.gov record does not report a formal test of the proportional-hazards assumption, so no such diagnostic conclusion is made here.

16. Results vs Statistical Interpretation

Reported result

Arm II vs Arm I produced HR 0.94, 95% CI 0.81–1.09, and P = 0.341 for progression-free survival.

Statistical meaning

The estimated relative progression hazard was close to 1, and the confidence interval included 1.

Reported result

Arm III vs Arm I produced HR 0.99, 95% CI 0.86–1.15, and P = 0.587 for progression-free survival.

Statistical meaning

The estimated relative progression hazard was also close to 1, with the confidence interval extending on both sides of 1.

The distinction between these two levels is important. The reported result is the numerical output of the registered analysis. The statistical interpretation explains what that number represents and what conclusions cannot safely be drawn from it.

17. Important Limitations and Interpretation Issues

18. Why This Trial Matters Statistically

GOG-0252 is a useful statistical teaching case because it combines randomized treatment allocation, a three-arm design, a time-to-event primary endpoint, stratified survival analysis, hazard-ratio estimation, intention-to-treat analysis, and a non-inferiority or equivalence framework.

ConceptHow it appears in GOG-0252
RandomizationThe trial uses randomized allocation.
Parallel designThree treatment arms are evaluated in a parallel design.
Intention-to-treatPrimary analyses use all enrolled patients.
Time-to-event endpointMedian progression-free survival is the registered primary endpoint.
RECISTProgression is defined using RECIST v1.0 criteria in the registry definition.
Log-rank testingPrimary treatment comparisons use log-rank methodology.
Stratified analysisAnalyses are stratified by stage of disease and size of residual disease.
Hazard ratioRelative progression hazards are reported for each primary comparison.
Confidence intervalsBoth primary HR estimates have two-sided 95% confidence intervals.
Non-inferiority / equivalenceThe registry identifies this hypothesis type for both primary analyses.
MultiplicityThe critical p-value accounts for correlation between the two primary hypotheses.
Safety by armSerious adverse events are reported as affected patients divided by patients at risk for each arm.

19. A Statistical Reading of the Primary Results

Viewed strictly as reported, both primary hazard-ratio estimates are close to the null value of 1. Arm II versus Arm I has an HR of 0.94 with a 95% CI of 0.81–1.09, while Arm III versus Arm I has an HR of 0.99 with a 95% CI of 0.86–1.15.

Both confidence intervals cross 1. That observation is descriptive; it should not be transformed into a claim of equivalence. The trial's stated hypothesis type makes the prespecified non-inferiority or equivalence margin central to the formal decision, and that numerical margin is not contained in the ClinicalTrials.gov record.

The two p-values, 0.341 and 0.587, should also be interpreted within the trial's multiple-hypothesis framework. The registry explicitly states that the critical p-value accounts for correlation between the two primary hypotheses. Consequently, replacing that design-specific framework with an assumed generic threshold would discard an important part of the registered statistical design.

There is also a conceptual distinction between saying that the estimated hazard ratios are close to 1 and saying that the treatments are equivalent. The first statement follows directly from the estimates. The second requires a prespecified equivalence or non-inferiority framework and its decision criterion.

20. Serious Adverse Events in Statistical Context

The serious-adverse-event counts provide an arm-level safety perspective alongside the primary efficacy analysis.

Treatment armSerious adverse events
Arm I156/511 affected/at risk
Arm II179/510 affected/at risk
Arm III215/508 affected/at risk

These counts should be kept separate from the hazard-ratio analysis. The progression-free survival HR describes a relative time-to-event treatment comparison, whereas the serious-adverse-event figures describe the number affected among those at risk. Without an event definition, observation period, formal comparison, or confidence interval beyond the ClinicalTrials.gov record, further quantitative inference would go beyond the ClinicalTrials.gov record.

21. What Is Not Reported in the Supplied Trial Data

The registry-reported GOG-0252 data provide formal statistical results for the primary progression-free survival endpoint, but they do not provide numerical results for other posted outcome measures. They also do not provide the numeric non-inferiority margin, Kaplan-Meier median estimates, event counts for the primary endpoint, subgroup hazard ratios, or formal safety p-values.

Accordingly, this analysis does not manufacture those quantities from the available HRs, confidence intervals, enrollment number, or serious-adverse-event counts. Keeping those boundaries explicit is part of a reproducible statistical analysis.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue through Clinical Biostats

Use the trial analysis as a practical entry point into survival analysis, hazard ratios, confidence intervals, intention-to-treat principles, and non-inferiority trial design.

25. Record Summary

GOG-0252 provides a useful example of how a three-arm randomized phase 3 trial can be analyzed when the primary endpoint is progression-free survival and the statistical framework includes stratification, intention-to-treat analysis, log-rank testing, hazard-ratio estimation, and a non-inferiority or equivalence hypothesis. The two posted primary comparisons are Arm II versus Arm I and Arm III versus Arm I.

The reported HR was 0.94 for Arm II versus Arm I, with a two-sided 95% CI of 0.81–1.09 and P = 0.341. For Arm III versus Arm I, the reported HR was 0.99, with a two-sided 95% CI of 0.86–1.15 and P = 0.587. Both analyses used the intention-to-treat population and stratified the comparison by stage of disease and size of residual disease.

The statistical story is more nuanced than simply comparing the two p-values. The registry explicitly describes a non-inferiority or equivalence framework, states that the study was designed around an 80% power assumption involving a 20% reduction in the progression-free survival event rate, and notes that the critical p-value accounts for correlation between the two primary hypotheses. Because the ClinicalTrials.gov record does not state the numerical non-inferiority margin, the formal non-inferiority decision cannot be reconstructed beyond the reported estimates and confidence intervals.

Clinical Biostats methodology: A trial-results page should distinguish the numerical result from the statistical conclusion that can legitimately be drawn from it. For GOG-0252, that means preserving the reported HRs, confidence intervals, p-values, analysis population, stratification factors, and hypothesis framework without converting absence of statistical evidence into a claim of equivalence.