← Clinical Trials
Ovarian Cancer Phase 3 Randomized NCT00976911

AURELIA: Complete Statistical Analysis of Bevacizumab in Platinum-Resistant Ovarian Cancer

An independent statistical review of the randomized phase 3 AURELIA trial evaluating bevacizumab added to chemotherapy in patients with platinum-resistant ovarian cancer, with emphasis on progression-free survival, objective response, overall survival, quality-of-life response, and the statistical methods used in the registry analyses.

Trial start: 29 October 2009  ·  Primary completion: 9 July 2014  ·  Sponsor: Hoffmann-La Roche
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record. The registry provides the official trial record.

Registry record: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View NCT00976911 on ClinicalTrials.gov.

1. Trial at a Glance

AURELIA was a randomized, open-label, parallel phase 3 trial in ovarian cancer. The trial compared chemotherapy alone with chemotherapy plus bevacizumab, using time-to-event and binary outcomes to evaluate progression, survival, objective response, and quality-of-life response.

361
Enrollment
Randomized trial
2
Arms
Parallel design
0.379
Stratified PFS HR
95% CI 0.296–0.485
<0.0001
PFS P-value
Two-sided log-rank analysis
FeatureAURELIA
PhasePhase 3
ConditionOvarian Cancer
DesignRandomized, parallel, open-label
AllocationRandomized
MaskingNone
Primary purposeTreatment
Enrollment361.0
Primary endpointsPercentage of Participants With Disease Progression or Death; Progression Free Survival
Primary endpoint typeTime-to-event
Statistical analyses posted17
ClinicalTrials.govNCT00976911

2. Clinical Question

The central statistical question was whether adding bevacizumab to chemotherapy changed outcomes compared with chemotherapy alone in the randomized trial population with ovarian cancer, particularly with respect to disease progression or death and progression-free survival.

Population

Participants with ovarian cancer enrolled in the phase 3 AURELIA trial. The ClinicalTrials.gov record identifies the condition as ovarian cancer and the intervention set as bevacizumab, liposomal doxorubicin, paclitaxel, and topotecan.

Intervention

Chemotherapy plus bevacizumab. The chemotherapy options identified in the registry data were liposomal doxorubicin, paclitaxel, and topotecan.

Comparator

Chemotherapy without bevacizumab, using the chemotherapy options identified in the trial data.

Primary question

Does adding bevacizumab to chemotherapy change the time-to-event outcomes measuring disease progression or death and progression-free survival?

3. Trial Design

01
Randomize361 participants
02
Two armsChemotherapy vs combination
03
AssessProgression / death / response
04
FollowTime-to-event outcomes
05
CompareStratified and unstratified analyses
ARM A · CHEMOTHERAPY

Chemotherapy

  • Liposomal doxorubicin
  • Paclitaxel
  • Topotecan
ARM B · CHEMOTHERAPY + BEVACIZUMAB

Chemotherapy plus bevacizumab

  • Bevacizumab
  • Liposomal doxorubicin
  • Paclitaxel
  • Topotecan

The registry classifies the trial as randomized, with a parallel design and no masking. The absence of masking is relevant to outcomes such as investigator-assessed tumor progression and patient-reported quality-of-life measures, because knowledge of treatment assignment can potentially affect assessments or reporting even when the statistical analysis itself is formally specified.

4. Endpoints

EndpointRegistry definition / time frameType
Percentage of Participants With Disease Progression or Death
(Data Cutoff 14 November 2011)
Progression free survival was defined as the time from the date of randomization to the first documented disease progression or death, whichever occurs first. Progression was based on tumour assessment made by the investigators according to the Response Evaluation Criteria In Solid Tumors (RECIST) criteria for participants with measurable disease, and for those with non-measurable disease presence or absence of lesions was noted. Assessment occurred at screening and every 8 weeks, or 9 weeks if receiving topotecan, until progression. Time-to-event
Progression Free Survival (PFS; Data Cutoff 14 November 2011) PFS was defined as the time from the date of randomization to the first documented disease progression or death, whichever occurs first. Progression was based on tumor assessment made by the investigators according to RECIST criteria for participants with measurable disease, with presence or absence of lesions noted for non-measurable disease. Time-to-event
Percentage of Participants With Best Overall Confirmed Objective Response of Complete Response (CR) or Partial Response (PR) Per Modified RECIST Screening Visit, Every 8 weeks (or 9 weeks if receiving topotecan) until progression. Binary
Duration of Objective Response Screening Visit, Every 8 weeks (or 9 weeks if receiving topotecan) until progression. Time-to-event
Overall Survival Screening Visit, Every 8 weeks (or 9 weeks if receiving topotecan) until progression; Data Cutoff 25 January 2013. Time-to-event
EORTC QLQ OV28 AB/GI Symptom Scale — Percentage of Responders Baseline and Weeks 8, 9, 16, 18, 24 and 30; Data Cutoff 14 November 2011. Binary
Endpoint distinction: the ClinicalTrials.gov record contains two registered primary endpoints that both concern progression or death/PFS. The formal comparative statistical estimates posted on ClinicalTrials.gov for the primary PFS endpoint are therefore the main primary efficacy results presented below. The ClinicalTrials.gov record does not contain a numerical result estimate for the separately named “Percentage of Participants With Disease Progression or Death” endpoint.

5. Analysis Populations and Stratification

The primary PFS analyses were identified as using the ITT Population, with the registry analysis field additionally stating that only participants with an event of progression or death were included in the analysis. The ClinicalTrials.gov record also identify stratified analyses using chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval.

Analysis featureRegistry information
Primary efficacy populationITT Population
PFS event definitionProgression or death
Stratification factorChemotherapy selected: paclitaxel, PLD, or topotecan
Stratification factorPrior anti-angiogenic therapy: yes or no
Stratification factorPlatinum-free interval: less than 3 or 3-6 months
Primary hypothesisSuperiority
Primary comparative methodLog-rank test

Stratification is important because the trial did not simply compare all observed event times without regard to these prespecified clinical factors. A stratified analysis asks whether the treatment groups differ in event experience while accounting for the specified strata. This can improve alignment between the analysis and the randomized design when those factors are prognostically relevant.

6. Statistical Methodology

Log-rank test

The principal formal analysis of PFS used a log-rank test. This is a standard method for comparing two time-to-event distributions. Rather than comparing only a single summary such as a median, the test uses the ordering of events throughout follow-up while accounting for patients who are censored.

Conceptual comparison
Observed events vs expected events under the null hypothesis

The log-rank framework evaluates whether the pattern of event occurrence differs systematically between randomized groups over the observed follow-up.

Cox regression and hazard ratios

The registry notes that a Cox regression model was used to determine the hazard ratio for the stratified PFS analysis. The registry-reported analysis notes identify the chemotherapy choice, prior anti-angiogenic therapy, and platinum-free interval as the stratification variables.

Hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the bevacizumab group

The hazard ratio is a relative time-to-event measure. It is not a probability, an absolute risk difference, or a statement that the same percentage of individual patients experienced a particular benefit.

Kaplan-Meier estimation

For a time-to-event endpoint such as PFS or overall survival, Kaplan-Meier estimation is the natural descriptive framework for representing the probability of remaining event-free over time while incorporating right-censored observations. The registry-reported statistical-analyses data identify log-rank and Cox methods for the formal comparisons, but do not provide Kaplan-Meier numerical estimates or median event times.

Peto-Peto-Prentice analysis

The registry also reports Peto-Peto-Prentice analyses for PFS, duration of objective response, and overall survival. This is a weighted survival-comparison approach related to the log-rank family. The ClinicalTrials.gov record contains the corresponding p-values for these analyses but do not provide hazard-ratio estimates or confidence intervals for those Peto-Peto-Prentice entries.

Categorical response analysis

Objective response was treated as a binary endpoint: whether a participant achieved a best overall confirmed objective response of complete response or partial response according to modified RECIST. The registry reports Pearson's chi-squared and Cochran-Mantel-Haenszel analyses for this endpoint, as well as a difference in response rates with a 95% confidence interval.

Fisher exact analysis of quality-of-life response

The EORTC QLQ OV28 abdominal/gastrointestinal symptom scale was analyzed as a binary responder outcome at specified follow-up visits. Fisher exact testing was used, and the registry reports confidence intervals approximated with a Hauck-Anderson continuity correction for the response-rate differences.

7. Primary Results: Progression-Free Survival

The formal primary PFS analysis used the ITT population and compared chemotherapy with chemotherapy plus bevacizumab. The registry reports both stratified and unstratified analyses. The primary stratified estimate was obtained using a log-rank comparison with a Cox regression model for the hazard ratio.

Stratified hazard ratio for progression or death

0.379

95% CI: 0.296–0.485   ·   P < 0.0001

Two-sided superiority analysis; log-rank test with stratification by chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval.

Primary PFS analysisEstimate95% CIP-value
Stratified log-rank / Cox analysisHR 0.3790.296–0.485<0.0001
Unstratified log-rank / Cox analysisHR 0.4600.366–0.577<0.0001
Unstratified Peto-Peto-Prentice analysisNot reportedNot reported<0.0001
Stratified Peto-Peto-Prentice analysisNot reportedNot reported<0.0001
Clinical Biostats interpretation

The HR of 0.379 means that, under the fitted stratified time-to-event model, the estimated instantaneous rate of progression or death in the chemotherapy-plus-bevacizumab group was 0.379 times that in the chemotherapy group over the analyzed follow-up. Expressed as a simple relative interpretation, 1 − 0.379 = 0.621, so the estimate corresponds to an approximately 62.1% lower estimated hazard of progression or death.

That does not mean that 62.1% of participants were prevented from progressing, that every patient experienced a 62.1% reduction in risk, or that the probability of progression or death was reduced by exactly 62.1% at every time point.

The 95% CI of 0.296–0.485 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient effects. The fact that the entire interval is below 1 is consistent with a lower estimated hazard in the bevacizumab group under this analysis.

The P-value < 0.0001 addresses evidence against the null hypothesis used for the comparison; it does not measure the size or clinical importance of the effect. The hazard ratio and confidence interval provide the effect-size and precision information.

Interpretation also depends on the time-to-event framework, censoring, the analysis population, and the assumptions underlying the Cox model. In particular, a single hazard ratio is most straightforward to interpret as a common relative effect when proportional hazards are a reasonable approximation over follow-up.

Why the stratified and unstratified estimates differ

The ClinicalTrials.gov record reports a stratified HR of 0.379 and an unstratified HR of 0.460. These are not contradictory results. They answer the same broad treatment-comparison question under different statistical specifications.

The stratified analysis incorporates the three specified strata: chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval. The unstratified analysis does not make that adjustment. A difference between the two estimates can therefore arise because the distribution of event information across these clinical strata affects the fitted comparison.

Important: the difference between 0.379 and 0.460 should not be interpreted as evidence that one analysis is “correct” and the other is “incorrect.” They are distinct prespecified-style statistical representations of the same randomized comparison, and the stratified result directly reflects the stratification factors identified in the registry analysis notes.

8. Primary Endpoint: Percentage of Participants With Disease Progression or Death

The registry lists Percentage of Participants With Disease Progression or Death as a primary endpoint, with a data cutoff of 14 November 2011. Its definition links the outcome to the first documented disease progression or death, whichever occurs first, and describes tumor assessment at screening and every 8 weeks, or 9 weeks for participants receiving topotecan.

What the registry-reported statistical record contains: ClinicalTrials.gov identifies this as a primary time-to-event endpoint, but the registry-reported statistical-analysis entries do not provide a separate numerical estimate, confidence interval, or p-value specifically for this endpoint under its exact registered title. The formal comparative PFS analyses reported above provide the available statistical estimates for the corresponding progression-or-death time-to-event framework.

For an endpoint of this type, a conventional analysis would use time-to-event methods such as Kaplan-Meier estimation for the event-time distribution, a log-rank test for comparison between randomized groups, and a Cox model for a hazard ratio. The ClinicalTrials.gov record specifically document those methods for the PFS primary endpoint.

9. Secondary Results: Objective Response

The secondary objective-response endpoint was the percentage of participants with a best overall confirmed objective response of complete response or partial response according to modified RECIST. The analysis population was the ITT population among participants with measurable disease at baseline.

Difference in response rates

15.7

95% CI: 6.5–24.8   ·   P-value not posted on ClinicalTrials.gov for this estimate

Two-sided superiority comparison; estimate reported as the difference in response rates.

Objective-response analysisEffect estimate95% CIP-value
Difference in response rates15.76.5–24.8Not reported for this estimate
Pearson's chi-squareNot reportedNot reported0.0010
Cochran-Mantel-HaenszelNot reportedNot reported0.0007
Clinical Biostats interpretation

The reported difference in response rates of 15.7 is a percentage-point difference between the two randomized groups. The 95% CI of 6.5–24.8 describes uncertainty around that difference under the reported analysis framework.

This is a different effect measure from the PFS hazard ratio. A response-rate difference asks about the proportion of participants achieving CR or PR, whereas the hazard ratio uses the timing of progression or death. Neither measure can be substituted for the other.

The chi-squared and Cochran-Mantel-Haenszel analyses provide p-values of 0.0010 and 0.0007, respectively. These p-values address the corresponding hypothesis tests; they do not quantify how large the response difference is. The response-rate difference and its confidence interval are the appropriate quantities for describing the magnitude and precision of the reported binary effect.

10. Secondary Results: Duration of Objective Response

Duration of objective response was analyzed among participants whose best overall confirmed response was CR or PR. The ClinicalTrials.gov record reports a log-rank analysis and an unstratified Cox regression hazard ratio.

Hazard ratio for duration of objective response

0.450

95% CI: 0.225–0.900   ·   P = 0.0202

Two-sided superiority analysis; hazard ratio estimated by unstratified Cox regression.

AnalysisEffect estimate95% CIP-value
Log-rank / unstratified CoxHR 0.4500.225–0.9000.0202
Peto-Peto-PrenticeNot reportedNot reported0.0081
Clinical Biostats interpretation

The HR of 0.450 indicates an estimated instantaneous rate of loss of response of 0.450 times that of the comparator group under the fitted unstratified Cox model. In simple relative terms, 1 − 0.450 = 0.550, corresponding to an approximately 55% lower estimated hazard of the response-duration event.

It does not mean that exactly 55% of responders retained their response, nor does it provide the median duration of response. The ClinicalTrials.gov record does not provide median duration of response.

The 95% CI of 0.225–0.900 indicates substantial uncertainty around the estimate. The interval remains below 1, while its breadth shows that the estimate is less precise than the primary PFS estimate. The p-value of 0.0202 addresses the statistical comparison and should not be interpreted as a measure of effect magnitude.

11. Secondary Results: Overall Survival

Overall survival was evaluated with a data cutoff of 25 January 2013. The analysis population was the ITT population, with the registry specifying participants who died as the events included in the analysis.

Overall survival hazard ratios

0.833   /   0.870

95% CI: 0.655–1.059 and 0.678–1.116

P-values: 0.1360 and 0.2711

Overall survival analysisEstimate95% CIP-value
Log-rank / Cox analysisHR 0.8330.655–1.0590.1360
Log-rank / Cox analysisHR 0.8700.678–1.1160.2711
Peto-Peto-Prentice analysisNot reportedNot reported0.0715
Peto-Peto-Prentice analysisNot reportedNot reported0.0890
Clinical Biostats interpretation

An HR of 0.833 corresponds to an estimated instantaneous mortality rate of 0.833 times that of the comparator group under the associated Cox analysis. The corresponding 95% CI, 0.655–1.059, crosses 1. The second reported HR of 0.870 has a 95% CI of 0.678–1.116, which also crosses 1.

These confidence intervals mean that the ClinicalTrials.gov record does not establish a precise mortality effect in the same way as the primary PFS estimate. The p-values of 0.1360 and 0.2711 are hypothesis-test quantities; they should not be used to infer that one treatment has a particular percentage advantage or disadvantage.

The contrast with PFS is statistically instructive: progression-free survival and overall survival measure different events and can produce different treatment-effect estimates. An apparent difference in PFS does not mechanically imply a particular OS hazard ratio.

12. Secondary Results: Quality-of-Life Response

The registry reports responder analyses for the EORTC QLQ OV28 abdominal/gastrointestinal symptom scale at specified follow-up visits. The analysis population consisted of ITT participants who completed the questionnaire at baseline and at the specified visit. Fisher exact testing was used.

ComparisonDifference in response rates95% CIP-value
Baseline versus Week 8/98.8-3.8–21.40.1859
Baseline versus Week 16/183.5-14–20.90.8309
Baseline versus Week 249.3-15–34.10.5790
Baseline versus Week 30-4.8-40–30.60.7339

The registry notes that the 95% confidence intervals were approximated using a Hauck-Anderson continuity correction. This is important because the intervals are not simply unqualified arithmetic transformations of the point estimates; the registry identifies the correction as part of the reported confidence-interval methodology.

Clinical Biostats interpretation

The response-rate differences vary across the specified visits, and all four registry-reported confidence intervals include zero. The p-values are likewise not small under the conventional two-sided hypothesis-testing framework. These results should be read as the registry-reported visit-specific comparisons rather than as a single longitudinal summary of quality of life.

The use of Fisher exact testing is appropriate for binary responder comparisons when the assumptions underlying large-sample chi-squared approximations may be less reliable. However, the ClinicalTrials.gov record does not provide the underlying responder counts for each visit, so the numerical results should not be reverse-engineered into counts.

13. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by participants at risk.

Safety measureChemotherapyChemotherapy + Bevacizumab
Serious adverse events49 / 18156 / 179
Serious adverse events — affected / at risk
Chemotherapy
49 / 181
+ Bevacizumab
56 / 179

The bar display above is a visual representation of the registry-reported affected/at-risk fractions. The underlying reported numbers remain 49/181 and 56/179; no additional safety event categories are inferred.

Safety interpretation: these are serious-adverse-event counts by arm, not a complete safety profile. The ClinicalTrials.gov record does not provide a formal between-arm p-value or confidence interval for this safety measure, so no comparative statistical significance claim should be attached to these counts alone.

14. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint because the analysis records not only whether progression or death occurred, but also when it occurred. Some participants may be censored before experiencing the event. The log-rank test is designed to compare two such survival-type distributions while incorporating event timing and censoring.

What does an HR of 0.379 mean?

It is a model-based relative measure of the instantaneous progression-or-death rate. Under the reported stratified Cox analysis, the estimated hazard in the bevacizumab group was 0.379 times the hazard in the chemotherapy group. It does not mean that 62.1% of patients avoided progression or death, and it is not an absolute risk reduction.

Why use stratified analysis?

The registry identifies three strata: chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval. Stratification allows the time-to-event comparison to account for these specified clinical factors rather than treating all participants as if they belonged to one homogeneous risk set.

Why is the unstratified HR different from the stratified HR?

The unstratified estimate of 0.460 and stratified estimate of 0.379 come from different statistical specifications. The stratified analysis incorporates the specified clinical strata; the unstratified analysis does not. Differences between them therefore do not automatically indicate a data problem.

Why use a Cochran-Mantel-Haenszel test for objective response?

The Cochran-Mantel-Haenszel framework can compare binary outcomes across treatment groups while accounting for defined strata. In AURELIA, the registry separately reports a Cochran-Mantel-Haenszel p-value of 0.0007 for the objective-response endpoint. The ClinicalTrials.gov record does not provide a corresponding effect estimate for that specific test entry.

Why use Fisher exact testing for the quality-of-life responder endpoint?

The quality-of-life outcome was converted into a binary responder measure and compared between treatment groups. Fisher exact testing calculates an exact probability under the null hypothesis and is useful for categorical comparisons when reliance on large-sample approximations may be undesirable.

Why can the PFS and OS results differ?

PFS and OS measure different events. PFS records progression or death, whereas OS records death. A treatment effect on disease progression does not impose a fixed mathematical relationship on the subsequent mortality comparison. The AURELIA registry data illustrate this distinction directly through the different reported PFS and OS hazard ratios.

15. Confidence Intervals and P-values

Confidence interval

A confidence interval describes uncertainty around an estimated effect under the specified statistical framework. For example, the primary PFS HR of 0.379 has a 95% CI of 0.296–0.485.

P-value

A p-value measures compatibility of the observed data with the null hypothesis used in the specified test. It does not quantify effect size, clinical importance, or the probability that the treatment works.

Relative effect

The hazard ratio is relative. It should not be translated into an absolute percentage of patients benefiting without additional absolute-risk information.

Binary effect

The objective-response estimate of 15.7 is a difference in response rates, which is naturally interpreted in percentage points rather than as a hazard ratio.

16. Primary PFS Result in Statistical Context

The primary PFS result is particularly useful for teaching because the registry provides several analyses of the same endpoint. The estimates are directionally consistent: both reported Cox hazard ratios are below 1, and all four reported PFS hypothesis tests have p-values <0.0001.

MethodStratificationEffect measureEstimate95% CIP-value
Log-rank / CoxStratifiedHazard ratio0.3790.296–0.485<0.0001
Log-rank / CoxUnstratifiedHazard ratio0.4600.366–0.577<0.0001
Peto-Peto-PrenticeUnstratifiedNot reportedNot reportedNot reported<0.0001
Peto-Peto-PrenticeStratifiedNot reportedNot reportedNot reported<0.0001

This is a useful example of why a clinical trial result should not be reduced to a single p-value. The analysis contains the endpoint definition, analysis population, treatment comparison, stratification structure, statistical test, effect measure, point estimate, and confidence interval. Each component answers a different part of the statistical question.

17. Understanding the Analysis Population

The registry identifies the primary PFS analysis population as ITT. The ITT principle preserves the treatment assignment created by randomization when comparing randomized groups. This is important because the principal causal contrast in a randomized trial is the difference associated with assignment to the treatment strategy rather than a comparison created after selectively removing participants.

At the same time, the registry's analysis-population field states that only participants with an event of progression or death were included in the analysis. This wording is specific to the ClinicalTrials.gov record and should not be silently rewritten into a different population definition.

Statistical reading rule: analysis-population labels matter. When interpreting a reported estimate, the denominator and eligibility criteria for that estimate are part of the result itself, not merely administrative metadata.

18. Time Frames and Data Cutoffs

29 October 2009

Trial start

The ClinicalTrials.gov record identifies this as the trial start date.

14 November 2011

PFS, response, and quality-of-life cutoff

The primary PFS and disease-progression/death endpoints, objective-response analyses, duration-of-response analyses, and EORTC QLQ OV28 responder analyses use this data cutoff in the registry-reported statistical records.

25 January 2013

Overall survival cutoff

The registry-reported overall-survival analyses use this separate data cutoff.

9 July 2014

Primary completion

the ClinicalTrials.gov record identifies this as the primary completion date.

Keeping the data cutoffs separate is essential. A PFS estimate based on 14 November 2011 should not be silently presented as though it came from the later overall-survival cutoff of 25 January 2013.

19. What the Hazard Ratio Does — and Does Not — Mean

Primary PFS interpretation

The reported stratified PFS HR of 0.379 indicates a lower estimated instantaneous rate of progression or death in the chemotherapy-plus-bevacizumab group under the fitted model.

Because 0.379 is below 1, the reciprocal comparison would point in the opposite direction if the reference group were reversed. The hazard ratio is therefore inherently tied to which treatment is listed first and which group is the reference.

Why the confidence interval matters

The 95% CI of 0.296–0.485 provides the uncertainty interval for the primary hazard-ratio estimate. Its lower and upper bounds are estimates of statistical precision, not predictions for individual patients.

Why the p-value is not the effect size

The primary p-value is <0.0001. That tells us about the evidence against the specified null hypothesis under the log-rank analysis. It does not tell us that the effect is “0.0001 large,” nor does it quantify the magnitude of treatment benefit.

20. Multiplicity and Multiple Analyses

The ClinicalTrials.gov record contains 17 statistical analyses, including multiple analyses of the same endpoints using different methods and specifications. The data identify the hypothesis type as superiority.

Multiple reported analyses create an important interpretive distinction. A p-value attached to an individual analysis is not automatically a familywise-error-adjusted p-value covering every statistical test reported for the trial. The ClinicalTrials.gov record does not provide a multiplicity-adjustment scheme or alpha-allocation procedure, so no additional correction should be inferred.

Endpoint / analysis familyMethods reportedEffect information reported
Primary PFSLog-rank; Peto-Peto-PrenticeHR estimates posted on ClinicalTrials.gov for log-rank/Cox analyses
Objective responsePearson's chi-square; Cochran-Mantel-HaenszelResponse-rate difference reported; p-values registry-reported separately
Duration of responseLog-rank; Peto-Peto-PrenticeHR 0.450 posted on ClinicalTrials.gov for log-rank/Cox analysis
Overall survivalLog-rank; Peto-Peto-PrenticeTwo HR estimates posted on ClinicalTrials.gov for log-rank/Cox analyses
Quality-of-life responder endpointFisher exactResponse-rate differences and confidence intervals reported

The appropriate educational lesson is not that multiple analyses invalidate the trial. Rather, each reported test must be interpreted according to its endpoint, analysis population, statistical method, and role in the trial's prespecified testing structure. The ClinicalTrials.gov record does not provide enough information to reconstruct an overall multiplicity hierarchy.

21. Missing Data, Censoring, and What Is Not Reported

Time-to-event analysis inherently involves censoring when participants have not experienced the specified event by the time their follow-up ends or becomes unavailable. The ClinicalTrials.gov record provides the endpoint definitions and statistical methods, but they do not provide a complete censoring table, missing-data strategy, imputation algorithm, or detailed rules for every potential censoring scenario.

Similarly, the quality-of-life endpoint specifies that the analysis population consists of participants who completed the questionnaire at baseline and at the specified visit. This means the analysis is conditioned on completion at those time points. The ClinicalTrials.gov record does not provide enough information to quantify how questionnaire completion differed between treatment groups or whether missingness was handled with imputation.

Data boundary: because the ClinicalTrials.gov record does not report a formal missing-data or imputation strategy, none is assumed here. The absence of a registry-reported method should not be filled with a generic claim that a particular imputation procedure was used.

22. Stratified Analysis and Clinical Interpretation

The AURELIA registry analyses provide a useful example of why stratification is more than a cosmetic statistical adjustment. The primary PFS analysis identifies three clinical variables for stratification:

These variables define the strata through which the primary stratified comparison is constructed. The purpose is to make the treatment comparison conditional on these specified categories rather than treating all participants as belonging to one undifferentiated risk set.

Educational principle
Stratified comparison = treatment contrast evaluated within the specified strata

The exact statistical implementation depends on the model and test used. The ClinicalTrials.gov record specifically identify stratified log-rank/Peto-Peto-Prentice analyses and a Cox regression model for the primary PFS hazard ratio.

23. Limitations

24. Why This Trial Matters Statistically

AURELIA is a useful teaching case because its registry results connect several core clinical-trial methods in one randomized phase 3 analysis. The trial uses time-to-event methodology for PFS, duration of response, and OS; categorical methods for objective response; exact testing for quality-of-life responder comparisons; and stratification to account for prespecified clinical factors.

ConceptHow it appears in AURELIA
Randomization361 participants allocated in a randomized parallel phase 3 trial
ITT analysisPrimary PFS and several secondary analyses identify the ITT population
Time-to-event endpointsPFS, progression/death, duration of response, and OS
Log-rank testingPrimary PFS, duration of response, and OS analyses
Hazard ratioPrimary PFS, duration of response, and OS effect measures
Cox regressionUsed to determine the reported hazard ratios
Stratified analysisPrimary PFS analysis accounts for chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval
Chi-squared testingObjective-response analysis
Cochran-Mantel-Haenszel testingObjective-response analysis
Fisher exact testingEORTC QLQ OV28 responder analyses
Confidence intervalsReported for hazard ratios, response-rate differences, and quality-of-life response differences
Multiple analyses17 statistical analyses posted in the ClinicalTrials.gov record

25. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through Clinical Biostats

Use the related tutorials and statistical calculators to explore the methods behind randomized clinical-trial endpoints, survival analysis, categorical comparisons, confidence intervals, and stratified testing.

28. Record Summary

AURELIA provides a compact example of how several statistical methods work together in a randomized oncology trial. The primary PFS analysis used an ITT population, a two-sided log-rank comparison, stratification by chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval, and a Cox regression model yielding a hazard ratio of 0.379 with a 95% CI of 0.296–0.485 and P < 0.0001. The corresponding unstratified HR was 0.460 with a 95% CI of 0.366–0.577 and P < 0.0001.

The secondary analyses illustrate why effect measures must match endpoint type. Objective response was summarized with a difference in response rates of 15.7 and a 95% CI of 6.5–24.8, with chi-squared and Cochran-Mantel-Haenszel p-values of 0.0010 and 0.0007. Duration of objective response produced an HR of 0.450 with a 95% CI of 0.225–0.900 and P = 0.0202. Overall survival analyses produced HR estimates of 0.833 and 0.870, with confidence intervals crossing 1 and p-values of 0.1360 and 0.2711.

The quality-of-life responder analyses further demonstrate how categorical methods can be used at multiple follow-up visits, while the serious-adverse-event data show why efficacy and safety should be examined as separate statistical questions. Most importantly, the page illustrates a central principle of clinical-trial interpretation: the estimate, confidence interval, p-value, endpoint definition, analysis population, statistical model, and data cutoff all belong to the result.

Clinical Biostats methodology: A trial-results page should not merely repeat a headline result. The goal is to reconstruct the statistical story using the reported endpoint definitions, analysis populations, methods, effect measures, uncertainty intervals, and data cutoffs while clearly separating documented evidence from statistical interpretation.