This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record. The registry provides the official trial record.
1. Trial at a Glance
AURELIA was a randomized, open-label, parallel phase 3 trial in ovarian cancer. The trial compared chemotherapy alone with chemotherapy plus bevacizumab, using time-to-event and binary outcomes to evaluate progression, survival, objective response, and quality-of-life response.
| Feature | AURELIA |
|---|---|
| Phase | Phase 3 |
| Condition | Ovarian Cancer |
| Design | Randomized, parallel, open-label |
| Allocation | Randomized |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 361.0 |
| Primary endpoints | Percentage of Participants With Disease Progression or Death; Progression Free Survival |
| Primary endpoint type | Time-to-event |
| Statistical analyses posted | 17 |
| ClinicalTrials.gov | NCT00976911 |
2. Clinical Question
The central statistical question was whether adding bevacizumab to chemotherapy changed outcomes compared with chemotherapy alone in the randomized trial population with ovarian cancer, particularly with respect to disease progression or death and progression-free survival.
Population
Participants with ovarian cancer enrolled in the phase 3 AURELIA trial. The ClinicalTrials.gov record identifies the condition as ovarian cancer and the intervention set as bevacizumab, liposomal doxorubicin, paclitaxel, and topotecan.
Intervention
Chemotherapy plus bevacizumab. The chemotherapy options identified in the registry data were liposomal doxorubicin, paclitaxel, and topotecan.
Comparator
Chemotherapy without bevacizumab, using the chemotherapy options identified in the trial data.
Primary question
Does adding bevacizumab to chemotherapy change the time-to-event outcomes measuring disease progression or death and progression-free survival?
3. Trial Design
Chemotherapy
- Liposomal doxorubicin
- Paclitaxel
- Topotecan
Chemotherapy plus bevacizumab
- Bevacizumab
- Liposomal doxorubicin
- Paclitaxel
- Topotecan
The registry classifies the trial as randomized, with a parallel design and no masking. The absence of masking is relevant to outcomes such as investigator-assessed tumor progression and patient-reported quality-of-life measures, because knowledge of treatment assignment can potentially affect assessments or reporting even when the statistical analysis itself is formally specified.
4. Endpoints
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Percentage of Participants With Disease Progression or Death (Data Cutoff 14 November 2011) |
Progression free survival was defined as the time from the date of randomization to the first documented disease progression or death, whichever occurs first. Progression was based on tumour assessment made by the investigators according to the Response Evaluation Criteria In Solid Tumors (RECIST) criteria for participants with measurable disease, and for those with non-measurable disease presence or absence of lesions was noted. Assessment occurred at screening and every 8 weeks, or 9 weeks if receiving topotecan, until progression. | Time-to-event |
| Progression Free Survival (PFS; Data Cutoff 14 November 2011) | PFS was defined as the time from the date of randomization to the first documented disease progression or death, whichever occurs first. Progression was based on tumor assessment made by the investigators according to RECIST criteria for participants with measurable disease, with presence or absence of lesions noted for non-measurable disease. | Time-to-event |
| Percentage of Participants With Best Overall Confirmed Objective Response of Complete Response (CR) or Partial Response (PR) Per Modified RECIST | Screening Visit, Every 8 weeks (or 9 weeks if receiving topotecan) until progression. | Binary |
| Duration of Objective Response | Screening Visit, Every 8 weeks (or 9 weeks if receiving topotecan) until progression. | Time-to-event |
| Overall Survival | Screening Visit, Every 8 weeks (or 9 weeks if receiving topotecan) until progression; Data Cutoff 25 January 2013. | Time-to-event |
| EORTC QLQ OV28 AB/GI Symptom Scale — Percentage of Responders | Baseline and Weeks 8, 9, 16, 18, 24 and 30; Data Cutoff 14 November 2011. | Binary |
5. Analysis Populations and Stratification
The primary PFS analyses were identified as using the ITT Population, with the registry analysis field additionally stating that only participants with an event of progression or death were included in the analysis. The ClinicalTrials.gov record also identify stratified analyses using chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval.
| Analysis feature | Registry information |
|---|---|
| Primary efficacy population | ITT Population |
| PFS event definition | Progression or death |
| Stratification factor | Chemotherapy selected: paclitaxel, PLD, or topotecan |
| Stratification factor | Prior anti-angiogenic therapy: yes or no |
| Stratification factor | Platinum-free interval: less than 3 or 3-6 months |
| Primary hypothesis | Superiority |
| Primary comparative method | Log-rank test |
Stratification is important because the trial did not simply compare all observed event times without regard to these prespecified clinical factors. A stratified analysis asks whether the treatment groups differ in event experience while accounting for the specified strata. This can improve alignment between the analysis and the randomized design when those factors are prognostically relevant.
6. Statistical Methodology
Log-rank test
The principal formal analysis of PFS used a log-rank test. This is a standard method for comparing two time-to-event distributions. Rather than comparing only a single summary such as a median, the test uses the ordering of events throughout follow-up while accounting for patients who are censored.
The log-rank framework evaluates whether the pattern of event occurrence differs systematically between randomized groups over the observed follow-up.
Cox regression and hazard ratios
The registry notes that a Cox regression model was used to determine the hazard ratio for the stratified PFS analysis. The registry-reported analysis notes identify the chemotherapy choice, prior anti-angiogenic therapy, and platinum-free interval as the stratification variables.
The hazard ratio is a relative time-to-event measure. It is not a probability, an absolute risk difference, or a statement that the same percentage of individual patients experienced a particular benefit.
Kaplan-Meier estimation
For a time-to-event endpoint such as PFS or overall survival, Kaplan-Meier estimation is the natural descriptive framework for representing the probability of remaining event-free over time while incorporating right-censored observations. The registry-reported statistical-analyses data identify log-rank and Cox methods for the formal comparisons, but do not provide Kaplan-Meier numerical estimates or median event times.
Peto-Peto-Prentice analysis
The registry also reports Peto-Peto-Prentice analyses for PFS, duration of objective response, and overall survival. This is a weighted survival-comparison approach related to the log-rank family. The ClinicalTrials.gov record contains the corresponding p-values for these analyses but do not provide hazard-ratio estimates or confidence intervals for those Peto-Peto-Prentice entries.
Categorical response analysis
Objective response was treated as a binary endpoint: whether a participant achieved a best overall confirmed objective response of complete response or partial response according to modified RECIST. The registry reports Pearson's chi-squared and Cochran-Mantel-Haenszel analyses for this endpoint, as well as a difference in response rates with a 95% confidence interval.
Fisher exact analysis of quality-of-life response
The EORTC QLQ OV28 abdominal/gastrointestinal symptom scale was analyzed as a binary responder outcome at specified follow-up visits. Fisher exact testing was used, and the registry reports confidence intervals approximated with a Hauck-Anderson continuity correction for the response-rate differences.
7. Primary Results: Progression-Free Survival
The formal primary PFS analysis used the ITT population and compared chemotherapy with chemotherapy plus bevacizumab. The registry reports both stratified and unstratified analyses. The primary stratified estimate was obtained using a log-rank comparison with a Cox regression model for the hazard ratio.
Stratified hazard ratio for progression or death
95% CI: 0.296–0.485 · P < 0.0001
Two-sided superiority analysis; log-rank test with stratification by chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval.
| Primary PFS analysis | Estimate | 95% CI | P-value |
|---|---|---|---|
| Stratified log-rank / Cox analysis | HR 0.379 | 0.296–0.485 | <0.0001 |
| Unstratified log-rank / Cox analysis | HR 0.460 | 0.366–0.577 | <0.0001 |
| Unstratified Peto-Peto-Prentice analysis | Not reported | Not reported | <0.0001 |
| Stratified Peto-Peto-Prentice analysis | Not reported | Not reported | <0.0001 |
The HR of 0.379 means that, under the fitted stratified time-to-event model, the estimated instantaneous rate of progression or death in the chemotherapy-plus-bevacizumab group was 0.379 times that in the chemotherapy group over the analyzed follow-up. Expressed as a simple relative interpretation, 1 − 0.379 = 0.621, so the estimate corresponds to an approximately 62.1% lower estimated hazard of progression or death.
That does not mean that 62.1% of participants were prevented from progressing, that every patient experienced a 62.1% reduction in risk, or that the probability of progression or death was reduced by exactly 62.1% at every time point.
The 95% CI of 0.296–0.485 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient effects. The fact that the entire interval is below 1 is consistent with a lower estimated hazard in the bevacizumab group under this analysis.
The P-value < 0.0001 addresses evidence against the null hypothesis used for the comparison; it does not measure the size or clinical importance of the effect. The hazard ratio and confidence interval provide the effect-size and precision information.
Interpretation also depends on the time-to-event framework, censoring, the analysis population, and the assumptions underlying the Cox model. In particular, a single hazard ratio is most straightforward to interpret as a common relative effect when proportional hazards are a reasonable approximation over follow-up.
Why the stratified and unstratified estimates differ
The ClinicalTrials.gov record reports a stratified HR of 0.379 and an unstratified HR of 0.460. These are not contradictory results. They answer the same broad treatment-comparison question under different statistical specifications.
The stratified analysis incorporates the three specified strata: chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval. The unstratified analysis does not make that adjustment. A difference between the two estimates can therefore arise because the distribution of event information across these clinical strata affects the fitted comparison.
8. Primary Endpoint: Percentage of Participants With Disease Progression or Death
The registry lists Percentage of Participants With Disease Progression or Death as a primary endpoint, with a data cutoff of 14 November 2011. Its definition links the outcome to the first documented disease progression or death, whichever occurs first, and describes tumor assessment at screening and every 8 weeks, or 9 weeks for participants receiving topotecan.
For an endpoint of this type, a conventional analysis would use time-to-event methods such as Kaplan-Meier estimation for the event-time distribution, a log-rank test for comparison between randomized groups, and a Cox model for a hazard ratio. The ClinicalTrials.gov record specifically document those methods for the PFS primary endpoint.
9. Secondary Results: Objective Response
The secondary objective-response endpoint was the percentage of participants with a best overall confirmed objective response of complete response or partial response according to modified RECIST. The analysis population was the ITT population among participants with measurable disease at baseline.
Difference in response rates
95% CI: 6.5–24.8 · P-value not posted on ClinicalTrials.gov for this estimate
Two-sided superiority comparison; estimate reported as the difference in response rates.
| Objective-response analysis | Effect estimate | 95% CI | P-value |
|---|---|---|---|
| Difference in response rates | 15.7 | 6.5–24.8 | Not reported for this estimate |
| Pearson's chi-square | Not reported | Not reported | 0.0010 |
| Cochran-Mantel-Haenszel | Not reported | Not reported | 0.0007 |
The reported difference in response rates of 15.7 is a percentage-point difference between the two randomized groups. The 95% CI of 6.5–24.8 describes uncertainty around that difference under the reported analysis framework.
This is a different effect measure from the PFS hazard ratio. A response-rate difference asks about the proportion of participants achieving CR or PR, whereas the hazard ratio uses the timing of progression or death. Neither measure can be substituted for the other.
The chi-squared and Cochran-Mantel-Haenszel analyses provide p-values of 0.0010 and 0.0007, respectively. These p-values address the corresponding hypothesis tests; they do not quantify how large the response difference is. The response-rate difference and its confidence interval are the appropriate quantities for describing the magnitude and precision of the reported binary effect.
10. Secondary Results: Duration of Objective Response
Duration of objective response was analyzed among participants whose best overall confirmed response was CR or PR. The ClinicalTrials.gov record reports a log-rank analysis and an unstratified Cox regression hazard ratio.
Hazard ratio for duration of objective response
95% CI: 0.225–0.900 · P = 0.0202
Two-sided superiority analysis; hazard ratio estimated by unstratified Cox regression.
| Analysis | Effect estimate | 95% CI | P-value |
|---|---|---|---|
| Log-rank / unstratified Cox | HR 0.450 | 0.225–0.900 | 0.0202 |
| Peto-Peto-Prentice | Not reported | Not reported | 0.0081 |
The HR of 0.450 indicates an estimated instantaneous rate of loss of response of 0.450 times that of the comparator group under the fitted unstratified Cox model. In simple relative terms, 1 − 0.450 = 0.550, corresponding to an approximately 55% lower estimated hazard of the response-duration event.
It does not mean that exactly 55% of responders retained their response, nor does it provide the median duration of response. The ClinicalTrials.gov record does not provide median duration of response.
The 95% CI of 0.225–0.900 indicates substantial uncertainty around the estimate. The interval remains below 1, while its breadth shows that the estimate is less precise than the primary PFS estimate. The p-value of 0.0202 addresses the statistical comparison and should not be interpreted as a measure of effect magnitude.
11. Secondary Results: Overall Survival
Overall survival was evaluated with a data cutoff of 25 January 2013. The analysis population was the ITT population, with the registry specifying participants who died as the events included in the analysis.
Overall survival hazard ratios
95% CI: 0.655–1.059 and 0.678–1.116
P-values: 0.1360 and 0.2711
| Overall survival analysis | Estimate | 95% CI | P-value |
|---|---|---|---|
| Log-rank / Cox analysis | HR 0.833 | 0.655–1.059 | 0.1360 |
| Log-rank / Cox analysis | HR 0.870 | 0.678–1.116 | 0.2711 |
| Peto-Peto-Prentice analysis | Not reported | Not reported | 0.0715 |
| Peto-Peto-Prentice analysis | Not reported | Not reported | 0.0890 |
An HR of 0.833 corresponds to an estimated instantaneous mortality rate of 0.833 times that of the comparator group under the associated Cox analysis. The corresponding 95% CI, 0.655–1.059, crosses 1. The second reported HR of 0.870 has a 95% CI of 0.678–1.116, which also crosses 1.
These confidence intervals mean that the ClinicalTrials.gov record does not establish a precise mortality effect in the same way as the primary PFS estimate. The p-values of 0.1360 and 0.2711 are hypothesis-test quantities; they should not be used to infer that one treatment has a particular percentage advantage or disadvantage.
The contrast with PFS is statistically instructive: progression-free survival and overall survival measure different events and can produce different treatment-effect estimates. An apparent difference in PFS does not mechanically imply a particular OS hazard ratio.
12. Secondary Results: Quality-of-Life Response
The registry reports responder analyses for the EORTC QLQ OV28 abdominal/gastrointestinal symptom scale at specified follow-up visits. The analysis population consisted of ITT participants who completed the questionnaire at baseline and at the specified visit. Fisher exact testing was used.
| Comparison | Difference in response rates | 95% CI | P-value |
|---|---|---|---|
| Baseline versus Week 8/9 | 8.8 | -3.8–21.4 | 0.1859 |
| Baseline versus Week 16/18 | 3.5 | -14–20.9 | 0.8309 |
| Baseline versus Week 24 | 9.3 | -15–34.1 | 0.5790 |
| Baseline versus Week 30 | -4.8 | -40–30.6 | 0.7339 |
The registry notes that the 95% confidence intervals were approximated using a Hauck-Anderson continuity correction. This is important because the intervals are not simply unqualified arithmetic transformations of the point estimates; the registry identifies the correction as part of the reported confidence-interval methodology.
The response-rate differences vary across the specified visits, and all four registry-reported confidence intervals include zero. The p-values are likewise not small under the conventional two-sided hypothesis-testing framework. These results should be read as the registry-reported visit-specific comparisons rather than as a single longitudinal summary of quality of life.
The use of Fisher exact testing is appropriate for binary responder comparisons when the assumptions underlying large-sample chi-squared approximations may be less reliable. However, the ClinicalTrials.gov record does not provide the underlying responder counts for each visit, so the numerical results should not be reverse-engineered into counts.
13. Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by participants at risk.
| Safety measure | Chemotherapy | Chemotherapy + Bevacizumab |
|---|---|---|
| Serious adverse events | 49 / 181 | 56 / 179 |
The bar display above is a visual representation of the registry-reported affected/at-risk fractions. The underlying reported numbers remain 49/181 and 56/179; no additional safety event categories are inferred.
14. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS is a time-to-event endpoint because the analysis records not only whether progression or death occurred, but also when it occurred. Some participants may be censored before experiencing the event. The log-rank test is designed to compare two such survival-type distributions while incorporating event timing and censoring.
What does an HR of 0.379 mean?
It is a model-based relative measure of the instantaneous progression-or-death rate. Under the reported stratified Cox analysis, the estimated hazard in the bevacizumab group was 0.379 times the hazard in the chemotherapy group. It does not mean that 62.1% of patients avoided progression or death, and it is not an absolute risk reduction.
Why use stratified analysis?
The registry identifies three strata: chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval. Stratification allows the time-to-event comparison to account for these specified clinical factors rather than treating all participants as if they belonged to one homogeneous risk set.
Why is the unstratified HR different from the stratified HR?
The unstratified estimate of 0.460 and stratified estimate of 0.379 come from different statistical specifications. The stratified analysis incorporates the specified clinical strata; the unstratified analysis does not. Differences between them therefore do not automatically indicate a data problem.
Why use a Cochran-Mantel-Haenszel test for objective response?
The Cochran-Mantel-Haenszel framework can compare binary outcomes across treatment groups while accounting for defined strata. In AURELIA, the registry separately reports a Cochran-Mantel-Haenszel p-value of 0.0007 for the objective-response endpoint. The ClinicalTrials.gov record does not provide a corresponding effect estimate for that specific test entry.
Why use Fisher exact testing for the quality-of-life responder endpoint?
The quality-of-life outcome was converted into a binary responder measure and compared between treatment groups. Fisher exact testing calculates an exact probability under the null hypothesis and is useful for categorical comparisons when reliance on large-sample approximations may be undesirable.
Why can the PFS and OS results differ?
PFS and OS measure different events. PFS records progression or death, whereas OS records death. A treatment effect on disease progression does not impose a fixed mathematical relationship on the subsequent mortality comparison. The AURELIA registry data illustrate this distinction directly through the different reported PFS and OS hazard ratios.
15. Confidence Intervals and P-values
Confidence interval
A confidence interval describes uncertainty around an estimated effect under the specified statistical framework. For example, the primary PFS HR of 0.379 has a 95% CI of 0.296–0.485.
P-value
A p-value measures compatibility of the observed data with the null hypothesis used in the specified test. It does not quantify effect size, clinical importance, or the probability that the treatment works.
Relative effect
The hazard ratio is relative. It should not be translated into an absolute percentage of patients benefiting without additional absolute-risk information.
Binary effect
The objective-response estimate of 15.7 is a difference in response rates, which is naturally interpreted in percentage points rather than as a hazard ratio.
16. Primary PFS Result in Statistical Context
The primary PFS result is particularly useful for teaching because the registry provides several analyses of the same endpoint. The estimates are directionally consistent: both reported Cox hazard ratios are below 1, and all four reported PFS hypothesis tests have p-values <0.0001.
| Method | Stratification | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Log-rank / Cox | Stratified | Hazard ratio | 0.379 | 0.296–0.485 | <0.0001 |
| Log-rank / Cox | Unstratified | Hazard ratio | 0.460 | 0.366–0.577 | <0.0001 |
| Peto-Peto-Prentice | Unstratified | Not reported | Not reported | Not reported | <0.0001 |
| Peto-Peto-Prentice | Stratified | Not reported | Not reported | Not reported | <0.0001 |
This is a useful example of why a clinical trial result should not be reduced to a single p-value. The analysis contains the endpoint definition, analysis population, treatment comparison, stratification structure, statistical test, effect measure, point estimate, and confidence interval. Each component answers a different part of the statistical question.
17. Understanding the Analysis Population
The registry identifies the primary PFS analysis population as ITT. The ITT principle preserves the treatment assignment created by randomization when comparing randomized groups. This is important because the principal causal contrast in a randomized trial is the difference associated with assignment to the treatment strategy rather than a comparison created after selectively removing participants.
At the same time, the registry's analysis-population field states that only participants with an event of progression or death were included in the analysis. This wording is specific to the ClinicalTrials.gov record and should not be silently rewritten into a different population definition.
18. Time Frames and Data Cutoffs
Trial start
The ClinicalTrials.gov record identifies this as the trial start date.
PFS, response, and quality-of-life cutoff
The primary PFS and disease-progression/death endpoints, objective-response analyses, duration-of-response analyses, and EORTC QLQ OV28 responder analyses use this data cutoff in the registry-reported statistical records.
Overall survival cutoff
The registry-reported overall-survival analyses use this separate data cutoff.
Primary completion
the ClinicalTrials.gov record identifies this as the primary completion date.
Keeping the data cutoffs separate is essential. A PFS estimate based on 14 November 2011 should not be silently presented as though it came from the later overall-survival cutoff of 25 January 2013.
19. What the Hazard Ratio Does — and Does Not — Mean
The reported stratified PFS HR of 0.379 indicates a lower estimated instantaneous rate of progression or death in the chemotherapy-plus-bevacizumab group under the fitted model.
Because 0.379 is below 1, the reciprocal comparison would point in the opposite direction if the reference group were reversed. The hazard ratio is therefore inherently tied to which treatment is listed first and which group is the reference.
The 95% CI of 0.296–0.485 provides the uncertainty interval for the primary hazard-ratio estimate. Its lower and upper bounds are estimates of statistical precision, not predictions for individual patients.
The primary p-value is <0.0001. That tells us about the evidence against the specified null hypothesis under the log-rank analysis. It does not tell us that the effect is “0.0001 large,” nor does it quantify the magnitude of treatment benefit.
20. Multiplicity and Multiple Analyses
The ClinicalTrials.gov record contains 17 statistical analyses, including multiple analyses of the same endpoints using different methods and specifications. The data identify the hypothesis type as superiority.
Multiple reported analyses create an important interpretive distinction. A p-value attached to an individual analysis is not automatically a familywise-error-adjusted p-value covering every statistical test reported for the trial. The ClinicalTrials.gov record does not provide a multiplicity-adjustment scheme or alpha-allocation procedure, so no additional correction should be inferred.
| Endpoint / analysis family | Methods reported | Effect information reported |
|---|---|---|
| Primary PFS | Log-rank; Peto-Peto-Prentice | HR estimates posted on ClinicalTrials.gov for log-rank/Cox analyses |
| Objective response | Pearson's chi-square; Cochran-Mantel-Haenszel | Response-rate difference reported; p-values registry-reported separately |
| Duration of response | Log-rank; Peto-Peto-Prentice | HR 0.450 posted on ClinicalTrials.gov for log-rank/Cox analysis |
| Overall survival | Log-rank; Peto-Peto-Prentice | Two HR estimates posted on ClinicalTrials.gov for log-rank/Cox analyses |
| Quality-of-life responder endpoint | Fisher exact | Response-rate differences and confidence intervals reported |
The appropriate educational lesson is not that multiple analyses invalidate the trial. Rather, each reported test must be interpreted according to its endpoint, analysis population, statistical method, and role in the trial's prespecified testing structure. The ClinicalTrials.gov record does not provide enough information to reconstruct an overall multiplicity hierarchy.
21. Missing Data, Censoring, and What Is Not Reported
Time-to-event analysis inherently involves censoring when participants have not experienced the specified event by the time their follow-up ends or becomes unavailable. The ClinicalTrials.gov record provides the endpoint definitions and statistical methods, but they do not provide a complete censoring table, missing-data strategy, imputation algorithm, or detailed rules for every potential censoring scenario.
Similarly, the quality-of-life endpoint specifies that the analysis population consists of participants who completed the questionnaire at baseline and at the specified visit. This means the analysis is conditioned on completion at those time points. The ClinicalTrials.gov record does not provide enough information to quantify how questionnaire completion differed between treatment groups or whether missingness was handled with imputation.
22. Stratified Analysis and Clinical Interpretation
The AURELIA registry analyses provide a useful example of why stratification is more than a cosmetic statistical adjustment. The primary PFS analysis identifies three clinical variables for stratification:
- Chemotherapy selected: paclitaxel, PLD, or topotecan.
- Prior anti-angiogenic therapy: yes or no.
- Platinum-free interval: less than 3 or 3-6 months.
These variables define the strata through which the primary stratified comparison is constructed. The purpose is to make the treatment comparison conditional on these specified categories rather than treating all participants as belonging to one undifferentiated risk set.
The exact statistical implementation depends on the model and test used. The ClinicalTrials.gov record specifically identify stratified log-rank/Peto-Peto-Prentice analyses and a Cox regression model for the primary PFS hazard ratio.
23. Limitations
- No masking: the registry classifies the trial as unmasked. Knowledge of treatment assignment can be relevant to investigator-assessed progression and patient-reported outcomes.
- Endpoint duplication: two registered primary endpoints concern disease progression or death/PFS, but the ClinicalTrials.gov record provides the detailed formal estimates for the PFS endpoint rather than a separate numerical comparison for the differently titled percentage endpoint.
- No median PFS or OS reported: the ClinicalTrials.gov record does not contain median survival estimates, so no median should be inferred from the hazard ratios.
- No underlying Kaplan-Meier data: the ClinicalTrials.gov record does not include event-time and censoring records needed to reconstruct a valid Kaplan-Meier curve.
- Hazard-ratio assumptions: Cox hazard ratios are model-based. Their simple interpretation as a single relative event-rate contrast is most direct when the proportional-hazards assumption is reasonable.
- Different data cutoffs: PFS and several secondary endpoints use 14 November 2011, whereas overall survival uses 25 January 2013. These analyses should not be merged into one undated result.
- Multiple analyses: the registry reports 17 statistical analyses, but the ClinicalTrials.gov record does not provide a complete multiplicity-adjustment framework.
- Quality-of-life completion: the quality-of-life analyses are restricted to participants who completed the questionnaire at baseline and the specified visit. The ClinicalTrials.gov record does not provide a detailed missing-data analysis.
- Safety detail: only serious adverse events by arm are reported here. A complete safety assessment would require the broader adverse-event dataset.
- No subgroup estimates reported: although stratified analysis is reported, the ClinicalTrials.gov record does not provide separate treatment-effect estimates for each stratum.
24. Why This Trial Matters Statistically
AURELIA is a useful teaching case because its registry results connect several core clinical-trial methods in one randomized phase 3 analysis. The trial uses time-to-event methodology for PFS, duration of response, and OS; categorical methods for objective response; exact testing for quality-of-life responder comparisons; and stratification to account for prespecified clinical factors.
| Concept | How it appears in AURELIA |
|---|---|
| Randomization | 361 participants allocated in a randomized parallel phase 3 trial |
| ITT analysis | Primary PFS and several secondary analyses identify the ITT population |
| Time-to-event endpoints | PFS, progression/death, duration of response, and OS |
| Log-rank testing | Primary PFS, duration of response, and OS analyses |
| Hazard ratio | Primary PFS, duration of response, and OS effect measures |
| Cox regression | Used to determine the reported hazard ratios |
| Stratified analysis | Primary PFS analysis accounts for chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval |
| Chi-squared testing | Objective-response analysis |
| Cochran-Mantel-Haenszel testing | Objective-response analysis |
| Fisher exact testing | EORTC QLQ OV28 responder analyses |
| Confidence intervals | Reported for hazard ratios, response-rate differences, and quality-of-life response differences |
| Multiple analyses | 17 statistical analyses posted in the ClinicalTrials.gov record |
25. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: AURELIA — NCT00976911. Official trial registry record and source for the trial data analyzed on this page.
- PubMed: PMID 37407274.
- PubMed: PMID 37185961.
- PubMed: PMID 28595285.
- PubMed: PMID 28481967.
- PubMed: PMID 27871723.
Continue through Clinical Biostats
Use the related tutorials and statistical calculators to explore the methods behind randomized clinical-trial endpoints, survival analysis, categorical comparisons, confidence intervals, and stratified testing.
28. Record Summary
AURELIA provides a compact example of how several statistical methods work together in a randomized oncology trial. The primary PFS analysis used an ITT population, a two-sided log-rank comparison, stratification by chemotherapy selected, prior anti-angiogenic therapy, and platinum-free interval, and a Cox regression model yielding a hazard ratio of 0.379 with a 95% CI of 0.296–0.485 and P < 0.0001. The corresponding unstratified HR was 0.460 with a 95% CI of 0.366–0.577 and P < 0.0001.
The secondary analyses illustrate why effect measures must match endpoint type. Objective response was summarized with a difference in response rates of 15.7 and a 95% CI of 6.5–24.8, with chi-squared and Cochran-Mantel-Haenszel p-values of 0.0010 and 0.0007. Duration of objective response produced an HR of 0.450 with a 95% CI of 0.225–0.900 and P = 0.0202. Overall survival analyses produced HR estimates of 0.833 and 0.870, with confidence intervals crossing 1 and p-values of 0.1360 and 0.2711.
The quality-of-life responder analyses further demonstrate how categorical methods can be used at multiple follow-up visits, while the serious-adverse-event data show why efficacy and safety should be examined as separate statistical questions. Most importantly, the page illustrates a central principle of clinical-trial interpretation: the estimate, confidence interval, p-value, endpoint definition, analysis population, statistical model, and data cutoff all belong to the result.