This page separates reported trial results from statistical interpretation. Numerical results and trial facts on this page are restricted to the ClinicalTrials.gov record for AGO-OVAR 12.
1. Trial at a Glance
AGO-OVAR 12 was a randomized, double-blind, parallel phase 3 treatment trial enrolling 1366 participants with ovarian or peritoneal neoplasms. The trial compared nintedanib with placebo, each administered in combination with paclitaxel and carboplatin.
| Feature | AGO-OVAR 12 |
|---|---|
| Trial name | AGO-OVAR 12 |
| Brief title | LUME-Ovar 1: Nintedanib (BIBF 1120) or Placebo in Combination With Paclitaxel and Carboplatin in First Line Treatment of Ovarian Cancer |
| Phase | Phase 3 |
| Status | Completed |
| Conditions | Ovarian Neoplasms; Peritoneal Neoplasms |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 1366 |
| Lead sponsor | Boehringer Ingelheim |
| Sponsor type | Industry |
| Primary endpoint type | Time-to-event |
| Registered primary endpoints | 2 |
| Outcome measures posted | 9 |
| Statistical analyses posted | 9 |
2. Clinical Question
The clinical question was whether adding nintedanib to paclitaxel and carboplatin changed progression-free survival compared with adding placebo to paclitaxel and carboplatin in first-line treatment of ovarian cancer.
Population
Participants enrolled in the phase 3 trial with ovarian neoplasms or peritoneal neoplasms receiving first-line treatment.
Intervention
Nintedanib (BIBF 1120) in combination with paclitaxel and carboplatin.
Comparator
Placebo in combination with paclitaxel and carboplatin.
Primary question
Does nintedanib plus paclitaxel and carboplatin improve the registered progression-free survival endpoint relative to placebo plus paclitaxel and carboplatin?
3. Trial Design
Nintedanib combination
- Nintedanib (BIBF 1120)
- Paclitaxel
- Carboplatin
Placebo combination
- Placebo
- Paclitaxel
- Carboplatin
The ClinicalTrials.gov record identifies the allocation as randomized, the design model as parallel, and the masking as double. These features are important statistically because randomization establishes the basis for a treatment-group comparison, while double masking is intended to reduce the influence of treatment knowledge on participant and investigator behavior and assessment.
4. Trial Timeline
Trial start
The registered trial start date was 2009-11-17.
Primary completion
The registered primary completion date was 2013-04-01.
Longer-term outcome analyses
The registry contains follow-up analyses for progression-free survival, overall survival, time to CA-125 tumour marker progression, objective response, symptoms and global health status/quality of life.
5. Endpoints
The registry lists two primary endpoints, both based on progression-free survival. The first is the primary PFS analysis, and the second is a follow-up PFS analysis conducted through final database lock.
| Role | Endpoint | Time frame | Type |
|---|---|---|---|
| Primary | PFS Based on Investigator Assessment According to Modified Response Evaluation Criteria in Solid Tumors, Version 1.1 (mRECIST), and Additional Clinical Criteria. | First drug administration to date of disease progression or death whichever occurs first, upto 29 months | Time-to-event |
| Primary | PFS Based on Investigator Assessment According to Modified Response Evaluation Criteria in Solid Tumors, Version 1.1 (mRECIST), and Additional Clinical Criteria (Follow up Analysis). | First drug administration to date of disease progression or death whichever occurs first until final Data Base Lock (DBL) 26September16, upto 62 months | Time-to-event |
The registry defines PFS as the time from randomisation to the date of disease progression or to the date of death, whichever occurs first, according to investigator assessment. For the primary PFS analysis, the registry states that the analysis was performed when approximately 753 patients had experienced a PFS event. The registry also describes median, 25th and 75th percentiles as being calculated from an unadjusted Kaplan-Meier curve for each treatment arm.
Secondary endpoints with formal analyses
| Endpoint | Time frame | Analysis type | Effect measure |
|---|---|---|---|
| PFS Based on Investigator Assessment According to mRECIST Version 1.1 (Key Secondary Endpoint) | First drug administration to date of disease progression or death whichever occurs first, upto 29 months | Stratified log-rank / proportional-hazards model | Hazard ratio |
| PFS Based on Investigator Assessment According to mRECIST Version 1.1 (Key Secondary Endpoint - Follow up Analysis) | First drug administration to date of disease progression or death whichever occurs first until final Data Base Lock (DBL) 26September16, upto 62 months | Stratified log-rank / proportional-hazards model | Hazard ratio |
| Overall Survival | First drug administration to date of death until final DBL 26September16, upto 62 months | Stratified log-rank / proportional-hazards model | Hazard ratio |
| Time to CA-125 Tumour Marker Progression | First drug administration until final DBL 26September16, upto 62 months | Stratified log-rank / proportional-hazards model | Hazard ratio |
| Objective Response Based on Investigator Assessment | First drug administration until final DBL 26September16, upto 62 months | Logistic regression | Odds ratio |
| Change in Abdominal/Gastro-intestinal Symptoms Over Time | First drug administration until final DBL 26September16, upto 62 months | Mixed effect growth curve model | Mean difference (Net) |
| Change in Global Health Status/Quality of Life (QoL) Scale Over Time | First drug administration until final DBL 26September16, upto 62 months | Mixed effect growth curve model | Mean difference (Net) |
6. Statistical Methodology
Kaplan-Meier estimation
The registry describes Kaplan-Meier estimation for PFS, with median, 25th and 75th percentiles calculated from an unadjusted Kaplan-Meier curve for each treatment arm. Kaplan-Meier estimation is appropriate for time-to-event outcomes because participants can have different follow-up times and some participants may not experience the event during observation.
Here, di represents the number of events at event time ti, while ni represents the number at risk immediately before that time.
Stratified log-rank test
The registry reports a log-rank method for the PFS and other time-to-event comparisons. The analysis text further identifies stratified analysis, with the proportional-hazards model using macroscopic residual postoperative tumour, FIGO stage and carboplatin level as stratification factors.
Stratified proportional-hazards model
For the time-to-event analyses, the reported hazard ratio, confidence interval and p-value were obtained from a proportional-hazards model stratified by macroscopic residual postoperative tumour (yes vs. no), FIGO stage (IIB-III vs. IV), and carboplatin level (AUC5 vs. AUC6). The Breslow method was used for handling ties.
The registry explicitly states that a hazard ratio below 1 favors nintedanib. The hazard ratio is a relative time-to-event measure; it is not a probability, an absolute risk difference, or a statement that every participant experiences the same proportional reduction.
Logistic regression
Objective response was analyzed using logistic regression. The model adjusted for macroscopic residual postoperative tumour, FIGO stage and carboplatin level. The reported effect measure was an odds ratio.
Constrained longitudinal data analysis
The registry reports a mixed effect growth curve model for change in abdominal/gastro-intestinal symptoms and for change in global health status/quality of life. The analysis text describes mixed-effects growth curve models with the average profile over time represented by a piecewise linear model. The registry-reported normalized methodology identifies this as constrained longitudinal data analysis (cLDA).
Covariate adjustment
Adjustment appears in two distinct statistical settings. The survival models were stratified by three prespecified factors, while the logistic regression model explicitly adjusted for those factors. The longitudinal models also adjusted for the same stratification factors. This distinction matters: stratification and covariate adjustment both account for important design variables, but they are not identical modeling operations.
7. Primary Results: Progression-Free Survival
The first primary analysis compared nintedanib with placebo in the Randomised Set (RS). The registry reports a stratified log-rank analysis and a hazard ratio from a proportional-hazards model.
Primary PFS analysis
Hazard ratio for progression or death
95% CI: 0.72–0.98 · P = 0.0239
Randomised Set (RS) · Nintedanib vs Placebo
| Feature | Reported result |
|---|---|
| Endpoint | PFS based on investigator assessment according to mRECIST and additional clinical criteria |
| Time frame | First drug administration to date of disease progression or death whichever occurs first, upto 29 months |
| Analysis population | Randomised Set (RS) |
| Comparison | Nintedanib vs Placebo |
| Method | Log-rank test; hazard ratio from a stratified proportional-hazards model |
| Hazard ratio | 0.84 |
| 95% CI | 0.72–0.98 |
| P-value | 0.0239 |
| Hypothesis | Superiority |
The estimated hazard ratio of 0.84 means that, under the proportional-hazards model used for this analysis, the estimated instantaneous rate of progression or death in the nintedanib group was approximately 84% of that in the placebo group. Expressed as a relative reduction in the estimated hazard, this corresponds to approximately 16% lower estimated hazard.
The HR does not mean that 16% of participants avoided progression, that every patient had a 16% reduction in risk, or that the absolute probability of progression or death was reduced by 16 percentage points. It is a model-based relative measure of the event rate over time.
The 95% confidence interval of 0.72–0.98 describes the statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual patient effects. Because the interval is relatively close to 1 at its upper boundary, the estimate should not be read as evidence of a very large treatment effect.
The p-value of 0.0239 addresses the statistical evidence against the null hypothesis under the specified analysis. It does not measure the size or clinical importance of the effect. The effect size is described by the hazard ratio and, ideally, by absolute event-time measures as well.
The analysis was stratified by macroscopic residual postoperative tumour, FIGO stage and carboplatin level, and the Breslow method was used to handle ties. As with other Cox-type analyses, interpretation of a single hazard ratio depends on the proportional-hazards modeling framework and on appropriate handling of censoring.
Primary PFS follow-up analysis
Follow-up hazard ratio for progression or death
95% CI: 0.75–0.98 · P = 0.0286
Randomised Set (RS) · Nintedanib vs Placebo
| Feature | Reported result |
|---|---|
| Endpoint | PFS based on investigator assessment according to mRECIST and additional clinical criteria, follow-up analysis |
| Time frame | First drug administration to date of disease progression or death whichever occurs first until final Data Base Lock (DBL) 26September16, upto 62 months |
| Analysis population | Randomised Set (RS) |
| Comparison | Nintedanib vs Placebo |
| Method | Log-rank test; hazard ratio from a stratified proportional-hazards model |
| Hazard ratio | 0.86 |
| 95% CI | 0.75–0.98 |
| P-value | 0.0286 |
| Hypothesis | Superiority |
The follow-up hazard ratio of 0.86 indicates an estimated instantaneous rate of progression or death approximately 86% of that in the placebo group under the fitted proportional-hazards model. In relative terms, that corresponds to approximately 14% lower estimated hazard.
The estimate should not be interpreted as a 14-percentage-point improvement in progression-free survival, nor as meaning that 14% more patients necessarily remained progression-free. Those interpretations require absolute survival probabilities or event-time summaries, which are not provided in the ClinicalTrials.gov record.
The 95% CI of 0.75–0.98 gives a measure of precision around the estimated HR. The interval remains below 1, but its upper bound is close to the null value. Thus, the numerical estimate is compatible with a smaller relative effect than the point estimate alone might suggest.
The p-value of 0.0286 is evidence against the null hypothesis within the reported superiority framework; it is not a measure of effect magnitude. Interpretation also requires attention to the analysis population, censoring, stratification and proportional-hazards assumptions.
8. Secondary Time-to-Event Results
Key secondary PFS endpoint
Hazard ratio for investigator-assessed PFS
95% CI: 0.72–0.97 · P = 0.0186
Key secondary endpoint · Randomised Set (RS)
The key secondary endpoint was PFS based on investigator assessment according to mRECIST Version 1.1. The registry reports a stratified log-rank analysis and a proportional-hazards model stratified by macroscopic residual postoperative tumour, FIGO stage and carboplatin level.
An HR of 0.83 corresponds to an estimated instantaneous event rate approximately 83% of the placebo group's rate under the fitted model, or approximately 17% lower estimated hazard. The 95% CI of 0.72–0.97 quantifies uncertainty around that estimate. The p-value of 0.0186 addresses evidence against the superiority null hypothesis; it does not quantify the size of the treatment effect.
Key secondary PFS follow-up analysis
Follow-up hazard ratio for investigator-assessed PFS
95% CI: 0.74–0.98 · P = 0.0256
Key secondary endpoint · Follow-up analysis
The follow-up analysis produced an HR of 0.85, with a 95% CI of 0.74–0.98 and P = 0.0256. Under the reported proportional-hazards model, this corresponds to approximately 15% lower estimated hazard for progression or death in the nintedanib group.
Overall survival
Hazard ratio for death
95% CI: 0.83–1.17 · P = 0.8653
Final DBL 26September16 · upto 62 months
Overall survival was a secondary time-to-event endpoint defined over the period from first drug administration to death, with the registry specifying final DBL 26September16 and upto 62 months. The analysis used the Randomised Set and a stratified log-rank approach with a hazard ratio from the stratified proportional-hazards model.
The OS hazard ratio of 0.99 is very close to 1. Under the fitted model, the estimated instantaneous rate of death was approximately 99% of the corresponding rate in the placebo group. The 95% CI of 0.83–1.17 spans 1, so the estimate is compatible with lower, similar, or higher hazard under the uncertainty represented by the interval.
The P-value of 0.8653 does not provide evidence against the null hypothesis in the reported superiority analysis. It should not be interpreted as proof that the two treatments are identical, because a nonsignificant p-value does not establish equivalence or non-inferiority.
Time to CA-125 tumour marker progression
Hazard ratio for CA-125 progression
95% CI: 0.77–1.01 · P = 0.0749
Final DBL 26September16 · upto 62 months
The time to CA-125 tumour marker progression analysis used the Randomised Set and the same general stratified time-to-event framework. The estimated HR was 0.88, corresponding to approximately 12% lower estimated hazard in the nintedanib group under the model.
The 95% CI of 0.77–1.01 crosses 1, and the reported P-value was 0.0749. This illustrates why the point estimate, confidence interval and p-value should be considered together rather than treating a hazard ratio below 1 as sufficient evidence of superiority.
9. Objective Response
Odds ratio for objective response
95% CI: 0.81–1.82 · P = 0.3490
Randomised Set for patients with at least 1 target lesion reported at baseline
Objective response was a binary endpoint assessed by investigators. The analysis population was the Randomised Set for patients with at least 1 target lesion reported at baseline. Logistic regression adjusted for macroscopic residual postoperative tumour, FIGO stage and carboplatin level.
| Feature | Reported result |
|---|---|
| Endpoint | Objective Response Based on Investigator Assessment |
| Time frame | First drug administration until final DBL 26September16, upto 62 months |
| Outcome unit | Percentage of participants |
| Analysis population | Randomised Set for patients with at least 1 target lesion reported at baseline |
| Method | Logistic regression |
| Covariate adjustment | Macroscopic residual postoperative tumour, FIGO stage and carboplatin level |
| Odds ratio | 1.22 |
| 95% CI | 0.81–1.82 |
| P-value | 0.3490 |
The odds ratio of 1.22 means that the modeled odds of objective response were estimated to be 1.22 times as high with nintedanib as with placebo after adjustment for the specified factors. The registry states that an odds ratio greater than 1 favors nintedanib.
An odds ratio is not the same as a risk ratio or percentage-point difference in response probability. For example, an OR of 1.22 cannot be read as a 22-percentage-point increase in the response rate.
The 95% CI of 0.81–1.82 is relatively broad and crosses 1. The P-value of 0.3490 does not provide evidence against the null hypothesis under the reported superiority analysis. The appropriate interpretation is therefore based on the estimated odds ratio together with its uncertainty, rather than on the point estimate alone.
10. Longitudinal Symptoms and Quality of Life
Change in abdominal/gastro-intestinal symptoms over time
Net mean difference
95% CI: 3.88–6.69 · P < 0.0001
Randomised Set for patients with abdominal/gastro-intestinal symptoms
The registry analyzed change in abdominal/gastro-intestinal symptoms over time using a mixed effect growth curve model. The analysis text describes a piecewise linear model for the average profile over time, adjusted for macroscopic residual postoperative tumour at baseline, FIGO stage and carboplatin level.
The reported net mean difference was 5.29 units. The confidence interval of 3.88–6.69 describes uncertainty around the estimated between-group difference in the modeled longitudinal outcome. The p-value of <0.0001 provides strong statistical evidence against the null hypothesis of no modeled mean difference within the reported analysis framework.
The direction of a score difference cannot be translated into "better" or "worse" without knowing the scoring direction of the underlying symptom scale. The ClinicalTrials.gov record reports the numerical mean difference but do not provide that scoring-direction information, so the estimate should not be assigned a clinical direction beyond the numerical comparison itself.
Change in global health status/quality of life scale over time
Net mean difference
95% CI: -3.35 to -0.41 · P = 0.0124
Randomised Set for patients with global health status/QoL
The global health status/quality of life endpoint was analyzed with the same general mixed-effects growth curve framework, including adjustment for the three stratification factors.
The modeled net mean difference was -1.88 units, with a 95% CI of -3.35 to -0.41. The interval does not include zero, and the reported P-value was 0.0124. Thus, the analysis provides statistical evidence of a difference in the modeled longitudinal mean outcome.
The negative sign is a mathematical direction, not automatically a clinical judgment. Whether a negative change represents improvement or worsening depends on how the underlying quality-of-life scale is scored. The ClinicalTrials.gov record does not provide enough information to assign that clinical meaning to the sign.
11. Results Summary
| Endpoint | Effect estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Primary PFS | HR 0.84 | 0.72–0.98 | 0.0239 | Stratified log-rank / proportional-hazards model |
| Primary PFS follow-up | HR 0.86 | 0.75–0.98 | 0.0286 | Stratified log-rank / proportional-hazards model |
| Key secondary PFS | HR 0.83 | 0.72–0.97 | 0.0186 | Stratified log-rank / proportional-hazards model |
| Key secondary PFS follow-up | HR 0.85 | 0.74–0.98 | 0.0256 | Stratified log-rank / proportional-hazards model |
| Overall survival | HR 0.99 | 0.83–1.17 | 0.8653 | Stratified log-rank / proportional-hazards model |
| Time to CA-125 progression | HR 0.88 | 0.77–1.01 | 0.0749 | Stratified log-rank / proportional-hazards model |
| Objective response | OR 1.22 | 0.81–1.82 | 0.3490 | Logistic regression |
| Abdominal/GI symptoms | Mean difference 5.29 | 3.88–6.69 | <0.0001 | cLDA / mixed effect growth curve model |
| Global health status/QoL | Mean difference -1.88 | -3.35 to -0.41 | 0.0124 | cLDA / mixed effect growth curve model |
The statistical pattern is not uniform across endpoints. The reported PFS analyses produced hazard ratios below 1 with confidence intervals below 1, whereas the overall survival estimate was close to 1 and its confidence interval crossed 1. The CA-125 analysis also had a confidence interval crossing 1. Objective response produced an odds ratio above 1, but with a confidence interval crossing 1. The longitudinal endpoints were analyzed on their continuous scales rather than as time-to-event or binary outcomes.
12. Statistical Methods Explained
Why was a log-rank test used?
PFS, overall survival and time to CA-125 progression are time-to-event endpoints. The log-rank test compares the observed and expected event patterns between treatment groups over follow-up while accounting for censoring. In this trial, the reported analysis was stratified to account for the specified design factors.
What does a hazard ratio of 0.84 mean?
An HR of 0.84 means that the fitted proportional-hazards model estimated the instantaneous progression-or-death rate in the nintedanib group at approximately 84% of the placebo-group rate. It can be described as approximately a 16% lower estimated hazard, but it is not a 16-percentage-point improvement in PFS and does not mean every patient experiences a 16% reduction.
Why were the analyses stratified?
The reported proportional-hazards models were stratified by macroscopic residual postoperative tumour, FIGO stage and carboplatin level. Stratification allows the baseline event hazard to differ across these strata rather than requiring one common baseline hazard structure. It also aligns the analysis with important factors used to structure the treatment comparison.
Why use logistic regression for objective response?
Objective response is a binary outcome: a participant either meets the response definition or does not. Logistic regression models the probability of a binary outcome and expresses treatment association through an odds ratio. In this analysis, the model adjusted for macroscopic residual postoperative tumour, FIGO stage and carboplatin level.
What is the difference between an odds ratio and a hazard ratio?
A hazard ratio is designed for a time-to-event outcome and incorporates the timing of events and censoring. An odds ratio compares odds of a binary outcome. Although both are relative measures, they answer different statistical questions and cannot be substituted for one another.
Why was a longitudinal model used for symptoms and quality of life?
Symptoms and global health status/quality of life were measured repeatedly over time. A mixed-effects growth curve model can use the longitudinal structure of these observations rather than reducing the entire follow-up trajectory to a single measurement. The registry describes a piecewise linear average profile over time and adjustment for the three stratification factors.
Why should the p-value not be treated as the effect size?
A p-value measures the compatibility of the observed data with a specified null hypothesis under the statistical model and testing framework. It does not tell the reader how large the treatment effect is. For this trial, the effect size is represented differently depending on the endpoint: hazard ratios for time-to-event outcomes, an odds ratio for objective response, and mean differences for longitudinal outcomes. Confidence intervals then provide information about uncertainty around those estimates.
13. Stratification and Covariate Adjustment
| Factor | Role reported in analysis |
|---|---|
| Macroscopic residual postoperative tumour | Stratification factor in proportional-hazards models; adjustment factor in logistic regression and longitudinal models |
| FIGO stage | Stratification factor in proportional-hazards models; adjustment factor in logistic regression and longitudinal models |
| Carboplatin level | AUC5 vs. AUC6; stratification factor in proportional-hazards models; adjustment factor in logistic regression and longitudinal models |
These factors appear repeatedly across the statistical analyses, which makes them an important part of understanding the trial's statistical architecture. They were not simply added to one isolated secondary model: the registry-reported analysis descriptions identify them in the survival, response and longitudinal analyses.
14. Understanding the Confidence Intervals
| Estimate | 95% confidence interval | Null value | What the interval shows |
|---|---|---|---|
| PFS HR 0.84 | 0.72–0.98 | 1 | Interval lies below 1 |
| PFS follow-up HR 0.86 | 0.75–0.98 | 1 | Interval lies below 1 |
| Key secondary PFS HR 0.83 | 0.72–0.97 | 1 | Interval lies below 1 |
| Key secondary PFS follow-up HR 0.85 | 0.74–0.98 | 1 | Interval lies below 1 |
| OS HR 0.99 | 0.83–1.17 | 1 | Interval crosses 1 |
| CA-125 progression HR 0.88 | 0.77–1.01 | 1 | Interval crosses 1 |
| Objective response OR 1.22 | 0.81–1.82 | 1 | Interval crosses 1 |
| Abdominal/GI symptoms mean difference 5.29 | 3.88–6.69 | 0 | Interval lies above 0 |
| Global health status/QoL mean difference -1.88 | -3.35 to -0.41 | 0 | Interval lies below 0 |
For hazard ratios and odds ratios, the null value is 1. For mean differences, the null value is 0. This distinction is fundamental when reading the table. A confidence interval crossing the appropriate null value indicates that the reported interval includes no difference on the corresponding effect-measure scale.
15. Multiplicity and the Meaning of Multiple Analyses
The ClinicalTrials.gov record identifies two registered primary endpoints and multiple secondary endpoints, with nine statistical analyses posted in total. Several of the endpoints have both an initial analysis and a follow-up analysis.
| Layer | Endpoints represented | Statistical issue |
|---|---|---|
| Primary | Two PFS analyses | Both are explicitly designated primary in the registry data |
| Key secondary | PFS and follow-up PFS | Additional efficacy comparisons beyond the primary endpoint records |
| Secondary | OS; CA-125 progression; objective response; symptoms; QoL | Multiple distinct estimands and outcome types |
| Follow-up analyses | PFS and key secondary PFS | Different follow-up windows should not be silently treated as identical analyses |
Multiplicity is important because a trial can generate many statistical tests. The ClinicalTrials.gov record identifies the hypothesis type as superiority but do not provide a detailed multiplicity-adjustment procedure, alpha allocation, or hierarchical testing sequence. Those features therefore should not be inferred from the posted p-values.
16. Interim Analysis and Data Cutoffs
The primary PFS definition states that the primary analysis was performed when approximately 753 patients had experienced a PFS event. This is an event-driven feature of the time-to-event analysis: the amount of information is linked to the number of observed events rather than simply to calendar time.
The ClinicalTrials.gov record does not provide a formal interim-analysis alpha-spending procedure, efficacy boundary, stopping rule, or detailed sample-size calculation. Those elements are therefore not assigned to the trial on this page.
Why event counts matter
For a time-to-event endpoint, information accumulates as participants experience the event and as follow-up develops. A trial can enroll a large number of participants while still having limited information if relatively few events have occurred. Conversely, a specified event target can provide a practical basis for determining when the primary time-to-event analysis is conducted.
17. Analysis Populations
| Endpoint | Analysis population |
|---|---|
| Primary PFS | Randomised Set (RS) |
| Primary PFS follow-up | Randomised Set (RS) |
| Key secondary PFS | Randomised Set (RS) |
| Key secondary PFS follow-up | Randomised Set (RS) |
| Overall survival | Randomised Set (RS) |
| Time to CA-125 tumour marker progression | Randomised set (RS) |
| Objective response | Randomised Set (RS) for patients with at least 1 target lesion reported at baseline |
| Abdominal/gastro-intestinal symptoms | Randomised Set (RS) for patients with abdominal/gastro-intestinal symptoms |
| Global health status/QoL | Randomised Set (RS) for patients with global health status/QoL |
The consistent use of the Randomised Set for the principal efficacy analyses preserves the treatment assignment established by randomization. The secondary outcomes appropriately narrow the analysis population when the endpoint requires an evaluable baseline characteristic or the presence of a particular symptom or measurement.
18. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm. These are presented as affected participants over participants at risk.
| Treatment arm | Serious adverse events | Affected / at risk |
|---|---|---|
| Placebo | Serious adverse events | 157 / 450 |
| Nintedanib | Serious adverse events | 379 / 902 |
The denominators reported for the serious-adverse-event summaries are different from the overall trial enrollment of 1366. They should therefore be interpreted as the registry-reported at-risk populations for this safety measure rather than assumed to represent all enrolled participants.
The ClinicalTrials.gov record does not provide a full table of individual adverse-event categories, grades, discontinuations or exposure-adjusted incidence rates. The safety interpretation is therefore limited to the reported serious-adverse-event counts.
19. A Deeper Look at the Longitudinal Analyses
The two longitudinal outcomes illustrate why the statistical model should match the structure of the endpoint. Symptoms and quality of life are not single event times. They are measurements that can change over repeated assessments.
Repeated observations
Each participant can contribute information at multiple time points, creating within-participant correlation that ordinary independent-observation methods would not appropriately represent.
Growth curves
The registry describes an average profile over time using a piecewise linear model, allowing the longitudinal trajectory to be represented rather than collapsing it to one observation.
Covariate adjustment
The longitudinal models adjusted for residual postoperative tumour, FIGO stage and carboplatin level.
Mean difference
The treatment comparison is expressed on the scale of the outcome as a net mean difference, not as a hazard ratio or odds ratio.
For abdominal/gastro-intestinal symptoms, the estimated net mean difference was 5.29 units. For global health status/quality of life, it was -1.88 units. The statistical significance of these estimates does not by itself establish clinical importance; that requires knowledge of the scale, its direction and a clinically meaningful difference threshold.
20. What the Time-to-Event Results Do — and Do Not — Say
The primary PFS HR of 0.84 is a relative time-to-event estimate. It describes the relationship between the event hazards under the fitted model rather than directly stating how many additional participants remained progression-free.
The ClinicalTrials.gov record does not provide arm-specific median PFS estimates or Kaplan-Meier survival probabilities. Those quantities should not be reconstructed from the hazard ratio alone.
Time-to-event analyses can include participants who have not experienced the event by their last observation. Their follow-up contributes information until censoring. Consequently, the analysis is not equivalent to simply comparing the proportions of participants who eventually experienced progression or death.
The reported hazard ratios come from proportional-hazards models. A single HR is most straightforward to interpret when the proportional-hazards structure is reasonable over the relevant follow-up. The ClinicalTrials.gov record does not provide a diagnostic assessment of that assumption.
21. Why the PFS and OS Results Should Be Read Separately
The primary PFS results and the secondary OS result address different endpoints. PFS counts progression or death as the event, while OS counts death. A treatment can therefore have a different estimated effect on PFS than on OS.
| Outcome | Reported estimate | 95% CI | P-value |
|---|---|---|---|
| Primary PFS | HR 0.84 | 0.72–0.98 | 0.0239 |
| Primary PFS follow-up | HR 0.86 | 0.75–0.98 | 0.0286 |
| Overall survival | HR 0.99 | 0.83–1.17 | 0.8653 |
The important statistical point is not to substitute one endpoint for another. The PFS estimates quantify the reported treatment association with progression or death, while the OS estimate quantifies the association with death. They have different event definitions and therefore can legitimately produce different estimates.
22. Limitations
- Limited absolute survival information in the ClinicalTrials.gov record: the registry analysis data provide hazard ratios and confidence intervals but do not provide arm-specific median PFS or OS values for these analyses.
- Hazard-ratio interpretation: the reported estimates rely on proportional-hazards modeling. The ClinicalTrials.gov record does not provide a formal assessment of the proportional-hazards assumption.
- Multiple endpoints: the trial has two registered primary endpoints and several secondary outcomes. The ClinicalTrials.gov record does not specify a complete multiplicity-adjustment strategy.
- Different estimands: PFS, OS, CA-125 progression, objective response, symptoms and quality of life measure different clinical or statistical constructs and should not be combined into one summary statistic.
- Safety denominators: serious adverse events are reported with arm-specific at-risk denominators of 450 and 902, which differ from total enrollment.
- Longitudinal scale interpretation: the ClinicalTrials.gov record reports numerical mean differences for symptoms and QoL but do not provide enough information to determine the clinical direction or meaningfulness of those differences.
- Subgroup information: the statistical analyses posted on ClinicalTrials.gov do not include treatment-effect estimates by individual stratification category. The stratification factors should therefore not be interpreted as subgroup efficacy results.
- Registry scope: this page uses the ClinicalTrials.gov record and does not add unpublished protocol, SAP or publication-derived numerical results.
23. What Is Not Appropriate to Infer From These Results
Not an absolute risk reduction
HR 0.84 cannot be converted directly into an absolute percentage-point reduction in progression or death.
Not a response probability
OR 1.22 is not a 22-percentage-point increase in objective response.
Not proof of equivalence
An OS HR of 0.99 with P = 0.8653 does not establish equivalence or non-inferiority.
Not a clinical-importance threshold
A statistically significant mean difference does not automatically establish that the magnitude is clinically meaningful.
24. Why This Trial Matters Statistically
AGO-OVAR 12 is a useful statistical teaching case because several major clinical-trial methods appear in one randomized phase 3 study. The principal efficacy endpoint is time-to-event, while other outcomes require binary and longitudinal methods.
| Concept | How it appears in AGO-OVAR 12 |
|---|---|
| Randomization | Randomized parallel-group phase 3 design |
| Blinding | Double-masked trial |
| Kaplan-Meier estimation | Used for PFS distribution summaries in the registry definition |
| Log-rank test | Reported for PFS, OS and CA-125 progression comparisons |
| Hazard ratio | Primary and secondary time-to-event treatment effect measure |
| Stratified analysis | Residual tumour, FIGO stage and carboplatin level used in survival modeling |
| Logistic regression | Objective response analysis |
| Odds ratio | Effect measure for objective response |
| Covariate adjustment | Used in response and longitudinal analyses |
| cLDA / mixed models | Longitudinal symptoms and global health status/QoL |
| Confidence intervals | Reported for all nine posted statistical analyses |
| Multiple endpoints | Two registered primary endpoints plus secondary outcomes |
| Event-driven analysis | Primary PFS analysis conducted when approximately 753 PFS events had occurred |
25. Statistical Interpretation of the Complete Evidence Pattern
The most informative way to read the registry results is endpoint by endpoint. The primary PFS analysis reports HR 0.84 with a 95% CI of 0.72–0.98 and P = 0.0239. Its follow-up analysis reports HR 0.86 with a 95% CI of 0.75–0.98 and P = 0.0286. These two estimates are directionally similar and both have confidence intervals below 1.
The key secondary PFS analyses are similarly below 1: HR 0.83 with a 95% CI of 0.72–0.97 and P = 0.0186, followed by HR 0.85 with a 95% CI of 0.74–0.98 and P = 0.0256. The consistency of these estimates is relevant descriptively, although the statistical interpretation of multiple analyses depends on the prespecified testing framework.
Other endpoints tell a different statistical story. Overall survival has an HR of 0.99 with a 95% CI of 0.83–1.17 and P = 0.8653. Time to CA-125 tumour marker progression has an HR of 0.88 with a 95% CI of 0.77–1.01 and P = 0.0749. Objective response has an OR of 1.22 with a 95% CI of 0.81–1.82 and P = 0.3490. These estimates illustrate why a trial should not be summarized by one favorable-looking effect measure.
The longitudinal outcomes use a different statistical language. The abdominal/gastro-intestinal symptom analysis reports a mean difference of 5.29 units with a 95% CI of 3.88–6.69 and P < 0.0001, while global health status/quality of life reports a mean difference of -1.88 units with a 95% CI of -3.35 to -0.41 and P = 0.0124. These are model-based longitudinal mean differences and require knowledge of the underlying scales before assigning clinical meaning to their direction or magnitude.
26. Related Tutorials
Learn more about the methods used in this trial:
27. Related Statistical Calculators
28. Sources
- ClinicalTrials.gov: AGO-OVAR 12 (NCT01015118).
- Linked publication: PubMed record for PMID 37185961.
- Linked publication: PubMed record for PMID 26590673.
Continue through Clinical Biostats
Explore the statistical methods behind randomized trials, survival analysis, regression, longitudinal models and clinical-trial interpretation.
29. Record Summary
AGO-OVAR 12 provides a compact example of how several statistical frameworks can coexist within one randomized phase 3 trial. The primary endpoints are time-to-event outcomes analyzed with log-rank testing and stratified proportional-hazards models. The trial also reports a binary objective-response endpoint analyzed with covariate-adjusted logistic regression and longitudinal symptom and quality-of-life outcomes analyzed using mixed-effects growth curve models identified in the registry-reported methodology as constrained longitudinal data analysis.
The primary PFS analysis reported an HR of 0.84 (95% CI 0.72–0.98; P = 0.0239), while the follow-up primary PFS analysis reported an HR of 0.86 (95% CI 0.75–0.98; P = 0.0286). The key secondary PFS analyses reported HRs of 0.83 and 0.85. By contrast, the overall survival HR was 0.99 (95% CI 0.83–1.17; P = 0.8653). The complete statistical picture therefore requires attention to the endpoint definition, effect measure, confidence interval, analysis population and model rather than relying on a single summary statistic.
The broader lesson is methodological: a hazard ratio describes a time-to-event association, an odds ratio describes a binary-outcome association, and a mean difference describes a difference on the outcome scale. Confidence intervals provide precision information around each estimate, while p-values address the specified hypothesis-testing framework. Reading all of these elements together is essential for a statistically responsible interpretation of randomized clinical-trial results.