← Clinical Trials
Ovarian Cancer Phase 3 Time-to-Event NCT01015118

AGO-OVAR 12: Complete Statistical Analysis of Nintedanib in Ovarian Cancer

An independent statistical analysis of the randomized phase 3 AGO-OVAR 12 trial evaluating nintedanib or placebo in combination with paclitaxel and carboplatin as first-line treatment for ovarian and peritoneal neoplasms.

Trial period: 2009-11-17 to 2013-04-01 primary completion  ·  Enrollment: 1366  ·  Industry sponsor: Boehringer Ingelheim
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial facts on this page are restricted to the ClinicalTrials.gov record for AGO-OVAR 12.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

AGO-OVAR 12 was a randomized, double-blind, parallel phase 3 treatment trial enrolling 1366 participants with ovarian or peritoneal neoplasms. The trial compared nintedanib with placebo, each administered in combination with paclitaxel and carboplatin.

1366
Enrollment
Randomized trial
2
Arms
Nintedanib vs placebo
0.84
Primary PFS HR
95% CI 0.72–0.98
0.86
Follow-up PFS HR
95% CI 0.75–0.98
FeatureAGO-OVAR 12
Trial nameAGO-OVAR 12
Brief titleLUME-Ovar 1: Nintedanib (BIBF 1120) or Placebo in Combination With Paclitaxel and Carboplatin in First Line Treatment of Ovarian Cancer
PhasePhase 3
StatusCompleted
ConditionsOvarian Neoplasms; Peritoneal Neoplasms
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment1366
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry
Primary endpoint typeTime-to-event
Registered primary endpoints2
Outcome measures posted9
Statistical analyses posted9

2. Clinical Question

The clinical question was whether adding nintedanib to paclitaxel and carboplatin changed progression-free survival compared with adding placebo to paclitaxel and carboplatin in first-line treatment of ovarian cancer.

Population

Participants enrolled in the phase 3 trial with ovarian neoplasms or peritoneal neoplasms receiving first-line treatment.

Intervention

Nintedanib (BIBF 1120) in combination with paclitaxel and carboplatin.

Comparator

Placebo in combination with paclitaxel and carboplatin.

Primary question

Does nintedanib plus paclitaxel and carboplatin improve the registered progression-free survival endpoint relative to placebo plus paclitaxel and carboplatin?

3. Trial Design

01
Randomize1366 enrolled
02
Double-blindNintedanib or placebo
03
CombinationPaclitaxel + carboplatin
04
AssessPFS, OS, response, symptoms and QoL
05
Follow-upFinal database lock analyses
ARM · NINTEDANIB

Nintedanib combination

  • Nintedanib (BIBF 1120)
  • Paclitaxel
  • Carboplatin
ARM · PLACEBO

Placebo combination

  • Placebo
  • Paclitaxel
  • Carboplatin

The ClinicalTrials.gov record identifies the allocation as randomized, the design model as parallel, and the masking as double. These features are important statistically because randomization establishes the basis for a treatment-group comparison, while double masking is intended to reduce the influence of treatment knowledge on participant and investigator behavior and assessment.

4. Trial Timeline

2009-11-17

Trial start

The registered trial start date was 2009-11-17.

2013-04-01

Primary completion

The registered primary completion date was 2013-04-01.

Final database lock

Longer-term outcome analyses

The registry contains follow-up analyses for progression-free survival, overall survival, time to CA-125 tumour marker progression, objective response, symptoms and global health status/quality of life.

5. Endpoints

The registry lists two primary endpoints, both based on progression-free survival. The first is the primary PFS analysis, and the second is a follow-up PFS analysis conducted through final database lock.

RoleEndpointTime frameType
Primary PFS Based on Investigator Assessment According to Modified Response Evaluation Criteria in Solid Tumors, Version 1.1 (mRECIST), and Additional Clinical Criteria. First drug administration to date of disease progression or death whichever occurs first, upto 29 months Time-to-event
Primary PFS Based on Investigator Assessment According to Modified Response Evaluation Criteria in Solid Tumors, Version 1.1 (mRECIST), and Additional Clinical Criteria (Follow up Analysis). First drug administration to date of disease progression or death whichever occurs first until final Data Base Lock (DBL) 26September16, upto 62 months Time-to-event

The registry defines PFS as the time from randomisation to the date of disease progression or to the date of death, whichever occurs first, according to investigator assessment. For the primary PFS analysis, the registry states that the analysis was performed when approximately 753 patients had experienced a PFS event. The registry also describes median, 25th and 75th percentiles as being calculated from an unadjusted Kaplan-Meier curve for each treatment arm.

Secondary endpoints with formal analyses

EndpointTime frameAnalysis typeEffect measure
PFS Based on Investigator Assessment According to mRECIST Version 1.1 (Key Secondary Endpoint) First drug administration to date of disease progression or death whichever occurs first, upto 29 months Stratified log-rank / proportional-hazards model Hazard ratio
PFS Based on Investigator Assessment According to mRECIST Version 1.1 (Key Secondary Endpoint - Follow up Analysis) First drug administration to date of disease progression or death whichever occurs first until final Data Base Lock (DBL) 26September16, upto 62 months Stratified log-rank / proportional-hazards model Hazard ratio
Overall Survival First drug administration to date of death until final DBL 26September16, upto 62 months Stratified log-rank / proportional-hazards model Hazard ratio
Time to CA-125 Tumour Marker Progression First drug administration until final DBL 26September16, upto 62 months Stratified log-rank / proportional-hazards model Hazard ratio
Objective Response Based on Investigator Assessment First drug administration until final DBL 26September16, upto 62 months Logistic regression Odds ratio
Change in Abdominal/Gastro-intestinal Symptoms Over Time First drug administration until final DBL 26September16, upto 62 months Mixed effect growth curve model Mean difference (Net)
Change in Global Health Status/Quality of Life (QoL) Scale Over Time First drug administration until final DBL 26September16, upto 62 months Mixed effect growth curve model Mean difference (Net)

6. Statistical Methodology

Kaplan-Meier estimation

The registry describes Kaplan-Meier estimation for PFS, with median, 25th and 75th percentiles calculated from an unadjusted Kaplan-Meier curve for each treatment arm. Kaplan-Meier estimation is appropriate for time-to-event outcomes because participants can have different follow-up times and some participants may not experience the event during observation.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at event time ti, while ni represents the number at risk immediately before that time.

Stratified log-rank test

The registry reports a log-rank method for the PFS and other time-to-event comparisons. The analysis text further identifies stratified analysis, with the proportional-hazards model using macroscopic residual postoperative tumour, FIGO stage and carboplatin level as stratification factors.

Stratified proportional-hazards model

For the time-to-event analyses, the reported hazard ratio, confidence interval and p-value were obtained from a proportional-hazards model stratified by macroscopic residual postoperative tumour (yes vs. no), FIGO stage (IIB-III vs. IV), and carboplatin level (AUC5 vs. AUC6). The Breslow method was used for handling ties.

Hazard-ratio interpretation
HR < 1  →  lower estimated instantaneous event rate in the nintedanib group

The registry explicitly states that a hazard ratio below 1 favors nintedanib. The hazard ratio is a relative time-to-event measure; it is not a probability, an absolute risk difference, or a statement that every participant experiences the same proportional reduction.

Logistic regression

Objective response was analyzed using logistic regression. The model adjusted for macroscopic residual postoperative tumour, FIGO stage and carboplatin level. The reported effect measure was an odds ratio.

Constrained longitudinal data analysis

The registry reports a mixed effect growth curve model for change in abdominal/gastro-intestinal symptoms and for change in global health status/quality of life. The analysis text describes mixed-effects growth curve models with the average profile over time represented by a piecewise linear model. The registry-reported normalized methodology identifies this as constrained longitudinal data analysis (cLDA).

Covariate adjustment

Adjustment appears in two distinct statistical settings. The survival models were stratified by three prespecified factors, while the logistic regression model explicitly adjusted for those factors. The longitudinal models also adjusted for the same stratification factors. This distinction matters: stratification and covariate adjustment both account for important design variables, but they are not identical modeling operations.

7. Primary Results: Progression-Free Survival

The first primary analysis compared nintedanib with placebo in the Randomised Set (RS). The registry reports a stratified log-rank analysis and a hazard ratio from a proportional-hazards model.

Primary PFS analysis

Hazard ratio for progression or death

0.84

95% CI: 0.72–0.98   ·   P = 0.0239

Randomised Set (RS) · Nintedanib vs Placebo

FeatureReported result
EndpointPFS based on investigator assessment according to mRECIST and additional clinical criteria
Time frameFirst drug administration to date of disease progression or death whichever occurs first, upto 29 months
Analysis populationRandomised Set (RS)
ComparisonNintedanib vs Placebo
MethodLog-rank test; hazard ratio from a stratified proportional-hazards model
Hazard ratio0.84
95% CI0.72–0.98
P-value0.0239
HypothesisSuperiority
Clinical Biostats interpretation

The estimated hazard ratio of 0.84 means that, under the proportional-hazards model used for this analysis, the estimated instantaneous rate of progression or death in the nintedanib group was approximately 84% of that in the placebo group. Expressed as a relative reduction in the estimated hazard, this corresponds to approximately 16% lower estimated hazard.

The HR does not mean that 16% of participants avoided progression, that every patient had a 16% reduction in risk, or that the absolute probability of progression or death was reduced by 16 percentage points. It is a model-based relative measure of the event rate over time.

The 95% confidence interval of 0.72–0.98 describes the statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual patient effects. Because the interval is relatively close to 1 at its upper boundary, the estimate should not be read as evidence of a very large treatment effect.

The p-value of 0.0239 addresses the statistical evidence against the null hypothesis under the specified analysis. It does not measure the size or clinical importance of the effect. The effect size is described by the hazard ratio and, ideally, by absolute event-time measures as well.

The analysis was stratified by macroscopic residual postoperative tumour, FIGO stage and carboplatin level, and the Breslow method was used to handle ties. As with other Cox-type analyses, interpretation of a single hazard ratio depends on the proportional-hazards modeling framework and on appropriate handling of censoring.

Primary PFS follow-up analysis

Follow-up hazard ratio for progression or death

0.86

95% CI: 0.75–0.98   ·   P = 0.0286

Randomised Set (RS) · Nintedanib vs Placebo

FeatureReported result
EndpointPFS based on investigator assessment according to mRECIST and additional clinical criteria, follow-up analysis
Time frameFirst drug administration to date of disease progression or death whichever occurs first until final Data Base Lock (DBL) 26September16, upto 62 months
Analysis populationRandomised Set (RS)
ComparisonNintedanib vs Placebo
MethodLog-rank test; hazard ratio from a stratified proportional-hazards model
Hazard ratio0.86
95% CI0.75–0.98
P-value0.0286
HypothesisSuperiority
Clinical Biostats interpretation

The follow-up hazard ratio of 0.86 indicates an estimated instantaneous rate of progression or death approximately 86% of that in the placebo group under the fitted proportional-hazards model. In relative terms, that corresponds to approximately 14% lower estimated hazard.

The estimate should not be interpreted as a 14-percentage-point improvement in progression-free survival, nor as meaning that 14% more patients necessarily remained progression-free. Those interpretations require absolute survival probabilities or event-time summaries, which are not provided in the ClinicalTrials.gov record.

The 95% CI of 0.75–0.98 gives a measure of precision around the estimated HR. The interval remains below 1, but its upper bound is close to the null value. Thus, the numerical estimate is compatible with a smaller relative effect than the point estimate alone might suggest.

The p-value of 0.0286 is evidence against the null hypothesis within the reported superiority framework; it is not a measure of effect magnitude. Interpretation also requires attention to the analysis population, censoring, stratification and proportional-hazards assumptions.

8. Secondary Time-to-Event Results

Key secondary PFS endpoint

Hazard ratio for investigator-assessed PFS

0.83

95% CI: 0.72–0.97   ·   P = 0.0186

Key secondary endpoint · Randomised Set (RS)

The key secondary endpoint was PFS based on investigator assessment according to mRECIST Version 1.1. The registry reports a stratified log-rank analysis and a proportional-hazards model stratified by macroscopic residual postoperative tumour, FIGO stage and carboplatin level.

Interpretation

An HR of 0.83 corresponds to an estimated instantaneous event rate approximately 83% of the placebo group's rate under the fitted model, or approximately 17% lower estimated hazard. The 95% CI of 0.72–0.97 quantifies uncertainty around that estimate. The p-value of 0.0186 addresses evidence against the superiority null hypothesis; it does not quantify the size of the treatment effect.

Key secondary PFS follow-up analysis

Follow-up hazard ratio for investigator-assessed PFS

0.85

95% CI: 0.74–0.98   ·   P = 0.0256

Key secondary endpoint · Follow-up analysis

The follow-up analysis produced an HR of 0.85, with a 95% CI of 0.74–0.98 and P = 0.0256. Under the reported proportional-hazards model, this corresponds to approximately 15% lower estimated hazard for progression or death in the nintedanib group.

Overall survival

Hazard ratio for death

0.99

95% CI: 0.83–1.17   ·   P = 0.8653

Final DBL 26September16 · upto 62 months

Overall survival was a secondary time-to-event endpoint defined over the period from first drug administration to death, with the registry specifying final DBL 26September16 and upto 62 months. The analysis used the Randomised Set and a stratified log-rank approach with a hazard ratio from the stratified proportional-hazards model.

Interpretation

The OS hazard ratio of 0.99 is very close to 1. Under the fitted model, the estimated instantaneous rate of death was approximately 99% of the corresponding rate in the placebo group. The 95% CI of 0.83–1.17 spans 1, so the estimate is compatible with lower, similar, or higher hazard under the uncertainty represented by the interval.

The P-value of 0.8653 does not provide evidence against the null hypothesis in the reported superiority analysis. It should not be interpreted as proof that the two treatments are identical, because a nonsignificant p-value does not establish equivalence or non-inferiority.

Time to CA-125 tumour marker progression

Hazard ratio for CA-125 progression

0.88

95% CI: 0.77–1.01   ·   P = 0.0749

Final DBL 26September16 · upto 62 months

The time to CA-125 tumour marker progression analysis used the Randomised Set and the same general stratified time-to-event framework. The estimated HR was 0.88, corresponding to approximately 12% lower estimated hazard in the nintedanib group under the model.

The 95% CI of 0.77–1.01 crosses 1, and the reported P-value was 0.0749. This illustrates why the point estimate, confidence interval and p-value should be considered together rather than treating a hazard ratio below 1 as sufficient evidence of superiority.

9. Objective Response

Odds ratio for objective response

1.22

95% CI: 0.81–1.82   ·   P = 0.3490

Randomised Set for patients with at least 1 target lesion reported at baseline

Objective response was a binary endpoint assessed by investigators. The analysis population was the Randomised Set for patients with at least 1 target lesion reported at baseline. Logistic regression adjusted for macroscopic residual postoperative tumour, FIGO stage and carboplatin level.

FeatureReported result
EndpointObjective Response Based on Investigator Assessment
Time frameFirst drug administration until final DBL 26September16, upto 62 months
Outcome unitPercentage of participants
Analysis populationRandomised Set for patients with at least 1 target lesion reported at baseline
MethodLogistic regression
Covariate adjustmentMacroscopic residual postoperative tumour, FIGO stage and carboplatin level
Odds ratio1.22
95% CI0.81–1.82
P-value0.3490
Clinical Biostats interpretation

The odds ratio of 1.22 means that the modeled odds of objective response were estimated to be 1.22 times as high with nintedanib as with placebo after adjustment for the specified factors. The registry states that an odds ratio greater than 1 favors nintedanib.

An odds ratio is not the same as a risk ratio or percentage-point difference in response probability. For example, an OR of 1.22 cannot be read as a 22-percentage-point increase in the response rate.

The 95% CI of 0.81–1.82 is relatively broad and crosses 1. The P-value of 0.3490 does not provide evidence against the null hypothesis under the reported superiority analysis. The appropriate interpretation is therefore based on the estimated odds ratio together with its uncertainty, rather than on the point estimate alone.

10. Longitudinal Symptoms and Quality of Life

Change in abdominal/gastro-intestinal symptoms over time

Net mean difference

5.29

95% CI: 3.88–6.69   ·   P < 0.0001

Randomised Set for patients with abdominal/gastro-intestinal symptoms

The registry analyzed change in abdominal/gastro-intestinal symptoms over time using a mixed effect growth curve model. The analysis text describes a piecewise linear model for the average profile over time, adjusted for macroscopic residual postoperative tumour at baseline, FIGO stage and carboplatin level.

Interpretation

The reported net mean difference was 5.29 units. The confidence interval of 3.88–6.69 describes uncertainty around the estimated between-group difference in the modeled longitudinal outcome. The p-value of <0.0001 provides strong statistical evidence against the null hypothesis of no modeled mean difference within the reported analysis framework.

The direction of a score difference cannot be translated into "better" or "worse" without knowing the scoring direction of the underlying symptom scale. The ClinicalTrials.gov record reports the numerical mean difference but do not provide that scoring-direction information, so the estimate should not be assigned a clinical direction beyond the numerical comparison itself.

Change in global health status/quality of life scale over time

Net mean difference

-1.88

95% CI: -3.35 to -0.41   ·   P = 0.0124

Randomised Set for patients with global health status/QoL

The global health status/quality of life endpoint was analyzed with the same general mixed-effects growth curve framework, including adjustment for the three stratification factors.

Interpretation

The modeled net mean difference was -1.88 units, with a 95% CI of -3.35 to -0.41. The interval does not include zero, and the reported P-value was 0.0124. Thus, the analysis provides statistical evidence of a difference in the modeled longitudinal mean outcome.

The negative sign is a mathematical direction, not automatically a clinical judgment. Whether a negative change represents improvement or worsening depends on how the underlying quality-of-life scale is scored. The ClinicalTrials.gov record does not provide enough information to assign that clinical meaning to the sign.

11. Results Summary

EndpointEffect estimate95% CIP-valueMethod
Primary PFSHR 0.840.72–0.980.0239Stratified log-rank / proportional-hazards model
Primary PFS follow-upHR 0.860.75–0.980.0286Stratified log-rank / proportional-hazards model
Key secondary PFSHR 0.830.72–0.970.0186Stratified log-rank / proportional-hazards model
Key secondary PFS follow-upHR 0.850.74–0.980.0256Stratified log-rank / proportional-hazards model
Overall survivalHR 0.990.83–1.170.8653Stratified log-rank / proportional-hazards model
Time to CA-125 progressionHR 0.880.77–1.010.0749Stratified log-rank / proportional-hazards model
Objective responseOR 1.220.81–1.820.3490Logistic regression
Abdominal/GI symptomsMean difference 5.293.88–6.69<0.0001cLDA / mixed effect growth curve model
Global health status/QoLMean difference -1.88-3.35 to -0.410.0124cLDA / mixed effect growth curve model

The statistical pattern is not uniform across endpoints. The reported PFS analyses produced hazard ratios below 1 with confidence intervals below 1, whereas the overall survival estimate was close to 1 and its confidence interval crossed 1. The CA-125 analysis also had a confidence interval crossing 1. Objective response produced an odds ratio above 1, but with a confidence interval crossing 1. The longitudinal endpoints were analyzed on their continuous scales rather than as time-to-event or binary outcomes.

12. Statistical Methods Explained

Why was a log-rank test used?

PFS, overall survival and time to CA-125 progression are time-to-event endpoints. The log-rank test compares the observed and expected event patterns between treatment groups over follow-up while accounting for censoring. In this trial, the reported analysis was stratified to account for the specified design factors.

What does a hazard ratio of 0.84 mean?

An HR of 0.84 means that the fitted proportional-hazards model estimated the instantaneous progression-or-death rate in the nintedanib group at approximately 84% of the placebo-group rate. It can be described as approximately a 16% lower estimated hazard, but it is not a 16-percentage-point improvement in PFS and does not mean every patient experiences a 16% reduction.

Why were the analyses stratified?

The reported proportional-hazards models were stratified by macroscopic residual postoperative tumour, FIGO stage and carboplatin level. Stratification allows the baseline event hazard to differ across these strata rather than requiring one common baseline hazard structure. It also aligns the analysis with important factors used to structure the treatment comparison.

Why use logistic regression for objective response?

Objective response is a binary outcome: a participant either meets the response definition or does not. Logistic regression models the probability of a binary outcome and expresses treatment association through an odds ratio. In this analysis, the model adjusted for macroscopic residual postoperative tumour, FIGO stage and carboplatin level.

What is the difference between an odds ratio and a hazard ratio?

A hazard ratio is designed for a time-to-event outcome and incorporates the timing of events and censoring. An odds ratio compares odds of a binary outcome. Although both are relative measures, they answer different statistical questions and cannot be substituted for one another.

Why was a longitudinal model used for symptoms and quality of life?

Symptoms and global health status/quality of life were measured repeatedly over time. A mixed-effects growth curve model can use the longitudinal structure of these observations rather than reducing the entire follow-up trajectory to a single measurement. The registry describes a piecewise linear average profile over time and adjustment for the three stratification factors.

Why should the p-value not be treated as the effect size?

A p-value measures the compatibility of the observed data with a specified null hypothesis under the statistical model and testing framework. It does not tell the reader how large the treatment effect is. For this trial, the effect size is represented differently depending on the endpoint: hazard ratios for time-to-event outcomes, an odds ratio for objective response, and mean differences for longitudinal outcomes. Confidence intervals then provide information about uncertainty around those estimates.

13. Stratification and Covariate Adjustment

FactorRole reported in analysis
Macroscopic residual postoperative tumourStratification factor in proportional-hazards models; adjustment factor in logistic regression and longitudinal models
FIGO stageStratification factor in proportional-hazards models; adjustment factor in logistic regression and longitudinal models
Carboplatin levelAUC5 vs. AUC6; stratification factor in proportional-hazards models; adjustment factor in logistic regression and longitudinal models

These factors appear repeatedly across the statistical analyses, which makes them an important part of understanding the trial's statistical architecture. They were not simply added to one isolated secondary model: the registry-reported analysis descriptions identify them in the survival, response and longitudinal analyses.

Stratification is not the same as subgroup analysis. The registry's stratified model uses the factors to account for differences in baseline hazard across strata. That does not mean the reported results constitute separate treatment-effect tests within each stratum.

14. Understanding the Confidence Intervals

Estimate95% confidence intervalNull valueWhat the interval shows
PFS HR 0.840.72–0.981Interval lies below 1
PFS follow-up HR 0.860.75–0.981Interval lies below 1
Key secondary PFS HR 0.830.72–0.971Interval lies below 1
Key secondary PFS follow-up HR 0.850.74–0.981Interval lies below 1
OS HR 0.990.83–1.171Interval crosses 1
CA-125 progression HR 0.880.77–1.011Interval crosses 1
Objective response OR 1.220.81–1.821Interval crosses 1
Abdominal/GI symptoms mean difference 5.293.88–6.690Interval lies above 0
Global health status/QoL mean difference -1.88-3.35 to -0.410Interval lies below 0

For hazard ratios and odds ratios, the null value is 1. For mean differences, the null value is 0. This distinction is fundamental when reading the table. A confidence interval crossing the appropriate null value indicates that the reported interval includes no difference on the corresponding effect-measure scale.

15. Multiplicity and the Meaning of Multiple Analyses

The ClinicalTrials.gov record identifies two registered primary endpoints and multiple secondary endpoints, with nine statistical analyses posted in total. Several of the endpoints have both an initial analysis and a follow-up analysis.

LayerEndpoints representedStatistical issue
PrimaryTwo PFS analysesBoth are explicitly designated primary in the registry data
Key secondaryPFS and follow-up PFSAdditional efficacy comparisons beyond the primary endpoint records
SecondaryOS; CA-125 progression; objective response; symptoms; QoLMultiple distinct estimands and outcome types
Follow-up analysesPFS and key secondary PFSDifferent follow-up windows should not be silently treated as identical analyses

Multiplicity is important because a trial can generate many statistical tests. The ClinicalTrials.gov record identifies the hypothesis type as superiority but do not provide a detailed multiplicity-adjustment procedure, alpha allocation, or hierarchical testing sequence. Those features therefore should not be inferred from the posted p-values.

Interpretation rule: each reported p-value should be tied to its specific endpoint, analysis population, time frame and model. A collection of nominal p-values across several endpoints should not automatically be treated as if each were an isolated confirmatory test.

16. Interim Analysis and Data Cutoffs

The primary PFS definition states that the primary analysis was performed when approximately 753 patients had experienced a PFS event. This is an event-driven feature of the time-to-event analysis: the amount of information is linked to the number of observed events rather than simply to calendar time.

The ClinicalTrials.gov record does not provide a formal interim-analysis alpha-spending procedure, efficacy boundary, stopping rule, or detailed sample-size calculation. Those elements are therefore not assigned to the trial on this page.

Why event counts matter

For a time-to-event endpoint, information accumulates as participants experience the event and as follow-up develops. A trial can enroll a large number of participants while still having limited information if relatively few events have occurred. Conversely, a specified event target can provide a practical basis for determining when the primary time-to-event analysis is conducted.

17. Analysis Populations

EndpointAnalysis population
Primary PFSRandomised Set (RS)
Primary PFS follow-upRandomised Set (RS)
Key secondary PFSRandomised Set (RS)
Key secondary PFS follow-upRandomised Set (RS)
Overall survivalRandomised Set (RS)
Time to CA-125 tumour marker progressionRandomised set (RS)
Objective responseRandomised Set (RS) for patients with at least 1 target lesion reported at baseline
Abdominal/gastro-intestinal symptomsRandomised Set (RS) for patients with abdominal/gastro-intestinal symptoms
Global health status/QoLRandomised Set (RS) for patients with global health status/QoL

The consistent use of the Randomised Set for the principal efficacy analyses preserves the treatment assignment established by randomization. The secondary outcomes appropriately narrow the analysis population when the endpoint requires an evaluable baseline characteristic or the presence of a particular symptom or measurement.

18. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These are presented as affected participants over participants at risk.

Treatment armSerious adverse eventsAffected / at risk
PlaceboSerious adverse events157 / 450
NintedanibSerious adverse events379 / 902

The denominators reported for the serious-adverse-event summaries are different from the overall trial enrollment of 1366. They should therefore be interpreted as the registry-reported at-risk populations for this safety measure rather than assumed to represent all enrolled participants.

The ClinicalTrials.gov record does not provide a full table of individual adverse-event categories, grades, discontinuations or exposure-adjusted incidence rates. The safety interpretation is therefore limited to the reported serious-adverse-event counts.

Do not compare the raw counts alone. Because the reported at-risk populations differ between the two arms, the numbers 157 and 379 cannot be interpreted as if they represented equal-sized groups. A proportion requires both the affected count and the corresponding at-risk denominator, and even then exposure duration and other safety-analysis conventions may matter.

19. A Deeper Look at the Longitudinal Analyses

The two longitudinal outcomes illustrate why the statistical model should match the structure of the endpoint. Symptoms and quality of life are not single event times. They are measurements that can change over repeated assessments.

Repeated observations

Each participant can contribute information at multiple time points, creating within-participant correlation that ordinary independent-observation methods would not appropriately represent.

Growth curves

The registry describes an average profile over time using a piecewise linear model, allowing the longitudinal trajectory to be represented rather than collapsing it to one observation.

Covariate adjustment

The longitudinal models adjusted for residual postoperative tumour, FIGO stage and carboplatin level.

Mean difference

The treatment comparison is expressed on the scale of the outcome as a net mean difference, not as a hazard ratio or odds ratio.

For abdominal/gastro-intestinal symptoms, the estimated net mean difference was 5.29 units. For global health status/quality of life, it was -1.88 units. The statistical significance of these estimates does not by itself establish clinical importance; that requires knowledge of the scale, its direction and a clinically meaningful difference threshold.

20. What the Time-to-Event Results Do — and Do Not — Say

Relative effect

The primary PFS HR of 0.84 is a relative time-to-event estimate. It describes the relationship between the event hazards under the fitted model rather than directly stating how many additional participants remained progression-free.

Absolute effect

The ClinicalTrials.gov record does not provide arm-specific median PFS estimates or Kaplan-Meier survival probabilities. Those quantities should not be reconstructed from the hazard ratio alone.

Censoring

Time-to-event analyses can include participants who have not experienced the event by their last observation. Their follow-up contributes information until censoring. Consequently, the analysis is not equivalent to simply comparing the proportions of participants who eventually experienced progression or death.

Model assumptions

The reported hazard ratios come from proportional-hazards models. A single HR is most straightforward to interpret when the proportional-hazards structure is reasonable over the relevant follow-up. The ClinicalTrials.gov record does not provide a diagnostic assessment of that assumption.

21. Why the PFS and OS Results Should Be Read Separately

The primary PFS results and the secondary OS result address different endpoints. PFS counts progression or death as the event, while OS counts death. A treatment can therefore have a different estimated effect on PFS than on OS.

OutcomeReported estimate95% CIP-value
Primary PFSHR 0.840.72–0.980.0239
Primary PFS follow-upHR 0.860.75–0.980.0286
Overall survivalHR 0.990.83–1.170.8653

The important statistical point is not to substitute one endpoint for another. The PFS estimates quantify the reported treatment association with progression or death, while the OS estimate quantifies the association with death. They have different event definitions and therefore can legitimately produce different estimates.

22. Limitations

23. What Is Not Appropriate to Infer From These Results

Not an absolute risk reduction

HR 0.84 cannot be converted directly into an absolute percentage-point reduction in progression or death.

Not a response probability

OR 1.22 is not a 22-percentage-point increase in objective response.

Not proof of equivalence

An OS HR of 0.99 with P = 0.8653 does not establish equivalence or non-inferiority.

Not a clinical-importance threshold

A statistically significant mean difference does not automatically establish that the magnitude is clinically meaningful.

24. Why This Trial Matters Statistically

AGO-OVAR 12 is a useful statistical teaching case because several major clinical-trial methods appear in one randomized phase 3 study. The principal efficacy endpoint is time-to-event, while other outcomes require binary and longitudinal methods.

ConceptHow it appears in AGO-OVAR 12
RandomizationRandomized parallel-group phase 3 design
BlindingDouble-masked trial
Kaplan-Meier estimationUsed for PFS distribution summaries in the registry definition
Log-rank testReported for PFS, OS and CA-125 progression comparisons
Hazard ratioPrimary and secondary time-to-event treatment effect measure
Stratified analysisResidual tumour, FIGO stage and carboplatin level used in survival modeling
Logistic regressionObjective response analysis
Odds ratioEffect measure for objective response
Covariate adjustmentUsed in response and longitudinal analyses
cLDA / mixed modelsLongitudinal symptoms and global health status/QoL
Confidence intervalsReported for all nine posted statistical analyses
Multiple endpointsTwo registered primary endpoints plus secondary outcomes
Event-driven analysisPrimary PFS analysis conducted when approximately 753 PFS events had occurred

25. Statistical Interpretation of the Complete Evidence Pattern

The most informative way to read the registry results is endpoint by endpoint. The primary PFS analysis reports HR 0.84 with a 95% CI of 0.72–0.98 and P = 0.0239. Its follow-up analysis reports HR 0.86 with a 95% CI of 0.75–0.98 and P = 0.0286. These two estimates are directionally similar and both have confidence intervals below 1.

The key secondary PFS analyses are similarly below 1: HR 0.83 with a 95% CI of 0.72–0.97 and P = 0.0186, followed by HR 0.85 with a 95% CI of 0.74–0.98 and P = 0.0256. The consistency of these estimates is relevant descriptively, although the statistical interpretation of multiple analyses depends on the prespecified testing framework.

Other endpoints tell a different statistical story. Overall survival has an HR of 0.99 with a 95% CI of 0.83–1.17 and P = 0.8653. Time to CA-125 tumour marker progression has an HR of 0.88 with a 95% CI of 0.77–1.01 and P = 0.0749. Objective response has an OR of 1.22 with a 95% CI of 0.81–1.82 and P = 0.3490. These estimates illustrate why a trial should not be summarized by one favorable-looking effect measure.

The longitudinal outcomes use a different statistical language. The abdominal/gastro-intestinal symptom analysis reports a mean difference of 5.29 units with a 95% CI of 3.88–6.69 and P < 0.0001, while global health status/quality of life reports a mean difference of -1.88 units with a 95% CI of -3.35 to -0.41 and P = 0.0124. These are model-based longitudinal mean differences and require knowledge of the underlying scales before assigning clinical meaning to their direction or magnitude.

26. Related Tutorials

Learn more about the methods used in this trial:

27. Related Statistical Calculators

28. Sources

Continue through Clinical Biostats

Explore the statistical methods behind randomized trials, survival analysis, regression, longitudinal models and clinical-trial interpretation.

29. Record Summary

AGO-OVAR 12 provides a compact example of how several statistical frameworks can coexist within one randomized phase 3 trial. The primary endpoints are time-to-event outcomes analyzed with log-rank testing and stratified proportional-hazards models. The trial also reports a binary objective-response endpoint analyzed with covariate-adjusted logistic regression and longitudinal symptom and quality-of-life outcomes analyzed using mixed-effects growth curve models identified in the registry-reported methodology as constrained longitudinal data analysis.

The primary PFS analysis reported an HR of 0.84 (95% CI 0.72–0.98; P = 0.0239), while the follow-up primary PFS analysis reported an HR of 0.86 (95% CI 0.75–0.98; P = 0.0286). The key secondary PFS analyses reported HRs of 0.83 and 0.85. By contrast, the overall survival HR was 0.99 (95% CI 0.83–1.17; P = 0.8653). The complete statistical picture therefore requires attention to the endpoint definition, effect measure, confidence interval, analysis population and model rather than relying on a single summary statistic.

The broader lesson is methodological: a hazard ratio describes a time-to-event association, an odds ratio describes a binary-outcome association, and a mean difference describes a difference on the outcome scale. Confidence intervals provide precision information around each estimate, while p-values address the specified hypothesis-testing framework. Reading all of these elements together is essential for a statistically responsible interpretation of randomized clinical-trial results.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. The purpose of this analysis is to explain how the registered endpoints and posted statistical analyses fit together without adding numerical results that are not contained in the ClinicalTrials.gov record.