This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the trial data posted on ClinicalTrials.gov for ARIEL4.
1. Trial at a Glance
ARIEL4 was a randomized, parallel-group, open-label phase 3 trial evaluating rucaparib versus chemotherapy in patients with BRCA-mutant ovarian, fallopian tube, or primary peritoneal cancer. The registered primary endpoints were both investigator-assessed progression-free survival (invPFS) endpoints evaluated with RECIST Version 1.1, one in the efficacy population and one in the intention-to-treat population.
| Feature | ARIEL4 |
|---|---|
| Phase | Phase 3 |
| Conditions | Ovarian Cancer; Epithelial Ovarian Cancer; Fallopian Tube Cancer; Peritoneal Cancer |
| Design | Randomized, parallel |
| Masking | None |
| Primary purpose | Treatment |
| Interventions | Chemotherapy; Rucaparib |
| Enrollment | 349 |
| Primary endpoints | Two investigator-assessed progression-free survival endpoints |
| Primary endpoint type | Time-to-event |
| Hypothesis type | Superiority |
| ClinicalTrials.gov | NCT02855944 |
2. Clinical Question
The central statistical question was whether rucaparib differed from chemotherapy with respect to investigator-assessed progression-free survival in patients with BRCA-mutant ovarian, fallopian tube, or primary peritoneal cancer. The registered primary analyses used a superiority framework and compared the treatment groups with hazard ratios from Cox regression.
Population
Patients in the treatment part of the study with a deleterious BRCA mutation. For the efficacy population, patients identified to have a BRCA reversion mutation were excluded.
Intervention
Rucaparib.
Comparator
Chemotherapy.
Primary question
Does rucaparib produce a different investigator-assessed progression-free survival experience than chemotherapy under the prespecified superiority comparison?
3. Trial Design
Rucaparib
- Rucaparib was one of the two randomized interventions.
- Serious adverse events were reported for 66 of 232 patients in the Treatment Part.
Chemotherapy
- Chemotherapy was the randomized comparator intervention.
- Serious adverse events were reported for 14 of 113 patients in the Treatment Part.
4. Endpoints
The two registered primary endpoints were both investigator-assessed progression-free survival endpoints by RECIST Version 1.1. The registry classifies them as time-to-event outcomes and reports superiority hypotheses.
| Endpoint | Definition / time frame | Population |
|---|---|---|
| Investigator Assessed Progression-Free Survival (invPFS) by RECIST Version 1.1 for Rucaparib Versus Chemotherapy (Efficacy Population) | Assessments every 8 weeks from Cycle 1 Day 1 (C1D1) until disease progression, death, or initiation of subsequent treatment. After 18 months on study, assessments every 16 weeks. Total follow-up was up to approximately 3.5 years. | All randomized patients in the treatment part of study with a deleterious BRCA mutation, excluding those identified to have a BRCA reversion mutation. |
| Investigator Assessed Progression-Free Survival (invPFS) by RECIST Version 1.1 for Rucaparib Versus Chemotherapy (ITT Population) | Assessments every 8 weeks from C1D1 until disease progression, death, or initiation of subsequent treatment. After 18 months on study, assessments every 16 weeks. Total follow-up was up to approximately 3.5 years. | All randomized patients in the treatment part of study. |
How invPFS was defined
The registry defines the primary efficacy endpoint as investigator-assessed progression-free survival for the Treatment Part. Time to invPFS is calculated in months from randomization to disease progression +1 day, as determined by RECIST v1.1 criteria, or death due to any cause, whichever occurs first. The registry describes progression using RECIST v1.1, including at least a 20% increase in the sum of measurements.
Because the endpoint is time-to-event, patients who have not experienced progression or death by their relevant follow-up can contribute information up to the point at which their outcome is censored under the study's analysis framework.
5. Analysis Populations
ARIEL4 illustrates an important distinction between the population used for a primary efficacy analysis and the broader intention-to-treat population. The ClinicalTrials.gov record explicitly defines both populations.
| Analysis population | Definition | Primary role |
|---|---|---|
| Efficacy Population | All randomized patients in the treatment part of study with a deleterious BRCA mutation, excluding those identified to have a BRCA reversion mutation. | Primary invPFS analysis. |
| Intent-to-treat Population | All randomized patients in the treatment part of study. | Second primary invPFS analysis. |
6. Statistical Methodology
Cox proportional-hazards regression
The registry reports "Regression, Cox" as the statistical method for both primary endpoints. The statistical method is a Cox proportional-hazards model, and the effect measure is a hazard ratio.
For a two-group comparison, the hazard ratio is the exponentiated treatment coefficient. It summarizes the relative instantaneous event rate represented by the fitted model over the analyzed time period.
Hazard ratio
A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the rucaparib group relative to chemotherapy under the fitted model. For example, an HR of 0.639 corresponds to an estimated instantaneous event rate approximately 36.1% lower than the comparator under that model.
The hazard ratio is not a probability, not a percentage of patients who benefit, and not the ratio of median survival times. It is a relative time-to-event measure produced by the Cox model.
Confidence intervals
Each primary analysis reports a two-sided 95% confidence interval around the hazard ratio. The interval quantifies statistical uncertainty around the estimated treatment effect under the model and sampling framework. A narrower interval generally indicates greater precision than a wider interval, although the width is also influenced by the amount of information available to the analysis.
P-values
The primary analyses report P-values of 0.0010 and 0.0017. These values address evidence against the null hypothesis under the specified statistical testing framework; they do not measure the magnitude of the treatment effect. Effect size is described by the hazard ratio, while the confidence interval describes uncertainty around that estimate.
Intention-to-treat analysis
The ITT analysis includes all randomized patients in the treatment part of the study. An ITT comparison retains the treatment assignment established by randomization and therefore provides a treatment-group comparison defined by the original randomized allocation rather than selectively restricting analysis to patients who remained on treatment or met later outcome criteria.
7. Primary Results: invPFS in the Efficacy Population
The first registered primary analysis evaluated investigator-assessed progression-free survival by RECIST Version 1.1 in the efficacy population. The analysis compared rucaparib with chemotherapy using Cox regression.
Hazard ratio for investigator-assessed progression-free survival
95% CI: 0.489–0.835 · P = 0.0010
Two-sided 95% confidence interval; superiority hypothesis.
| Primary endpoint | Analysis population | Method | Effect | P-value |
|---|---|---|---|---|
| invPFS by RECIST Version 1.1 | Efficacy Population | Cox proportional-hazards model | HR 0.639 (95% CI 0.489–0.835) | 0.0010 |
The hazard ratio of 0.639 means that, under the fitted Cox model, the estimated instantaneous rate of progression or death in the rucaparib group was approximately 63.9% of that in the chemotherapy group, corresponding to an approximately 36.1% lower estimated hazard.
It does not mean that 36.1% of patients avoided progression, that progression was delayed by exactly 36.1%, or that every patient experienced the same proportional reduction in risk. A hazard ratio is a population-level model-based summary of the time-to-event comparison.
The 95% CI of 0.489–0.835 describes uncertainty around the estimated hazard ratio. It does not describe the range of treatment effects that individual patients experienced.
The P-value of 0.0010 addresses the statistical evidence against the relevant null hypothesis. It does not tell us that the effect is "0.0010 large," nor does it provide a measure of clinical magnitude. The hazard ratio and its confidence interval are the appropriate quantities for describing effect size and precision.
Because this is a Cox proportional-hazards analysis, interpretation of a single hazard ratio also depends on the proportional-hazards assumption. If the relative hazards vary substantially over time, one summary hazard ratio can obscure important features of the underlying survival experience.
8. Primary Results: invPFS in the ITT Population
The second registered primary analysis evaluated the same investigator-assessed progression-free survival endpoint in the intention-to-treat population. The registry again reports Cox regression and a superiority hypothesis.
Hazard ratio for investigator-assessed progression-free survival
95% CI: 0.516–0.858 · P = 0.0017
Two-sided 95% confidence interval; superiority hypothesis.
| Primary endpoint | Analysis population | Method | Effect | P-value |
|---|---|---|---|---|
| invPFS by RECIST Version 1.1 | Intent-to-treat Population | Cox proportional-hazards model | HR 0.665 (95% CI 0.516–0.858) | 0.0017 |
The hazard ratio of 0.665 means that, under the fitted Cox model, the estimated instantaneous rate of progression or death in the rucaparib group was approximately 66.5% of that in the chemotherapy group, corresponding to an approximately 33.5% lower estimated hazard.
This estimate is not a statement that 33.5% of patients benefited, nor does it mean that each patient had exactly a 33.5% reduction in progression or death. The result summarizes the randomized time-to-event comparison at the population level.
The 95% CI of 0.516–0.858 gives the uncertainty range for the estimated hazard ratio under the analysis framework. The confidence interval is therefore important alongside the point estimate rather than being treated as a secondary decoration.
The P-value of 0.0017 provides evidence against the relevant null hypothesis within the reported superiority analysis. It is not an effect-size statistic and should not be used to compare the magnitude of this result with the P-value from another endpoint.
The analysis uses the ITT population, so its interpretation is tied to the randomized treatment groups. As with the efficacy-population analysis, the Cox interpretation depends on the proportional-hazards model and on the censoring and event-assessment framework underlying the registry endpoint.
9. Comparing the Two Primary invPFS Analyses
The two primary analyses answer closely related questions in different analysis populations. That distinction is statistically important: the efficacy-population estimate applies to the registry-defined population with a deleterious BRCA mutation after exclusion of identified BRCA reversion mutations, whereas the ITT estimate applies to all randomized patients in the treatment part of the study.
| Feature | Efficacy Population | ITT Population |
|---|---|---|
| Endpoint | Investigator-assessed invPFS by RECIST Version 1.1 | Investigator-assessed invPFS by RECIST Version 1.1 |
| Method | Cox proportional-hazards model | Cox proportional-hazards model |
| Hazard ratio | 0.639 | 0.665 |
| 95% CI | 0.489–0.835 | 0.516–0.858 |
| P-value | 0.0010 | 0.0017 |
| Hypothesis | Superiority | Superiority |
The two estimates are not identical, which is expected when the analysis populations differ. The important statistical point is not to average, combine, or otherwise replace the two estimates with a single number. Each estimate describes its specified population and analysis.
10. Secondary Endpoint Results: Duration of Response
The registry also reports formal statistical analyses for investigator-assessed duration of response (DOR) by RECIST v1.1. DOR is a time-to-event endpoint among patients meeting response-related eligibility criteria, so Cox regression is again used to compare the treatment groups.
| Endpoint | Analysis population | Hazard ratio | 95% CI | P-value |
|---|---|---|---|---|
| Investigator Assessed Duration of Response (DOR) by RECIST v1.1 | Efficacy population with measurable disease at baseline and a confirmed response | 0.589 | 0.356–0.976 | 0.0401 |
| Investigator Assessed Duration of Response (DOR) by RECIST v1.1 | ITT population with measurable disease at baseline and a confirmed response | 0.564 | 0.343–0.927 | 0.0240 |
DOR: Efficacy Population
95% CI: 0.356–0.976 · P = 0.0401
The DOR hazard ratio of 0.589 represents the estimated relative hazard of the DOR event in the rucaparib group compared with chemotherapy among the specified efficacy-population responders. Under the Cox model, this corresponds to an approximately 41.1% lower estimated hazard.
The DOR population is inherently narrower than the overall randomized population because it requires measurable disease at baseline and a confirmed response. Consequently, this analysis should not be interpreted as a treatment comparison among every randomized patient.
The 95% CI of 0.356–0.976 indicates substantial uncertainty around the point estimate. The interval is notably wider than the confidence intervals reported for the two primary invPFS analyses, illustrating why precision must be considered separately from the direction of an estimate.
The P-value of 0.0401 is a statistical testing result, not a measure of effect magnitude. Because DOR is a secondary endpoint, its interpretation should also remain distinct from the prespecified primary invPFS analyses.
DOR: ITT Population
95% CI: 0.343–0.927 · P = 0.0240
The ITT-response-population DOR hazard ratio of 0.564 corresponds, under the Cox model, to an approximately 43.6% lower estimated instantaneous hazard of the DOR event in the rucaparib group relative to chemotherapy.
Again, the denominator is not all randomized patients: the analysis population is the ITT population with measurable disease at baseline and a confirmed response. DOR therefore describes durability among responders rather than the probability that a randomly assigned patient will respond.
The 95% CI of 0.343–0.927 expresses uncertainty around the estimate. The P-value of 0.0240 addresses the corresponding statistical test and does not quantify the size or clinical importance of the DOR difference.
As with the primary invPFS analyses, the Cox model is a time-to-event model. Interpretation therefore requires attention to censoring and the proportional-hazards assumption rather than treating the hazard ratio as a simple ratio of event probabilities.
11. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment part and, separately, for the crossover part. The available safety measure is the number of patients affected divided by the number at risk.
| Study part / group | Serious adverse events |
|---|---|
| Rucaparib (Treatment Part) | 66/232 |
| Chemotherapy (Treatment Part) | 14/113 |
| Rucaparib (Crossover Part) | 24/80 |
12. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
Both primary endpoints and both DOR analyses are time-to-event outcomes. Cox regression is designed to compare event rates over follow-up while producing a hazard ratio that summarizes the relative event rate between treatment groups. This is more appropriate for a time-to-event endpoint than treating the outcome as a simple binary response at a fixed time point.
What does an HR of 0.639 mean?
It means that, under the fitted model, the estimated instantaneous rate of progression or death in the rucaparib group was 0.639 times the corresponding rate in the chemotherapy group. Equivalently, the estimated hazard was approximately 36.1% lower. It does not mean that 36.1% of patients were protected from progression or death.
Why are there two primary invPFS analyses?
The registry identifies two primary endpoints with different analysis populations. One uses the efficacy population, which excludes patients identified to have a BRCA reversion mutation, while the other uses the ITT population of all randomized patients in the treatment part. These analyses therefore provide related but distinct estimates.
Why is the confidence interval important?
A point estimate such as 0.639 is only one estimate of the underlying treatment effect. The 95% confidence interval of 0.489–0.835 communicates the statistical uncertainty surrounding that estimate. A confidence interval also helps readers judge how precisely the treatment effect has been estimated.
Why doesn't the P-value measure effect size?
The P-value quantifies evidence against a null hypothesis under the specified statistical model and testing framework. It is affected by both the magnitude of an observed effect and the amount of information available. The hazard ratio is the effect-size measure here, while the confidence interval provides information about precision.
What is the difference between invPFS and DOR?
invPFS is defined from randomization until disease progression or death, whichever occurs first. DOR is evaluated among patients with measurable disease at baseline and a confirmed response. Thus, invPFS addresses the randomized population-level time-to-event experience, whereas DOR describes how long confirmed responses persist among a response-defined population.
13. Understanding the Hazard Ratio More Precisely
The ARIEL4 results are reported primarily as hazard ratios. A hazard ratio is a relative measure, whereas an absolute treatment effect would describe an actual difference in event probability or survival at a specified time. The ClinicalTrials.gov record does not provide median invPFS, Kaplan-Meier survival probabilities at specific time points, or an absolute risk difference, so those quantities cannot be reconstructed from the ClinicalTrials.gov record.
An HR of 0.665 does not imply that the probability of progression or death is 0.665 times lower at every particular time point. It is a model-based relative hazard measure. If hazards are not proportional over time, a single Cox hazard ratio may compress a more complicated time-varying treatment effect into one summary statistic.
The efficacy-population HR of 0.639 and ITT HR of 0.665 should not be treated as duplicate measurements of exactly the same quantity. Their populations differ by definition. The correct interpretation is to report each estimate with its corresponding population, confidence interval, and P-value.
14. Time-to-Event Analysis in ARIEL4
Progression-free survival and duration of response are naturally analyzed as time-to-event endpoints because the timing of progression, death, or loss of observable event-free follow-up contains information. A simple proportion of patients who progressed would discard much of that timing information.
Kaplan-Meier estimation is a standard way to estimate the probability of remaining event-free over time while accommodating right-censored observations. The registry-reported ARIEL4 statistical-analysis data identify Cox regression as the posted method; they do not provide Kaplan-Meier estimates or survival-curve coordinates.
The absence of posted Kaplan-Meier summary estimates in the ClinicalTrials.gov record is important. It means the reported hazard ratios can be interpreted directly, but median event times or time-specific event-free percentages should not be inferred from the hazard ratios alone.
15. Randomization and the Meaning of the Comparison
ARIEL4 is explicitly described as randomized. Randomization is central to the causal interpretation of a treatment comparison because treatment assignment is determined by the trial design rather than by a patient's subsequent response or outcome.
Before the outcome
Randomization assigns participants to the rucaparib or chemotherapy comparison groups.
During follow-up
Progression, death, and subsequent treatment can occur at different times, producing a time-to-event dataset rather than a single binary endpoint.
At analysis
Cox regression converts the observed event-time information into a relative hazard estimate with a confidence interval.
For interpretation
The analysis population must remain attached to the estimate because the efficacy and ITT populations are defined differently.
16. What the Results Establish Statistically
The statistical analyses posted on ClinicalTrials.gov show hazard ratios below 1 for both primary invPFS comparisons and both secondary DOR comparisons. The primary invPFS analyses have two-sided 95% confidence intervals and P-values of 0.0010 and 0.0017, respectively, while the DOR analyses report P-values of 0.0401 and 0.0240.
| Endpoint | Population | HR | 95% CI | P-value | Role |
|---|---|---|---|---|---|
| invPFS | Efficacy | 0.639 | 0.489–0.835 | 0.0010 | Primary |
| invPFS | ITT | 0.665 | 0.516–0.858 | 0.0017 | Primary |
| DOR | Efficacy responders | 0.589 | 0.356–0.976 | 0.0401 | Secondary |
| DOR | ITT responders | 0.564 | 0.343–0.927 | 0.0240 | Secondary |
The table illustrates why effect estimates, uncertainty intervals, analysis populations, and endpoint roles should be displayed together. A P-value without the corresponding hazard ratio and confidence interval would provide an incomplete description of the statistical evidence.
17. Multiplicity and Endpoint Interpretation
The ClinicalTrials.gov record identifies two registered primary endpoints and two secondary DOR analyses, all tested under superiority hypotheses. The data do not provide an alpha-allocation strategy, hierarchical testing procedure, multiplicity-adjustment method, or interim-analysis plan.
18. Stratification, Interim Analysis, and Other Design Features
The ClinicalTrials.gov record identifies Cox regression, hazard ratios, superiority hypotheses, and the relevant analysis populations. They do not provide a documented stratification scheme, non-inferiority margin, factorial structure, interim-analysis boundary, alpha-spending method, Bayesian model, or missing-data imputation method.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Randomization | Yes; allocation is listed as randomized. |
| Parallel design | Yes; design model is parallel. |
| Masking | None. |
| Superiority | Yes; reported for all four statistical analyses. |
| Non-inferiority margin | Not provided in the ClinicalTrials.gov record. |
| Factorial design | Not provided in the ClinicalTrials.gov record. |
| Interim analysis | Not provided in the ClinicalTrials.gov record. |
| Alpha spending | Not provided in the ClinicalTrials.gov record. |
| Bayesian methods | Not provided in the ClinicalTrials.gov record. |
| Missing-data / imputation method | Not provided in the ClinicalTrials.gov record. |
| Stratification factors | Not provided in the ClinicalTrials.gov record. |
19. Limitations and Interpretation Issues
- Analysis-population differences: the efficacy population excludes patients identified to have a BRCA reversion mutation, while the ITT population includes all randomized patients in the treatment part.
- Time-to-event interpretation: hazard ratios are model-based relative measures and should not be interpreted as simple risk ratios or absolute differences.
- Proportional-hazards assumption: the Cox model's single hazard ratio is most straightforward to interpret when the relative hazards are reasonably represented by a proportional-hazards structure.
- Censoring: progression-free survival analyses necessarily involve follow-up that can end before an observed event. The ClinicalTrials.gov record does not provide the detailed censoring rules or counts needed to evaluate the censoring pattern.
- Secondary endpoint interpretation: DOR is evaluated in a response-defined population, so it does not estimate the probability that an arbitrary randomized patient will experience a durable response.
- Multiplicity: the ClinicalTrials.gov record identifies multiple primary and secondary analyses but do not specify the multiplicity-control strategy.
- Safety scope: only serious adverse-event counts by the reported study parts are available in the ClinicalTrials.gov record; a complete safety profile cannot be inferred.
- Missing statistical detail: the ClinicalTrials.gov record does not specify stratification factors, interim monitoring, imputation methods, or a detailed censoring algorithm.
- Limited absolute-effect information: median progression-free survival and time-specific Kaplan-Meier estimates are not contained in the ClinicalTrials.gov record and therefore are not reported here.
20. Why This Trial Matters Statistically
ARIEL4 is a useful teaching example because it demonstrates how a randomized clinical trial can use the same general time-to-event framework across endpoints while changing the analysis population according to the scientific question.
| Concept | How it appears in ARIEL4 |
|---|---|
| Randomization | The trial is randomized with a parallel design. |
| ITT analysis | One primary invPFS analysis is explicitly conducted in the ITT population. |
| Restricted efficacy population | The other primary invPFS analysis excludes patients identified to have a BRCA reversion mutation. |
| Time-to-event endpoints | Both primary endpoints and both DOR analyses are classified as time-to-event outcomes. |
| Cox regression | All four posted statistical analyses use Cox regression. |
| Hazard ratio | The treatment effect is reported as a Cox proportional-hazards estimate. |
| Confidence interval | All four analyses report two-sided 95% confidence intervals. |
| P-values | Each posted analysis includes a P-value for its superiority comparison. |
| Response-defined analysis | DOR is analyzed among patients with measurable disease at baseline and a confirmed response. |
| Safety population distinction | Serious adverse events are reported separately for treatment and crossover parts. |
21. A Practical Reading Strategy for ARIEL4
Step 1 · Identify the endpoint
Start with invPFS rather than the P-value. The endpoint determines the appropriate statistical framework.
Step 2 · Identify the population
Check whether the estimate belongs to the efficacy population or the ITT population before comparing estimates.
Step 3 · Read the HR
Interpret the hazard ratio as a relative time-to-event measure rather than as an absolute risk reduction.
Step 4 · Read the CI
Use the 95% confidence interval to understand the precision and uncertainty of the treatment-effect estimate.
Step 5 · Read the P-value
Use the P-value to understand statistical evidence against the null hypothesis, not to quantify effect size.
Step 6 · Check endpoint hierarchy
Keep the primary invPFS analyses separate from the secondary DOR analyses and avoid assuming an undocumented multiplicity procedure.
22. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: ARIEL4, NCT02855944. Official trial registry record and source for the registered endpoints and posted statistical analyses used on this page.
- PubMed: PMID 39914419.
- PubMed: PMID 35298906.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to explore the survival-analysis methods that underpin randomized time-to-event comparisons such as ARIEL4.
25. Record Summary
ARIEL4 provides a clear example of randomized time-to-event analysis in a phase 3 clinical trial. The study compares rucaparib with chemotherapy using investigator-assessed progression-free survival as two registered primary endpoints, analyzed in an efficacy population and an ITT population. Both primary analyses use Cox proportional-hazards regression and report hazard ratios below 1 with two-sided 95% confidence intervals and P-values of 0.0010 and 0.0017.
The statistical story extends to duration of response, where the registry also reports Cox-model hazard ratios of 0.589 in the efficacy population and 0.564 in the ITT response population. These secondary analyses require a narrower interpretation because DOR is assessed only among patients with measurable disease at baseline and a confirmed response.
The most important statistical lesson is to keep the endpoint, analysis population, effect measure, confidence interval, and P-value connected. A hazard ratio is not an absolute risk difference, a confidence interval is not a range of individual patient outcomes, and a P-value is not a measure of treatment-effect magnitude. Reading all of these quantities together provides a more complete understanding of what the ARIEL4 analyses actually estimate.