← Clinical Trials
Ovarian Cancer Phase 3 Time-to-Event NCT02470585

VELIA: Complete Statistical Analysis of Veliparib in Newly Diagnosed Advanced Ovarian Cancer

An independent statistical analysis of the randomized, double-blind phase 3 VELIA trial evaluating veliparib with carboplatin and paclitaxel and as continuation maintenance therapy in adults with newly diagnosed stage III or IV, high-grade serous, epithelial ovarian, fallopian tube, or primary peritoneal cancer.

Phase 3  ·  Randomized  ·  Double-blind  ·  Enrollment 1140
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics are restricted to the ClinicalTrials.gov record.

Registry context: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

VELIA was a randomized, double-blind, parallel phase 3 trial evaluating veliparib in combination with carboplatin and paclitaxel, followed by continuation maintenance therapy, in adults with newly diagnosed stage III or IV high-grade serous epithelial ovarian, fallopian tube, or primary peritoneal cancer.

1140
Enrollment
3 treatment arms
3
Primary endpoints
All time-to-event
0.435
BRCA-deficient PFS HR
95% CI 0.277–0.683
0.683
ITT PFS HR
95% CI 0.562–0.831
FeatureVELIA
Trial nameVELIA
PhasePhase 3
StatusTerminated
Start2015-07-14
Primary completion2019-05-03
PopulationAdults with newly diagnosed stage III or IV, high-grade serous, epithelial ovarian, fallopian tube, or primary peritoneal cancer
DesignRandomized, double-blind, parallel
AllocationRandomized
Arms3
Primary purposeTreatment
Lead sponsorAbbVie
Sponsor typeIndustry
ClinicalTrials.govNCT02470585

2. Clinical Question

The central statistical question was whether adding veliparib to carboplatin and paclitaxel, with veliparib continued as maintenance therapy, improved investigator-assessed progression-free survival compared with the corresponding placebo regimen. The registry evaluated this question sequentially in the BRCA-deficient population, the homologous recombination deficiency cohort, and the intention-to-treat population.

Population

Adults with newly diagnosed stage III or IV, high-grade serous, epithelial ovarian, fallopian tube, or primary peritoneal cancer. The primary analyses used sequentially inclusive BRCA-deficient, HRD, and ITT populations.

Intervention

Veliparib with carboplatin and paclitaxel, followed by veliparib continuation maintenance therapy.

Comparator

Placebo with carboplatin and paclitaxel, followed by placebo continuation maintenance therapy.

Primary question

Does the Arm 3 veliparib strategy improve PFS relative to Arm 1 placebo in the prespecified sequential populations?

3. Trial Design

01
Randomize1140 enrolled
02
3 armsParallel randomized design
03
CombinationCarboplatin + paclitaxel ± veliparib/placebo
04
MaintenanceVeliparib or placebo continuation
05
AssessPFS and OS time-to-event endpoints
ARM 1 · CONTROL

Placebo continuation strategy

  • Placebo to veliparib
  • Carboplatin
  • Paclitaxel
  • Continuation maintenance with placebo to veliparib
ARM 2 · INTERMEDIATE STRATEGY

Veliparib during combination treatment

  • Veliparib
  • Carboplatin
  • Paclitaxel
  • Continuation maintenance with placebo to veliparib
ARM 3 · CONTINUATION STRATEGY

Veliparib continuation strategy

  • Veliparib
  • Carboplatin
  • Paclitaxel
  • Continuation maintenance with veliparib
PRIMARY COMPARISON

Arm 3 versus Arm 1

  • Veliparib + carboplatin + paclitaxel → veliparib
  • versus placebo + carboplatin + paclitaxel → placebo
  • Primary efficacy comparison

4. Endpoints

The registry lists three primary endpoints, all based on progression-free survival. Each primary analysis compares Arm 3 with Arm 1 and uses a sequentially inclusive analysis strategy.

Primary endpointRegistry definitionTime frame
Progression-Free Survival (PFS) in the BRCA-deficient Population (Arm 3 vs Arm 1) PFS was defined as the time from the date that the participant was randomized to the date the participant experienced an event of disease progression, according to RECIST version 1.1 as determined by the investigator, or to the date of death if disease progression was not reached. From randomization until the primary analysis data cut-off date of 03 May 2019; the median duration of follow-up was 28 months as reported in the registry time frame.
Progression-Free Survival (PFS) in the Homologous Recombination Deficiency Cohort (Arm 3 vs Arm 1) PFS was defined as the time from the date that the participant was randomized to the date the participant experienced an event of disease progression, according to RECIST version 1.1 as determined by the investigator, or to the date of death if disease progression was not reached. From randomization until the primary analysis data cut-off date of 03 May 2019; the median duration of follow-up was 28 months as reported in the registry time frame.
Progression-Free Survival (PFS) in the Intention-to-treat Population (Arm 3 vs Arm 1) PFS was defined as the time from the date the participant was randomized to the date of disease progression according to RECIST version 1.1 as determined by the investigator, or to the date of death from all causes if disease progression was not reached. From randomization until the primary analysis data cut-off date of 03 May 2019; the median duration of follow-up was 28 months as reported in the registry time frame.

Secondary time-to-event endpoints

The registry also reports PFS comparisons of Arm 2 versus Arm 1 in the same three sequentially inclusive populations and overall survival comparisons in the BRCA-deficient, HRD, and whole populations. The OS time frame was from randomization to the end of the study, up to 98 months.

5. Statistical Methodology

Sequentially inclusive analysis populations

The primary efficacy analyses were conducted in three sequentially inclusive populations. The first was the BRCA-mutation cohort, the second expanded to the HRD cohort, and the third was the ITT population. The ITT population included all randomized participants.

PopulationRegistry descriptionStatistical role
BRCA-deficientParticipants with either a germline and/or tissue deleterious or suspected deleterious BRCA1/2 mutation.First primary PFS analysis.
HRDParticipants in the BRCA-deficient population and those determined to have HRD tumors based on HRD score.Second primary PFS analysis.
ITTAll randomized participants.Third primary PFS analysis.

Log-rank testing

The registry reports a log-rank test as the primary statistical method for the PFS and OS comparisons. Because PFS and OS are time-to-event outcomes, the log-rank test evaluates whether the observed event-time experience differs between treatment groups over follow-up.

Stratified analysis

The primary efficacy comparisons were stratified according to residual disease status and disease stage. For the ITT analysis, the registry additionally reports stratification according to the choice of the paclitaxel regimen and BRCA-mutation status. The Cox proportional-hazards model was stratified according to the same factors used in the corresponding log-rank test.

Hazard ratio

The treatment effect was expressed as a hazard ratio. For the primary PFS analyses, the registry reports hazard ratios from Cox proportional-hazards models alongside two-sided 95% confidence intervals.

Conceptual interpretation
HR = estimated instantaneous event rate in Arm 3 ÷ estimated instantaneous event rate in Arm 1

An HR below 1 indicates a lower estimated instantaneous event rate in Arm 3 than in Arm 1 under the fitted model. It is a relative time-to-event measure, not a direct statement about absolute survival probability or the percentage of participants who benefit.

Superiority framework

The registry identifies the hypothesis type as superiority. Thus, the relevant question is whether the observed time-to-event distributions provide evidence of a treatment difference favoring Arm 3, rather than whether Arm 3 merely satisfies a non-inferiority margin.

Multiplicity control

The registry explicitly identifies several sources of multiplicity: three treatment arms, two pairwise comparisons, three sequentially inclusive populations, and multiple endpoints. A fixed-sequence testing procedure was used to control the Type I error rate at 0.05.

Why this matters: the three primary PFS analyses should not be treated as three unrelated hypothesis tests. The fixed-sequence procedure establishes an order for formal testing and is part of the trial's Type I error control. A nominal P-value should therefore be interpreted in the context of the prespecified testing strategy rather than in isolation.

6. Results: Primary Progression-Free Survival Analyses

The registry contains formal statistical analyses for all three primary endpoints. All three compare the veliparib continuation strategy in Arm 3 with the placebo continuation strategy in Arm 1.

BRCA-deficient Population

Progression-free survival hazard ratio

0.435

95% CI: 0.277–0.683   ·   P < 0.001

Two-sided 95% confidence interval · Superiority hypothesis

The analysis used a log-rank test, with the comparison stratified according to residual disease status and disease stage. The registry also reports a Cox proportional-hazards model stratified according to the same factors.

Clinical Biostats interpretation

An HR of 0.435 means that, under the fitted time-to-event model, the estimated instantaneous rate of progression or death in Arm 3 was about 43.5% of that in Arm 1. Expressed as a relative hazard difference, this corresponds to an estimated 56.5% lower hazard for Arm 3 under the model.

This does not mean that 56.5% of participants avoided progression, that 56.5% of participants benefited, or that each participant experienced the same reduction. The hazard ratio summarizes a relative comparison of event rates over follow-up.

The 95% CI of 0.277–0.683 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of treatment effects that individual patients can experience.

The P-value of <0.001 addresses evidence against the null hypothesis in the prespecified testing framework; it does not measure the magnitude or clinical importance of the treatment effect. Because the analysis is based on a Cox model, the proportional-hazards assumption is also relevant when interpreting a single HR as a summary over time.

Homologous Recombination Deficiency Cohort

Progression-free survival hazard ratio

0.572

95% CI: 0.433–0.756   ·   P < 0.001

Two-sided 95% confidence interval · Superiority hypothesis

The HRD analysis expanded the first population to include the BRCA-mutation cohort and participants determined to have HRD tumors based on HRD score. The primary comparison remained Arm 3 versus Arm 1 and was analyzed using a stratified log-rank test and stratified Cox proportional-hazards model.

Clinical Biostats interpretation

An HR of 0.572 corresponds to an estimated instantaneous progression-or-death rate in Arm 3 that was about 57.2% of the Arm 1 rate under the fitted model. Equivalently, the estimated hazard was about 42.8% lower in Arm 3.

The 95% CI of 0.433–0.756 indicates uncertainty around that relative effect. Because the interval remains below 1, the point estimate and interval are consistent with a lower hazard in Arm 3 under the model.

The P-value of <0.001 indicates strong statistical evidence under the trial's testing framework, but a P-value is not an effect-size measure. The magnitude of the effect is better communicated by the HR together with its confidence interval.

The HRD analysis also illustrates why analysis populations matter. This is not simply another independent subgroup: the registry describes the populations as sequentially inclusive, and the fixed-sequence procedure accounts for that multiplicity structure.

Intention-to-Treat Population

Progression-free survival hazard ratio

0.683

95% CI: 0.562–0.831   ·   P < 0.001

Two-sided 95% confidence interval · Superiority hypothesis

The ITT analysis included all randomized participants. In addition to residual disease status and disease stage, the registry reports stratification according to the choice of the paclitaxel regimen and BRCA-mutation status for this analysis.

Clinical Biostats interpretation

An HR of 0.683 means that the estimated instantaneous rate of progression or death in Arm 3 was approximately 68.3% of the corresponding rate in Arm 1 under the fitted Cox model, or approximately 31.7% lower in relative terms.

The 95% CI of 0.562–0.831 provides the uncertainty interval for the estimated relative hazard. It is narrower than the BRCA-deficient estimate's interval in absolute width, consistent with the broader ITT analysis population containing more participants than the more restricted biomarker-defined population.

The P-value of <0.001 provides evidence against the null hypothesis in the prespecified superiority framework. It does not tell us that the effect is 31.7%, nor does it provide a probability that the treatment is effective.

The ITT result is particularly important statistically because randomization defines the comparison. The result should still be interpreted as a time-to-event estimate subject to censoring and the assumptions of the Cox model rather than as an absolute risk reduction.

Primary PFS populationComparisonHR95% CIP-value
BRCA-deficientArm 3 vs Arm 10.4350.277–0.683<0.001
HRDArm 3 vs Arm 10.5720.433–0.756<0.001
ITTArm 3 vs Arm 10.6830.562–0.831<0.001
Educational note: a Kaplan-Meier curve cannot be reconstructed reliably from hazard ratios, confidence intervals, and P-values alone. The ClinicalTrials.gov record does not provide the participant-level event and censoring information required for an independent curve reconstruction.

7. Secondary Progression-Free Survival Analyses

The registry also reports PFS comparisons of Arm 2 versus Arm 1 in the same three sequentially inclusive populations. These analyses use the same general log-rank and stratified Cox framework.

PopulationComparisonHR95% CIP-value
BRCA-deficientArm 2 vs Arm 11.2150.821–1.7990.335
HRDArm 2 vs Arm 11.1000.855–1.4140.462
ITTArm 2 vs Arm 11.0730.895–1.2870.450

These estimates are descriptive of the registry-reported comparisons. For all three analyses, the confidence interval includes 1, and the registry reports P-values of 0.335, 0.462, and 0.450, respectively. The absence of a statistically significant comparison should not be converted into a claim that the two treatment strategies are identical; the confidence interval is the more informative description of the range of relative effects compatible with the data.

8. Secondary Overall Survival Results

The registry reports overall survival analyses from randomization to the end of the study, up to 98 months. These are time-to-event analyses using log-rank tests, with hazard ratios from Cox proportional-hazards models.

BRCA-deficient Population: Arm 3 vs Arm 1

Overall survival hazard ratio

0.900

95% CI: 0.567–1.429   ·   P = 0.328

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The HR of 0.900 corresponds to an estimated instantaneous death rate in Arm 3 about 90.0% of the Arm 1 rate under the fitted model. The estimate alone therefore suggests a relative hazard below 1, but the confidence interval extends from 0.567 to 1.429, spanning both lower and higher hazards.

The P-value of 0.328 does not provide evidence against the null hypothesis at conventional significance levels. It should not be interpreted as proof of no survival difference, because the confidence interval permits a range of effects in either direction.

BRCA-deficient Population: Arm 2 vs Arm 1

Overall survival hazard ratio

1.218

95% CI: 0.780–1.903   ·   P = 0.808

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The HR of 1.218 is above 1, corresponding to an estimated instantaneous death rate in Arm 2 approximately 21.8% higher than Arm 1 under the fitted model. However, the 95% CI of 0.780–1.903 includes 1 and spans a broad range of possible relative hazards.

The P-value of 0.808 does not provide evidence of a statistically detectable difference in the reported comparison. It does not establish equivalence between Arm 2 and Arm 1.

HRD Population: Arm 3 vs Arm 1

Overall survival hazard ratio

0.844

95% CI: 0.640–1.114   ·   P = 0.116

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The HR of 0.844 corresponds to an estimated instantaneous death rate approximately 15.6% lower in Arm 3 than Arm 1 under the model. The 95% CI of 0.640–1.114 crosses 1, so the interval remains compatible with no difference as well as with a lower or higher hazard.

The P-value of 0.116 should be interpreted as evidence insufficient to reject the relevant null hypothesis in this reported comparison, rather than as evidence that the two strategies have identical survival.

HRD Population: Arm 2 vs Arm 1

Overall survival hazard ratio

0.949

95% CI: 0.726–1.242   ·   P = 0.352

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The HR of 0.949 is close to 1, but the confidence interval of 0.726–1.242 includes both values below and above 1. The P-value of 0.352 does not provide evidence of a statistically detectable difference in this comparison.

Whole Population: Arm 3 vs Arm 1

Overall survival hazard ratio

0.946

95% CI: 0.782–1.144   ·   P = 0.283

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The HR of 0.946 represents an estimated instantaneous death rate approximately 5.4% lower in Arm 3 than Arm 1 under the fitted model. The 95% CI of 0.782–1.144 crosses 1, so the estimate is compatible with both lower and higher hazards.

The P-value of 0.283 does not provide evidence against the null hypothesis in this reported comparison. Because this is a whole-population OS analysis, it should not be substituted for the primary PFS analyses or treated as though it were the same endpoint.

Whole Population: Arm 2 vs Arm 1

Overall survival hazard ratio

1.034

95% CI: 0.859–1.244   ·   P = 0.638

Two-sided 95% confidence interval · Superiority hypothesis

Clinical Biostats interpretation

The HR of 1.034 corresponds to an estimated instantaneous death rate about 3.4% higher in Arm 2 than Arm 1 under the fitted model. The 95% CI of 0.859–1.244 includes 1, and the P-value of 0.638 does not provide evidence of a statistically detectable difference in this comparison.

OS populationComparisonHR95% CIP-value
BRCA-deficientArm 3 vs Arm 10.9000.567–1.4290.328
BRCA-deficientArm 2 vs Arm 11.2180.780–1.9030.808
HRDArm 3 vs Arm 10.8440.640–1.1140.116
HRDArm 2 vs Arm 10.9490.726–1.2420.352
Whole populationArm 3 vs Arm 10.9460.782–1.1440.283
Whole populationArm 2 vs Arm 11.0340.859–1.2440.638

9. Safety Results

The registry reports serious adverse events by treatment arm using affected participants divided by participants at risk.

ArmSerious adverse events affected / at risk
Arm 1: Placebo + Carboplatin + Paclitaxel → Placebo143/375
Arm 3: Veliparib + Carboplatin + Paclitaxel → Veliparib130/383
Arm 2: Veliparib + Carboplatin + Paclitaxel → Placebo146/382
Serious adverse events: affected participants / at risk
Arm 1
143/375
Arm 3
130/383
Arm 2
146/382
Statistical caution: the displayed bar widths are simple descriptive representations of the reported affected/at-risk counts. The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates, so these figures should not be interpreted as evidence of a statistically significant difference between arms.

10. Statistical Methods Explained

Why was a log-rank test used?

PFS and OS are time-to-event endpoints, so the exact timing of progression, death, and censoring matters. A log-rank test compares the event experience between groups across follow-up rather than reducing each participant to a simple yes/no outcome at a single fixed time.

What does a hazard ratio of 0.435 mean?

An HR of 0.435 means that the estimated instantaneous event rate in Arm 3 was 43.5% of the corresponding rate in Arm 1 under the fitted Cox model. It can also be described as an estimated 56.5% lower hazard. It does not mean a 56.5% absolute reduction in the probability of progression or death.

Why were the BRCA, HRD, and ITT populations analysed sequentially?

The registry describes these populations as sequentially inclusive. The BRCA-mutation cohort forms the first analysis population; the HRD cohort adds participants identified through the HRD definition; and the ITT population includes all randomized participants. This structure lets the trial examine treatment effects from a biomarker-defined population through a broader population while incorporating the prespecified multiplicity strategy.

Why does stratification matter?

Stratification allows the time-to-event comparison to account for prespecified factors that may influence prognosis or treatment balance. For the primary analyses, residual disease status and disease stage were used. The ITT analysis additionally used the choice of paclitaxel regimen and BRCA-mutation status.

Why is the confidence interval more informative than the P-value alone?

The P-value addresses the strength of evidence against a null hypothesis, whereas the confidence interval describes uncertainty around the effect estimate. For example, the ITT PFS HR of 0.683 has a 95% CI of 0.562–0.831. Showing both communicates both the estimated relative effect and the uncertainty around it.

Why should an HR below 1 not be read as a percentage of patients helped?

A hazard ratio is a relative model-based measure of event rates over time. It does not directly provide the proportion of patients who benefit, the absolute risk reduction, or a patient's individual probability of remaining progression-free. Those questions require different estimands and, when available, absolute survival estimates or other clinically interpretable measures.

Why does multiplicity matter here?

The trial involved three arms, two pairwise comparisons, three sequentially inclusive analysis populations, and multiple endpoints. The registry states that a fixed-sequence testing procedure was used to control the Type I error rate at 0.05. Therefore, statistical evidence should be interpreted within that prespecified sequence rather than treating every reported P-value as an isolated test.

11. How to Read the Primary PFS Pattern

The three primary PFS estimates form a useful statistical teaching example because the point estimates change as the analysis population broadens. The BRCA-deficient analysis reports an HR of 0.435, the HRD analysis reports 0.572, and the ITT analysis reports 0.683. These are not contradictory results: they answer the same broad treatment question in progressively broader populations.

Primary PFS hazard ratios
BRCA-deficient
0.435
HRD
0.572
ITT
0.683

The visual scale above is simply the reported HR expressed relative to 1.0; it is not a confidence-interval plot and does not replace the formal estimates and intervals. In all three primary analyses, the reported 95% confidence interval lies below 1 and the P-value is <0.001. The statistical interpretation therefore depends on both the point estimate and uncertainty interval, together with the prespecified fixed-sequence testing procedure.

Important distinction: a change in the HR across populations does not by itself demonstrate biological effect modification. Formal claims that treatment effects differ between populations require an appropriate interaction or heterogeneity analysis. The ClinicalTrials.gov record does not provide such an interaction test.

12. Intention-to-Treat Analysis

The ITT population is explicitly defined in the registry as all randomized participants. This is important because randomization creates the basis for the treatment comparison. Analysing participants according to their randomized group preserves that assignment rather than redefining groups according to subsequent treatment exposure.

For VELIA, the ITT primary PFS analysis produced an HR of 0.683 with a two-sided 95% CI of 0.562–0.831 and P < 0.001. The analysis was stratified according to residual disease status, disease stage, choice of the paclitaxel regimen, and BRCA-mutation status.

ITT analysis does not eliminate all statistical complications. Time-to-event analyses still require appropriate handling of censoring, and the Cox model relies on assumptions about the relationship between treatment and the hazard over time. The ClinicalTrials.gov record does not report a separate missing-data or imputation procedure for the primary PFS analyses.

13. Time-to-Event Analysis and Censoring

PFS and OS are fundamentally different from binary endpoints measured at a fixed time. A participant may be followed for different lengths of time, may experience the event during follow-up, or may be censored before an event is observed. Kaplan-Meier methods are designed to estimate the survival function under right censoring, while the log-rank test and Cox model compare groups using the observed event-time information.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

where di is the number of events at time ti and ni is the number at risk immediately before that time. The registry data posted on ClinicalTrials.gov for VELIA report formal log-rank and Cox analyses but do not provide participant-level information needed to reconstruct the complete Kaplan-Meier curves.

The primary PFS endpoint uses investigator-assessed progression according to RECIST version 1.1 or death if progression was not reached. This definition means that death is treated as a PFS event rather than as a separate competing endpoint.

14. Hazard Ratios and Proportional Hazards

The Cox proportional-hazards model is the principal model underlying the reported hazard ratios. The model summarizes the relative hazard between treatment groups while accounting for the stratification factors specified for the analysis.

Hazard ratio scale
HR = 1  →  no relative hazard difference
HR < 1  →  lower estimated hazard in Arm 3
HR > 1  →  higher estimated hazard in Arm 3

The interpretation assumes the comparison is defined as Arm 3 relative to Arm 1. The secondary analyses involving Arm 2 use Arm 2 relative to Arm 1.

A single HR is most straightforward to interpret when the proportional-hazards assumption is reasonable. If relative hazards vary substantially over time, a single HR may compress a changing treatment effect into one summary number. The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic or an alternative time-varying effect estimate.

15. Multiplicity and Fixed-Sequence Testing

Multiplicity is one of the central statistical features of VELIA. The registry explicitly identifies three treatment arms, two pairwise comparisons, three sequentially inclusive populations, and multiple endpoints as sources of multiplicity.

Multiplicity sourceWhat it means
Three treatment armsThere are three randomized strategies rather than a single two-arm comparison.
Two pairwise comparisonsThe registry reports Arm 3 vs Arm 1 and Arm 2 vs Arm 1 comparisons.
Three sequential populationsBRCA-deficient, HRD, and ITT populations were analysed in sequence.
Multiple endpointsThe registry reports three primary PFS endpoints as well as secondary PFS and OS endpoints.
Fixed-sequence procedureThe registry states that a fixed-sequence testing procedure was used to control the Type I error rate at 0.05.

A fixed-sequence strategy is useful because it specifies an ordered testing hierarchy. The statistical interpretation of a later test can depend on whether the preceding test has met its criterion. This is different from simply running several tests independently and treating each P-value as though no other hypotheses existed.

What the registry does not provide here: the ClinicalTrials.gov record does not give the complete numerical alpha-allocation details or the full statistical analysis plan. The appropriate conclusion is therefore limited to the documented statement that a fixed-sequence procedure was used to control Type I error at 0.05.

16. What the P-Values Do — and Do Not — Mean

The three primary PFS analyses all report P < 0.001. That is strong statistical evidence under the relevant testing framework, but it should not be mistaken for a measure of treatment magnitude.

What a P-value addresses

It quantifies how incompatible the observed data are with the specified null hypothesis, under the statistical model and testing procedure.

What it does not address

It does not tell us the probability that the null hypothesis is true, the probability that the treatment works for an individual, or the clinical importance of an effect.

Why the HR matters

The HR describes the estimated relative treatment effect. For example, the ITT PFS HR is 0.683.

Why the CI matters

The 95% CI describes uncertainty around the estimated HR. For the ITT PFS analysis it is 0.562–0.831.

17. Overall Survival Versus Progression-Free Survival

The registry's primary endpoints are PFS endpoints, while OS is reported as a secondary endpoint. These outcomes answer related but different questions.

EndpointWhat it measuresVELIA statistical role
PFSTime from randomization to disease progression or death according to the registered definition.Three primary endpoints for Arm 3 vs Arm 1.
OSTime from randomization to death.Secondary endpoint analyses in BRCA-deficient, HRD, and whole populations.

The distinction matters because an intervention can affect progression timing without producing the same magnitude of effect on overall survival. OS is also influenced by events and treatments occurring after progression. The ClinicalTrials.gov record does not provide subsequent-treatment information that would allow a more detailed decomposition of the OS results.

18. Results Summary

EndpointComparisonPopulationHR95% CIP-value
Primary PFSArm 3 vs Arm 1BRCA-deficient0.4350.277–0.683<0.001
Primary PFSArm 3 vs Arm 1HRD0.5720.433–0.756<0.001
Primary PFSArm 3 vs Arm 1ITT0.6830.562–0.831<0.001
Secondary PFSArm 2 vs Arm 1BRCA-deficient1.2150.821–1.7990.335
Secondary PFSArm 2 vs Arm 1HRD1.1000.855–1.4140.462
Secondary PFSArm 2 vs Arm 1ITT1.0730.895–1.2870.450
Secondary OSArm 3 vs Arm 1BRCA-deficient0.9000.567–1.4290.328
Secondary OSArm 2 vs Arm 1BRCA-deficient1.2180.780–1.9030.808
Secondary OSArm 3 vs Arm 1HRD0.8440.640–1.1140.116
Secondary OSArm 2 vs Arm 1HRD0.9490.726–1.2420.352
Secondary OSArm 3 vs Arm 1Whole population0.9460.782–1.1440.283
Secondary OSArm 2 vs Arm 1Whole population1.0340.859–1.2440.638

19. Important Limitations

20. Why This Trial Matters Statistically

VELIA is a useful teaching case because the registry results bring together several central principles of clinical-trial statistics: randomized allocation, double blinding, multiple treatment arms, sequentially inclusive biomarker-defined populations, time-to-event endpoints, stratified log-rank testing, Cox proportional-hazards models, intention-to-treat analysis, hazard ratios, confidence intervals, P-values, and multiplicity control.

ConceptHow it appears in VELIA
RandomizationThe study used randomized allocation in a parallel phase 3 design.
BlindingThe study was double-blind.
Three-arm designThree treatment strategies were evaluated.
Time-to-event endpointsAll three registered primary endpoints were PFS endpoints, and OS was also analysed.
Log-rank testUsed for the reported PFS and OS treatment comparisons.
Hazard ratioUsed to express the relative treatment effect in the Cox proportional-hazards analysis.
Stratified analysisPrimary comparisons were stratified by prespecified clinical and treatment factors.
ITT analysisThe third primary analysis included all randomized participants.
MultiplicityThree arms, pairwise comparisons, sequential populations, and multiple endpoints were addressed with a fixed-sequence procedure.
Confidence intervalsAll primary analyses report two-sided 95% confidence intervals around the HR.

The particularly instructive feature is the relationship between analysis population and effect estimate. The primary PFS HR moves from 0.435 in the BRCA-deficient population to 0.572 in the HRD population and 0.683 in the ITT population. That progression demonstrates why the population attached to an effect estimate is as important as the estimate itself.

21. Planned Statistical Interpretation Framework

For a trial with this design, the statistical analysis can be understood as a sequence of linked questions:

01
RandomizePreserve treatment comparability
02
Define PFSProgression or death
03
Compare curvesLog-rank framework
04
Estimate HRCox model
05
Control errorFixed sequence

That framework prevents a common statistical mistake: interpreting a single P-value without identifying the endpoint, analysis population, treatment contrast, estimand, model, confidence interval, and multiplicity structure that generated it.

22. Limitations of the Reported Evidence for Secondary Analyses

The secondary PFS and OS analyses illustrate why a complete clinical-trial interpretation should preserve the distinction between primary confirmatory questions and additional analyses. The ClinicalTrials.gov record reports formal statistical analyses for 12 outcome measures, but only three are identified as primary endpoints.

For the secondary OS analyses, all six reported confidence intervals include 1. The corresponding P-values range from 0.116 to 0.808. These results are useful for describing the registry-reported estimates, but they should not be transformed into a binary statement that the treatment has either "no effect" or a proven survival benefit. The confidence intervals communicate that the observed estimates remain uncertain.

Likewise, the secondary Arm 2 versus Arm 1 PFS estimates are 1.215, 1.100, and 1.073 across the BRCA-deficient, HRD, and ITT populations. Each confidence interval includes 1. These results should be interpreted as estimates with uncertainty rather than as evidence of equivalence.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Calculators

25. Sources

Continue through the Clinical Biostats statistical methods library

Explore the statistical concepts behind randomized clinical trials, survival analysis, confidence intervals, and treatment-effect estimation.

26. Record Summary

VELIA provides a compact teaching example of how a randomized phase 3 trial can combine a three-arm design with sequentially inclusive biomarker-defined analysis populations and time-to-event methodology. The three primary PFS analyses compare Arm 3 with Arm 1 and report HRs of 0.435 in the BRCA-deficient population, 0.572 in the HRD cohort, and 0.683 in the ITT population, with two-sided 95% confidence intervals and P-values <0.001 for all three analyses.

The same registry record reports secondary PFS analyses comparing Arm 2 with Arm 1 and secondary OS analyses across BRCA-deficient, HRD, and whole populations. These estimates should be interpreted with attention to their analysis populations, confidence intervals, and the trial's multiplicity structure rather than by P-value alone.

Clinical Biostats methodology: The statistical story of a trial is more than its headline P-value. A rigorous interpretation identifies the population, treatment contrast, endpoint definition, time frame, analysis method, effect estimate, confidence interval, multiplicity strategy, and assumptions that determine what the result actually tells us.