← Clinical Trials
Metastatic Colorectal Cancer Phase 3 Time-to-Event Analysis NCT00561470

VELOUR: Complete Statistical Analysis of Aflibercept in Metastatic Colorectal Cancer

An independent statistical review of the randomized phase 3 VELOUR trial evaluating aflibercept plus FOLFIRI versus placebo plus FOLFIRI in patients with metastatic colorectal cancer after failure of an oxaliplatin-based regimen.

Trial status: Completed  ·  Enrollment: 1226  ·  Primary completion: 2011-02
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results and trial characteristics presented here are limited to the ClinicalTrials.gov record data. Where the registry does not provide an estimate, confidence interval, or additional result, this page does not infer one.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

VELOUR was a randomized, parallel-group, triple-masked phase 3 treatment trial enrolling 1226 participants with metastatic colorectal cancer after failure of an oxaliplatin-based regimen. The registered comparison was aflibercept plus FOLFIRI versus placebo plus FOLFIRI.

1226
Enrollment
Phase 3
2
Arms
Parallel design
0.817
OS HR
95.34% CI 0.713–0.937
0.758
PFS HR
99.99% CI 0.578–0.995
FeatureVELOUR
Trial nameVELOUR
NCT identifierNCT00561470
PhasePhase 3
StatusCompleted
Therapeutic areaOncology
ConditionsColorectal Neoplasms; Neoplasm Metastasis
Enrollment1226
AllocationRandomized
Design modelParallel
MaskingTriple
Primary purposeTreatment
Lead sponsorSanofi
Sponsor typeIndustry
Start2007-11
Primary completion2011-02
Results postedYes

2. Clinical Question

The central statistical question was whether adding aflibercept to FOLFIRI improved time-to-event outcomes compared with placebo plus FOLFIRI in patients with metastatic colorectal cancer after failure of an oxaliplatin-based regimen.

Population

Patients with metastatic colorectal cancer after failure of an oxaliplatin-based regimen.

Intervention

Aflibercept (ziv-aflibercept, AVE0005, VEGF trap, ZALTRAP®) in combination with FOLFIRI.

Comparator

Placebo in combination with FOLFIRI.

Primary question

Does aflibercept plus FOLFIRI improve overall survival relative to placebo plus FOLFIRI?

3. Trial Design

01
Randomize1226 participants
02
Two armsParallel treatment groups
03
Triple maskedRegistered masking designation
04
Follow-upTime-to-event assessment
05
AnalysisITT efficacy analysis
Allocation
Randomized
Participants were assigned to randomized treatment groups.
Design model
Parallel
The registry identifies a two-arm parallel design.
Masking
Triple
The trial's registry designation is triple masked.
Primary purpose
Treatment
The registered primary purpose is treatment.
TREATMENT ARM

Aflibercept / FOLFIRI

  • Aflibercept (ziv-aflibercept, AVE0005, VEGF trap, ZALTRAP®)
  • FOLFIRI: irinotecan, 5-fluorouracil, and leucovorin
CONTROL ARM

Placebo / FOLFIRI

  • Placebo
  • FOLFIRI: irinotecan, 5-fluorouracil, and leucovorin

4. Endpoints

The registry identifies one primary endpoint, Overall Survival (OS), and reports statistical analyses for OS, progression-free survival assessed by an Independent Review Committee, and overall objective response rate.

EndpointTypeRegistry time frameAnalysis reported
Overall Survival (OS) Time-to-event From the date of the first randomization until the study data cut-off date, 07 February 2011 (approximately three years) Stratified log-rank test; stratified hazard ratio
Progression-free Survival (PFS) Assessed by Independent Review Committee (IRC) Time-to-event From the date of the first randomization until the occurrence of 561 OS events, 06 May 2010 (approximately 30 months) Stratified log-rank test; stratified hazard ratio
Overall Objective Response Rate (ORR) Based on the Tumor Assessment by the Independent Review Committee (IRC) as Per Response Evaluation Criteria in Solid Tumours (RECIST) Criteria Binary From the date of the first randomization until the study data cut-off date, 06 May 2010 (approximately 30 months) Stratified Cochran-Mantel-Haenszel test

Primary endpoint definition: Overall Survival

ClinicalTrials.gov defines overall survival as the time interval from the date of randomization to the date of death due to any cause. Once disease progression was documented, participants were followed every 2 months for survival status, until death or until the study cutoff date, whichever came first. The final data cutoff date for the analysis of OS was the date when 863 deaths had occurred, 07 February 2011.

Endpoint hierarchy: OS is the single registered primary endpoint in the ClinicalTrials.gov record. PFS and ORR are represented in the posted statistical analyses as secondary endpoints.

5. Statistical Methodology

Intention-to-treat analysis

The OS analysis population was the intent-to-treat population, defined in the registry as all participants who gave informed consent and were randomized. The PFS analysis likewise used an ITT population consisting of all participants who gave informed consent and were randomized.

This distinction is important because the treatment comparison remains anchored to randomized assignment rather than being restricted to participants who completed treatment. In a randomized trial, preserving the randomized population helps maintain the comparability created by randomization.

Stratified log-rank test

The primary OS comparison used a stratified log-rank test. The PFS comparison also used a stratified log-rank test. Both analyses were stratified on ECOG Performance Status (0 vs 1 vs 2) and prior Bevacizumab (yes vs no) according to IVRS.

Stratification allows the time-to-event comparison to account for prespecified factors used to structure the randomized comparison. Rather than treating all participants as belonging to one undifferentiated risk set, the analysis compares treatment groups within the specified strata and combines the information across them.

Stratified Cox proportional-hazards model

The OS analysis notes specify a Cox Proportional Hazard Model stratified on ECOG Performance Status and prior Bevacizumab. The reported effect measure was a stratified hazard ratio.

The PFS analysis similarly used a stratified Cox Proportional Hazard Model, with the same two stratification factors. The hazard ratio is therefore a relative time-to-event measure rather than a direct comparison of proportions at a single time point.

Cochran-Mantel-Haenszel analysis

The ORR analysis used a stratified Cochran-Mantel-Haenszel test. The registry identifies ORR as a binary endpoint and states that the analysis was stratified on ECOG Performance Status (0 vs 1 vs 2) and prior Bevacizumab (yes vs no) according to IVRS.

The ClinicalTrials.gov record reports a P-value and confidence-interval level for this analysis but does not provide an ORR effect estimate or its confidence limits. Accordingly, this page does not construct an unreported effect measure.

Interim analysis and alpha spending

The OS analysis notes specify a significance threshold of 0.0466 using the O'Brien-Fleming alpha spending function. The statistical-analysis metadata also identifies interim analysis / alpha spending as an analysis concept.

This matters because a trial that permits an interim efficacy assessment cannot generally be interpreted as though the data were examined only once at the end. Alpha spending provides a framework for controlling the prespecified type I error while allowing information to be evaluated during the study.

6. Results: Overall Survival

Overall survival was the registered primary endpoint. The analysis was performed in the ITT population and compared placebo/FOLFIRI with aflibercept/FOLFIRI using a stratified log-rank test.

Stratified hazard ratio for overall survival

0.817

95.34% CI: 0.713–0.937   ·   P = 0.0032

Two-sided confidence interval · Superiority hypothesis

OS analysis characteristicReported value
Analysis populationIntent-to-treat population
Groups comparedPlacebo/FOLFIRI vs Aflibercept/FOLFIRI
MethodStratified log-rank test
Effect measureStratified Hazard Ratio
Estimate0.817
Confidence interval95.34% CI 0.713–0.937
P-value0.0032
HypothesisSuperiority
OS data cutoff07 February 2011
Event basis for final OS cutoff863 deaths
Clinical Biostats interpretation

The estimated OS hazard ratio of 0.817 means that, under the stratified Cox model used for the analysis, the estimated instantaneous hazard of death in the aflibercept/FOLFIRI group was about 18.3% lower than in the placebo/FOLFIRI group. This is a relative model-based comparison of event hazards, not a statement that 18.3% of participants benefited or that each participant experienced an 18.3% reduction in their individual probability of death.

The 95.34% confidence interval of 0.713–0.937 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of effects that individual patients would experience.

The P-value of 0.0032 addresses evidence against the null hypothesis under the specified statistical test; it does not measure the magnitude or clinical importance of the treatment effect. Effect size is conveyed by the hazard ratio and its confidence interval.

The OS analysis was stratified on ECOG Performance Status and prior Bevacizumab, and the registry notes a significance threshold of 0.0466 using an O'Brien-Fleming alpha spending function. The interpretation therefore belongs to the prespecified group-sequential framework rather than to an unadjusted single-look test.

Finally, the hazard-ratio interpretation relies on the Cox proportional-hazards modeling framework. A single hazard ratio summarizes a relative event-rate relationship over the analyzed period; it should not be read as an absolute survival difference or as proof that the hazards are identical in relative proportion at every point in time.

What the OS result does not tell us

7. Results: Progression-Free Survival

Progression-free survival assessed by an Independent Review Committee was reported as a secondary time-to-event endpoint. The analysis used the ITT population and a stratified log-rank test.

Stratified hazard ratio for PFS

0.758

99.99% CI: 0.578–0.995   ·   P = 0.00007

Two-sided confidence interval · Superiority hypothesis

PFS analysis characteristicReported value
Analysis populationIntent-to-treat population
Groups comparedPlacebo/FOLFIRI vs Aflibercept/FOLFIRI
AssessmentIndependent Review Committee (IRC)
MethodStratified log-rank test
Effect measureStratified Hazard ratio
Estimate0.758
Confidence interval99.99% CI 0.578–0.995
P-value0.00007
HypothesisSuperiority
Analysis cutoff06 May 2010
Event/time basisUntil occurrence of 561 OS events; approximately 30 months
Clinical Biostats interpretation

The PFS hazard ratio of 0.758 corresponds to an estimated instantaneous hazard of progression or death about 24.2% lower with aflibercept/FOLFIRI than with placebo/FOLFIRI under the reported stratified Cox model.

The 99.99% confidence interval of 0.578–0.995 is the registry-reported uncertainty interval for this estimate. The unusually high confidence level should be retained exactly as reported rather than silently replacing it with a conventional 95% interval.

The P-value of 0.00007 quantifies evidence under the specified statistical testing framework. It is not an effect-size measure. The hazard ratio and its confidence interval provide the information about the estimated relative magnitude and precision of the treatment comparison.

The PFS analysis was stratified on ECOG Performance Status and prior Bevacizumab. Because PFS is a time-to-event endpoint, censoring and the assumptions underlying the survival-analysis framework remain relevant to interpretation.

Educational note: the ClinicalTrials.gov record provides summary hazard-ratio results but not the underlying individual event and censoring times. A Kaplan-Meier curve should therefore not be reconstructed from these summary statistics alone.

8. Results: Overall Objective Response Rate

Overall objective response rate was reported as a secondary binary endpoint based on tumor assessment by the Independent Review Committee according to RECIST criteria.

ORR analysis characteristicReported value
Analysis populationEvaluable patient population for tumor response
Population definitionAll randomized participants with measurable disease at study entry, as per IRC evaluation, and with at least one valid post-baseline assessment
Groups comparedPlacebo/FOLFIRI vs Aflibercept/FOLFIRI
Endpoint typeBinary
MethodStratified Cochran-Mantel-Haenszel
Confidence interval level95%
P-value0.0001
HypothesisSuperiority
Analysis cutoff06 May 2010
Clinical Biostats interpretation

The registry reports a P-value of 0.0001 for the stratified Cochran-Mantel-Haenszel comparison of ORR. The ClinicalTrials.gov record does not provide the corresponding response-rate estimates or a confidence interval for the treatment effect.

Therefore, the appropriate interpretation is limited to the reported statistical comparison. It would be inappropriate to reconstruct an ORR difference, risk ratio, odds ratio, or response-rate percentage from the P-value alone.

The analysis population also differs from the ITT population used for OS and PFS: the response analysis used an evaluable patient population with measurable disease at study entry and at least one valid post-baseline assessment. That distinction matters when comparing the interpretation of response with the primary survival analysis.

9. Secondary Endpoint Analysis and Population Differences

VELOUR illustrates an important principle in clinical-trial statistics: different endpoints can legitimately use different analysis populations and statistical procedures when their measurement requirements differ.

EndpointPopulationData typeStatistical method
Overall SurvivalITTTime-to-eventStratified log-rank test; stratified Cox model for HR
PFS by IRCITTTime-to-eventStratified log-rank test; stratified Cox model for HR
ORR by IRC / RECISTEvaluable patient populationBinaryStratified Cochran-Mantel-Haenszel test

The survival analyses preserve the randomized treatment assignment through ITT analysis. The response analysis instead requires measurable disease and a valid post-baseline tumor assessment, which is why the registry specifies a different evaluable population.

This difference does not make the ORR analysis invalid. It means that ORR and survival answer different questions and are estimated in populations defined by different information requirements.

10. Stratification

The reported OS and PFS analyses were stratified on two factors: ECOG Performance Status (0 vs 1 vs 2) and prior Bevacizumab (yes vs no), according to IVRS. The ORR analysis used the same stratification factors.

ECOG Performance Status

The registered analysis distinguishes ECOG Performance Status categories 0, 1, and 2. The analysis accounts for these categories through stratification.

Prior Bevacizumab

Prior Bevacizumab was categorized as yes versus no and incorporated into the reported stratified analyses.

Stratification is particularly useful when an important prognostic factor could influence the event distribution. The objective is not to create a separate treatment effect for every subgroup, but to perform the primary comparison while accounting for the prespecified strata.

Conceptual structure
Treatment comparison = information combined across prespecified strata

The stratified analysis compares the randomized groups while conditioning on the registered stratification factors rather than ignoring them.

11. Interim Analysis and Alpha Spending

The registry-reported OS analysis notes specify that the significance threshold was set to 0.0466 using the O'Brien-Fleming alpha spending function. Interim analysis / alpha spending is also identified among the concepts in the analysis text.

Why alpha spending matters

When accumulating trial data are examined during the study, the statistical framework must account for those repeated opportunities to evaluate efficacy. Alpha spending allocates the available type I error across information times.

O'Brien-Fleming principle

An O'Brien-Fleming approach places a stringent evidentiary requirement on an early look and permits a less stringent boundary as information accumulates toward the final analysis.

The important practical point is that the reported OS P-value should be interpreted in the context of the prespecified group-sequential framework. The threshold of 0.0466 is part of that framework and should not be replaced by an assumption that the only relevant threshold was an unadjusted 0.05.

Do not infer an unreported stopping rule: the ClinicalTrials.gov record identifies the O'Brien-Fleming alpha spending function and significance threshold but does not provide enough information here to reconstruct the complete interim boundary schedule. This page therefore does not infer additional interim thresholds or stopping decisions.

12. What the Hazard Ratio Means in VELOUR

OS hazard ratio

The OS HR of 0.817 means the fitted model estimated a lower instantaneous hazard of death in the aflibercept/FOLFIRI group relative to placebo/FOLFIRI. Numerically, 1 − 0.817 gives approximately an 18.3% lower estimated hazard.

PFS hazard ratio

The PFS HR of 0.758 means the fitted model estimated a lower instantaneous hazard of progression or death in the aflibercept/FOLFIRI group. Numerically, 1 − 0.758 gives approximately a 24.2% lower estimated hazard.

What neither HR means

Neither hazard ratio is a probability that an individual patient will experience an event. Neither is an absolute risk reduction, a median-survival difference, or a statement that the same proportional reduction applies identically to every participant.

The two hazard ratios also describe different endpoints. OS concerns death from any cause, whereas the registered PFS analysis concerns progression-free survival assessed by an Independent Review Committee. They should therefore not be combined into one generic measure of "benefit."

13. Confidence Intervals and Precision

Confidence intervals provide information that a point estimate alone cannot. The OS estimate of 0.817 is accompanied by a 95.34% CI of 0.713–0.937, while the PFS estimate of 0.758 is accompanied by a 99.99% CI of 0.578–0.995.

EndpointEstimateConfidence levelConfidence interval
Overall Survival0.81795.34%0.713–0.937
Progression-free Survival0.75899.99%0.578–0.995

The confidence interval is about uncertainty in the estimated treatment effect under the statistical model and sampling framework. It is not a range containing the true treatment effect for every individual participant, nor does it represent the distribution of individual responses.

The different confidence levels should also be preserved exactly as reported. A confidence interval cannot be compared solely by looking at its numerical width without considering the confidence level at which it was constructed.

14. Statistical Methods Explained

Why was a stratified log-rank test used for OS and PFS?

OS and PFS are time-to-event endpoints, so the analysis must account for both event timing and censoring. The stratified log-rank test provides a way to compare the event-time distributions between randomized treatment groups while accounting for the prespecified stratification factors.

What does an OS hazard ratio of 0.817 mean?

It is a relative estimate from the stratified Cox model. Under that model, the aflibercept/FOLFIRI group had an estimated instantaneous hazard of death approximately 18.3% lower than the placebo/FOLFIRI group. It is not an 18.3-percentage-point reduction in mortality.

Why does the PFS analysis have a different confidence level?

The registry reports a 99.99% confidence interval for the PFS hazard ratio. The page retains that exact level rather than converting it to another conventional confidence level. Confidence intervals are inseparable from their stated confidence level.

Why was the Cochran-Mantel-Haenszel test used for ORR?

ORR is a binary endpoint rather than a time-to-event endpoint. A stratified Cochran-Mantel-Haenszel analysis is designed to compare categorical outcomes across treatment groups while accounting for prespecified strata.

Why is the ORR analysis population different from the OS population?

Survival can be analyzed for randomized participants through the ITT principle, whereas tumor response requires evaluable tumor information. The registry therefore defines an evaluable patient population with measurable disease and at least one valid post-baseline assessment for the ORR analysis.

Why does the O'Brien-Fleming alpha-spending approach matter?

Interim examination of efficacy can increase the chance of a false-positive finding if repeated looks are treated as independent tests. Alpha spending provides a prespecified mechanism for allocating type I error over the accumulating information, and the registry specifically reports an O'Brien-Fleming alpha spending function.

Why should the P-value not be treated as the effect size?

A P-value measures the strength of evidence against a null hypothesis under the specified testing framework. It does not quantify how large the treatment effect is. The hazard ratio provides the relative effect estimate, while its confidence interval describes uncertainty around that estimate.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. The counts are presented as affected participants over participants at risk.

Safety measurePlacebo/FOLFIRIAflibercept/FOLFIRI
Serious adverse events198 / 605294 / 611

Placebo/FOLFIRI

Serious adverse events affected 198 participants among 605 participants at risk.

Aflibercept/FOLFIRI

Serious adverse events affected 294 participants among 611 participants at risk.

These safety counts should be interpreted separately from the efficacy hazard ratios. An OS hazard ratio does not incorporate serious adverse events, and a serious-adverse-event count does not measure overall survival. Efficacy and safety are distinct statistical domains that must be considered on their respective scales.

Safety denominator matters: the ClinicalTrials.gov record reports serious adverse events as affected participants divided by participants at risk. These values should not be converted into a different safety population or compared using an unreported statistical test.

16. ClinicalTrials.gov Result Structure

The ClinicalTrials.gov record identifies 5 posted outcome measures and 3 posted statistical analyses. One of those analyses is for the registered primary endpoint, OS; two are secondary analyses for PFS and ORR.

Posted analysisEndpoint roleAnalysis familyEffect measure / result
Overall SurvivalPrimarySurvival analysisHR 0.817; 95.34% CI 0.713–0.937; P = 0.0032
Progression-free Survival by IRCSecondarySurvival analysisHR 0.758; 99.99% CI 0.578–0.995; P = 0.00007
Overall Objective Response RateSecondaryCategorical data95% CI level; P = 0.0001; no effect estimate reported

This structure is useful for statistical interpretation because it distinguishes the primary endpoint from secondary outcomes and also shows that the trial did not use one universal statistical method for every endpoint.

17. Limitations

18. Why This Trial Matters Statistically

VELOUR is a useful statistical teaching case because it combines randomized treatment allocation, triple masking, stratified time-to-event analysis, a binary response endpoint, ITT efficacy analysis, an evaluable response population, and an interim-analysis framework.

ConceptHow it appears in VELOUR
RandomizationThe trial is registered as randomized with two parallel arms.
BlindingThe registry identifies triple masking.
ITT analysisOS and PFS analyses used participants who gave informed consent and were randomized.
Time-to-event endpointsOS and IRC-assessed PFS were analyzed as time-to-event outcomes.
Stratified log-rank testUsed for the OS and PFS comparisons.
Hazard ratioUsed as the reported effect measure for OS and PFS.
Cox modelUsed to obtain the stratified hazard ratio for OS and PFS.
StratificationECOG Performance Status and prior Bevacizumab were used as analysis strata.
Cochran-Mantel-Haenszel testUsed for the binary ORR analysis.
Interim analysisThe OS analysis notes identify interim analysis / alpha spending.
O'Brien-Fleming alpha spendingThe OS significance threshold was reported as 0.0466 using an O'Brien-Fleming alpha spending function.
Different analysis populationsSurvival endpoints used ITT; ORR used an evaluable patient population.

19. A Statistical Reading of the VELOUR Evidence

The most important feature of the VELOUR statistical record is the alignment between endpoint type and analysis method. OS and PFS are time-to-event outcomes, so the registry uses stratified log-rank tests and Cox-model hazard ratios. ORR is binary, so the registry uses a stratified Cochran-Mantel-Haenszel test.

The OS result provides the most complete primary-endpoint statistical record: a hazard ratio of 0.817, a 95.34% confidence interval of 0.713–0.937, and a P-value of 0.0032. The analysis was conducted in the ITT population and incorporated the registered stratification factors.

The PFS result provides a similar structure, with an HR of 0.758, a 99.99% confidence interval of 0.578–0.995, and a P-value of 0.00007. The PFS analysis was also stratified and used an ITT population.

ORR demonstrates why statistical interpretation cannot be reduced to P-values. Although the registry reports P = 0.0001, it does not supply the treatment-group response estimates in the provided statistical-analysis record. The correct statistical response is therefore to report the available evidence without manufacturing an effect estimate.

Key statistical lesson: a complete clinical-trial interpretation requires the endpoint definition, analysis population, statistical method, effect measure, confidence interval, P-value, and design framework to be considered together. A single numerical result without that context can be misleading.

20. Primary Endpoint vs Secondary Endpoints

FeatureOverall SurvivalPFS by IRCORR by IRC / RECIST
RolePrimarySecondarySecondary
Endpoint typeTime-to-eventTime-to-eventBinary
PopulationITTITTEvaluable patient population
MethodStratified log-rankStratified log-rankStratified Cochran-Mantel-Haenszel
Effect measureHazard ratioHazard ratioNot reported
Estimate0.8170.758Not reported
Confidence interval95.34%: 0.713–0.93799.99%: 0.578–0.99595% confidence level; limits not reported
P-value0.00320.000070.0001

The hierarchy matters. OS is the registered primary endpoint, while PFS and ORR are secondary analyses in the ClinicalTrials.gov record. Statistical evidence for a secondary endpoint should not automatically be described as though it were the result of the registered primary analysis.

21. Time-to-Event Analysis: Why Kaplan-Meier Methods Matter

Although the ClinicalTrials.gov record does not provide Kaplan-Meier estimates, OS and PFS are time-to-event endpoints for which Kaplan-Meier estimation is the standard descriptive framework accompanying comparative survival analysis.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at time ti and ni represents participants at risk immediately before that time.

The value of Kaplan-Meier estimation is that it preserves the timing of events and appropriately handles right-censored observations. A participant who has not experienced the event by the last available follow-up can still contribute information up to the censoring time.

The ClinicalTrials.gov record does not provide the individual event and censoring times necessary to independently reproduce a Kaplan-Meier curve. Accordingly, no curve is fabricated here.

22. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The randomized ITT comparison for OS produced a stratified HR of 0.817 with a 95.34% CI of 0.713–0.937 and P = 0.0032. The secondary PFS analysis produced an HR of 0.758 with a 99.99% CI of 0.578–0.995 and P = 0.00007.

Clinical interpretation

Clinical interpretation requires more than the hazard ratios. It must consider the endpoint definitions, absolute event probabilities when available, treatment exposure, safety, follow-up, and the characteristics of the population represented by the trial. The ClinicalTrials.gov record does not provide all of those quantities.

This distinction prevents an important statistical error: a statistically detectable difference is not itself a complete description of clinical value. Conversely, the absence of an effect estimate for ORR in the ClinicalTrials.gov record does not justify assuming that no effect existed. It simply limits what can responsibly be concluded from the available registry result.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats knowledge graph

Connect the endpoints and statistical methods used in VELOUR with deeper tutorials and practical statistical tools.

26. Record Summary

VELOUR provides a compact teaching example of how a randomized phase 3 trial can use different statistical methods for different endpoint types. The registered primary endpoint, OS, was analyzed in the ITT population with a stratified log-rank test and a stratified Cox proportional-hazards model. The reported HR was 0.817, with a 95.34% CI of 0.713–0.937 and P = 0.0032.

The secondary PFS analysis used the same general time-to-event framework and reported an HR of 0.758, with a 99.99% CI of 0.578–0.995 and P = 0.00007. ORR was analyzed as a binary endpoint using a stratified Cochran-Mantel-Haenszel test and had a reported P-value of 0.0001, but the ClinicalTrials.gov record does not provide the corresponding treatment-group response estimates.

The statistical design also illustrates the importance of stratification and interim monitoring. ECOG Performance Status and prior Bevacizumab were incorporated into the reported stratified analyses, while the OS analysis used an O'Brien-Fleming alpha spending function with a reported significance threshold of 0.0466.

Finally, the trial demonstrates why effect estimates should always be read alongside their confidence intervals, P-values, analysis populations, endpoint definitions, and statistical methods. A hazard ratio describes a relative time-to-event effect; it does not replace absolute outcome measures, and a P-value does not quantify effect size.

Clinical Biostats methodology: A trial-results page should not merely repeat a registry result. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and avoiding unsupported estimates.