This page separates reported trial results from statistical interpretation. The numerical results and trial characteristics presented here are limited to the ClinicalTrials.gov record data. Where the registry does not provide an estimate, confidence interval, or additional result, this page does not infer one.
1. Trial at a Glance
VELOUR was a randomized, parallel-group, triple-masked phase 3 treatment trial enrolling 1226 participants with metastatic colorectal cancer after failure of an oxaliplatin-based regimen. The registered comparison was aflibercept plus FOLFIRI versus placebo plus FOLFIRI.
| Feature | VELOUR |
|---|---|
| Trial name | VELOUR |
| NCT identifier | NCT00561470 |
| Phase | Phase 3 |
| Status | Completed |
| Therapeutic area | Oncology |
| Conditions | Colorectal Neoplasms; Neoplasm Metastasis |
| Enrollment | 1226 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Triple |
| Primary purpose | Treatment |
| Lead sponsor | Sanofi |
| Sponsor type | Industry |
| Start | 2007-11 |
| Primary completion | 2011-02 |
| Results posted | Yes |
2. Clinical Question
The central statistical question was whether adding aflibercept to FOLFIRI improved time-to-event outcomes compared with placebo plus FOLFIRI in patients with metastatic colorectal cancer after failure of an oxaliplatin-based regimen.
Population
Patients with metastatic colorectal cancer after failure of an oxaliplatin-based regimen.
Intervention
Aflibercept (ziv-aflibercept, AVE0005, VEGF trap, ZALTRAP®) in combination with FOLFIRI.
Comparator
Placebo in combination with FOLFIRI.
Primary question
Does aflibercept plus FOLFIRI improve overall survival relative to placebo plus FOLFIRI?
3. Trial Design
Participants were assigned to randomized treatment groups.
The registry identifies a two-arm parallel design.
The trial's registry designation is triple masked.
The registered primary purpose is treatment.
Aflibercept / FOLFIRI
- Aflibercept (ziv-aflibercept, AVE0005, VEGF trap, ZALTRAP®)
- FOLFIRI: irinotecan, 5-fluorouracil, and leucovorin
Placebo / FOLFIRI
- Placebo
- FOLFIRI: irinotecan, 5-fluorouracil, and leucovorin
4. Endpoints
The registry identifies one primary endpoint, Overall Survival (OS), and reports statistical analyses for OS, progression-free survival assessed by an Independent Review Committee, and overall objective response rate.
| Endpoint | Type | Registry time frame | Analysis reported |
|---|---|---|---|
| Overall Survival (OS) | Time-to-event | From the date of the first randomization until the study data cut-off date, 07 February 2011 (approximately three years) | Stratified log-rank test; stratified hazard ratio |
| Progression-free Survival (PFS) Assessed by Independent Review Committee (IRC) | Time-to-event | From the date of the first randomization until the occurrence of 561 OS events, 06 May 2010 (approximately 30 months) | Stratified log-rank test; stratified hazard ratio |
| Overall Objective Response Rate (ORR) Based on the Tumor Assessment by the Independent Review Committee (IRC) as Per Response Evaluation Criteria in Solid Tumours (RECIST) Criteria | Binary | From the date of the first randomization until the study data cut-off date, 06 May 2010 (approximately 30 months) | Stratified Cochran-Mantel-Haenszel test |
Primary endpoint definition: Overall Survival
ClinicalTrials.gov defines overall survival as the time interval from the date of randomization to the date of death due to any cause. Once disease progression was documented, participants were followed every 2 months for survival status, until death or until the study cutoff date, whichever came first. The final data cutoff date for the analysis of OS was the date when 863 deaths had occurred, 07 February 2011.
5. Statistical Methodology
Intention-to-treat analysis
The OS analysis population was the intent-to-treat population, defined in the registry as all participants who gave informed consent and were randomized. The PFS analysis likewise used an ITT population consisting of all participants who gave informed consent and were randomized.
This distinction is important because the treatment comparison remains anchored to randomized assignment rather than being restricted to participants who completed treatment. In a randomized trial, preserving the randomized population helps maintain the comparability created by randomization.
Stratified log-rank test
The primary OS comparison used a stratified log-rank test. The PFS comparison also used a stratified log-rank test. Both analyses were stratified on ECOG Performance Status (0 vs 1 vs 2) and prior Bevacizumab (yes vs no) according to IVRS.
Stratification allows the time-to-event comparison to account for prespecified factors used to structure the randomized comparison. Rather than treating all participants as belonging to one undifferentiated risk set, the analysis compares treatment groups within the specified strata and combines the information across them.
Stratified Cox proportional-hazards model
The OS analysis notes specify a Cox Proportional Hazard Model stratified on ECOG Performance Status and prior Bevacizumab. The reported effect measure was a stratified hazard ratio.
The PFS analysis similarly used a stratified Cox Proportional Hazard Model, with the same two stratification factors. The hazard ratio is therefore a relative time-to-event measure rather than a direct comparison of proportions at a single time point.
Cochran-Mantel-Haenszel analysis
The ORR analysis used a stratified Cochran-Mantel-Haenszel test. The registry identifies ORR as a binary endpoint and states that the analysis was stratified on ECOG Performance Status (0 vs 1 vs 2) and prior Bevacizumab (yes vs no) according to IVRS.
The ClinicalTrials.gov record reports a P-value and confidence-interval level for this analysis but does not provide an ORR effect estimate or its confidence limits. Accordingly, this page does not construct an unreported effect measure.
Interim analysis and alpha spending
The OS analysis notes specify a significance threshold of 0.0466 using the O'Brien-Fleming alpha spending function. The statistical-analysis metadata also identifies interim analysis / alpha spending as an analysis concept.
This matters because a trial that permits an interim efficacy assessment cannot generally be interpreted as though the data were examined only once at the end. Alpha spending provides a framework for controlling the prespecified type I error while allowing information to be evaluated during the study.
6. Results: Overall Survival
Overall survival was the registered primary endpoint. The analysis was performed in the ITT population and compared placebo/FOLFIRI with aflibercept/FOLFIRI using a stratified log-rank test.
Stratified hazard ratio for overall survival
95.34% CI: 0.713–0.937 · P = 0.0032
Two-sided confidence interval · Superiority hypothesis
| OS analysis characteristic | Reported value |
|---|---|
| Analysis population | Intent-to-treat population |
| Groups compared | Placebo/FOLFIRI vs Aflibercept/FOLFIRI |
| Method | Stratified log-rank test |
| Effect measure | Stratified Hazard Ratio |
| Estimate | 0.817 |
| Confidence interval | 95.34% CI 0.713–0.937 |
| P-value | 0.0032 |
| Hypothesis | Superiority |
| OS data cutoff | 07 February 2011 |
| Event basis for final OS cutoff | 863 deaths |
The estimated OS hazard ratio of 0.817 means that, under the stratified Cox model used for the analysis, the estimated instantaneous hazard of death in the aflibercept/FOLFIRI group was about 18.3% lower than in the placebo/FOLFIRI group. This is a relative model-based comparison of event hazards, not a statement that 18.3% of participants benefited or that each participant experienced an 18.3% reduction in their individual probability of death.
The 95.34% confidence interval of 0.713–0.937 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of effects that individual patients would experience.
The P-value of 0.0032 addresses evidence against the null hypothesis under the specified statistical test; it does not measure the magnitude or clinical importance of the treatment effect. Effect size is conveyed by the hazard ratio and its confidence interval.
The OS analysis was stratified on ECOG Performance Status and prior Bevacizumab, and the registry notes a significance threshold of 0.0466 using an O'Brien-Fleming alpha spending function. The interpretation therefore belongs to the prespecified group-sequential framework rather than to an unadjusted single-look test.
Finally, the hazard-ratio interpretation relies on the Cox proportional-hazards modeling framework. A single hazard ratio summarizes a relative event-rate relationship over the analyzed period; it should not be read as an absolute survival difference or as proof that the hazards are identical in relative proportion at every point in time.
What the OS result does not tell us
- It does not provide the median OS because a median survival estimate is not included in the ClinicalTrials.gov record.
- It does not provide survival probabilities at a particular calendar time because those estimates are not reported in the ClinicalTrials.gov record.
- It does not establish that every participant experienced the same relative treatment effect.
- It does not by itself describe toxicity, quality of life, or individual patient benefit.
7. Results: Progression-Free Survival
Progression-free survival assessed by an Independent Review Committee was reported as a secondary time-to-event endpoint. The analysis used the ITT population and a stratified log-rank test.
Stratified hazard ratio for PFS
99.99% CI: 0.578–0.995 · P = 0.00007
Two-sided confidence interval · Superiority hypothesis
| PFS analysis characteristic | Reported value |
|---|---|
| Analysis population | Intent-to-treat population |
| Groups compared | Placebo/FOLFIRI vs Aflibercept/FOLFIRI |
| Assessment | Independent Review Committee (IRC) |
| Method | Stratified log-rank test |
| Effect measure | Stratified Hazard ratio |
| Estimate | 0.758 |
| Confidence interval | 99.99% CI 0.578–0.995 |
| P-value | 0.00007 |
| Hypothesis | Superiority |
| Analysis cutoff | 06 May 2010 |
| Event/time basis | Until occurrence of 561 OS events; approximately 30 months |
The PFS hazard ratio of 0.758 corresponds to an estimated instantaneous hazard of progression or death about 24.2% lower with aflibercept/FOLFIRI than with placebo/FOLFIRI under the reported stratified Cox model.
The 99.99% confidence interval of 0.578–0.995 is the registry-reported uncertainty interval for this estimate. The unusually high confidence level should be retained exactly as reported rather than silently replacing it with a conventional 95% interval.
The P-value of 0.00007 quantifies evidence under the specified statistical testing framework. It is not an effect-size measure. The hazard ratio and its confidence interval provide the information about the estimated relative magnitude and precision of the treatment comparison.
The PFS analysis was stratified on ECOG Performance Status and prior Bevacizumab. Because PFS is a time-to-event endpoint, censoring and the assumptions underlying the survival-analysis framework remain relevant to interpretation.
8. Results: Overall Objective Response Rate
Overall objective response rate was reported as a secondary binary endpoint based on tumor assessment by the Independent Review Committee according to RECIST criteria.
| ORR analysis characteristic | Reported value |
|---|---|
| Analysis population | Evaluable patient population for tumor response |
| Population definition | All randomized participants with measurable disease at study entry, as per IRC evaluation, and with at least one valid post-baseline assessment |
| Groups compared | Placebo/FOLFIRI vs Aflibercept/FOLFIRI |
| Endpoint type | Binary |
| Method | Stratified Cochran-Mantel-Haenszel |
| Confidence interval level | 95% |
| P-value | 0.0001 |
| Hypothesis | Superiority |
| Analysis cutoff | 06 May 2010 |
The registry reports a P-value of 0.0001 for the stratified Cochran-Mantel-Haenszel comparison of ORR. The ClinicalTrials.gov record does not provide the corresponding response-rate estimates or a confidence interval for the treatment effect.
Therefore, the appropriate interpretation is limited to the reported statistical comparison. It would be inappropriate to reconstruct an ORR difference, risk ratio, odds ratio, or response-rate percentage from the P-value alone.
The analysis population also differs from the ITT population used for OS and PFS: the response analysis used an evaluable patient population with measurable disease at study entry and at least one valid post-baseline assessment. That distinction matters when comparing the interpretation of response with the primary survival analysis.
9. Secondary Endpoint Analysis and Population Differences
VELOUR illustrates an important principle in clinical-trial statistics: different endpoints can legitimately use different analysis populations and statistical procedures when their measurement requirements differ.
| Endpoint | Population | Data type | Statistical method |
|---|---|---|---|
| Overall Survival | ITT | Time-to-event | Stratified log-rank test; stratified Cox model for HR |
| PFS by IRC | ITT | Time-to-event | Stratified log-rank test; stratified Cox model for HR |
| ORR by IRC / RECIST | Evaluable patient population | Binary | Stratified Cochran-Mantel-Haenszel test |
The survival analyses preserve the randomized treatment assignment through ITT analysis. The response analysis instead requires measurable disease and a valid post-baseline tumor assessment, which is why the registry specifies a different evaluable population.
This difference does not make the ORR analysis invalid. It means that ORR and survival answer different questions and are estimated in populations defined by different information requirements.
10. Stratification
The reported OS and PFS analyses were stratified on two factors: ECOG Performance Status (0 vs 1 vs 2) and prior Bevacizumab (yes vs no), according to IVRS. The ORR analysis used the same stratification factors.
ECOG Performance Status
The registered analysis distinguishes ECOG Performance Status categories 0, 1, and 2. The analysis accounts for these categories through stratification.
Prior Bevacizumab
Prior Bevacizumab was categorized as yes versus no and incorporated into the reported stratified analyses.
Stratification is particularly useful when an important prognostic factor could influence the event distribution. The objective is not to create a separate treatment effect for every subgroup, but to perform the primary comparison while accounting for the prespecified strata.
The stratified analysis compares the randomized groups while conditioning on the registered stratification factors rather than ignoring them.
11. Interim Analysis and Alpha Spending
The registry-reported OS analysis notes specify that the significance threshold was set to 0.0466 using the O'Brien-Fleming alpha spending function. Interim analysis / alpha spending is also identified among the concepts in the analysis text.
Why alpha spending matters
When accumulating trial data are examined during the study, the statistical framework must account for those repeated opportunities to evaluate efficacy. Alpha spending allocates the available type I error across information times.
O'Brien-Fleming principle
An O'Brien-Fleming approach places a stringent evidentiary requirement on an early look and permits a less stringent boundary as information accumulates toward the final analysis.
The important practical point is that the reported OS P-value should be interpreted in the context of the prespecified group-sequential framework. The threshold of 0.0466 is part of that framework and should not be replaced by an assumption that the only relevant threshold was an unadjusted 0.05.
12. What the Hazard Ratio Means in VELOUR
The OS HR of 0.817 means the fitted model estimated a lower instantaneous hazard of death in the aflibercept/FOLFIRI group relative to placebo/FOLFIRI. Numerically, 1 − 0.817 gives approximately an 18.3% lower estimated hazard.
The PFS HR of 0.758 means the fitted model estimated a lower instantaneous hazard of progression or death in the aflibercept/FOLFIRI group. Numerically, 1 − 0.758 gives approximately a 24.2% lower estimated hazard.
Neither hazard ratio is a probability that an individual patient will experience an event. Neither is an absolute risk reduction, a median-survival difference, or a statement that the same proportional reduction applies identically to every participant.
The two hazard ratios also describe different endpoints. OS concerns death from any cause, whereas the registered PFS analysis concerns progression-free survival assessed by an Independent Review Committee. They should therefore not be combined into one generic measure of "benefit."
13. Confidence Intervals and Precision
Confidence intervals provide information that a point estimate alone cannot. The OS estimate of 0.817 is accompanied by a 95.34% CI of 0.713–0.937, while the PFS estimate of 0.758 is accompanied by a 99.99% CI of 0.578–0.995.
| Endpoint | Estimate | Confidence level | Confidence interval |
|---|---|---|---|
| Overall Survival | 0.817 | 95.34% | 0.713–0.937 |
| Progression-free Survival | 0.758 | 99.99% | 0.578–0.995 |
The confidence interval is about uncertainty in the estimated treatment effect under the statistical model and sampling framework. It is not a range containing the true treatment effect for every individual participant, nor does it represent the distribution of individual responses.
The different confidence levels should also be preserved exactly as reported. A confidence interval cannot be compared solely by looking at its numerical width without considering the confidence level at which it was constructed.
14. Statistical Methods Explained
Why was a stratified log-rank test used for OS and PFS?
OS and PFS are time-to-event endpoints, so the analysis must account for both event timing and censoring. The stratified log-rank test provides a way to compare the event-time distributions between randomized treatment groups while accounting for the prespecified stratification factors.
What does an OS hazard ratio of 0.817 mean?
It is a relative estimate from the stratified Cox model. Under that model, the aflibercept/FOLFIRI group had an estimated instantaneous hazard of death approximately 18.3% lower than the placebo/FOLFIRI group. It is not an 18.3-percentage-point reduction in mortality.
Why does the PFS analysis have a different confidence level?
The registry reports a 99.99% confidence interval for the PFS hazard ratio. The page retains that exact level rather than converting it to another conventional confidence level. Confidence intervals are inseparable from their stated confidence level.
Why was the Cochran-Mantel-Haenszel test used for ORR?
ORR is a binary endpoint rather than a time-to-event endpoint. A stratified Cochran-Mantel-Haenszel analysis is designed to compare categorical outcomes across treatment groups while accounting for prespecified strata.
Why is the ORR analysis population different from the OS population?
Survival can be analyzed for randomized participants through the ITT principle, whereas tumor response requires evaluable tumor information. The registry therefore defines an evaluable patient population with measurable disease and at least one valid post-baseline assessment for the ORR analysis.
Why does the O'Brien-Fleming alpha-spending approach matter?
Interim examination of efficacy can increase the chance of a false-positive finding if repeated looks are treated as independent tests. Alpha spending provides a prespecified mechanism for allocating type I error over the accumulating information, and the registry specifically reports an O'Brien-Fleming alpha spending function.
Why should the P-value not be treated as the effect size?
A P-value measures the strength of evidence against a null hypothesis under the specified testing framework. It does not quantify how large the treatment effect is. The hazard ratio provides the relative effect estimate, while its confidence interval describes uncertainty around that estimate.
15. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. The counts are presented as affected participants over participants at risk.
| Safety measure | Placebo/FOLFIRI | Aflibercept/FOLFIRI |
|---|---|---|
| Serious adverse events | 198 / 605 | 294 / 611 |
Placebo/FOLFIRI
Serious adverse events affected 198 participants among 605 participants at risk.
Aflibercept/FOLFIRI
Serious adverse events affected 294 participants among 611 participants at risk.
These safety counts should be interpreted separately from the efficacy hazard ratios. An OS hazard ratio does not incorporate serious adverse events, and a serious-adverse-event count does not measure overall survival. Efficacy and safety are distinct statistical domains that must be considered on their respective scales.
16. ClinicalTrials.gov Result Structure
The ClinicalTrials.gov record identifies 5 posted outcome measures and 3 posted statistical analyses. One of those analyses is for the registered primary endpoint, OS; two are secondary analyses for PFS and ORR.
| Posted analysis | Endpoint role | Analysis family | Effect measure / result |
|---|---|---|---|
| Overall Survival | Primary | Survival analysis | HR 0.817; 95.34% CI 0.713–0.937; P = 0.0032 |
| Progression-free Survival by IRC | Secondary | Survival analysis | HR 0.758; 99.99% CI 0.578–0.995; P = 0.00007 |
| Overall Objective Response Rate | Secondary | Categorical data | 95% CI level; P = 0.0001; no effect estimate reported |
This structure is useful for statistical interpretation because it distinguishes the primary endpoint from secondary outcomes and also shows that the trial did not use one universal statistical method for every endpoint.
17. Limitations
- Registry-level detail: this page is limited to the ClinicalTrials.gov record. It does not add results from publications or other external sources.
- Incomplete effect reporting for ORR: the ClinicalTrials.gov record reports a P-value and confidence-level information but not the corresponding ORR estimates or confidence limits.
- No median survival estimates reported: median OS and median PFS are not included in the ClinicalTrials.gov record and therefore are not reported here.
- No subgroup estimates reported: the registry identifies ECOG Performance Status and prior Bevacizumab as stratification factors, but no subgroup-specific treatment-effect estimates are reported.
- Hazard-ratio assumptions: Cox proportional-hazards models rely on a model structure in which the treatment effect is represented through a hazard ratio. A single HR does not directly describe absolute risk at every time point.
- Censoring: time-to-event methods rely on the available event and censoring information. The summary registry record does not provide the individual event/censoring data needed for independent reconstruction of survival curves.
- Different analysis populations: OS and PFS used ITT populations, whereas ORR used an evaluable patient population requiring measurable disease and a valid post-baseline assessment.
- Interim analysis: the OS analysis uses an O'Brien-Fleming alpha-spending framework. The ClinicalTrials.gov record does not provide the complete sequence of interim information fractions and boundaries.
- Safety interpretation: serious adverse-event counts are provided by arm, but no formal comparative safety test or effect estimate is reported.
18. Why This Trial Matters Statistically
VELOUR is a useful statistical teaching case because it combines randomized treatment allocation, triple masking, stratified time-to-event analysis, a binary response endpoint, ITT efficacy analysis, an evaluable response population, and an interim-analysis framework.
| Concept | How it appears in VELOUR |
|---|---|
| Randomization | The trial is registered as randomized with two parallel arms. |
| Blinding | The registry identifies triple masking. |
| ITT analysis | OS and PFS analyses used participants who gave informed consent and were randomized. |
| Time-to-event endpoints | OS and IRC-assessed PFS were analyzed as time-to-event outcomes. |
| Stratified log-rank test | Used for the OS and PFS comparisons. |
| Hazard ratio | Used as the reported effect measure for OS and PFS. |
| Cox model | Used to obtain the stratified hazard ratio for OS and PFS. |
| Stratification | ECOG Performance Status and prior Bevacizumab were used as analysis strata. |
| Cochran-Mantel-Haenszel test | Used for the binary ORR analysis. |
| Interim analysis | The OS analysis notes identify interim analysis / alpha spending. |
| O'Brien-Fleming alpha spending | The OS significance threshold was reported as 0.0466 using an O'Brien-Fleming alpha spending function. |
| Different analysis populations | Survival endpoints used ITT; ORR used an evaluable patient population. |
19. A Statistical Reading of the VELOUR Evidence
The most important feature of the VELOUR statistical record is the alignment between endpoint type and analysis method. OS and PFS are time-to-event outcomes, so the registry uses stratified log-rank tests and Cox-model hazard ratios. ORR is binary, so the registry uses a stratified Cochran-Mantel-Haenszel test.
The OS result provides the most complete primary-endpoint statistical record: a hazard ratio of 0.817, a 95.34% confidence interval of 0.713–0.937, and a P-value of 0.0032. The analysis was conducted in the ITT population and incorporated the registered stratification factors.
The PFS result provides a similar structure, with an HR of 0.758, a 99.99% confidence interval of 0.578–0.995, and a P-value of 0.00007. The PFS analysis was also stratified and used an ITT population.
ORR demonstrates why statistical interpretation cannot be reduced to P-values. Although the registry reports P = 0.0001, it does not supply the treatment-group response estimates in the provided statistical-analysis record. The correct statistical response is therefore to report the available evidence without manufacturing an effect estimate.
20. Primary Endpoint vs Secondary Endpoints
| Feature | Overall Survival | PFS by IRC | ORR by IRC / RECIST |
|---|---|---|---|
| Role | Primary | Secondary | Secondary |
| Endpoint type | Time-to-event | Time-to-event | Binary |
| Population | ITT | ITT | Evaluable patient population |
| Method | Stratified log-rank | Stratified log-rank | Stratified Cochran-Mantel-Haenszel |
| Effect measure | Hazard ratio | Hazard ratio | Not reported |
| Estimate | 0.817 | 0.758 | Not reported |
| Confidence interval | 95.34%: 0.713–0.937 | 99.99%: 0.578–0.995 | 95% confidence level; limits not reported |
| P-value | 0.0032 | 0.00007 | 0.0001 |
The hierarchy matters. OS is the registered primary endpoint, while PFS and ORR are secondary analyses in the ClinicalTrials.gov record. Statistical evidence for a secondary endpoint should not automatically be described as though it were the result of the registered primary analysis.
21. Time-to-Event Analysis: Why Kaplan-Meier Methods Matter
Although the ClinicalTrials.gov record does not provide Kaplan-Meier estimates, OS and PFS are time-to-event endpoints for which Kaplan-Meier estimation is the standard descriptive framework accompanying comparative survival analysis.
Here, di represents events at time ti and ni represents participants at risk immediately before that time.
The value of Kaplan-Meier estimation is that it preserves the timing of events and appropriately handles right-censored observations. A participant who has not experienced the event by the last available follow-up can still contribute information up to the censoring time.
The ClinicalTrials.gov record does not provide the individual event and censoring times necessary to independently reproduce a Kaplan-Meier curve. Accordingly, no curve is fabricated here.
22. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The randomized ITT comparison for OS produced a stratified HR of 0.817 with a 95.34% CI of 0.713–0.937 and P = 0.0032. The secondary PFS analysis produced an HR of 0.758 with a 99.99% CI of 0.578–0.995 and P = 0.00007.
Clinical interpretation
Clinical interpretation requires more than the hazard ratios. It must consider the endpoint definitions, absolute event probabilities when available, treatment exposure, safety, follow-up, and the characteristics of the population represented by the trial. The ClinicalTrials.gov record does not provide all of those quantities.
This distinction prevents an important statistical error: a statistically detectable difference is not itself a complete description of clinical value. Conversely, the absence of an effect estimate for ORR in the ClinicalTrials.gov record does not justify assuming that no effect existed. It simply limits what can responsibly be concluded from the available registry result.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NCT00561470 — VELOUR.
- PubMed: PMID 41530374.
- PubMed: PMID 39207600.
- PubMed: PMID 34789774.
- PubMed: PMID 32168980.
- PubMed: PMID 28807738.
Continue through the Clinical Biostats knowledge graph
Connect the endpoints and statistical methods used in VELOUR with deeper tutorials and practical statistical tools.
26. Record Summary
VELOUR provides a compact teaching example of how a randomized phase 3 trial can use different statistical methods for different endpoint types. The registered primary endpoint, OS, was analyzed in the ITT population with a stratified log-rank test and a stratified Cox proportional-hazards model. The reported HR was 0.817, with a 95.34% CI of 0.713–0.937 and P = 0.0032.
The secondary PFS analysis used the same general time-to-event framework and reported an HR of 0.758, with a 99.99% CI of 0.578–0.995 and P = 0.00007. ORR was analyzed as a binary endpoint using a stratified Cochran-Mantel-Haenszel test and had a reported P-value of 0.0001, but the ClinicalTrials.gov record does not provide the corresponding treatment-group response estimates.
The statistical design also illustrates the importance of stratification and interim monitoring. ECOG Performance Status and prior Bevacizumab were incorporated into the reported stratified analyses, while the OS analysis used an O'Brien-Fleming alpha spending function with a reported significance threshold of 0.0466.
Finally, the trial demonstrates why effect estimates should always be read alongside their confidence intervals, P-values, analysis populations, endpoint definitions, and statistical methods. A hazard ratio describes a relative time-to-event effect; it does not replace absolute outcome measures, and a P-value does not quantify effect size.