← Clinical Trials
Non-Small Cell Lung Cancer Phase 3 Time-to-Event NCT00257608

ATLAS: Complete Statistical Analysis of Bevacizumab With or Without Erlotinib in Non-Small Cell Lung Cancer

An independent statistical analysis of the randomized, double-blind phase 3 ATLAS trial comparing bevacizumab plus erlotinib with bevacizumab plus placebo for first-line treatment of non-small cell lung cancer, with progression-free survival as the registered primary endpoint.

Trial status: Completed  ·  Enrollment: 1145  ·  Primary completion: 2014-11
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical trial results are taken only from the ClinicalTrials.gov data posted on ClinicalTrials.gov for ATLAS. The registry record is the official source for the registered design, endpoints, and posted statistical analyses.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ATLAS was a randomized, double-blind, parallel phase 3 trial evaluating bevacizumab therapy with or without erlotinib for first-line treatment of non-small cell lung cancer. The registered primary endpoint was progression-free survival (PFS), a time-to-event outcome assessed over approximately 3 years.

1145
Enrollment
Randomized trial
2
Treatment arms
Parallel design
0.708
PFS HR
95% CI 0.580–0.864
0.0006
PFS P-value
Two-sided
FeatureATLAS
PhasePhase 3
ConditionNon-Small Cell Lung Cancer
DesignRandomized, double-blind, parallel
AllocationRandomized
Primary purposeTreatment
Enrollment1145
Arms2
Primary endpointProgression-free Survival (PFS)
Primary endpoint typeTime-to-event
Primary hypothesisSuperiority
Statistical methodLog-rank test
Effect measureHazard ratio
ClinicalTrials.govNCT00257608

2. Clinical Question

The trial's statistical question was whether treatment with bevacizumab plus erlotinib produced a different time-to-progression-or-death profile from bevacizumab plus placebo in the randomized trial population. The registered hypothesis type was superiority.

Population

Participants enrolled in the phase 3 ATLAS study for first-line treatment of non-small cell lung cancer.

Intervention

Bevacizumab + erlotinib.

Comparator

Bevacizumab + placebo.

Primary question

Does bevacizumab plus erlotinib improve progression-free survival compared with bevacizumab plus placebo?

3. Trial Design

01
Randomize1145 participants
02
Two armsBevacizumab + placebo or erlotinib
03
Double-blindBlinded treatment assignment
04
Follow-upTime-to-event outcomes
05
AnalysisLog-rank and hazard ratio
ARM A

Bevacizumab + Placebo

  • Bevacizumab
  • Placebo
ARM B

Bevacizumab + Erlotinib

  • Bevacizumab
  • Erlotinib HCl

The registry describes ATLAS as a randomized, double-blind, parallel phase 3 treatment trial. The intervention list contains bevacizumab, placebo, and erlotinib HCl. The design therefore creates a direct randomized comparison between the two treatment combinations while maintaining blinding.

Why the placebo matters statistically: the placebo component helps preserve the blinded comparison between the two randomized groups. Blinding can reduce the potential for treatment knowledge to influence participant behavior, clinical management, assessment, or reporting, depending on the endpoint and how the trial was conducted.

4. Trial Timeline

2006-01

Trial start

The ATLAS trial began in January 2006 according to the registry data.

Phase 3

Randomized comparison

The study used randomized allocation, a parallel design, double masking, and a treatment-focused primary purpose.

Approximately 3 years

Primary PFS assessment

The registered primary endpoint was progression-free survival over approximately 3 years.

2014-11

Primary completion

The registry data identify November 2014 as the primary completion date.

5. Primary Endpoint

EndpointRegistered definitionTime frameType
Progression-free Survival (PFS) PFS was defined as the length of time from randomization until documented disease progression or death from any cause, whichever occurred earlier. Progression is defined using Response Evaluation Criteria In Solid Tumors Criteria (RECIST v1.0), as a 20% increase in the sum of the longest diameter of target lesions, or a measurable increase in a non-target lesion, or the appearance of new lesions. Data presented until cut-off date 18 July 2008. Approximately 3 years Time-to-event

This endpoint combines two possible events: documented disease progression and death from any cause. The first of these events to occur determines the PFS event time. This is important because a patient who dies before documented progression still experiences a PFS event under the registered definition.

6. Secondary Endpoint

The posted statistical analyses also include overall survival. The registry classifies this analysis as secondary rather than primary.

EndpointRoleTime frameTypeAnalysis
Overall Survival Secondary Approximately 3.5 years Time-to-event Log-rank test; hazard ratio

7. Statistical Analysis Population

The registry specifies an intent-to-treat (ITT) population for both the primary PFS analysis and the secondary overall-survival analysis. The ITT population included all participants who were randomized during the post-chemotherapy phase.

Why ITT is important

Analyzing participants according to their randomized assignment preserves the treatment comparison created by randomization, even when subsequent events affect the amount of observed follow-up.

What ITT does not guarantee

ITT does not eliminate missing follow-up, censoring, treatment discontinuation, or other practical complications. Those issues still have to be handled appropriately in the time-to-event analysis.

8. Primary Results: Progression-Free Survival

The posted primary analysis compares bevacizumab + placebo with bevacizumab + erlotinib in the ITT population. The reported method was a log-rank test, with the treatment effect expressed as a hazard ratio.

Hazard ratio for progression or death

0.708

95% CI: 0.580–0.864   ·   P = 0.0006

Two-sided confidence interval  ·  Superiority hypothesis

Primary endpointComparisonMethodHazard ratio95% CIP-value
Progression-free Survival Bevacizumab + Placebo vs Bevacizumab + Erlotinib Log-rank test 0.708 0.580–0.864 0.0006
Clinical Biostats interpretation

The reported hazard ratio of 0.708 means that the estimated instantaneous rate of a PFS event was about 70.8% for the comparison represented by the reported hazard ratio, or equivalently that the estimated hazard was approximately 29.2% lower for bevacizumab plus erlotinib relative to bevacizumab plus placebo under the fitted time-to-event comparison.

This does not mean that 29.2% of patients avoided progression, that every patient experienced a 29.2% reduction in risk, or that the probability of progression or death at a particular time point was reduced by exactly 29.2%.

The 95% confidence interval of 0.580–0.864 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It is an interval for the treatment-effect estimate, not a range containing the individual treatment effects experienced by patients.

The P-value of 0.0006 addresses evidence against the null hypothesis within the reported statistical framework. It does not measure the size, clinical importance, or certainty of the treatment effect. The hazard ratio and its confidence interval are needed to describe the magnitude and precision of the estimated effect.

Because PFS is a time-to-event endpoint, interpretation also depends on censoring and the assumptions underlying the survival-analysis framework. A hazard ratio is a relative rate measure over follow-up; it is not interchangeable with a risk ratio or an absolute difference in PFS probability.

Educational note: a Kaplan-Meier curve cannot be reconstructed from the reported hazard ratio, confidence interval, and P-value alone. A valid reconstruction would require underlying event and censoring information or sufficiently detailed digitized source data.

9. Secondary Results: Overall Survival

The posted secondary analysis evaluated overall survival in the same ITT population. The registry reports a log-rank comparison and a hazard ratio with a two-sided 95% confidence interval.

Hazard ratio for overall survival

0.917

95% CI: 0.698–1.205   ·   P = 0.5341

Approximately 3.5-year time frame  ·  Superiority hypothesis

Secondary endpointComparisonMethodHazard ratio95% CIP-value
Overall Survival Bevacizumab + Placebo vs Bevacizumab + Erlotinib Log-rank test 0.917 0.698–1.205 0.5341
Clinical Biostats interpretation

The reported hazard ratio of 0.917 is below 1, corresponding to an estimated instantaneous death hazard approximately 8.3% lower for the comparison represented by the reported treatment effect. This is a relative model-based estimate, not an 8.3% reduction in the proportion of patients who died.

The 95% confidence interval of 0.698–1.205 is substantially wider than the PFS interval and crosses 1. This indicates considerable uncertainty around the estimated relative effect and includes values corresponding to lower, similar, and higher hazards under the model.

The P-value of 0.5341 does not measure the magnitude of the observed hazard ratio. It describes the compatibility of the observed data with the statistical null hypothesis under the specified testing framework. A nonsmall P-value should not be translated into proof that the treatments are identical.

Overall survival can also be affected by events and treatments occurring after the initial randomized treatment. The ClinicalTrials.gov record does not provide additional information about subsequent treatment, crossover, or other post-progression factors, so those mechanisms should not be inferred here.

10. Comparing PFS and Overall Survival Statistically

ATLAS provides a useful example of why different time-to-event endpoints should not be treated as interchangeable.

FeatureProgression-free SurvivalOverall Survival
RolePrimary endpointSecondary endpoint
Time frameApproximately 3 yearsApproximately 3.5 years
Event definitionDocumented progression or death, whichever occurs earlierDeath from any cause
Analysis populationITTITT
MethodLog-rank testLog-rank test
Hazard ratio0.7080.917
95% CI0.580–0.8640.698–1.205
P-value0.00060.5341

The two endpoints measure different clinical events. PFS incorporates disease progression as well as death, whereas overall survival uses death from any cause. Consequently, a treatment effect on PFS need not have the same numerical magnitude as its effect on OS.

11. Statistical Methodology

Time-to-event analysis

Both posted statistical analyses concern outcomes for which the timing of an event matters. In a time-to-event analysis, a participant contributes information not only through whether an event occurs, but also through the observed time until the event or the point at which follow-up ends.

Conceptual survival function
S(t) = P(T > t)

The survival function represents the probability that the event time T exceeds time t. For PFS, the event is progression or death according to the registered definition. For OS, the event is death from any cause.

Log-rank test

The registry reports the log-rank test for both PFS and overall survival. The log-rank test compares the observed pattern of event occurrence between randomized groups across follow-up, accounting for the time at which participants experience events or are censored.

The test is therefore fundamentally different from a simple comparison of proportions. A participant who remains event-free for a longer period contributes information about the treatment comparison even if that participant has not experienced the event by the end of available follow-up.

Hazard ratio

The treatment effect was expressed as a hazard ratio. Conceptually, a hazard ratio compares the instantaneous event rates between groups within a time-to-event framework.

Interpretation
HR < 1  →  lower estimated instantaneous event rate in the numerator treatment group

The exact direction of interpretation depends on which treatment group is represented in the numerator of the reported comparison. A hazard ratio should not be interpreted as an absolute probability difference.

Confidence intervals

A confidence interval provides information about the precision of the estimated treatment effect. Narrower intervals generally indicate greater statistical precision, while wider intervals indicate more uncertainty. For a hazard ratio, the value 1 represents equal hazards between the compared groups.

Intention-to-treat analysis

The ATLAS registry analysis specifies an ITT population consisting of all participants randomized during the post-chemotherapy phase. This preserves the randomized treatment assignment as the basis of the efficacy comparison and avoids redefining treatment groups according to later treatment behavior.

Superiority hypothesis

The registry identifies the hypothesis type for both posted analyses as superiority. This means the statistical question is whether the randomized treatment groups differ in the favorable direction specified by the trial hypothesis, rather than whether one treatment is merely not unacceptably worse than another under a non-inferiority margin.

12. Statistical Methods Explained

Why was a log-rank test used?

PFS and OS are time-to-event endpoints, so the timing of events and censoring matters. The log-rank test is designed to compare survival experience between randomized groups while using information across the follow-up period rather than reducing each patient to a simple event/no-event indicator.

What does a PFS hazard ratio of 0.708 mean?

A hazard ratio of 0.708 indicates an estimated event rate ratio below 1 for the reported comparison. Expressed descriptively, 0.708 corresponds to an estimated hazard about 29.2% lower for the treatment represented as the numerator relative to the comparator. It does not mean that exactly 29.2% fewer patients progressed or died.

Why is the confidence interval important?

The point estimate is only one estimate of the treatment effect. The 95% CI of 0.580–0.864 shows the statistical uncertainty around the PFS hazard ratio. Looking only at 0.708 would conceal information about how precisely the effect was estimated.

Why does the P-value not measure effect size?

A P-value is a measure of evidence against a specified null hypothesis under the statistical model and testing procedure. It is affected by both the observed effect and the amount of information in the study. The hazard ratio describes relative effect size, while its confidence interval describes precision.

Why is PFS different from overall survival?

PFS records progression or death, whichever occurs first, whereas OS records death from any cause. These endpoints therefore capture different clinical processes. A treatment can have different estimated effects on the two endpoints without the statistical analysis being contradictory.

Why does randomization matter for the statistical comparison?

Randomization establishes the treatment assignment before subsequent outcomes occur. When an ITT analysis retains participants according to that assignment, the treatment comparison remains anchored to the randomized design rather than being reconstructed from post-randomization treatment choices.

13. Understanding Censoring in ATLAS

Time-to-event trials commonly include participants whose event has not been observed by the end of their available follow-up. Such observations are typically treated as censored at the appropriate follow-up time rather than being treated as if an event never occurred.

Event observed

For PFS, a documented progression or death establishes the event time under the registered definition. For OS, death establishes the event.

Event not observed

A participant without an observed event by the relevant end of follow-up contributes information up to the point at which the observation is censored.

This is one reason a simple proportion of participants with events is not equivalent to a survival-analysis result. Time-to-event methods use the available timing information and account for different lengths of observed follow-up.

14. Primary Endpoint Interpretation in Context

The PFS result combines three pieces of statistical information that should be read together: the hazard ratio, the confidence interval, and the P-value.

ComponentATLAS PFS resultWhat it tells the reader
Hazard ratio0.708Magnitude and direction of the estimated relative time-to-event effect
95% CI0.580–0.864Statistical uncertainty around the hazard-ratio estimate
P-value0.0006Evidence against the relevant null hypothesis under the reported test
HypothesisSuperiorityThe trial tested for a treatment difference rather than non-inferiority

The three quantities answer different questions. The hazard ratio asks how large is the estimated relative effect? The confidence interval asks how precisely is that effect estimated? The P-value asks how compatible are the data with the null hypothesis under the specified testing framework?

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These figures should be interpreted separately from the efficacy analyses because safety and efficacy address different outcomes and use different conceptual denominators.

Treatment groupParticipants with serious adverse eventsParticipants at risk
Bevacizumab + Placebo63367
Bevacizumab + Erlotinib86368
Serious adverse events: affected participants
Bevacizumab + Placebo
63
Bevacizumab + Erlotinib
86

The ClinicalTrials.gov record reports affected and at-risk counts but do not provide a formal statistical comparison for serious adverse events. Therefore, the counts should not be converted into an inferential treatment-effect claim. Safety interpretation would ordinarily consider the denominator, event definitions, exposure time, severity, timing, and clinical attribution.

16. What the Hazard Ratio Does — and Does Not — Mean

PFS hazard ratio

The ATLAS PFS hazard ratio of 0.708 indicates a lower estimated instantaneous rate of PFS events for the treatment represented by the numerator of the reported comparison. The complementary calculation, 1 − 0.708, gives 0.292, or approximately a 29.2% lower estimated hazard.

It does not mean that 29.2% of participants benefited, that every participant experienced a 29.2% reduction, or that the absolute probability of progression or death was reduced by 29.2 percentage points.

PFS confidence interval

The 95% CI of 0.580–0.864 gives a range for the hazard-ratio estimate under the statistical framework used for the analysis. Because the entire interval lies below 1, the interval is consistent with a lower estimated hazard for the treatment represented by the numerator of the comparison.

OS hazard ratio

The OS hazard ratio of 0.917 is closer to 1 than the PFS estimate, and its 95% CI of 0.698–1.205 crosses 1. This means the posted OS estimate is less precise and does not establish a statistically distinguishable difference under the reported superiority test.

17. Limitations and Interpretation Issues

18. Why This Trial Matters Statistically

ATLAS is a useful teaching example because it illustrates the complete statistical chain for a randomized time-to-event endpoint: randomization establishes the comparison, an ITT population preserves that randomized assignment for efficacy analysis, the endpoint is defined as a time until progression or death, the groups are compared with a log-rank test, and the treatment effect is summarized with a hazard ratio and confidence interval.

ConceptHow it appears in ATLAS
RandomizationRandomized allocation in a parallel phase 3 design
BlindingDouble masking
ITT analysisPrimary and secondary efficacy analyses use the ITT population
Time-to-event endpointPFS is the registered primary endpoint
RECIST-based progressionProgression is defined using RECIST v1.0 criteria in the registry definition
Log-rank testReported for PFS and overall survival
Hazard ratioReported as the treatment-effect measure
Confidence intervalTwo-sided 95% CIs reported for both hazard ratios
Superiority testingBoth posted analyses are identified as superiority hypotheses

19. Statistical Methods Explained: A Deeper View

Why not compare only the number of patients who progressed?

Because the timing of progression matters. Two trials could have the same proportion of participants experiencing an event while having very different distributions of when those events occurred. A time-to-event method uses the follow-up time rather than discarding it.

Why does randomization support causal interpretation?

Randomization creates the treatment groups before post-randomization outcomes occur. In expectation, this balances measured and unmeasured prognostic factors across treatment assignments. The resulting comparison is therefore fundamentally different from an observational comparison in which treatment choice may be associated with baseline risk.

Why use the ITT population?

The ITT approach maintains the treatment groups created by randomization. If participants are removed from their randomized groups because of what happens after assignment, the original randomized comparison can be weakened or distorted.

What does a confidence interval crossing 1 mean for a hazard ratio?

For a hazard ratio, 1 represents equal hazards. A confidence interval that crosses 1 therefore includes values compatible with equal hazards as well as values on both sides of the null value. In ATLAS, the OS interval of 0.698–1.205 has this property.

Can the PFS and OS hazard ratios be directly compared as if they were the same outcome?

No. The PFS event includes progression or death, while the OS event is death from any cause. Different event definitions produce different estimands, so the numerical hazard ratios describe different treatment effects.

20. PFS Versus OS: An Important Statistical Distinction

PFS

The primary endpoint asks how long participants remain alive without documented disease progression. Its event definition combines progression and death.

Overall survival

The secondary endpoint asks about time to death from any cause. It does not count documented disease progression as an event by itself.

This distinction is especially important when reading a trial report. A favorable PFS result and a different OS estimate are not inherently inconsistent. They are estimates for different endpoints, with different event definitions and potentially different sources of information over follow-up.

21. Results Summary

EndpointRoleTime frameHR95% CIP-value
Progression-free SurvivalPrimaryApproximately 3 years0.7080.580–0.8640.0006
Overall SurvivalSecondaryApproximately 3.5 years0.9170.698–1.2050.5341

The two results illustrate why a complete statistical interpretation should report the point estimate, uncertainty interval, endpoint definition, analysis population, and testing framework together. The PFS analysis reports a hazard ratio below 1 with a 95% confidence interval entirely below 1 and a P-value of 0.0006. The OS analysis reports a hazard ratio below 1, but its 95% confidence interval crosses 1 and its P-value is 0.5341.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue with the statistical methods

Explore the broader Clinical Biostats collection of biostatistics tutorials, statistical calculators, and clinical-trial analyses.

25. Record Summary

ATLAS is a phase 3 randomized, double-blind, parallel clinical trial with a registered primary endpoint of progression-free survival. The primary PFS analysis used an ITT population, a log-rank test, and a hazard ratio as the effect measure. The reported PFS hazard ratio was 0.708, with a two-sided 95% CI of 0.580–0.864 and a P-value of 0.0006.

The posted secondary overall-survival analysis used the same general time-to-event framework and reported an HR of 0.917, with a two-sided 95% CI of 0.698–1.205 and a P-value of 0.5341. The contrast between these estimates demonstrates why endpoint definition, effect measure, confidence interval, and hypothesis test must all be considered together rather than relying on a single statistic.

The ClinicalTrials.gov record also report serious adverse events affecting 63/367 participants in the bevacizumab + placebo group and 86/368 participants in the bevacizumab + erlotinib group. These safety counts are descriptive in the ClinicalTrials.gov record and do not by themselves establish a statistical difference between groups.

Clinical Biostats methodology: A trial-results page should not merely repeat a numerical result. The goal is to explain what the endpoint measures, why the statistical method fits the endpoint, how the effect estimate and uncertainty should be interpreted, and which conclusions are supported by the reported analysis.