← Clinical Trials
HER2+ Metastatic Breast Cancer Phase 3 Completed NCT01808573

NALA: Complete Statistical Analysis of Neratinib Plus Capecitabine in HER2+ Metastatic Breast Cancer

An independent statistical review of the randomized phase 3 NALA trial comparing neratinib plus capecitabine with lapatinib plus capecitabine in patients with HER2+ metastatic breast cancer who had received two or more prior HER2-directed regimens in the metastatic setting.

Randomized phase 3  ·  Enrollment 621  ·  Primary completion September 28, 2018
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

NALA was a randomized, open-label, parallel phase 3 trial comparing neratinib plus capecitabine with lapatinib plus capecitabine in patients with HER2+ metastatic breast cancer who had received two or more prior HER2-directed regimens in the metastatic setting.

621
Enrolled
Randomized phase 3 trial
2
Treatment arms
Parallel design
0.762
PFS HR
95% CI 0.626–0.926
0.881
OS HR
95% CI 0.723–1.073
FeatureNALA
PhasePhase 3
ConditionHER2+ metastatic breast cancer
PopulationPatients who had received two or more prior HER2-directed regimens in the metastatic setting
DesignRandomized, parallel, open-label
AllocationRandomized
Primary endpointsCentrally assessed progression-free survival and overall survival
Primary endpoint typeTime-to-event
Enrollment621
Lead sponsorPuma Biotechnology, Inc.
Sponsor typeIndustry
ClinicalTrials.govNCT01808573

2. Clinical Question

The central statistical question was whether neratinib plus capecitabine produced superior time-to-event outcomes compared with lapatinib plus capecitabine in patients with HER2+ metastatic breast cancer who had received two or more prior HER2-directed regimens in the metastatic setting.

Population

Patients with HER2+ metastatic breast cancer who had received two or more prior HER2-directed regimens in the metastatic setting.

Intervention

Neratinib plus capecitabine.

Comparator

Lapatinib plus capecitabine.

Primary question

Does the neratinib-plus-capecitabine regimen provide superior centrally assessed PFS and overall survival compared with lapatinib plus capecitabine?

3. Trial Design

01
Randomize621 enrolled
02
Two armsNeratinib or lapatinib
03
Combination therapyEach regimen included capecitabine
04
Follow-upTime-to-event endpoints
05
AnalysisStratified survival methods
INTERVENTION

Neratinib Plus Capecitabine

  • Neratinib
  • Capecitabine
  • Randomized comparison arm
REFERENCE ARM

Lapatinib Plus Capecitabine

  • Lapatinib
  • Capecitabine
  • Reference treatment group for the reported hazard ratios
Open-label design: the registry describes masking as none. Thus, the treatment assignment was not masked. This is an important design distinction from a blinded randomized trial because knowledge of treatment assignment can potentially affect aspects of treatment delivery, assessment, or patient-reported experience. The primary PFS endpoint was nevertheless based on central assessment as defined in the registry.

4. Trial Timeline and Registry Status

MilestoneRegistry information
Trial startMarch 29, 2013
Primary completionSeptember 28, 2018
Trial statusCompleted
Enrollment621
Phase3
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment

The registry reports statistical results for seven posted outcome measures, including six formal statistical analyses. Two of those analyses correspond to the registered primary endpoints.

5. Primary Endpoints

EndpointRegistered definition and time frameAnalysis
Centrally Assessed Progression Free Survival Progression Free Survival (PFS), measured in months, for randomized subjects of the central assessment. The time interval from the date of randomization until the first date on which recurrence, progression, or death due to any cause is documented. The registry time frame is from randomization to recurrence, progression or death, assessed up to 38 months. Stratified log-rank test; hazard ratio
Overall Survival Time from randomization to death due to any cause, censored at the last date known alive on or prior to the data cutoff employed for the analysis. The registry reports the result as restricted mean survival time, defined as the area under the survival function up to 48 months. The registry time frame is from randomization to death, assessed up to 59 months. Stratified log-rank test; hazard ratio

Both primary endpoints are time-to-event endpoints. That means the analysis incorporates not only whether an event occurred but also the amount of follow-up available for each randomized participant.

6. Statistical Methodology

Intention-to-treat analysis

The primary PFS analysis was performed in the Randomized ITT Population. An intention-to-treat framework analyzes randomized participants according to the treatment assignment generated by randomization. This preserves the central comparison created by the randomized design rather than redefining treatment groups according to subsequent treatment exposure.

Log-rank testing

The registry reports a log-rank test for both primary endpoints. The log-rank test is designed for comparing time-to-event distributions between treatment groups while accounting for censored observations.

Conceptual survival comparison
H0: the treatment groups have the same event-time distribution

The log-rank framework compares the observed pattern of events with the pattern expected under the null hypothesis across the follow-up period. It is therefore different from a simple comparison of proportions at one fixed time point.

Stratified analysis

The primary analyses were stratified by hormone receptor status, number of prior HER2-directed regimens in the metastatic setting, and visceral disease versus non-visceral disease.

Stratification allows the time-to-event comparison to account for these prespecified factors rather than treating the randomized population as if every participant had the same baseline risk. The same stratification factors are reported for the primary PFS and OS analyses and for the secondary CNS-disease analysis.

Hazard ratio

The registry reports hazard ratios as the effect measure for the primary PFS and OS analyses. With lapatinib plus capecitabine as the reference group, a hazard ratio below 1 indicates a lower estimated instantaneous event rate in the neratinib-plus-capecitabine group under the fitted time-to-event comparison.

Interpretive framework
HR = estimated event hazard with neratinib + capecitabine ÷ estimated event hazard with lapatinib + capecitabine

The hazard ratio is a relative time-to-event measure. It is not a probability, not a percentage of patients who benefit, and not a statement that every participant experiences the same proportional change in risk.

Confidence intervals

Each primary endpoint has a two-sided 95% confidence interval around its hazard ratio. The interval describes statistical uncertainty around the estimated relative treatment effect under the analysis framework. A narrower interval generally indicates greater precision than a wider interval, although precision and clinical importance are separate questions.

P-values

The P-value addresses the statistical evidence against the null hypothesis under the specified testing framework. It does not measure the size of the treatment effect. For that reason, the P-value should be interpreted together with the hazard ratio and its confidence interval.

7. Primary Results: Centrally Assessed Progression-Free Survival

Hazard ratio for progression, recurrence, or death

0.762

95% CI: 0.626–0.926   ·   P = 0.0059

Randomized ITT population  ·  Stratified log-rank analysis

ElementReported result
EndpointCentrally Assessed Progression Free Survival
PopulationRandomized ITT Population
ComparisonNeratinib Plus Capecitabine vs Lapatinib Plus Capecitabine
AnalysisLog-rank test
Effect measureHazard ratio
Estimate0.762
95% CI0.626–0.926
P-value0.0059
HypothesisSuperiority
Time frameFrom randomization date to recurrence, progression or death, assessed up to 38 months
Clinical Biostats interpretation

The reported HR of 0.762 means that, under the reported time-to-event analysis, the estimated instantaneous rate of the PFS event was about 76.2% of the corresponding rate in the lapatinib-plus-capecitabine reference group. Equivalently, 1 − 0.762 = 0.238, so the estimate corresponds to an approximately 23.8% lower estimated hazard for the PFS event in the neratinib group.

This does not mean that 23.8% of patients avoided progression, that 23.8% of patients were cured, or that every patient experienced exactly a 23.8% reduction in risk. It is a relative hazard estimate from a time-to-event analysis.

The 95% CI of 0.626–0.926 describes uncertainty around the estimated hazard ratio. It does not describe the range of outcomes for individual patients. Because the interval is below 1 throughout, the reported estimate is consistent with a lower event hazard for the neratinib-plus-capecitabine group under this analysis.

The P-value of 0.0059 is evidence against the null hypothesis used for the reported superiority comparison. It does not tell us that the treatment effect is "0.59%" or that there is a 0.59% probability that the treatment has no effect. The magnitude and precision of the effect are better conveyed by the HR and its confidence interval.

As with other hazard-ratio analyses, interpretation also depends on the time-to-event model and censoring structure. A single HR summarizes a relative event-rate comparison; it is not an absolute difference in survival probabilities at a particular time.

What the PFS endpoint actually measures

The registered endpoint begins at randomization and ends at the first documented recurrence, progression according to RECIST v1.1, or death from any cause. Participants without recurrence, progression, or death are censored according to the registry definition.

This makes PFS different from a binary response endpoint. A patient who remains progression-free contributes information for as long as that patient remains under observation without the event, rather than simply being classified as "progressed" or "not progressed" at a single time point.

8. Primary Results: Overall Survival

Hazard ratio for death

0.881

95% CI: 0.723–1.073   ·   P = 0.2086

Stratified log-rank analysis  ·  Lapatinib Plus Capecitabine as reference

ElementReported result
EndpointOverall Survival
ComparisonNeratinib Plus Capecitabine vs Lapatinib Plus Capecitabine
AnalysisLog-rank test
Effect measureHazard ratio
Estimate0.881
95% CI0.723–1.073
P-value0.2086
HypothesisSuperiority
Time frameFrom randomization date to death, assessed up to 59 months
Additional registry definitionTime to death due to any cause, censored at the last date known alive on or prior to the data cutoff. The result was reported as restricted mean survival time, defined as the area under the survival function up to 48 months.
Clinical Biostats interpretation

The reported OS HR of 0.881 corresponds to an estimated instantaneous rate of death approximately 88.1% of that in the lapatinib-plus-capecitabine reference group under the reported analysis. Expressed as a relative hazard difference, 1 − 0.881 = 0.119, corresponding to an approximately 11.9% lower estimated hazard of death.

That estimate should not be translated into an 11.9% increase in survival probability or an 11.9% absolute improvement in survival. The hazard ratio is a relative time-to-event measure, not an absolute survival difference.

The 95% CI is 0.723–1.073. Unlike the PFS interval, this interval crosses 1.00. Consequently, the reported estimate is compatible with both a lower and a higher instantaneous mortality hazard for the neratinib group relative to the reference group within the uncertainty represented by this interval.

The P-value of 0.2086 does not measure the size of the observed effect. It indicates that the reported superiority test did not provide strong statistical evidence against its null hypothesis under the specified analysis framework. A non-small P-value should not be interpreted as proof that the two treatments are identical.

The registry also describes the OS result using restricted mean survival time, with the restricted mean defined as the area under the survival curve up to 48 months. That is conceptually different from a hazard ratio because it summarizes survival time accumulated through a specified time horizon rather than expressing an instantaneous relative event rate.

9. Understanding the Two Primary Endpoints Together

Progression-Free Survival

HR 0.762 with a 95% CI of 0.626–0.926 and P = 0.0059. The reported analysis provides evidence against the null hypothesis for the superiority comparison.

Overall Survival

HR 0.881 with a 95% CI of 0.723–1.073 and P = 0.2086. The confidence interval includes 1.00 and the reported superiority comparison has a larger P-value.

The two endpoints answer different questions. PFS evaluates the time from randomization to recurrence, progression, or death, whereas OS evaluates the time from randomization to death from any cause. A treatment effect on one endpoint does not mathematically guarantee the same magnitude of effect on another.

The numerical contrast between the two reported hazard ratios is therefore itself statistically informative: the estimated relative effect is smaller for PFS than for OS, and the uncertainty around the OS estimate is sufficiently broad to include 1.00.

Do not reverse the interpretation: an OS P-value of 0.2086 does not mean that the probability of treatment benefit is 20.86%, nor does it establish equivalence of the two regimens. Similarly, the PFS P-value of 0.0059 does not quantify the clinical size of the PFS benefit.

10. Secondary Endpoint Results

Duration of Response

Hazard ratio for duration of response

0.495

95% CI: 0.332–0.736   ·   P = 0.0004

Central assessment among the population that had a response with measurable disease at screening

The registry defines this endpoint as the time from the start date of response after randomization to first progressive disease, with assessment up to 33 months. The analysis used a log-rank test and reported lapatinib plus capecitabine as the reference group.

Clinical Biostats interpretation

The HR of 0.495 indicates an estimated instantaneous rate of the duration-of-response event of about 49.5% of that in the reference group. Equivalently, the estimated hazard is approximately 50.5% lower under the reported model.

The confidence interval of 0.332–0.736 quantifies uncertainty around that estimate. This endpoint is conditional on having a response with measurable disease at screening, so it is not a substitute for the ITT PFS analysis.

The P-value of 0.0004 provides evidence against the reported superiority null hypothesis. It does not indicate the probability that a particular patient will have a durable response.

Intervention for Symptomatic Metastatic Central Nervous System Disease

ElementReported result
EndpointIntervention for Symptomatic Metastatic Central Nervous System Disease
Time frameFrom randomization date to first intervention for symptomatic metastatic CNS disease, assessed up to 59 months
AnalysisGray's test
Effect measureNo hazard ratio reported in the registry statistical analysis
P-value0.043
HypothesisSuperiority
StratificationHormone receptor status, number of prior HER2-directed regimens in the metastatic setting, and visceral disease vs. non-visceral disease
Clinical Biostats interpretation

The registry reports a Gray's test P-value of 0.043 for this endpoint but does not provide a hazard ratio or corresponding confidence interval in the registry-reported statistical analysis. Therefore, the registry result supports a statistical comparison of the cumulative incidence framework but does not provide a numerical relative effect estimate that can be interpreted like the PFS or OS hazard ratios.

Gray's test is used for comparing cumulative incidence functions when competing risks are relevant. The key distinction is that a competing event can prevent the event of interest from subsequently occurring, so treating every competing event as ordinary censoring can produce an inappropriate risk comparison.

Objective Response Rate

ElementReported result
EndpointObjective Response Rate (ORR) — Central Assessment
PopulationITT Population With Measurable Disease at Screening
Time frameFrom randomization date to first confirmed Complete or Partial Response, whichever came earlier, up to 42 months
AnalysisCochran-Mantel-Haenszel test
Effect measureNo effect estimate reported in the registry statistical analysis
P-value0.1201
HypothesisSuperiority
StratificationHormone receptor status, number of prior HER2-directed regimens in the metastatic setting, and visceral disease vs. non-visceral disease

The registry reports a Cochran-Mantel-Haenszel P-value of 0.1201, but it does not provide the response percentages, an effect estimate, or a confidence interval in the registry-reported statistical analysis. The formal method is therefore reported without adding unprovided response rates.

Clinical Benefit Rate

ElementReported result
EndpointClinical Benefit Rate (CBR) — Central Assessment
PopulationITT Population With Measurable Disease at Screening
Time frameFrom randomization date to either first confirmed CR or PR or Stable Disease, whichever came earlier, up to 42 months
AnalysisCochran-Mantel-Haenszel test
Effect measureNo effect estimate reported in the registry statistical analysis
P-value0.0328
HypothesisSuperiority
StratificationHormone receptor status, number of prior HER2-directed regimens in the metastatic setting, and visceral disease vs. non-visceral disease
Clinical Biostats interpretation

The reported P = 0.0328 indicates evidence against the null hypothesis for the specified superiority comparison under the Cochran-Mantel-Haenszel framework. However, because the ClinicalTrials.gov record does not report the underlying percentages or an effect estimate, the size of the difference cannot be quantified from this result alone.

This illustrates why a P-value should not be presented as a stand-alone measure of clinical effect. Without the corresponding group estimates and confidence interval, the numerical magnitude and precision of the CBR difference remain unspecified in the ClinicalTrials.gov record.

11. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These figures are presented as affected participants over participants at risk.

Safety measureNeratinib Plus CapecitabineLapatinib Plus Capecitabine
Serious adverse events103/30393/311

Neratinib Plus Capecitabine

103/303 participants were affected by serious adverse events among the reported number at risk.

Lapatinib Plus Capecitabine

93/311 participants were affected by serious adverse events among the reported number at risk.

The ClinicalTrials.gov record does not provide a formal statistical comparison or confidence interval for these serious-adverse-event counts. Accordingly, the page reports the registry values without constructing an unreported hypothesis test.

Safety interpretation: the denominators reported for serious adverse events are 303 and 311, respectively, rather than the overall enrollment of 621. These figures should therefore not be silently converted into percentages using 621 as the denominator.

12. Statistical Methods Explained

Why was a log-rank test used for PFS and OS?

PFS and OS are time-to-event outcomes. Some participants may not experience the event during the observation period, creating right-censored observations. The log-rank test is designed for comparing survival distributions while incorporating that censoring rather than treating all participants as if they had identical follow-up.

What does a PFS hazard ratio of 0.762 mean?

With lapatinib plus capecitabine as the reference, the reported HR of 0.762 means that the estimated instantaneous event rate for the PFS endpoint was approximately 76.2% of the reference rate under the fitted analysis. It corresponds to an approximately 23.8% lower estimated hazard, but it is not a 23.8-percentage-point improvement in the probability of remaining progression-free.

Why does the confidence interval matter?

The point estimate is only one possible estimate supported by the data. The 95% CI of 0.626–0.926 for PFS communicates uncertainty around the HR. For OS, the interval is wider relative to the estimate and spans 1.00, which means the reported data are compatible with a range of relative effects that includes no difference on the hazard-ratio scale.

Why is the OS P-value not a measure of treatment effect?

The OS P-value of 0.2086 summarizes evidence against the specified null hypothesis under the reported testing framework. It does not describe how large the treatment effect is. Effect magnitude is communicated by the HR, while uncertainty is communicated by its confidence interval.

Why was Gray's test used for the CNS endpoint?

The endpoint concerns the first intervention for symptomatic metastatic CNS disease, and the registry identifies Gray's test as the analysis method. Gray's test is used when cumulative incidence must be compared in the presence of competing risks. This is different from a standard Kaplan-Meier comparison that treats competing events as ordinary censoring.

Why was the Cochran-Mantel-Haenszel test used for ORR and CBR?

ORR and CBR are categorical endpoints rather than continuous time-to-event measurements. The Cochran-Mantel-Haenszel framework allows a treatment comparison while accounting for the reported stratification factors. In this trial those factors included hormone receptor status, number of prior HER2-directed regimens in the metastatic setting, and visceral disease versus non-visceral disease.

Why does the analysis population matter?

The primary PFS analysis is explicitly identified as using the randomized ITT population. By contrast, the duration-of-response analysis is restricted to the population that had a response with measurable disease at screening. These populations answer different questions. A response-duration analysis cannot be interpreted as though it represented every randomized participant.

13. Stratification and the Cochran-Mantel-Haenszel Framework

The registry identifies three stratification factors for the reported stratified analyses:

Stratification factorCategories as reported
Hormone receptor statusHormone receptor status
Prior HER2-directed treatmentNumber of prior HER2-directed regimens in the metastatic setting
Visceral diseaseVisceral disease vs. non-visceral disease

Stratification is useful when important baseline factors are expected to influence event rates. Instead of ignoring those factors, the analysis compares treatment groups within the relevant strata and combines information across strata using the prespecified statistical framework.

The same general principle applies to the Cochran-Mantel-Haenszel analyses of ORR and CBR: the categorical treatment comparison is adjusted for the specified stratification factors rather than being reduced to an unstratified comparison.

Conceptual distinction
Unstratified comparison → one overall comparison
Stratified comparison → comparisons within prespecified strata, combined across strata

The purpose is not to manufacture a larger treatment effect. It is to account for the trial's prespecified structure and the prognostic factors used in the analysis.

14. Time-to-Event Endpoints and Censoring

NALA provides several examples of why time-to-event analysis requires more information than a simple event percentage.

EndpointStarting pointEventReported statistical framework
Centrally Assessed PFSRandomizationRecurrence, progression, or deathLog-rank test; HR
Overall SurvivalRandomizationDeath due to any causeLog-rank test; HR
Duration of ResponseStart of response after randomizationFirst progressive diseaseLog-rank test; HR
Symptomatic metastatic CNS interventionRandomizationFirst intervention for symptomatic metastatic CNS diseaseGray's test

For PFS, the registry definition states that participants without recurrence, progression, or death are censored at the relevant end of follow-up. For OS, participants are censored at the last date known alive on or prior to the analysis data cutoff.

Censoring allows participants with incomplete event information to contribute the follow-up that is actually observed. The validity of the resulting survival analysis depends on the censoring and analysis assumptions, so a hazard ratio should always be understood as an estimate generated from a particular time-to-event framework rather than as a simple observed percentage.

15. Restricted Mean Survival Time in the OS Definition

The registry includes an additional description of the OS result: time to event was reported as the restricted mean survival time, with the restricted mean defined as the area under the survival function up to 48 months.

Restricted mean survival time
RMST(τ) = ∫0τ S(t) dt

Here, the registry defines the relevant restriction at 48 months. Conceptually, the area under the survival curve represents average survival time accumulated through the specified horizon.

RMST is useful because it expresses treatment differences on a time scale rather than only as a relative hazard. It does not require the reader to interpret a hazard ratio as an instantaneous event-rate ratio. In this registry record, however, the registry-reported statistical analysis reports the primary OS comparison as a hazard ratio and does not provide a numerical RMST difference.

16. Primary Endpoint Interpretation: Relative vs Absolute Evidence

The NALA registry results illustrate why statistical interpretation should not rely on a single number.

Relative effect

The PFS HR is 0.762 and the OS HR is 0.881. These summarize relative event-rate differences under the reported time-to-event analyses.

Precision

The PFS 95% CI is 0.626–0.926, while the OS 95% CI is 0.723–1.073. The intervals communicate uncertainty around the point estimates.

Statistical evidence

The PFS P-value is 0.0059 and the OS P-value is 0.2086. These address evidence against the respective null hypotheses rather than clinical magnitude.

Endpoint definition

PFS and OS measure different events. A difference in one endpoint cannot simply be substituted for the other.

17. Multiplicity and the Two Primary Endpoints

The registry identifies two registered primary endpoints: centrally assessed PFS and overall survival. Both are superiority hypotheses and both have formal statistical analyses posted.

EndpointRoleHypothesisFormal analysis
Centrally Assessed PFSPrimarySuperiorityLog-rank test
Overall SurvivalPrimarySuperiorityLog-rank test

The ClinicalTrials.gov record does not provide an alpha-allocation scheme, multiplicity-adjustment procedure, or hierarchical testing sequence for the two primary endpoints. Accordingly, this page does not infer one.

Statistical caution: when a trial has multiple primary endpoints, the interpretation of nominal P-values depends on the prespecified multiplicity strategy. Because that strategy is not included in the ClinicalTrials.gov record, the two reported P-values are presented exactly as reported without assigning an unreported familywise-error interpretation.

18. Interim Analysis and Other Design Features

The ClinicalTrials.gov record does not report an interim-analysis plan, alpha-spending procedure, Bayesian method, non-inferiority margin, crossover rule, or missing-data/imputation method. Those design topics are therefore not characterized here.

This distinction matters because the absence of a field in the ClinicalTrials.gov record is not evidence that a particular method was or was not used in the full protocol or statistical analysis plan. The appropriate conclusion from the available data is simply that the ClinicalTrials.gov record does not specify the method.

19. Statistical Interpretation of the Secondary Analyses

Secondary endpointMethodReported estimate95% CIP-value
Duration of ResponseLog-rankHR 0.4950.332–0.7360.0004
Intervention for Symptomatic Metastatic CNS DiseaseGray's testNot reportedNot reported0.043
Objective Response RateCochran-Mantel-HaenszelNot reportedNot reported0.1201
Clinical Benefit RateCochran-Mantel-HaenszelNot reportedNot reported0.0328

The four secondary analyses demonstrate four distinct reporting situations. Duration of response has a hazard ratio and confidence interval; the CNS endpoint has a competing-risks test without a reported effect estimate; ORR and CBR have stratified categorical tests without reported effect estimates in the registry-reported statistical analysis.

That difference in reporting is important. A P-value without an effect estimate cannot tell the reader how large the between-group difference was. Conversely, an effect estimate without its uncertainty would provide an incomplete picture of precision.

20. Important Limitations and Interpretation Issues

21. Why This Trial Matters Statistically

NALA is a useful teaching case because it combines several major clinical-trial statistical concepts in one randomized phase 3 study: two primary time-to-event endpoints, stratified log-rank testing, hazard ratios, confidence intervals, an ITT efficacy population, a response-duration analysis restricted to responders, a competing-risks analysis, and stratified categorical analyses.

ConceptHow it appears in NALA
RandomizationRandomized phase 3 parallel-group comparison
Intention-to-treat analysisPrimary PFS analysis uses the Randomized ITT Population
Time-to-event endpointsPFS, OS, duration of response, and symptomatic CNS intervention
Log-rank testUsed for the two primary endpoints and duration of response
Hazard ratioReported for PFS, OS, and duration of response
Confidence interval95% two-sided intervals reported for the primary HR estimates and duration of response
Stratified analysisHormone receptor status, prior HER2-directed regimens, and visceral disease status
Gray's testUsed for the symptomatic metastatic CNS intervention endpoint
Cochran-Mantel-Haenszel testUsed for ORR and CBR
Competing risksRelevant to the Gray's-test analysis of symptomatic CNS intervention
Restricted mean survival timeOS definition describes RMST as area under the survival function through 48 months

22. What the Primary Hazard Ratios Do — and Do Not — Mean

PFS

A PFS HR of 0.762 means that the estimated instantaneous event rate was approximately 76.2% of the reference-group rate under the reported survival analysis. It corresponds to an approximately 23.8% lower estimated hazard.

It does not mean that 23.8% more patients remained progression-free, that 23.8% of patients were cured, or that each patient had exactly the same proportional reduction in risk.

OS

An OS HR of 0.881 means that the estimated instantaneous mortality rate was approximately 88.1% of the reference-group rate under the reported analysis. It corresponds to an approximately 11.9% lower estimated hazard.

The 95% CI of 0.723–1.073 includes 1.00, so the uncertainty interval encompasses the possibility of no relative difference on the hazard-ratio scale.

Duration of response

A duration-of-response HR of 0.495 corresponds to an estimated event hazard approximately 49.5% of the reference-group hazard, or an approximately 50.5% lower estimated hazard under the reported model.

This estimate applies to the response-defined analysis population, not to all randomized participants.

23. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The randomized comparison produced a PFS HR of 0.762 and an OS HR of 0.881. The PFS confidence interval lies below 1.00, whereas the OS confidence interval spans 1.00.

Endpoint interpretation

PFS, OS, duration of response, CNS intervention, ORR, and CBR are distinct outcomes. Their analyses should not be collapsed into a single numerical measure of treatment effect.

Uncertainty interpretation

Confidence intervals describe uncertainty around estimates. The OS interval is particularly important because it extends above 1.00.

Reporting interpretation

For secondary endpoints without reported effect estimates, the P-values provide evidence from the specified tests but do not quantify the size of the treatment difference.

24. Related Tutorials

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue through Clinical Biostats

Connect this trial's endpoints and statistical methods to deeper biostatistics tutorials, statistical calculators, and other clinical trial analyses.

27. Record Summary

NALA provides a useful example of how a randomized phase 3 oncology trial can combine several statistical frameworks according to endpoint type. The primary PFS and OS analyses use stratified log-rank testing and hazard ratios, while duration of response uses another time-to-event comparison, symptomatic CNS intervention uses Gray's test for competing risks, and ORR and CBR use Cochran-Mantel-Haenszel analyses.

The primary PFS result is reported as HR 0.762 (95% CI 0.626–0.926; P = 0.0059), while the primary OS result is HR 0.881 (95% CI 0.723–1.073; P = 0.2086). The statistical interpretation of these results depends on the distinction between relative effect, absolute outcome, uncertainty, and the specific endpoint being analyzed.

The secondary analyses further demonstrate why the method must match the outcome. Duration of response is analyzed as a time-to-event endpoint, the CNS intervention endpoint uses a competing-risks method, and ORR and CBR are categorical outcomes analyzed with the Cochran-Mantel-Haenszel test. For several secondary endpoints, the ClinicalTrials.gov record reports P-values without effect estimates, so the size of those differences cannot be inferred from the P-values alone.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. Where the ClinicalTrials.gov record provides an estimate, confidence interval, and P-value, all three are presented together. Where only a statistical test and P-value are reported, the analysis is described without inventing an effect size or confidence interval.