← Clinical Trials
Breast Cancer Phase 3 Completed NCT01358877

APHINITY: Complete Statistical Analysis of Pertuzumab in HER2-Positive Primary Breast Cancer

An independent statistical review of the randomized, double-blind phase 3 APHINITY trial evaluating pertuzumab added to trastuzumab and chemotherapy as adjuvant therapy in participants with HER2-positive primary breast cancer.

Trial start: 8 November 2011  ·  Primary completion: 19 December 2016  ·  Enrollment: 4804
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the trial data posted on ClinicalTrials.gov for APHINITY. Where the ClinicalTrials.gov record does not provide a formal estimate or analysis method, that limitation is stated rather than supplemented from outside sources.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

APHINITY was a randomized, double-blind, parallel-group phase 3 trial in breast cancer. The registry describes a two-arm comparison of pertuzumab plus trastuzumab and chemotherapy versus placebo plus trastuzumab and chemotherapy, with a total enrollment of 4804 participants.

4804
Enrollment
Participants
2
Arms
Parallel design
0.81
Primary HR
95% CI 0.66–1.00
0.0446
Primary P-value
Two-sided
FeatureAPHINITY
PhasePhase 3
ConditionBreast Cancer
Brief titleA Study of Pertuzumab in Addition to Chemotherapy and Trastuzumab as Adjuvant Therapy in Participants With Human Epidermal Growth Receptor 2 (HER2)-Positive Primary Breast Cancer
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment4804
Trial statusCompleted
Lead sponsorHoffmann-La Roche
Sponsor typeIndustry
ClinicalTrials.govNCT01358877

2. Clinical Question

The statistical question is whether adding pertuzumab to trastuzumab plus chemotherapy changes the time-to-event outcomes recorded as primary endpoints compared with placebo plus trastuzumab plus chemotherapy in participants with HER2-positive primary breast cancer receiving adjuvant therapy.

Population

Participants with HER2-positive primary breast cancer receiving adjuvant therapy, according to the registered trial population.

Intervention

Pertuzumab plus trastuzumab plus chemotherapy.

Comparator

Placebo plus trastuzumab plus chemotherapy.

Primary question

Does the pertuzumab-containing regimen alter invasive disease-free survival relative to the placebo-containing regimen?

The registry lists 10 interventions: 5-Fluorouracil, carboplatin, cyclophosphamide, docetaxel, doxorubicin, epirubicin, paclitaxel, pertuzumab, placebo, and trastuzumab. The ClinicalTrials.gov record does not provide a detailed treatment schedule or dosing sequence, so this page does not infer one.

3. Trial Design

01
Randomize4804 enrolled
02
Two armsPertuzumab or placebo
03
Double blindMasked comparison
04
FollowTime-to-event outcomes
05
AnalyzeITT efficacy analysis
Allocation
Randomized allocation to two parallel treatment arms.
Masking
Double masking.
Primary purpose
Treatment.
Primary endpoint type
Time-to-event.
ARM A

Pertuzumab-containing regimen

  • Pertuzumab
  • Trastuzumab
  • Chemotherapy
ARM B

Placebo-containing regimen

  • Placebo
  • Trastuzumab
  • Chemotherapy

The design is important statistically because randomization creates the basis for comparing outcomes between treatment assignments, while double masking is intended to reduce the influence of knowledge of treatment assignment on trial conduct and outcome assessment. The ClinicalTrials.gov record does not report the randomization ratio, so no ratio is assigned here.

4. Endpoints

The registry lists two primary endpoints, both concerning invasive disease-free survival (IDFS), and identifies the primary endpoint type as time-to-event.

EndpointRegistry time frameTypeFormal analysis reported
Percentage of Participants With Invasive Disease-Free Survival (IDFS) Event (Excluding Second Primary Non-Breast Cancer [SPNBC]), as Assessed Using Radiologic, Histologic Examinations or Laboratory Findings Randomization to the first occurrence of IDFS event (excluding SPNBC) until data cut-off date 19 December 2016 Time-to-event Yes
Kaplan-Meier Estimate of the Percentage of Participants Who Were IDFS Event-Free (Excluding SPNBC) at 3 Years, as Assessed Using Radiologic, Histologic Examinations or Laboratory Findings 3 years Time-to-event No formal comparative analysis reported

Why the endpoint is time-to-event rather than simply a proportion

Although the outcome measure is phrased as a “percentage of participants with” an IDFS event, the registered endpoint is explicitly classified as time-to-event and is defined from randomization to the first occurrence of an event. That distinction matters. A simple percentage ignores when events occur, whereas a time-to-event analysis uses both event occurrence and the amount of observed follow-up.

Participants who have not experienced an IDFS event by their last available assessment can contribute information up to that point and may subsequently be censored. This is one reason survival-analysis methods are useful for clinical trials with unequal follow-up.

5. Analysis Populations and Statistical Structure

AnalysisPopulation reportedRole
Primary IDFS analysis ITT population Primary efficacy comparison
Secondary time-to-event analyses ITT population Efficacy comparisons
Cardiac event analyses Safety population Safety assessment
LVEF analysis Safety population; evaluable participants for the outcome Safety assessment

The primary IDFS analysis was performed in the intention-to-treat population. This means the efficacy comparison is anchored to randomized treatment assignment rather than changing a participant's analysis group according to later treatment exposure. The registry analysis explicitly identifies ITT as an analysis concept.

The safety analyses use a different framework. For primary cardiac events, the registry defines the safety population as participants who received any amount of study medication—chemotherapy, pertuzumab/placebo, or trastuzumab—according to the treatment actually received. This distinction is important because efficacy and safety answer different statistical questions.

Why ITT matters: the strength of randomization comes from comparing groups as assigned. If post-randomization events such as treatment discontinuation were allowed to determine who remained in an efficacy group, the original randomized comparison could be distorted. ITT analysis therefore protects the treatment comparison even when actual treatment exposure is imperfect.

6. Statistical Methodology

Stratified log-rank test

The primary IDFS comparison used a stratified log-rank test. The registry analysis states that the stratification factors were nodal status, protocol version, central hormone receptor status, and adjuvant chemotherapy regimen. These factors were included as stratification factors in the randomization and in the time-to-event comparison.

What the log-rank test asks
H0: the event-time distributions are equivalent between treatment groups

The log-rank test compares the observed and expected pattern of events between randomized groups over follow-up. A stratified version performs that comparison while accounting for prespecified strata.

The test is therefore not comparing only the final percentage of participants with an event. It uses the timing of events throughout the follow-up period. This is appropriate for an endpoint defined from randomization to the first occurrence of an IDFS event.

Cox regression for the hazard ratio

The registry states that the hazard ratio was estimated by Cox regression. The same stratification factors used for the stratified log-rank analysis were incorporated into the analysis framework.

Hazard-ratio interpretation
HR = estimated hazard in pertuzumab group ÷ estimated hazard in placebo group

An HR below 1 indicates a lower estimated instantaneous event rate in the pertuzumab-containing group under the fitted model. An HR of 1 corresponds to no relative difference in hazard.

The hazard ratio is a relative time-to-event measure. It does not directly tell us the absolute probability of an event by a particular time, the number needed to treat, or the fraction of individual participants who benefit. Those quantities require additional information.

Kaplan-Meier estimation

The second registered primary endpoint is explicitly a Kaplan-Meier estimate of the percentage of participants who were IDFS event-free at 3 years. Kaplan-Meier estimation is designed for time-to-event data with censoring.

Kaplan-Meier survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di is the number of events at event time ti, while ni is the number of participants at risk immediately before that time.

The ClinicalTrials.gov record confirms that results were posted for the 3-year IDFS event-free endpoint, but they do not provide a formal statistical analysis entry or a numerical Kaplan-Meier estimate for that endpoint. The analysis therefore cannot responsibly supply a 3-year event-free percentage from the information provided.

Confidence intervals

The reported hazard ratios use two-sided 95% confidence intervals. A confidence interval describes the statistical uncertainty around an estimated treatment effect under the relevant analysis framework. It is not a range containing 95% of individual patient outcomes.

P-values

The primary analysis reports a two-sided P-value of 0.0446. A P-value measures the compatibility of the observed data with the null hypothesis under the specified testing framework. It does not measure the size of the treatment effect and does not provide the probability that the null hypothesis is true.

7. Results: Primary IDFS Analysis

The primary formal analysis reported in the registry compares the pertuzumab plus trastuzumab plus chemotherapy group with the placebo plus trastuzumab plus chemotherapy group for the percentage of participants with an IDFS event excluding SPNBC. The analysis population was the ITT population.

Hazard ratio for IDFS event

0.81

95% CI: 0.66–1.00   ·   P = 0.0446

Two-sided stratified log-rank test; HR estimated by Cox regression.

Primary endpointResult
OutcomeIDFS event excluding SPNBC
Analysis populationITT population
ComparisonPertuzumab + Trastuzumab + Chemotherapy vs Placebo + Trastuzumab + Chemotherapy
MethodStratified log-rank test
Effect measureHazard ratio
HR0.81
95% CI0.66–1.00
P-value0.0446
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 0.81 means that, under the Cox model used for the analysis, the estimated instantaneous rate of an IDFS event in the pertuzumab-containing group was approximately 19% lower than the corresponding estimated hazard in the placebo-containing group.

That does not mean that 19% fewer participants experienced an IDFS event, nor does it mean that each participant had exactly a 19% reduction in risk. A hazard ratio is a model-based relative measure of event rates over time.

The 95% CI of 0.66–1.00 indicates uncertainty around the HR estimate. Its upper boundary reaches 1.00, so the reported estimate is relatively close to the conventional no-difference value. The interval is therefore important context alongside the point estimate.

The P-value of 0.0446 does not quantify the magnitude of the treatment effect. It addresses the statistical test under the prespecified framework. The size of the effect is conveyed by the HR, while the CI communicates precision.

Because the analysis is time-to-event based, interpretation also depends on the assumptions and censoring structure underlying the survival analysis. In particular, a single Cox HR is most straightforward to interpret when the relative hazards are reasonably represented by a proportional-hazards framework over the relevant follow-up.

What the primary result establishes statistically

The registry analysis reports a superiority hypothesis, a stratified log-rank test, a hazard ratio of 0.81, a two-sided 95% CI of 0.66–1.00, and a P-value of 0.0446. These quantities form a coherent statistical description of the reported primary comparison.

The result should nevertheless be read as a time-to-event treatment comparison, not as a simple comparison of percentages. The registry's endpoint wording identifies the first occurrence of an IDFS event after randomization, and the statistical method explicitly accounts for event timing.

8. Planned Analysis for the 3-Year IDFS Event-Free Endpoint

The second registered primary endpoint is the Kaplan-Meier estimate of the percentage of participants who were IDFS event-free at 3 years. Results are posted for this endpoint, but the ClinicalTrials.gov record does not provide a formal comparative analysis, numerical 3-year estimate, confidence interval, or P-value for it.

For an endpoint of this type, the natural descriptive analysis is a Kaplan-Meier estimate evaluated at 3 years. If the objective is to formally compare treatment groups, a time-to-event comparison such as a log-rank test and an effect estimate such as a Cox-model hazard ratio would ordinarily be considered, subject to the prespecified statistical analysis plan.

Data boundary: this page does not supply a numerical 3-year IDFS event-free estimate because that number is not present in the ClinicalTrials.gov record. Reporting a value from another source would violate the trial-data rule for this analysis.

9. Secondary Time-to-Event Results

The registry provides several secondary time-to-event analyses. Each of the efficacy analyses below uses the ITT population and the same general stratified log-rank framework: nodal status, protocol version, central hormone receptor status, and adjuvant chemotherapy regimen were used as stratification factors, with the hazard ratio estimated by Cox regression.

Secondary endpointHR95% CIP-value
IDFS event including SPNBC 0.82 0.68–0.99 0.0430
DFS event 0.81 0.67–0.98 0.0327
Death, first interim overall survival analysis 0.89 0.66–1.21 0.4673
Death, final overall survival analysis 0.83 0.69–1.00 0.0441
Recurrence-free interval event 0.79 0.63–0.99 0.0430
Distant recurrence-free interval event 0.82 0.64–1.04 0.1007

IDFS including second primary non-breast cancer recurrence (i.e., an invasive breast cancer in the axilla, regional lymph nodes, chest wall, and/or skin of the ipsilateral breast); distant recurrence (i.e., evidence of breast cancer in any anatomic site - other than the two above mentioned sites); death attributable to any cause; contralateral invasive breast cancer. All SPNBCs and in situ carcinomas (including DCIS and LCIS) and non-melanoma skin cancer were excluded as an event.

Hazard ratio

0.82

95% CI: 0.68–0.99   ·   P = 0.0430

This endpoint broadens the IDFS definition by including SPNBC. The HR of 0.82 corresponds to an approximately 18% lower estimated hazard of the defined event in the pertuzumab-containing group under the Cox model.

Disease-free survival

Hazard ratio

0.81

95% CI: 0.67–0.98   ·   P = 0.0327

The DFS analysis has an HR of 0.81, corresponding to an approximately 19% lower estimated event hazard under the fitted model. The confidence interval lies below 1.00, although the interval itself is the appropriate measure of statistical precision rather than the P-value alone.

Overall survival — first interim analysis

Hazard ratio for death

0.89

95% CI: 0.66–1.21   ·   P = 0.4673

The first interim overall survival analysis estimated an HR of 0.89. Because the confidence interval extends from 0.66 to 1.21, the ClinicalTrials.gov record allows for considerable uncertainty around the relative hazard. The P-value of 0.4673 is not a measure of the size of the HR and should not be used as a substitute for examining the confidence interval.

Overall survival — final analysis

Hazard ratio for death

0.83

95% CI: 0.69–1.00   ·   P = 0.0441

Median [range] follow-up: 11.3 [0-12.9] years.

The final overall survival analysis reports an HR of 0.83. Under the Cox model, this represents an approximately 17% lower estimated hazard of death in the pertuzumab-containing group. The 95% CI of 0.69–1.00 again places the upper boundary at the conventional no-difference value.

Recurrence-free interval

Hazard ratio

0.79

95% CI: 0.63–0.99   ·   P = 0.0430

An HR of 0.79 corresponds to an approximately 21% lower estimated hazard of a recurrence-free interval event under the fitted model. The confidence interval provides the uncertainty around that estimate and does not describe the range of individual patient experiences.

Distant recurrence-free interval

Hazard ratio

0.82

95% CI: 0.64–1.04   ·   P = 0.1007

The estimated HR is 0.82, but the 95% CI extends above 1.00 to 1.04. This illustrates why a point estimate alone is insufficient. The estimate suggests a lower hazard in the pertuzumab-containing group, while the interval communicates uncertainty that includes the no-difference value.

Clinical Biostats interpretation

Across the registry-reported secondary time-to-event analyses, the HR estimates range from 0.79 to 0.89. These are relative measures of estimated event hazard, not direct absolute risk reductions.

The confidence intervals also differ meaningfully in their precision. Some exclude 1.00, while others include 1.00. A confidence interval should be read together with the prespecified endpoint hierarchy and multiplicity strategy rather than treating each P-value as an isolated test.

The two overall survival analyses are particularly useful for illustrating the role of follow-up. The first interim analysis reports HR 0.89 with a 95% CI of 0.66–1.21, whereas the final analysis reports HR 0.83 with a 95% CI of 0.69–1.00. These are distinct analyses and should not be silently combined into one estimate.

10. Cardiac Safety Results

The ClinicalTrials.gov record includes primary and secondary cardiac-event analyses as well as changes in LVEF. These analyses use the safety population rather than the ITT population used for the efficacy analyses.

Primary cardiac event

AnalysisEstimate95% CIFollow-up
Primary analysis Treatment difference 0.4 0.0–0.8 Baseline until data cut-off 19 December 2016; median [range] follow-up 3.8 [0.1-4.9] years
Final analysis Treatment difference 0.4 -0.1–0.9 Baseline until end of follow-up; median [range] follow-up 11.3 [0.1-12.9] years

The registry describes the effect measure as a treatment difference, with the difference in percentage of participants experiencing a primary cardiac event between the pertuzumab and placebo arms. The 95% confidence interval was estimated using the Hauck-Anderson correction.

The ClinicalTrials.gov record does not report a separate inferential method field for this endpoint. Therefore, the page does not assign a method beyond the stated Hauck-Anderson confidence-interval correction.

Secondary cardiac event

AnalysisEstimate95% CIFollow-up
Primary analysis Treatment difference -0.1 -1.0–0.9 Baseline until data cut-off 19 December 2016; median [range] follow-up 3.8 [0.1-4.9] years
Final analysis Treatment difference -0.1 -1.1–0.9 Baseline until end of follow-up; median [range] follow-up 11.3 [0.1-12.9] years

The secondary cardiac-event treatment difference is -0.1 in both registry-reported analyses. The corresponding confidence intervals extend on both sides of zero, illustrating substantial uncertainty around this small estimated difference.

How to interpret a treatment difference

Difference rather than hazard ratio
Treatment difference = percentage in pertuzumab group − percentage in placebo group

For these cardiac endpoints, the registry reports an “Other difference” rather than a hazard ratio. The confidence interval is centered on the reported treatment difference.

This is an important statistical distinction. A treatment difference on a percentage scale is not interchangeable with a hazard ratio. It is also not automatically equivalent to a risk ratio or odds ratio.

11. Serious Adverse Events by Arm

the ClinicalTrials.gov record reports serious adverse events by randomized treatment arm as affected participants divided by participants at risk.

Treatment groupParticipants with serious AEsParticipants at risk
Pertuzumab + Trastuzumab + Chemotherapy 792 2364
Placebo + Trastuzumab + Chemotherapy 686 2405

The ClinicalTrials.gov record therefore identify 792/2364 affected participants in the pertuzumab-containing group and 686/2405 in the placebo-containing group.

Interpretation boundary: the ClinicalTrials.gov record does not provide a formal comparative statistical analysis for serious adverse events. Accordingly, the page reports the affected and at-risk counts without calculating an unreported effect measure or P-value.

It is also important not to convert these safety counts into an efficacy-style hazard ratio. Serious adverse events are being presented as an affected/at-risk safety summary, whereas the primary IDFS analysis was explicitly a time-to-event analysis in the ITT population.

12. LVEF Results

The registry reports change from baseline in LVEF to the worst post-baseline value at both the primary and final analyses. The outcome is listed in the ClinicalTrials.gov record with the outcome unit “percentage of blood pumped out.”

AnalysisTreatment difference95% CIFollow-up
Primary analysis 0.1 -0.3–0.5 Baseline until data cut-off 19 December 2016; median [range] follow-up 3.8 [0.1-4.9] years
Final analysis 0.0 -0.4–0.4 Baseline until end of follow-up; median [range] follow-up 11.3 [0.1-12.9] years

The analysis population was the safety population, with the overall number analyzed defined as the participants evaluable for the outcome measure. The ClinicalTrials.gov record does not provide the corresponding numbers analyzed, so they are not added here.

The registry does not report a normalized inferential method for the LVEF endpoint. It does report that the 95% confidence interval was estimated using Hauck-Anderson correction.

Clinical Biostats interpretation

The primary-analysis treatment difference is 0.1 with a 95% CI of -0.3 to 0.5. The final-analysis treatment difference is 0.0 with a 95% CI of -0.4 to 0.4.

The key statistical feature is the width and location of the confidence intervals. Both intervals span zero, so the ClinicalTrials.gov record does not support treating the point estimate alone as evidence of a nonzero treatment difference in this endpoint.

These are changes to the worst post-baseline LVEF value, not hazard ratios or time-to-event estimates. They therefore require a different interpretation from the IDFS and overall survival analyses.

13. Statistical Methods Explained

Why was a stratified log-rank test used?

The primary endpoint is time-to-event, so the analysis needs to account for both whether an IDFS event occurred and when it occurred. The log-rank test is designed for comparing event-time distributions. Stratification allows the comparison to account for prespecified factors used in the randomization and analysis: nodal status, protocol version, central hormone receptor status, and adjuvant chemotherapy regimen.

What does an HR of 0.81 mean?

An HR of 0.81 means that the fitted Cox model estimates the instantaneous event hazard in the pertuzumab-containing group at approximately 81% of the corresponding hazard in the placebo-containing group. Expressed as a relative reduction in estimated hazard, that is approximately 19%.

It does not mean that 19% of participants avoided an event, that the absolute event probability fell by 19 percentage points, or that every participant experienced the same relative change.

Why is the confidence interval important?

The point estimate is only one summary of the observed data. The 95% CI of 0.66–1.00 for the primary HR describes the uncertainty around the estimated relative hazard. A narrower interval would generally indicate greater statistical precision; a wider interval indicates greater uncertainty. The CI also reveals information that a P-value alone cannot: the range of effect estimates reasonably compatible with the analysis framework.

Why doesn't the P-value measure effect size?

The P-value answers a hypothesis-testing question. It does not tell us whether an effect is clinically large or small. For APHINITY, the primary P-value is 0.0446, while the effect estimate is HR 0.81 and the 95% CI is 0.66–1.00. Those three quantities answer different questions and should be interpreted together.

Why is ITT important for the primary analysis?

The ITT population maintains the comparison created by randomization. An efficacy analysis based on randomized assignment helps preserve the causal interpretation of the treatment contrast even when participants do not remain perfectly adherent to their assigned therapy. The registry-reported primary analysis explicitly identifies the ITT population.

Why are the safety and efficacy populations different?

Efficacy asks what happened under randomized treatment assignment, whereas exposure-related safety is often summarized according to treatment actually received. APHINITY's the ClinicalTrials.gov record illustrate this distinction: the primary IDFS analysis uses the ITT population, while cardiac safety analyses use the safety population.

Why should the two primary endpoints not be collapsed into one statistic?

The registry lists two primary endpoints with different statistical summaries: the first concerns IDFS events from randomization until the first event, while the second is a Kaplan-Meier estimate of being IDFS event-free at 3 years. They are related, but they are not the same estimand. One should therefore not substitute a hazard ratio for the 3-year Kaplan-Meier estimate or vice versa.

14. Understanding the Primary Confidence Interval

Point estimate

HR 0.81 is the central estimate of the relative event hazard from the fitted Cox model.

Precision

The 95% CI of 0.66–1.00 communicates uncertainty around that point estimate.

Hypothesis test

P = 0.0446 is the reported two-sided result from the stratified log-rank framework.

Estimand

The effect is a relative time-to-event comparison, not an absolute percentage-point difference.

A useful way to read the primary result is therefore:

Four-part reading
Effect size → HR 0.81   |   Precision → 95% CI 0.66–1.00   |   Test → P = 0.0446   |   Population → ITT

Each component answers a different statistical question. Omitting any one of them makes the result harder to interpret correctly.

The same principle applies to the secondary endpoints. For example, the final overall survival estimate is HR 0.83 with a 95% CI of 0.69–1.00 and P = 0.0441. The hazard ratio communicates the estimated relative event rate, while the confidence interval describes its uncertainty and the P-value describes the reported hypothesis-test result.

15. Interpreting Time-to-Event Results Correctly

Time-to-event endpoints require several layers of interpretation that are easy to lose when a clinical trial is reduced to a single HR.

QuestionRelevant statistical quantity
How large is the estimated relative treatment effect?Hazard ratio
How uncertain is that estimate?Confidence interval
How compatible are the data with the null hypothesis?P-value
When did the endpoint occur?Event-time information
How is censoring handled?Kaplan-Meier / survival-analysis framework
Which participants form the primary efficacy comparison?ITT population
Were important randomization factors incorporated?Stratified analysis

This framework is particularly relevant for APHINITY because the registry reports several time-to-event outcomes, all analyzed using the same broad survival-analysis family. The individual endpoints nevertheless have different definitions and should not be treated as interchangeable.

Hazard ratio versus absolute event probability

An HR is inherently relative. Suppose two groups have different baseline event risks: the same hazard ratio could correspond to very different absolute differences depending on the underlying event rates and follow-up. The HR therefore cannot, by itself, answer how many additional participants avoid an event.

Hazard ratio versus median survival

A hazard ratio also should not be interpreted as a ratio of median event times. The registry-reported APHINITY data do not provide median IDFS or overall survival values, so this page does not infer them.

Hazard ratio versus risk ratio

A risk ratio compares probabilities over a specified period. A hazard ratio compares instantaneous event rates through a time-to-event model. They are mathematically and conceptually different quantities.

16. Secondary Endpoint Interpretation

The secondary analyses demonstrate why the statistical analysis of a clinical trial cannot be reduced to a single P-value. The registry-reported estimates include HR 0.82 for IDFS including SPNBC, HR 0.81 for DFS, HR 0.89 for the first interim overall survival analysis, HR 0.83 for final overall survival, HR 0.79 for recurrence-free interval, and HR 0.82 for distant recurrence-free interval.

EndpointWhat the HR describes95% CI relationship to 1.00
IDFS including SPNBCRelative hazard of the defined IDFS eventUpper bound 0.99
DFSRelative hazard of the defined DFS eventUpper bound 0.98
First interim OSRelative hazard of deathIncludes 1.00
Final OSRelative hazard of deathUpper bound 1.00
RFIRelative hazard of breast cancer recurrenceUpper bound 0.99
DRFIRelative hazard of distant breast cancer recurrenceIncludes 1.00

This table is descriptive rather than a ranking of endpoints. A confidence interval crossing 1.00 is not itself proof that there is no treatment effect; it indicates that the interval includes the no-difference value under the stated analysis. Conversely, an interval not crossing 1.00 does not by itself establish that an endpoint is clinically important.

17. Multiplicity and Multiple Endpoints

APHINITY has two registered primary endpoints and multiple secondary analyses. That creates an important statistical issue: a clinical trial can generate many numerical comparisons, and the interpretation of each test depends on the prespecified hierarchy and multiplicity-control strategy.

The ClinicalTrials.gov record identifies the hypothesis type for the reported analyses as superiority, but they do not provide a detailed multiplicity-adjustment strategy or alpha-allocation scheme. This page therefore does not infer one.

Why this matters: a collection of P-values should not automatically be interpreted as a collection of independent confirmatory findings. When multiple endpoints or analyses are evaluated, the statistical design determines how Type I error is controlled and which tests carry confirmatory status.

The primary IDFS analysis is clearly identified in the ClinicalTrials.gov record as a primary analysis. The secondary endpoint results should retain their secondary status even when an individual P-value is below a conventional threshold.

18. Stratification and Why It Matters

The primary analysis used stratification factors that included nodal status, protocol version, central hormone receptor status, and adjuvant chemotherapy regimen. These factors were used in the randomization and incorporated into the stratified log-rank analysis.

Improved alignment with randomization

Stratified analysis respects important factors that were built into the trial's allocation structure.

Control of prognostic imbalance

Stratification can reduce the influence of differences in event experience across prespecified strata.

Not a subgroup claim

Using a variable for stratification is different from claiming that treatment effects differ across that variable.

Interpretation

The reported HR is a stratified treatment comparison, not an unadjusted comparison of raw event percentages.

This distinction is often overlooked. A stratification factor does not automatically become a treatment-effect subgroup. To claim that treatment effects differ between categories, an interaction analysis or another formal heterogeneity assessment is generally required.

19. Follow-Up and Analysis Timing

8 November 2011

Trial start

The APHINITY trial began according to the registry profile.

19 December 2016

Primary completion / principal data cutoff

The registry-reported primary endpoint analyses use a data cut-off date of 19 December 2016. Several cardiac and time-to-event endpoints explicitly reference this date.

11.3 years

Final overall survival follow-up

The final overall survival analysis reports median [range] follow-up of 11.3 [0-12.9] years.

Follow-up duration matters because time-to-event estimates depend on how long participants are observed. The ClinicalTrials.gov record contains both an earlier first interim overall survival analysis and a later final overall survival analysis, illustrating that an HR is tied to a particular analysis time and data cutoff.

The primary IDFS analysis is tied to the data cutoff of 19 December 2016. The final overall survival analysis instead reports median [range] follow-up of 11.3 [0-12.9] years. These should be treated as distinct analysis contexts rather than merged into a single generic “long-term” estimate.

20. Safety Versus Efficacy: Different Statistical Questions

APHINITY provides a useful teaching example of why safety and efficacy analyses should be separated.

FeatureEfficacySafety
Primary population reportedITTSafety population
Primary IDFS measureHazard ratioNot applicable
Cardiac event measureNot applicableTreatment difference
Serious AEsNot applicableAffected / at risk counts
LVEFNot applicableTreatment difference
Primary time-to-event methodStratified log-rank + Cox HRNot reported for cardiac/LVEF endpoints

This separation prevents an important analytical mistake: an efficacy HR should not be compared numerically with a safety treatment difference as though both were measuring the same quantity. Each effect measure has its own estimand and interpretation.

21. Primary Result: A Complete Statistical Reading

Clinical Biostats interpretation

The primary APHINITY analysis reports a hazard ratio of 0.81 for IDFS events excluding SPNBC, based on the ITT population and a stratified log-rank test, with the HR estimated by Cox regression.

The estimate corresponds to an approximately 19% lower estimated instantaneous event hazard in the pertuzumab-containing group under the fitted model. This is a relative hazard interpretation, not an absolute risk reduction.

The 95% CI of 0.66–1.00 describes uncertainty around the HR. Its upper boundary reaches 1.00, so the precision of the estimated effect is an important part of the interpretation.

The P-value of 0.0446 is the reported two-sided hypothesis-test result. It should not be interpreted as the probability that the treatment has no effect, nor as a measure of how large the treatment effect is.

The analysis was stratified by nodal status, protocol version, central hormone receptor status, and adjuvant chemotherapy regimen. Consequently, the HR should be understood as the result of the specified stratified survival-analysis framework rather than as a simple unadjusted ratio of event percentages.

22. Important Limitations and Interpretation Issues

23. Why This Trial Matters Statistically

APHINITY is a useful statistical teaching case because it connects a randomized phase 3 design with several core concepts in clinical-trial analysis: time-to-event endpoints, Kaplan-Meier estimation, stratified log-rank testing, Cox regression, hazard ratios, confidence intervals, ITT analysis, secondary endpoints, and distinct safety estimands.

ConceptHow it appears in APHINITY
RandomizationRandomized, parallel-group phase 3 design
BlindingDouble masking
ITT analysisPrimary and secondary efficacy time-to-event analyses
Time-to-event endpointsIDFS, DFS, overall survival, RFI, and DRFI
Kaplan-MeierRegistered 3-year IDFS event-free endpoint
Log-rank testingPrimary and secondary time-to-event comparisons
Stratified analysisNodal status, protocol version, central hormone receptor status, and adjuvant chemotherapy regimen
Cox regressionEstimation of reported hazard ratios
Confidence intervals95% two-sided intervals for treatment effects
Superiority testingHypothesis type reported for the analyses
Safety populationCardiac and LVEF analyses
Treatment differenceCardiac-event and LVEF analyses
Multiple endpointsTwo registered primary endpoints plus multiple secondary analyses

24. Statistical Methods in Context

The statistical story of APHINITY is easiest to understand as a sequence rather than as a collection of isolated formulas.

01
RandomizationCreates treatment comparison
02
Follow-upObserve event times
03
Kaplan-MeierEstimate event-free experience
04
Log-rankCompare groups
05
Cox HRQuantify relative effect

Randomization establishes the comparison. Follow-up generates event and censoring information. Kaplan-Meier methods describe event-free experience over time. The stratified log-rank test compares the event-time distributions. Cox regression then supplies a model-based hazard ratio that summarizes the relative event hazard.

Confidence intervals and P-values add complementary inferential information. Finally, the analysis population and endpoint hierarchy determine what population and hypothesis the numerical result actually represents.

25. What the Primary HR Does — and Does Not — Mean

What HR 0.81 means

Under the reported Cox model, the estimated instantaneous rate of an IDFS event was approximately 19% lower in the pertuzumab-containing group than in the placebo-containing group.

What HR 0.81 does not mean

It does not mean that 19% of participants benefited, that the absolute event probability fell by 19 percentage points, that participants lived 19% longer, or that every participant experienced the same proportional change.

What the CI adds

The 95% CI of 0.66–1.00 communicates uncertainty around the estimated HR. It does not represent the distribution of individual treatment effects.

What the P-value adds

The P-value of 0.0446 describes the reported hypothesis-test result under the specified two-sided stratified log-rank framework. It is not an effect-size measure.

26. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary ITT analysis reports HR 0.81, 95% CI 0.66–1.00, and P = 0.0446 using a stratified log-rank test with Cox regression for HR estimation.

Endpoint interpretation

The effect concerns the time from randomization to the first occurrence of the registered IDFS event excluding SPNBC.

Uncertainty interpretation

The confidence interval should be considered alongside the point estimate and P-value; the upper boundary reaches 1.00.

Safety interpretation

Cardiac events, serious adverse events, and LVEF are separate safety questions and use different analysis populations and effect measures.

Keeping these layers separate prevents the common mistake of turning a trial into a single “positive” or “negative” number. The statistical evidence is better represented by the combination of endpoint definition, analysis population, effect measure, confidence interval, hypothesis test, and design structure.

27. Sources

The PubMed links above are provided as the linked publication records reported with the trial data. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

28. Related Tutorials

Learn more about the methods used in this trial:

29. Related Calculators

Continue through Clinical Biostats

Connect the statistical methods in this trial to tutorials and practical statistical tools for clinical-trial analysis.

30. Record Summary

APHINITY provides a compact example of a randomized phase 3 clinical-trial analysis centered on time-to-event methodology. The primary formal analysis used the ITT population, a stratified log-rank test, and Cox regression to estimate a hazard ratio of 0.81 with a two-sided 95% CI of 0.66–1.00 and P = 0.0446 for IDFS events excluding SPNBC.

The broader statistical record includes secondary analyses of IDFS including SPNBC, DFS, overall survival, recurrence-free interval, and distant recurrence-free interval. Their reported HR estimates range from 0.79 to 0.89. The final overall survival analysis reports HR 0.83 with a 95% CI of 0.69–1.00 and median [range] follow-up of 11.3 [0-12.9] years.

The safety analyses illustrate a different statistical structure. Cardiac events and LVEF are summarized using treatment differences and confidence intervals, while serious adverse events are reported as affected participants divided by participants at risk. These analyses use the safety population rather than the ITT population used for efficacy.

The most important lesson is therefore methodological: the effect measure must be interpreted in the context of its endpoint, population, analysis method, confidence interval, and hypothesis-testing framework. A hazard ratio is not a risk ratio, a treatment difference is not a hazard ratio, and a P-value is not an effect size.

Clinical Biostats methodology: A trial-results page should not merely repeat a headline result. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and avoiding unsupported numerical reconstruction.