← Clinical Trials
Stage III Unresectable NSCLC Phase 3 Completed NCT02125461

PACIFIC: Complete Statistical Analysis of MEDI4736 in Stage III Unresectable Non-Small Cell Lung Cancer

An independent statistical review of the randomized phase 3 PACIFIC trial evaluating MEDI4736 following concurrent chemoradiation in patients with stage III unresectable non-small cell lung cancer, with emphasis on time-to-event methodology, hazard ratios, confidence intervals, and the interpretation of primary and secondary endpoint results.

Trial start: 07 May 2014  ·  Primary completion: 13 Feb 2017  ·  Enrollment: 713
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View NCT02125461 on ClinicalTrials.gov.

1. Trial at a Glance

PACIFIC was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating MEDI4736 versus placebo in patients with stage III unresectable non-small cell lung cancer following concurrent chemoradiation. The registry reports 713 enrolled patients, two arms, two primary time-to-event endpoints, and formal statistical analyses for both primary endpoints.

713
Enrollment
2 randomized arms
2
Primary endpoints
Both time-to-event
0.52
PFS HR
95% CI 0.42–0.65
0.68
OS HR
95% CI 0.53–0.87
FeaturePACIFIC
Trial namePACIFIC
PhasePhase 3
ConditionNon-Small Cell Lung Cancer
Population descriptionPatients with Stage III Unresectable Non-Small Cell Lung Cancer following concurrent chemoradiation
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment713
Arms2
InterventionsMEDI4736 (drug); Placebo (other)
Primary endpoints2
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
Lead sponsorAstraZeneca
Sponsor typeIndustry
StatusCompleted
Trial start2014-05-07
Primary completion2017-02-13

2. Clinical Question

The central statistical question was whether MEDI4736 following concurrent chemoradiation improved the time-to-event outcomes defined as progression-free survival and overall survival compared with placebo.

Population

Patients with Stage III Unresectable Non-Small Cell Lung Cancer following concurrent chemoradiation.

Intervention

MEDI4736.

Comparator

Placebo.

Primary question

Does MEDI4736 improve progression-free survival and overall survival relative to placebo?

The registry identifies the primary hypotheses as superiority. This is important because the statistical question is whether the MEDI4736 group differs favorably from the placebo group under the prespecified time-to-event framework, rather than whether the two groups are sufficiently similar to support an equivalence or non-inferiority conclusion.

3. Trial Design

01
Randomize713 enrolled
02
Two armsMEDI4736 vs placebo
03
FollowTime-to-event outcomes
04
AssessPFS / OS / secondary outcomes
05
AnalyzeStratified survival methods
Allocation
Randomized allocation in a parallel-group design.
Masking
Quadruple masking.
Primary purpose
Treatment.
Statistical hypothesis
Superiority.
ARM A

MEDI4736

  • Intervention classified in the registry as a drug.
  • Compared with placebo in the randomized parallel-group design.
  • Primary efficacy analyses were based on the FAS and analyzed on an ITT basis.
ARM B

Placebo

  • Comparator classified in the registry as an other intervention.
  • Compared with MEDI4736 in the randomized parallel-group design.
  • Primary efficacy analyses were based on the FAS and analyzed on an ITT basis.

The registry does not provide an allocation ratio in the ClinicalTrials.gov record. Accordingly, this page does not infer one from the enrollment total or the safety denominators.

4. Trial Timeline and Registry Status

07 May 2014

Trial start

The registry lists 2014-05-07 as the study start date.

13 February 2017

Primary completion

The registry lists 2017-02-13 as the primary completion date.

22 March 2018

OS data cutoff

The registered overall-survival endpoint was assessed until the 22 Mar 2018 data cutoff, with a maximum of approximately 4 years.

Completed

Registry status

The trial is listed as completed. The registry caveat states that interim PFS and interim OS analyses are considered the final PFS and OS analyses, respectively.

Registry analysis caveat: Results of the interim PFS analysis are considered as final PFS analysis, and results of the interim OS analysis are considered as final OS analysis. Patients were followed up for long-term survival until approximately 5 years after the last patient enrolled.

5. Primary Endpoints

EndpointRegistry definitionTime framePrimary analysis
Progression Free Survival Based on Blinded Independent Central Review (BICR) According to Response Evaluation Criteria in Solid Tumors (RECIST 1.1) PFS was defined as the time from randomization until the date of objective disease progression (RECIST 1.1) or death (by any cause in the absence of progression). Progression was defined using RECIST 1.1 as a 20% increase in the sum of the longest diameter of target lesions, or a measurable increase in a non-target lesion, or the appearance of new lesions. PFS was calculated using the Kaplan-Meier technique. Tumor scans performed at baseline then every ~8 weeks up to 48 weeks, then every ~12 weeks thereafter until confirmed disease progression. Assessed until 13 Feb 2017 DCO; up to a maximum of approximately 3 years. Stratified log-rank test; hazard ratio
Overall Survival OS was defined as the time from the date of randomization until death due to any cause. OS was calculated using the Kaplan-Meier technique. From baseline until death due to any cause. Assessed until 22 Mar 2018 DCO; up to a maximum of approximately 4 years. Stratified log-rank test; hazard ratio

Both primary endpoints are time-to-event outcomes. This means the analysis must account not only for whether an event occurred, but also for the amount of follow-up available for each randomized patient and for patients who remain event-free at the time of analysis.

6. Analysis Population and Stratification

For both primary endpoints, the registry states that the FAS included all randomized patients and that patients were analyzed on an intent-to-treat (ITT) basis. The reported primary analyses used a stratified log-rank test.

FeatureRegistry-supported approach
Analysis populationFAS including all randomized patients
Analysis principleIntent-to-treat
Groups comparedDurvalumab (MEDI4736) vs Placebo
Primary comparisonStratified log-rank test
Effect measureHazard ratio
Confidence interval95%, two-sided
HypothesisSuperiority
Tie handlingBreslow approach
Stratification factorsAge at randomization (<65 vs ≥65), sex (male vs female), and smoking history (smoker vs non-smoker)

The same age, sex, and smoking-history factors were used in the stratified primary analyses. This matters because the treatment comparison is not simply an unadjusted comparison of two survival curves: the analysis explicitly incorporates these factors into the time-to-event comparison.

7. Statistical Methodology

Kaplan-Meier estimation

The registry states that PFS and OS were calculated using the Kaplan-Meier technique. Kaplan-Meier estimation is appropriate for time-to-event data because it can incorporate right-censored observations. A patient who has not experienced progression or death by the last available assessment does not simply disappear from the analysis; the patient's observed follow-up contributes information up to the censoring time.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di is the number of events at event time ti, while ni is the number at risk immediately before that time. The resulting survival function estimates the probability of remaining event-free beyond time t.

Stratified log-rank testing

The primary PFS and OS comparisons used a stratified log-rank test. Rather than treating all randomized patients as belonging to one homogeneous risk set, the analysis adjusted the comparison through strata defined by age at randomization, sex, and smoking history.

Conceptually, the log-rank test asks whether the observed pattern of events over follow-up differs between treatment groups under the null hypothesis of no difference in the survival experience. The stratified version performs this comparison while accounting for the specified stratification factors.

Hazard ratio

The principal effect measure reported for the primary analyses was the hazard ratio (HR). An HR compares the estimated instantaneous event rates between groups over the analyzed follow-up. An HR below 1 indicates a lower estimated instantaneous event rate in the MEDI4736 group relative to placebo under the fitted analysis.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous event rate with MEDI4736

The HR is not a probability, not a percentage of patients who benefit, and not the same quantity as an absolute risk difference. Its interpretation is tied to the time-to-event model and its assumptions.

Covariate adjustment and stratification

The primary analysis notes explicitly identify covariate adjustment and stratified analysis. The stratification factors were age at randomization, sex, and smoking history. Ties were handled using the Breslow approach.

Confidence intervals

Each primary endpoint has a two-sided 95% confidence interval for the hazard ratio. The interval provides information about statistical precision around the estimated treatment effect. It is not a range containing 95% of individual patient effects, and it does not mean that there is a 95% probability that the true hazard ratio lies inside this particular calculated interval.

Intent-to-treat analysis

The FAS included all randomized patients and the analyses were performed on an ITT basis. The key principle is that randomized patients remain associated with the treatment group to which they were assigned for the efficacy comparison. This preserves the treatment contrast created by randomization and reduces the risk that post-randomization treatment behavior determines the primary comparison population.

8. Primary Result: Progression-Free Survival

The registry reports a formal primary analysis for progression-free survival based on BICR according to RECIST 1.1. The comparison was between MEDI4736 and placebo in the FAS, analyzed on an ITT basis.

Progression-free survival hazard ratio

0.52

95% CI: 0.42–0.65   ·   P < 0.0001

Stratified log-rank analysis; superiority hypothesis.

Primary PFS analysis featureReported result
Groups comparedDurvalumab (MEDI4736) vs Placebo
Analysis populationFAS; all randomized patients; ITT
Endpoint typeTime-to-event
Effect measureHazard ratio
Estimate0.52
95% CI0.42–0.65
P-value<0.0001
MethodStratified log-rank test
StratificationAge at randomization (<65 vs ≥65), sex (male vs female), smoking history (smoker vs non-smoker)
TiesBreslow approach
Clinical Biostats interpretation

An HR of 0.52 means that, under the reported time-to-event analysis, the estimated instantaneous rate of progression or death in the MEDI4736 group was approximately 52% of that in the placebo group. Equivalently, 1 − 0.52 = 0.48, so the estimated hazard was approximately 48% lower with MEDI4736 under this model.

This does not mean that 48% of patients avoided progression, that every patient experienced a 48% reduction in risk, or that the absolute probability of progression or death was reduced by 48 percentage points. The hazard ratio is a relative time-to-event measure.

The 95% CI of 0.42–0.65 describes the statistical uncertainty around the estimated HR. Its width also shows why reporting the interval is important: the point estimate alone does not communicate the precision of the treatment-effect estimate.

The P-value of <0.0001 addresses evidence against the null hypothesis within the specified testing framework. It does not measure the magnitude of the treatment effect. Effect magnitude is described by the HR and its confidence interval.

The analysis is also dependent on censoring rules and the assumptions underlying the time-to-event framework. In particular, a hazard ratio is not automatically equivalent to a constant relative risk over the entire follow-up period; proportional-hazards assumptions should be considered when interpreting a single HR as a summary of treatment effect.

The registry provides the Kaplan-Meier methodology and summary hazard-ratio result, but the ClinicalTrials.gov record does not provide the underlying patient-level event and censoring times required to reconstruct a valid Kaplan-Meier curve independently.

9. Primary Result: Overall Survival

Overall survival was the second primary endpoint. The registry defines OS as the time from randomization until death due to any cause and states that OS was calculated using the Kaplan-Meier technique.

Overall survival hazard ratio

0.68

95% CI: 0.53–0.87   ·   P = 0.00251

Stratified log-rank analysis; superiority hypothesis.

Primary OS analysis featureReported result
Groups comparedDurvalumab (MEDI4736) vs Placebo
Analysis populationFAS; all randomized patients; ITT
Endpoint typeTime-to-event
Effect measureHazard ratio
Estimate0.68
95% CI0.53–0.87
P-value0.00251
MethodStratified log-rank test
StratificationAge at randomization (<65 vs ≥65), sex (male vs female), smoking history (smoker vs non-smoker)
TiesBreslow approach
Data cutoff22 Mar 2018
Clinical Biostats interpretation

An HR of 0.68 indicates that, under the reported time-to-event model, the estimated instantaneous rate of death in the MEDI4736 group was approximately 68% of that in the placebo group. Equivalently, 1 − 0.68 = 0.32, corresponding to an estimated 32% lower instantaneous hazard of death under the model.

This is not the same as saying that mortality was reduced by 32 percentage points, that 32% of patients were saved, or that every individual patient experienced the same relative reduction. The HR summarizes a randomized group comparison over the analyzed follow-up.

The 95% CI of 0.53–0.87 gives the statistical uncertainty around the estimated HR. Because the interval is not centered on the null value of 1, the interval is consistent with a treatment effect below the null under the stated two-sided confidence-interval framework.

The P-value of 0.00251 measures the compatibility of the observed result with the relevant null hypothesis under the analysis framework. It does not tell us how large the treatment effect is; the HR and CI provide that information.

As with PFS, interpretation of a single HR requires attention to censoring and the proportional-hazards model used for the treatment effect. The registry's use of stratification also means that the reported comparison should not be interpreted as a simple unadjusted ratio calculated from two crude event proportions.

10. Primary Endpoint Results Together

The two primary endpoints give complementary views of the randomized treatment comparison. PFS captures the time to objective disease progression or death, whereas OS captures time to death from any cause.

Primary endpointHR95% CIP-valueAnalysis
Progression Free Survival based on BICR according to RECIST 1.10.520.42–0.65<0.0001Stratified log-rank
Overall Survival0.680.53–0.870.00251Stratified log-rank

Both reported HR estimates are below 1, and both confidence intervals exclude 1. The registry classifies both analyses as superiority analyses. The PFS and OS results should nevertheless be understood as separate endpoints with different clinical meanings rather than as two measurements of exactly the same outcome.

PFS asks

How does the time from randomization to objective progression or death compare between the randomized groups?

OS asks

How does the time from randomization to death from any cause compare between the randomized groups?

11. Secondary Time-to-Event Results

The registry contains formal statistical analyses for several secondary time-to-event endpoints. These analyses use the same broad framework of FAS/ITT analysis, comparison of MEDI4736 with placebo, and stratified time-to-event methods unless otherwise noted in the endpoint-specific analysis text.

Secondary endpointHR95% CIP-valueMethod
Time to Death or Distant Metastasis (TTDM) based on BICR assessments according to RECIST 1.1 0.530.41–0.68<0.0001Stratified log-rank
Time to Second Progression or Death (PFS2) 0.580.46–0.73<0.0001Stratified log-rank
Time to Deterioration of Global Health Status / HRQoL using EORTC QLQ-C30 0.950.77–1.180.664Stratified log-rank
Time to deterioration of PRO symptom: dyspnea 1.060.88–1.290.522Stratified log-rank
Time to deterioration of PRO symptom: cough 0.910.74–1.120.380Stratified log-rank
Time to deterioration of PRO symptom: hemoptysis 0.750.56–1.000.048Stratified log-rank
Time to deterioration of PRO symptom: chest pain 0.940.75–1.190.626Stratified log-rank

These secondary endpoints illustrate why a trial-results page should distinguish the direction and size of an estimate from its statistical significance. The HRs range from 0.53 to 1.06, and the corresponding confidence intervals range from relatively precise intervals well below 1 to intervals that include 1.

Time to Death or Distant Metastasis

TTDM hazard ratio

0.53

95% CI: 0.41–0.68   ·   P < 0.0001

Stratified log-rank analysis.

The TTDM analysis estimates the relative time-to-event experience for death or distant metastasis. The HR of 0.53 corresponds to an estimated 47% lower instantaneous event rate under the reported model. This interpretation concerns the composite event definition itself; it should not be interpreted as a 47% reduction in each component considered separately.

Time to Second Progression or Death

PFS2 hazard ratio

0.58

95% CI: 0.46–0.73   ·   P < 0.0001

Stratified log-rank analysis.

PFS2 extends the time-to-event framework beyond the first progression. The registry states that, following confirmed progression, patients were assessed every ~12 weeks until second disease progression. The HR of 0.58 corresponds to an estimated 42% lower instantaneous rate of the PFS2 event under the reported model.

Global Health Status / HRQoL

MeasureHR95% CIP-value
Time to deterioration of global health status / HRQoL, EORTC QLQ-C300.950.77–1.180.664
Dyspnea1.060.88–1.290.522
Cough0.910.74–1.120.380
Hemoptysis0.750.56–1.000.048
Chest pain0.940.75–1.190.626

For the global-health-status/HRQoL analysis, only patients with baseline scores ≥ 10 were included. For the listed EORTC QLQ-LC13 symptom analyses, only patients with baseline scores ≤ 90 were included. These eligibility rules matter because the denominator for these analyses is therefore defined by the endpoint-specific baseline-score requirement rather than simply by all 713 enrolled patients.

Interpretation caution: the registry supplies several secondary endpoint P-values, but the ClinicalTrials.gov record does not provide a multiplicity-adjustment scheme for interpreting these secondary tests as one family of independent confirmatory claims. A nominal P-value should therefore not be treated as a universal measure of clinical importance or as proof that an isolated secondary finding is independently confirmatory.

12. Objective Response Rate

Objective Response Rate (ORR) based on BICR assessments according to RECIST 1.1 was a secondary binary endpoint. The analysis population was the FAS, which included all randomized patients, with measurable disease at baseline, analyzed on an ITT basis.

Fisher exact comparison

P < 0.001

Secondary endpoint: Objective Response Rate

Fisher's exact test with mid P-value modification.

FeatureReported analysis
EndpointObjective Response Rate based on BICR assessments according to RECIST 1.1
Endpoint typeBinary
Outcome unitPercentage of patients
PopulationFAS with measurable disease at baseline; ITT
Groups comparedDurvalumab (MEDI4736) vs Placebo
MethodFisher exact test
P-value<0.001
ModificationMid P-value modification by subtracting half of the probability of the observed table from Fisher's P-value

The ClinicalTrials.gov record does not provide the response percentages, response counts, or a confidence interval for ORR. Therefore, this page reports the formal comparison that is available without reconstructing an unreported effect estimate.

Why Fisher's exact test is appropriate here

ORR reduces the response assessment to a binary outcome for each evaluable patient: the statistical analysis compares the resulting treatment-group response classifications. Fisher's exact test is designed for categorical contingency tables and calculates the exact probability under the specified null framework rather than relying on a large-sample approximation.

The registry used a mid P-value modification. This subtracts half the probability of the observed table from Fisher's P-value. The resulting value should be identified as a mid-P result rather than silently treating it as an ordinary unmodified Fisher exact P-value.

13. Percentage of Patients Alive at 24 Months

The registry also reports a secondary analysis of the Percentage of Patients Alive at 24 Months (OS24). This is a binary/time-point summary rather than a continuous measurement of survival time.

OS24 comparison

P = 0.005

Secondary endpoint: Percentage of Patients Alive at 24 Months.

Wald / z-test; variance estimated using the delta method and Greenwood's formula.

FeatureReported analysis
EndpointPercentage of Patients Alive at 24 Months (OS24)
Time frameFrom baseline until death due to any cause. Assessed until 22 Mar 2018 DCO; up to a maximum of approximately 4 years.
Endpoint typeBinary
Outcome unitPercentage of patients
PopulationFAS including all randomized patients; ITT
Groups comparedDurvalumab (MEDI4736) vs Placebo
MethodWald / z-test
P-value0.005
Variance estimationDelta method and Greenwood's formula

The ClinicalTrials.gov record does not provide the actual 24-month survival percentages by treatment group. Accordingly, the P-value is reported without inventing the corresponding group estimates.

Statistically, this endpoint is useful because it provides a fixed time-point summary, while the primary OS analysis uses the full time-to-event information. A fixed-time survival percentage and a hazard ratio are therefore complementary rather than interchangeable summaries.

14. Statistical Methods Explained

Why was a stratified log-rank test used?

The primary endpoints were time-to-event outcomes, so the treatment groups needed to be compared over the entire follow-up rather than using a simple comparison of event proportions. The registry reports a stratified log-rank test that accounts for age at randomization, sex, and smoking history. Stratification can reduce the influence of these specified factors on the treatment comparison while preserving the randomized-group framework.

What does an HR of 0.52 mean for PFS?

An HR of 0.52 means that the estimated instantaneous rate of progression or death in the MEDI4736 group was 52% of that in the placebo group under the reported analysis. The corresponding 48% figure is a relative reduction in estimated hazard, not a 48-percentage-point reduction in the probability of progression or death.

Why is the confidence interval important?

The point estimate is only one summary of the data. The 95% CI of 0.42–0.65 for PFS and 0.53–0.87 for OS shows the statistical precision surrounding the respective HR estimates. A narrower interval generally conveys more precision than a wider interval, while the location of the interval relative to the null value of 1 informs the uncertainty about the direction of the relative hazard.

Why doesn't the P-value measure effect size?

A P-value quantifies how compatible the observed data are with a specified null hypothesis under the statistical model and testing procedure. It is affected by both the magnitude of an observed effect and the amount of information in the analysis. The HR describes relative effect magnitude, while the confidence interval describes uncertainty around that estimate. These should not be replaced by the P-value alone.

Why is the Breslow method mentioned?

In time-to-event analyses, multiple events can occur at the same recorded time. These are tied event times. The registry states that ties were handled using the Breslow approach. Identifying the tie-handling method is part of describing the actual statistical model rather than assuming a default implementation.

Why does the ITT population matter?

The registry defines the FAS as including all randomized patients and states that the primary analyses were conducted on an ITT basis. This keeps the efficacy comparison aligned with randomized assignment. Removing patients after randomization because of subsequent treatment behavior can compromise the balance produced by randomization and change the clinical question being answered.

Why are the HRQoL analyses based on restricted baseline scores?

The HRQoL analysis included only patients with baseline scores ≥ 10, while the listed PRO symptom analyses included only patients with baseline scores ≤ 90. These rules define the relevant analysis populations for those endpoints. They also mean that the results should not be interpreted as though every randomized patient necessarily contributed to every PRO analysis.

15. Primary Analysis Interpretation: What the Numbers Do and Do Not Mean

PFS HR 0.52

The estimated instantaneous rate of progression or death was approximately 48% lower in the MEDI4736 group under the reported model. This does not mean that 48% of patients were protected from progression, nor does it specify the absolute probability of being progression-free at any particular time.

OS HR 0.68

The estimated instantaneous rate of death was approximately 32% lower in the MEDI4736 group under the reported model. This does not mean that the absolute mortality probability was reduced by 32 percentage points or that every patient experienced the same relative reduction.

Confidence intervals

The PFS 95% CI of 0.42–0.65 and OS 95% CI of 0.53–0.87 quantify statistical uncertainty around the respective hazard-ratio estimates. They do not describe the range of outcomes that individual patients might experience.

P-values

The PFS P-value of <0.0001 and OS P-value of 0.00251 provide evidence against the corresponding null hypotheses under the reported testing framework. They do not rank the size or clinical importance of the two treatment effects.

16. Time-to-Event Endpoints and Censoring

PFS and OS are fundamentally different from ordinary binary endpoints because patients can have different amounts of observed follow-up. Some patients experience the event during follow-up, while others remain event-free at their last assessment. Those latter observations are censored rather than treated as if the event never occurred.

For PFS, the event is objective disease progression or death in the absence of progression. For OS, the event is death due to any cause. The registry therefore defines different event processes even though both endpoints use the same broad Kaplan-Meier and stratified survival-analysis framework.

PFS event

Objective disease progression according to RECIST 1.1 or death by any cause in the absence of progression.

OS event

Death due to any cause.

17. Safety: Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected patients over patients at risk. This provides a direct arm-specific safety count, although the ClinicalTrials.gov record does not include a formal statistical comparison or a complete adverse-event profile.

Safety measureMEDI4736Placebo
Serious adverse events, affected / at risk138 / 47554 / 234
Serious adverse events: affected patients
MEDI4736
138
Placebo
54

The denominators are 475 and 234, respectively. The ClinicalTrials.gov record does not provide a formal P-value, confidence interval, risk ratio, or odds ratio for serious adverse events. Such measures should therefore not be inferred from the affected/at-risk counts when the task is to reproduce reported trial results exactly.

Safety and efficacy answer different questions. A lower or higher event rate for a safety endpoint does not mathematically cancel or validate a time-to-event efficacy estimate. The two evidence streams should be presented separately.

18. Secondary Endpoint Analysis Methods

Endpoint familyEndpoint typeMethodEffect measure / output
PFSTime-to-eventStratified log-rankHazard ratio
OSTime-to-eventStratified log-rankHazard ratio
TTDMTime-to-eventStratified log-rankHazard ratio
PFS2Time-to-eventStratified log-rankHazard ratio
Global health status / HRQoL deteriorationTime-to-eventStratified log-rankHazard ratio
PRO symptom deteriorationTime-to-eventStratified log-rankHazard ratio
ORRBinaryFisher exact testP-value
OS24BinaryWald / z-testP-value

This distribution of methods is statistically coherent with the endpoint types. Time-to-event outcomes require methods that account for follow-up and censoring; binary response outcomes can be analyzed using contingency-table methods; and a fixed-time survival percentage can be compared using a variance-based test, as reported here.

19. Stratified Cox Interpretation for Secondary PRO Analyses

For the global health status / HRQoL endpoint and the listed PRO symptom endpoints, the registry specifies that the HR and CI were estimated from a stratified Cox proportional hazards model. The Breslow method was used to control for ties, and the strata statement included age at randomization (<65 vs ≥65), sex (male vs female), and smoking history (smoker vs non-smoker). The confidence interval was calculated using a profile likelihood approach.

EndpointHR estimation details
Global health status / HRQoL deteriorationStratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI.
Dyspnea deteriorationStratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI.
Cough deteriorationStratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI.
Hemoptysis deteriorationStratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI.
Chest pain deteriorationStratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI.

This is an important methodological distinction: the registry method field lists the comparison as a log-rank analysis, while the endpoint-specific analysis text describes the model used to estimate the HR and confidence interval. A complete statistical description should preserve both pieces of information rather than reducing the analysis to a single method label.

20. Limitations and Interpretation Issues

What is deliberately not included: the ClinicalTrials.gov record does not report baseline characteristics, median PFS, median OS, subgroup estimates, forest plots, treatment crossover information, a factorial design, a non-inferiority margin, or Bayesian analyses. These topics are not reconstructed from outside knowledge.

21. Why This Trial Matters Statistically

PACIFIC is a useful statistical teaching case because the ClinicalTrials.gov record bring together several common methods in modern randomized clinical-trial analysis: ITT efficacy analysis, Kaplan-Meier estimation, stratified log-rank testing, hazard ratios, Cox proportional-hazards modeling, exact categorical testing, fixed-time survival comparisons, and endpoint-specific analysis populations.

ConceptHow it appears in PACIFIC
RandomizationRandomized, parallel-group phase 3 design.
MaskingQuadruple masking.
ITT analysisPrimary efficacy analyses use the FAS including all randomized patients, analyzed on an ITT basis.
Kaplan-Meier estimationUsed for PFS and OS.
Hazard ratioPrimary and several secondary time-to-event effects are reported as HRs.
Confidence intervalPrimary HRs have two-sided 95% CIs.
Stratified log-rank testUsed for the primary PFS and OS comparisons and several secondary time-to-event analyses.
Stratified Cox modelUsed to estimate HRs and CIs for the reported PRO deterioration analyses.
Breslow methodUsed for handling ties in the reported survival analyses.
Fisher exact testUsed for ORR, with a mid-P modification.
Wald / z-testUsed for the 24-month overall-survival percentage analysis.
Delta method / Greenwood's formulaUsed for variance estimation in the OS24 analysis.
Endpoint-specific populationsPRO analyses restrict eligibility according to baseline score thresholds.
Superiority testingThe registry identifies the primary hypotheses as superiority.

22. A Practical Reading of the PACIFIC Statistical Results

A useful way to read the primary results is to separate four questions that are often collapsed into one:

1. What was estimated?

The principal effect measure was a hazard ratio comparing MEDI4736 with placebo for time-to-event endpoints.

2. How precise was it?

Precision is communicated by the 95% confidence interval: 0.42–0.65 for PFS and 0.53–0.87 for OS.

3. What was the statistical evidence?

The primary P-values were <0.0001 for PFS and 0.00251 for OS under the reported testing framework.

4. What does it not establish?

These statistics do not give an individual patient's probability of benefit, do not provide an absolute risk difference, and do not replace clinical or safety interpretation.

The distinction is particularly important in survival analysis. A hazard ratio compresses a potentially complex pattern of event and censoring times into one relative measure. Kaplan-Meier estimates preserve more information about the survival experience, while fixed-time estimates such as OS24 provide an absolute perspective at a particular time point.

23. Primary Endpoint Comparison Without Overinterpretation

QuestionPFSOS
What is the event?Objective progression or death in the absence of progressionDeath due to any cause
Analysis typeTime-to-eventTime-to-event
HR0.520.68
95% CI0.42–0.650.53–0.87
P-value<0.00010.00251
Relative interpretationEstimated instantaneous progression/death hazard approximately 48% lowerEstimated instantaneous death hazard approximately 32% lower
Primary methodStratified log-rankStratified log-rank

The two estimates should not be compared as though an HR of 0.52 is inherently “better” than an HR of 0.68. They measure different event processes. The appropriate interpretation is that each endpoint provides its own estimate of the treatment comparison under its own event definition.

24. Related Tutorials

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue through Clinical Biostats

Connect the endpoints and statistical methods in this trial to deeper biostatistics tutorials and statistical calculators.

27. Record Summary

PACIFIC provides a useful example of a randomized phase 3 time-to-event analysis. The registry describes a randomized, parallel-group, quadruple-masked trial with 713 enrolled patients and two primary endpoints: progression-free survival based on BICR according to RECIST 1.1 and overall survival. Both primary analyses used the FAS, included all randomized patients, followed an ITT principle, and used stratified log-rank testing with age at randomization, sex, and smoking history as stratification factors and the Breslow approach for ties.

The primary PFS analysis reported an HR of 0.52 with a two-sided 95% CI of 0.42–0.65 and P < 0.0001. The primary OS analysis reported an HR of 0.68 with a two-sided 95% CI of 0.53–0.87 and P = 0.00251. Secondary analyses extended the time-to-event framework to TTDM, PFS2, global health status / HRQoL deterioration, and individual PRO symptoms, while ORR used Fisher's exact test with a mid-P modification and OS24 used a Wald / z-test with variance estimated using the delta method and Greenwood's formula.

The most important statistical lesson is that these results should be read as a collection of endpoint-specific estimates rather than as one number. Hazard ratios describe relative time-to-event effects; confidence intervals describe statistical precision; P-values address evidence against specified null hypotheses; Kaplan-Meier methods describe survival distributions; and endpoint-specific populations and censoring rules determine exactly which patients and observations contribute to each analysis.

Clinical Biostats methodology: This page separates reported registry results from statistical interpretation. Numbers are reproduced from the registry-reported PACIFIC trial data, while explanatory material describes what the reported statistical methods mean without adding unreported efficacy estimates or clinical conclusions.