This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View NCT02125461 on ClinicalTrials.gov.
1. Trial at a Glance
PACIFIC was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating MEDI4736 versus placebo in patients with stage III unresectable non-small cell lung cancer following concurrent chemoradiation. The registry reports 713 enrolled patients, two arms, two primary time-to-event endpoints, and formal statistical analyses for both primary endpoints.
| Feature | PACIFIC |
|---|---|
| Trial name | PACIFIC |
| Phase | Phase 3 |
| Condition | Non-Small Cell Lung Cancer |
| Population description | Patients with Stage III Unresectable Non-Small Cell Lung Cancer following concurrent chemoradiation |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 713 |
| Arms | 2 |
| Interventions | MEDI4736 (drug); Placebo (other) |
| Primary endpoints | 2 |
| Primary endpoint type | Time-to-event |
| Hypothesis type | Superiority |
| Lead sponsor | AstraZeneca |
| Sponsor type | Industry |
| Status | Completed |
| Trial start | 2014-05-07 |
| Primary completion | 2017-02-13 |
2. Clinical Question
The central statistical question was whether MEDI4736 following concurrent chemoradiation improved the time-to-event outcomes defined as progression-free survival and overall survival compared with placebo.
Population
Patients with Stage III Unresectable Non-Small Cell Lung Cancer following concurrent chemoradiation.
Intervention
MEDI4736.
Comparator
Placebo.
Primary question
Does MEDI4736 improve progression-free survival and overall survival relative to placebo?
The registry identifies the primary hypotheses as superiority. This is important because the statistical question is whether the MEDI4736 group differs favorably from the placebo group under the prespecified time-to-event framework, rather than whether the two groups are sufficiently similar to support an equivalence or non-inferiority conclusion.
3. Trial Design
MEDI4736
- Intervention classified in the registry as a drug.
- Compared with placebo in the randomized parallel-group design.
- Primary efficacy analyses were based on the FAS and analyzed on an ITT basis.
Placebo
- Comparator classified in the registry as an other intervention.
- Compared with MEDI4736 in the randomized parallel-group design.
- Primary efficacy analyses were based on the FAS and analyzed on an ITT basis.
The registry does not provide an allocation ratio in the ClinicalTrials.gov record. Accordingly, this page does not infer one from the enrollment total or the safety denominators.
4. Trial Timeline and Registry Status
Trial start
The registry lists 2014-05-07 as the study start date.
Primary completion
The registry lists 2017-02-13 as the primary completion date.
OS data cutoff
The registered overall-survival endpoint was assessed until the 22 Mar 2018 data cutoff, with a maximum of approximately 4 years.
Registry status
The trial is listed as completed. The registry caveat states that interim PFS and interim OS analyses are considered the final PFS and OS analyses, respectively.
5. Primary Endpoints
| Endpoint | Registry definition | Time frame | Primary analysis |
|---|---|---|---|
| Progression Free Survival Based on Blinded Independent Central Review (BICR) According to Response Evaluation Criteria in Solid Tumors (RECIST 1.1) | PFS was defined as the time from randomization until the date of objective disease progression (RECIST 1.1) or death (by any cause in the absence of progression). Progression was defined using RECIST 1.1 as a 20% increase in the sum of the longest diameter of target lesions, or a measurable increase in a non-target lesion, or the appearance of new lesions. PFS was calculated using the Kaplan-Meier technique. | Tumor scans performed at baseline then every ~8 weeks up to 48 weeks, then every ~12 weeks thereafter until confirmed disease progression. Assessed until 13 Feb 2017 DCO; up to a maximum of approximately 3 years. | Stratified log-rank test; hazard ratio |
| Overall Survival | OS was defined as the time from the date of randomization until death due to any cause. OS was calculated using the Kaplan-Meier technique. | From baseline until death due to any cause. Assessed until 22 Mar 2018 DCO; up to a maximum of approximately 4 years. | Stratified log-rank test; hazard ratio |
Both primary endpoints are time-to-event outcomes. This means the analysis must account not only for whether an event occurred, but also for the amount of follow-up available for each randomized patient and for patients who remain event-free at the time of analysis.
6. Analysis Population and Stratification
For both primary endpoints, the registry states that the FAS included all randomized patients and that patients were analyzed on an intent-to-treat (ITT) basis. The reported primary analyses used a stratified log-rank test.
| Feature | Registry-supported approach |
|---|---|
| Analysis population | FAS including all randomized patients |
| Analysis principle | Intent-to-treat |
| Groups compared | Durvalumab (MEDI4736) vs Placebo |
| Primary comparison | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Confidence interval | 95%, two-sided |
| Hypothesis | Superiority |
| Tie handling | Breslow approach |
| Stratification factors | Age at randomization (<65 vs ≥65), sex (male vs female), and smoking history (smoker vs non-smoker) |
The same age, sex, and smoking-history factors were used in the stratified primary analyses. This matters because the treatment comparison is not simply an unadjusted comparison of two survival curves: the analysis explicitly incorporates these factors into the time-to-event comparison.
7. Statistical Methodology
Kaplan-Meier estimation
The registry states that PFS and OS were calculated using the Kaplan-Meier technique. Kaplan-Meier estimation is appropriate for time-to-event data because it can incorporate right-censored observations. A patient who has not experienced progression or death by the last available assessment does not simply disappear from the analysis; the patient's observed follow-up contributes information up to the censoring time.
Here, di is the number of events at event time ti, while ni is the number at risk immediately before that time. The resulting survival function estimates the probability of remaining event-free beyond time t.
Stratified log-rank testing
The primary PFS and OS comparisons used a stratified log-rank test. Rather than treating all randomized patients as belonging to one homogeneous risk set, the analysis adjusted the comparison through strata defined by age at randomization, sex, and smoking history.
Conceptually, the log-rank test asks whether the observed pattern of events over follow-up differs between treatment groups under the null hypothesis of no difference in the survival experience. The stratified version performs this comparison while accounting for the specified stratification factors.
Hazard ratio
The principal effect measure reported for the primary analyses was the hazard ratio (HR). An HR compares the estimated instantaneous event rates between groups over the analyzed follow-up. An HR below 1 indicates a lower estimated instantaneous event rate in the MEDI4736 group relative to placebo under the fitted analysis.
The HR is not a probability, not a percentage of patients who benefit, and not the same quantity as an absolute risk difference. Its interpretation is tied to the time-to-event model and its assumptions.
Covariate adjustment and stratification
The primary analysis notes explicitly identify covariate adjustment and stratified analysis. The stratification factors were age at randomization, sex, and smoking history. Ties were handled using the Breslow approach.
Confidence intervals
Each primary endpoint has a two-sided 95% confidence interval for the hazard ratio. The interval provides information about statistical precision around the estimated treatment effect. It is not a range containing 95% of individual patient effects, and it does not mean that there is a 95% probability that the true hazard ratio lies inside this particular calculated interval.
Intent-to-treat analysis
The FAS included all randomized patients and the analyses were performed on an ITT basis. The key principle is that randomized patients remain associated with the treatment group to which they were assigned for the efficacy comparison. This preserves the treatment contrast created by randomization and reduces the risk that post-randomization treatment behavior determines the primary comparison population.
8. Primary Result: Progression-Free Survival
The registry reports a formal primary analysis for progression-free survival based on BICR according to RECIST 1.1. The comparison was between MEDI4736 and placebo in the FAS, analyzed on an ITT basis.
Progression-free survival hazard ratio
95% CI: 0.42–0.65 · P < 0.0001
Stratified log-rank analysis; superiority hypothesis.
| Primary PFS analysis feature | Reported result |
|---|---|
| Groups compared | Durvalumab (MEDI4736) vs Placebo |
| Analysis population | FAS; all randomized patients; ITT |
| Endpoint type | Time-to-event |
| Effect measure | Hazard ratio |
| Estimate | 0.52 |
| 95% CI | 0.42–0.65 |
| P-value | <0.0001 |
| Method | Stratified log-rank test |
| Stratification | Age at randomization (<65 vs ≥65), sex (male vs female), smoking history (smoker vs non-smoker) |
| Ties | Breslow approach |
An HR of 0.52 means that, under the reported time-to-event analysis, the estimated instantaneous rate of progression or death in the MEDI4736 group was approximately 52% of that in the placebo group. Equivalently, 1 − 0.52 = 0.48, so the estimated hazard was approximately 48% lower with MEDI4736 under this model.
This does not mean that 48% of patients avoided progression, that every patient experienced a 48% reduction in risk, or that the absolute probability of progression or death was reduced by 48 percentage points. The hazard ratio is a relative time-to-event measure.
The 95% CI of 0.42–0.65 describes the statistical uncertainty around the estimated HR. Its width also shows why reporting the interval is important: the point estimate alone does not communicate the precision of the treatment-effect estimate.
The P-value of <0.0001 addresses evidence against the null hypothesis within the specified testing framework. It does not measure the magnitude of the treatment effect. Effect magnitude is described by the HR and its confidence interval.
The analysis is also dependent on censoring rules and the assumptions underlying the time-to-event framework. In particular, a hazard ratio is not automatically equivalent to a constant relative risk over the entire follow-up period; proportional-hazards assumptions should be considered when interpreting a single HR as a summary of treatment effect.
9. Primary Result: Overall Survival
Overall survival was the second primary endpoint. The registry defines OS as the time from randomization until death due to any cause and states that OS was calculated using the Kaplan-Meier technique.
Overall survival hazard ratio
95% CI: 0.53–0.87 · P = 0.00251
Stratified log-rank analysis; superiority hypothesis.
| Primary OS analysis feature | Reported result |
|---|---|
| Groups compared | Durvalumab (MEDI4736) vs Placebo |
| Analysis population | FAS; all randomized patients; ITT |
| Endpoint type | Time-to-event |
| Effect measure | Hazard ratio |
| Estimate | 0.68 |
| 95% CI | 0.53–0.87 |
| P-value | 0.00251 |
| Method | Stratified log-rank test |
| Stratification | Age at randomization (<65 vs ≥65), sex (male vs female), smoking history (smoker vs non-smoker) |
| Ties | Breslow approach |
| Data cutoff | 22 Mar 2018 |
An HR of 0.68 indicates that, under the reported time-to-event model, the estimated instantaneous rate of death in the MEDI4736 group was approximately 68% of that in the placebo group. Equivalently, 1 − 0.68 = 0.32, corresponding to an estimated 32% lower instantaneous hazard of death under the model.
This is not the same as saying that mortality was reduced by 32 percentage points, that 32% of patients were saved, or that every individual patient experienced the same relative reduction. The HR summarizes a randomized group comparison over the analyzed follow-up.
The 95% CI of 0.53–0.87 gives the statistical uncertainty around the estimated HR. Because the interval is not centered on the null value of 1, the interval is consistent with a treatment effect below the null under the stated two-sided confidence-interval framework.
The P-value of 0.00251 measures the compatibility of the observed result with the relevant null hypothesis under the analysis framework. It does not tell us how large the treatment effect is; the HR and CI provide that information.
As with PFS, interpretation of a single HR requires attention to censoring and the proportional-hazards model used for the treatment effect. The registry's use of stratification also means that the reported comparison should not be interpreted as a simple unadjusted ratio calculated from two crude event proportions.
10. Primary Endpoint Results Together
The two primary endpoints give complementary views of the randomized treatment comparison. PFS captures the time to objective disease progression or death, whereas OS captures time to death from any cause.
| Primary endpoint | HR | 95% CI | P-value | Analysis |
|---|---|---|---|---|
| Progression Free Survival based on BICR according to RECIST 1.1 | 0.52 | 0.42–0.65 | <0.0001 | Stratified log-rank |
| Overall Survival | 0.68 | 0.53–0.87 | 0.00251 | Stratified log-rank |
Both reported HR estimates are below 1, and both confidence intervals exclude 1. The registry classifies both analyses as superiority analyses. The PFS and OS results should nevertheless be understood as separate endpoints with different clinical meanings rather than as two measurements of exactly the same outcome.
PFS asks
How does the time from randomization to objective progression or death compare between the randomized groups?
OS asks
How does the time from randomization to death from any cause compare between the randomized groups?
11. Secondary Time-to-Event Results
The registry contains formal statistical analyses for several secondary time-to-event endpoints. These analyses use the same broad framework of FAS/ITT analysis, comparison of MEDI4736 with placebo, and stratified time-to-event methods unless otherwise noted in the endpoint-specific analysis text.
| Secondary endpoint | HR | 95% CI | P-value | Method |
|---|---|---|---|---|
| Time to Death or Distant Metastasis (TTDM) based on BICR assessments according to RECIST 1.1 | 0.53 | 0.41–0.68 | <0.0001 | Stratified log-rank |
| Time to Second Progression or Death (PFS2) | 0.58 | 0.46–0.73 | <0.0001 | Stratified log-rank |
| Time to Deterioration of Global Health Status / HRQoL using EORTC QLQ-C30 | 0.95 | 0.77–1.18 | 0.664 | Stratified log-rank |
| Time to deterioration of PRO symptom: dyspnea | 1.06 | 0.88–1.29 | 0.522 | Stratified log-rank |
| Time to deterioration of PRO symptom: cough | 0.91 | 0.74–1.12 | 0.380 | Stratified log-rank |
| Time to deterioration of PRO symptom: hemoptysis | 0.75 | 0.56–1.00 | 0.048 | Stratified log-rank |
| Time to deterioration of PRO symptom: chest pain | 0.94 | 0.75–1.19 | 0.626 | Stratified log-rank |
These secondary endpoints illustrate why a trial-results page should distinguish the direction and size of an estimate from its statistical significance. The HRs range from 0.53 to 1.06, and the corresponding confidence intervals range from relatively precise intervals well below 1 to intervals that include 1.
Time to Death or Distant Metastasis
TTDM hazard ratio
95% CI: 0.41–0.68 · P < 0.0001
Stratified log-rank analysis.
The TTDM analysis estimates the relative time-to-event experience for death or distant metastasis. The HR of 0.53 corresponds to an estimated 47% lower instantaneous event rate under the reported model. This interpretation concerns the composite event definition itself; it should not be interpreted as a 47% reduction in each component considered separately.
Time to Second Progression or Death
PFS2 hazard ratio
95% CI: 0.46–0.73 · P < 0.0001
Stratified log-rank analysis.
PFS2 extends the time-to-event framework beyond the first progression. The registry states that, following confirmed progression, patients were assessed every ~12 weeks until second disease progression. The HR of 0.58 corresponds to an estimated 42% lower instantaneous rate of the PFS2 event under the reported model.
Global Health Status / HRQoL
| Measure | HR | 95% CI | P-value |
|---|---|---|---|
| Time to deterioration of global health status / HRQoL, EORTC QLQ-C30 | 0.95 | 0.77–1.18 | 0.664 |
| Dyspnea | 1.06 | 0.88–1.29 | 0.522 |
| Cough | 0.91 | 0.74–1.12 | 0.380 |
| Hemoptysis | 0.75 | 0.56–1.00 | 0.048 |
| Chest pain | 0.94 | 0.75–1.19 | 0.626 |
For the global-health-status/HRQoL analysis, only patients with baseline scores ≥ 10 were included. For the listed EORTC QLQ-LC13 symptom analyses, only patients with baseline scores ≤ 90 were included. These eligibility rules matter because the denominator for these analyses is therefore defined by the endpoint-specific baseline-score requirement rather than simply by all 713 enrolled patients.
12. Objective Response Rate
Objective Response Rate (ORR) based on BICR assessments according to RECIST 1.1 was a secondary binary endpoint. The analysis population was the FAS, which included all randomized patients, with measurable disease at baseline, analyzed on an ITT basis.
Fisher exact comparison
Secondary endpoint: Objective Response Rate
Fisher's exact test with mid P-value modification.
| Feature | Reported analysis |
|---|---|
| Endpoint | Objective Response Rate based on BICR assessments according to RECIST 1.1 |
| Endpoint type | Binary |
| Outcome unit | Percentage of patients |
| Population | FAS with measurable disease at baseline; ITT |
| Groups compared | Durvalumab (MEDI4736) vs Placebo |
| Method | Fisher exact test |
| P-value | <0.001 |
| Modification | Mid P-value modification by subtracting half of the probability of the observed table from Fisher's P-value |
The ClinicalTrials.gov record does not provide the response percentages, response counts, or a confidence interval for ORR. Therefore, this page reports the formal comparison that is available without reconstructing an unreported effect estimate.
Why Fisher's exact test is appropriate here
ORR reduces the response assessment to a binary outcome for each evaluable patient: the statistical analysis compares the resulting treatment-group response classifications. Fisher's exact test is designed for categorical contingency tables and calculates the exact probability under the specified null framework rather than relying on a large-sample approximation.
The registry used a mid P-value modification. This subtracts half the probability of the observed table from Fisher's P-value. The resulting value should be identified as a mid-P result rather than silently treating it as an ordinary unmodified Fisher exact P-value.
13. Percentage of Patients Alive at 24 Months
The registry also reports a secondary analysis of the Percentage of Patients Alive at 24 Months (OS24). This is a binary/time-point summary rather than a continuous measurement of survival time.
OS24 comparison
Secondary endpoint: Percentage of Patients Alive at 24 Months.
Wald / z-test; variance estimated using the delta method and Greenwood's formula.
| Feature | Reported analysis |
|---|---|
| Endpoint | Percentage of Patients Alive at 24 Months (OS24) |
| Time frame | From baseline until death due to any cause. Assessed until 22 Mar 2018 DCO; up to a maximum of approximately 4 years. |
| Endpoint type | Binary |
| Outcome unit | Percentage of patients |
| Population | FAS including all randomized patients; ITT |
| Groups compared | Durvalumab (MEDI4736) vs Placebo |
| Method | Wald / z-test |
| P-value | 0.005 |
| Variance estimation | Delta method and Greenwood's formula |
The ClinicalTrials.gov record does not provide the actual 24-month survival percentages by treatment group. Accordingly, the P-value is reported without inventing the corresponding group estimates.
Statistically, this endpoint is useful because it provides a fixed time-point summary, while the primary OS analysis uses the full time-to-event information. A fixed-time survival percentage and a hazard ratio are therefore complementary rather than interchangeable summaries.
14. Statistical Methods Explained
Why was a stratified log-rank test used?
The primary endpoints were time-to-event outcomes, so the treatment groups needed to be compared over the entire follow-up rather than using a simple comparison of event proportions. The registry reports a stratified log-rank test that accounts for age at randomization, sex, and smoking history. Stratification can reduce the influence of these specified factors on the treatment comparison while preserving the randomized-group framework.
What does an HR of 0.52 mean for PFS?
An HR of 0.52 means that the estimated instantaneous rate of progression or death in the MEDI4736 group was 52% of that in the placebo group under the reported analysis. The corresponding 48% figure is a relative reduction in estimated hazard, not a 48-percentage-point reduction in the probability of progression or death.
Why is the confidence interval important?
The point estimate is only one summary of the data. The 95% CI of 0.42–0.65 for PFS and 0.53–0.87 for OS shows the statistical precision surrounding the respective HR estimates. A narrower interval generally conveys more precision than a wider interval, while the location of the interval relative to the null value of 1 informs the uncertainty about the direction of the relative hazard.
Why doesn't the P-value measure effect size?
A P-value quantifies how compatible the observed data are with a specified null hypothesis under the statistical model and testing procedure. It is affected by both the magnitude of an observed effect and the amount of information in the analysis. The HR describes relative effect magnitude, while the confidence interval describes uncertainty around that estimate. These should not be replaced by the P-value alone.
Why is the Breslow method mentioned?
In time-to-event analyses, multiple events can occur at the same recorded time. These are tied event times. The registry states that ties were handled using the Breslow approach. Identifying the tie-handling method is part of describing the actual statistical model rather than assuming a default implementation.
Why does the ITT population matter?
The registry defines the FAS as including all randomized patients and states that the primary analyses were conducted on an ITT basis. This keeps the efficacy comparison aligned with randomized assignment. Removing patients after randomization because of subsequent treatment behavior can compromise the balance produced by randomization and change the clinical question being answered.
Why are the HRQoL analyses based on restricted baseline scores?
The HRQoL analysis included only patients with baseline scores ≥ 10, while the listed PRO symptom analyses included only patients with baseline scores ≤ 90. These rules define the relevant analysis populations for those endpoints. They also mean that the results should not be interpreted as though every randomized patient necessarily contributed to every PRO analysis.
15. Primary Analysis Interpretation: What the Numbers Do and Do Not Mean
The estimated instantaneous rate of progression or death was approximately 48% lower in the MEDI4736 group under the reported model. This does not mean that 48% of patients were protected from progression, nor does it specify the absolute probability of being progression-free at any particular time.
The estimated instantaneous rate of death was approximately 32% lower in the MEDI4736 group under the reported model. This does not mean that the absolute mortality probability was reduced by 32 percentage points or that every patient experienced the same relative reduction.
The PFS 95% CI of 0.42–0.65 and OS 95% CI of 0.53–0.87 quantify statistical uncertainty around the respective hazard-ratio estimates. They do not describe the range of outcomes that individual patients might experience.
The PFS P-value of <0.0001 and OS P-value of 0.00251 provide evidence against the corresponding null hypotheses under the reported testing framework. They do not rank the size or clinical importance of the two treatment effects.
16. Time-to-Event Endpoints and Censoring
PFS and OS are fundamentally different from ordinary binary endpoints because patients can have different amounts of observed follow-up. Some patients experience the event during follow-up, while others remain event-free at their last assessment. Those latter observations are censored rather than treated as if the event never occurred.
For PFS, the event is objective disease progression or death in the absence of progression. For OS, the event is death due to any cause. The registry therefore defines different event processes even though both endpoints use the same broad Kaplan-Meier and stratified survival-analysis framework.
PFS event
Objective disease progression according to RECIST 1.1 or death by any cause in the absence of progression.
OS event
Death due to any cause.
17. Safety: Serious Adverse Events by Arm
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected patients over patients at risk. This provides a direct arm-specific safety count, although the ClinicalTrials.gov record does not include a formal statistical comparison or a complete adverse-event profile.
| Safety measure | MEDI4736 | Placebo |
|---|---|---|
| Serious adverse events, affected / at risk | 138 / 475 | 54 / 234 |
The denominators are 475 and 234, respectively. The ClinicalTrials.gov record does not provide a formal P-value, confidence interval, risk ratio, or odds ratio for serious adverse events. Such measures should therefore not be inferred from the affected/at-risk counts when the task is to reproduce reported trial results exactly.
Safety and efficacy answer different questions. A lower or higher event rate for a safety endpoint does not mathematically cancel or validate a time-to-event efficacy estimate. The two evidence streams should be presented separately.
18. Secondary Endpoint Analysis Methods
| Endpoint family | Endpoint type | Method | Effect measure / output |
|---|---|---|---|
| PFS | Time-to-event | Stratified log-rank | Hazard ratio |
| OS | Time-to-event | Stratified log-rank | Hazard ratio |
| TTDM | Time-to-event | Stratified log-rank | Hazard ratio |
| PFS2 | Time-to-event | Stratified log-rank | Hazard ratio |
| Global health status / HRQoL deterioration | Time-to-event | Stratified log-rank | Hazard ratio |
| PRO symptom deterioration | Time-to-event | Stratified log-rank | Hazard ratio |
| ORR | Binary | Fisher exact test | P-value |
| OS24 | Binary | Wald / z-test | P-value |
This distribution of methods is statistically coherent with the endpoint types. Time-to-event outcomes require methods that account for follow-up and censoring; binary response outcomes can be analyzed using contingency-table methods; and a fixed-time survival percentage can be compared using a variance-based test, as reported here.
19. Stratified Cox Interpretation for Secondary PRO Analyses
For the global health status / HRQoL endpoint and the listed PRO symptom endpoints, the registry specifies that the HR and CI were estimated from a stratified Cox proportional hazards model. The Breslow method was used to control for ties, and the strata statement included age at randomization (<65 vs ≥65), sex (male vs female), and smoking history (smoker vs non-smoker). The confidence interval was calculated using a profile likelihood approach.
| Endpoint | HR estimation details |
|---|---|
| Global health status / HRQoL deterioration | Stratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI. |
| Dyspnea deterioration | Stratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI. |
| Cough deterioration | Stratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI. |
| Hemoptysis deterioration | Stratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI. |
| Chest pain deterioration | Stratified Cox proportional hazards model; Breslow method for ties; age, sex, and smoking history in strata statement; profile likelihood CI. |
This is an important methodological distinction: the registry method field lists the comparison as a log-rank analysis, while the endpoint-specific analysis text describes the model used to estimate the HR and confidence interval. A complete statistical description should preserve both pieces of information rather than reducing the analysis to a single method label.
20. Limitations and Interpretation Issues
- Registry-level detail: the ClinicalTrials.gov record provides summary estimates and statistical methods, but not every element of the underlying statistical analysis plan.
- Interim analyses treated as final: the registry explicitly states that the interim PFS and OS analyses are considered the final PFS and OS analyses.
- No baseline table reported: the ClinicalTrials.gov record does not contain baseline demographic or disease-characteristic results, so none are reproduced here.
- No subgroup estimates reported: the ClinicalTrials.gov record does not contain subgroup HRs or interaction tests. Therefore, no subgroup conclusions are presented.
- No median PFS or OS reported: the ClinicalTrials.gov record provides HRs, confidence intervals, and P-values but no median survival estimates. Medians are therefore not inferred.
- Hazard-ratio assumptions: a single HR is a model-based relative measure. Its interpretation depends on the time-to-event modeling framework and should not automatically be read as a constant risk ratio at every time point.
- Multiple endpoints: the trial contains two primary endpoints and numerous secondary analyses. Individual secondary P-values should be interpreted in the context of the overall testing strategy; the ClinicalTrials.gov record does not specify a complete multiplicity-adjustment hierarchy.
- Endpoint-specific populations: the PRO analyses impose baseline-score eligibility criteria, so their populations are narrower than the all-randomized FAS.
- Safety denominator: serious adverse-event counts are reported as affected/at-risk counts, but the ClinicalTrials.gov record does not provide a formal comparative safety analysis.
- Long-term follow-up: the registry states that patients were followed for long-term survival until approximately 5 years after the last patient enrolled, but the ClinicalTrials.gov record does not provide the corresponding long-term survival estimates.
21. Why This Trial Matters Statistically
PACIFIC is a useful statistical teaching case because the ClinicalTrials.gov record bring together several common methods in modern randomized clinical-trial analysis: ITT efficacy analysis, Kaplan-Meier estimation, stratified log-rank testing, hazard ratios, Cox proportional-hazards modeling, exact categorical testing, fixed-time survival comparisons, and endpoint-specific analysis populations.
| Concept | How it appears in PACIFIC |
|---|---|
| Randomization | Randomized, parallel-group phase 3 design. |
| Masking | Quadruple masking. |
| ITT analysis | Primary efficacy analyses use the FAS including all randomized patients, analyzed on an ITT basis. |
| Kaplan-Meier estimation | Used for PFS and OS. |
| Hazard ratio | Primary and several secondary time-to-event effects are reported as HRs. |
| Confidence interval | Primary HRs have two-sided 95% CIs. |
| Stratified log-rank test | Used for the primary PFS and OS comparisons and several secondary time-to-event analyses. |
| Stratified Cox model | Used to estimate HRs and CIs for the reported PRO deterioration analyses. |
| Breslow method | Used for handling ties in the reported survival analyses. |
| Fisher exact test | Used for ORR, with a mid-P modification. |
| Wald / z-test | Used for the 24-month overall-survival percentage analysis. |
| Delta method / Greenwood's formula | Used for variance estimation in the OS24 analysis. |
| Endpoint-specific populations | PRO analyses restrict eligibility according to baseline score thresholds. |
| Superiority testing | The registry identifies the primary hypotheses as superiority. |
22. A Practical Reading of the PACIFIC Statistical Results
A useful way to read the primary results is to separate four questions that are often collapsed into one:
1. What was estimated?
The principal effect measure was a hazard ratio comparing MEDI4736 with placebo for time-to-event endpoints.
2. How precise was it?
Precision is communicated by the 95% confidence interval: 0.42–0.65 for PFS and 0.53–0.87 for OS.
3. What was the statistical evidence?
The primary P-values were <0.0001 for PFS and 0.00251 for OS under the reported testing framework.
4. What does it not establish?
These statistics do not give an individual patient's probability of benefit, do not provide an absolute risk difference, and do not replace clinical or safety interpretation.
The distinction is particularly important in survival analysis. A hazard ratio compresses a potentially complex pattern of event and censoring times into one relative measure. Kaplan-Meier estimates preserve more information about the survival experience, while fixed-time estimates such as OS24 provide an absolute perspective at a particular time point.
23. Primary Endpoint Comparison Without Overinterpretation
| Question | PFS | OS |
|---|---|---|
| What is the event? | Objective progression or death in the absence of progression | Death due to any cause |
| Analysis type | Time-to-event | Time-to-event |
| HR | 0.52 | 0.68 |
| 95% CI | 0.42–0.65 | 0.53–0.87 |
| P-value | <0.0001 | 0.00251 |
| Relative interpretation | Estimated instantaneous progression/death hazard approximately 48% lower | Estimated instantaneous death hazard approximately 32% lower |
| Primary method | Stratified log-rank | Stratified log-rank |
The two estimates should not be compared as though an HR of 0.52 is inherently “better” than an HR of 0.68. They measure different event processes. The appropriate interpretation is that each endpoint provides its own estimate of the treatment comparison under its own event definition.
24. Related Tutorials
Learn more about the methods used in this trial:
25. Related Statistical Calculators
26. Sources
- ClinicalTrials.gov: NCT02125461 — PACIFIC.
- PubMed: PMID 42398520.
- PubMed: PMID 36841540.
- PubMed: PMID 35247871.
- PubMed: PMID 35245844.
- PubMed: PMID 35108059.
Continue through Clinical Biostats
Connect the endpoints and statistical methods in this trial to deeper biostatistics tutorials and statistical calculators.
27. Record Summary
PACIFIC provides a useful example of a randomized phase 3 time-to-event analysis. The registry describes a randomized, parallel-group, quadruple-masked trial with 713 enrolled patients and two primary endpoints: progression-free survival based on BICR according to RECIST 1.1 and overall survival. Both primary analyses used the FAS, included all randomized patients, followed an ITT principle, and used stratified log-rank testing with age at randomization, sex, and smoking history as stratification factors and the Breslow approach for ties.
The primary PFS analysis reported an HR of 0.52 with a two-sided 95% CI of 0.42–0.65 and P < 0.0001. The primary OS analysis reported an HR of 0.68 with a two-sided 95% CI of 0.53–0.87 and P = 0.00251. Secondary analyses extended the time-to-event framework to TTDM, PFS2, global health status / HRQoL deterioration, and individual PRO symptoms, while ORR used Fisher's exact test with a mid-P modification and OS24 used a Wald / z-test with variance estimated using the delta method and Greenwood's formula.
The most important statistical lesson is that these results should be read as a collection of endpoint-specific estimates rather than as one number. Hazard ratios describe relative time-to-event effects; confidence intervals describe statistical precision; P-values address evidence against specified null hypotheses; Kaplan-Meier methods describe survival distributions; and endpoint-specific populations and censoring rules determine exactly which patients and observations contribute to each analysis.