← Clinical Trials
Non-Small Cell Lung Cancer Phase 3 Active, Not Recruiting NCT03164616

POSEIDON: Complete Statistical Analysis of Durvalumab-Based Therapy in Non-Small Cell Lung Cancer

An independent statistical review of the randomized phase 3 POSEIDON trial comparing durvalumab plus tremelimumab with chemotherapy, durvalumab with chemotherapy, and chemotherapy alone in patients with non-small cell lung cancer.

Trial start: June 1, 2017  ·  Primary completion: March 12, 2021  ·  Enrollment: 1,186
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record for NCT03164616 and the statistical analyses posted there.

Registry context: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

POSEIDON was a randomized, open-label, parallel-group phase 3 trial evaluating durvalumab plus tremelimumab with chemotherapy, durvalumab with chemotherapy, or chemotherapy alone in patients with non-small cell lung cancer. The registry reports 1,186 enrolled participants across three arms.

1,186
Enrolled
Global trial record
3
Arms
Parallel design
0.74
Primary PFS HR
95% CI 0.620–0.885
0.86
Primary OS HR
95% CI 0.724–1.016
FeaturePOSEIDON
PhasePhase 3
ConditionNon Small Cell Lung Cancer NSCLC
DesignRandomized, parallel-group, unmasked
AllocationRandomized
Primary purposeTreatment
Enrollment1,186
Arms3
Primary endpointsProgression-Free Survival (PFS) and Overall Survival (OS)
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted14
Statistical analyses posted8
Lead sponsorAstraZeneca
Sponsor typeIndustry
ClinicalTrials.govNCT03164616

2. Clinical Question

The central statistical question was whether treatment strategies incorporating durvalumab, with or without tremelimumab, and chemotherapy could improve time-to-event outcomes compared with chemotherapy alone. The registry's two primary comparisons were specifically durvalumab plus standard of care (D + SoC) versus SoC alone for PFS and OS.

Population

Patients with non-small cell lung cancer enrolled in the phase 3 POSEIDON trial.

Intervention strategy

Durvalumab plus chemotherapy, with tremelimumab plus durvalumab plus chemotherapy forming a separate randomized arm.

Comparator

Standard-of-care chemotherapy alone, represented in the registry as SoC Alone.

Primary question

Does D + SoC improve PFS and OS compared with SoC alone under the prespecified superiority framework?

3. Trial Design

01
Randomize 1,186 enrolled
02
Three arms D + SoC; T + D + SoC; SoC Alone
03
Treatment Randomized treatment assignment
04
Assessment PFS, OS, response, and safety
05
Analysis Survival and categorical methods
ARM 1

T + D + SoC

  • Tremelimumab
  • Durvalumab
  • Standard-of-care chemotherapy
ARM 2

D + SoC

  • Durvalumab
  • Standard-of-care chemotherapy
ARM 3

SoC Alone

  • Standard-of-care chemotherapy

The registry identifies the allocation as RANDOMIZED, the design model as PARALLEL, and masking as NONE. Thus, the treatment comparison is based on randomized assignment, but the trial was not described as blinded.

Important design distinction: three randomized treatment arms do not mean that every possible pairwise comparison is automatically a primary confirmatory comparison. The posted primary analyses specifically evaluate D + SoC versus SoC Alone for PFS and OS. The registry also reports secondary comparisons involving T + D + SoC.

4. Endpoints

EndpointRegistry definitionTime frame
Progression-Free Survival (PFS); D + SoC Compared With SoC Alone PFS (per RECIST version 1.1 [RECIST 1.1] using Blinded Independent Central Review [BICR] assessments) was defined as time from date of randomization until date of objective disease progression or death (by any cause in the absence of progression), regardless of whether the patient withdrew from randomized therapy or received another anticancer therapy prior to progression. Tumor scans performed at baseline, Week 6, Week 12 and then every 8 weeks relative to date of randomization until radiological progression. Assessed until global cohort DCO of 24 July 2019 (maximum of approximately 25 months).
Overall Survival (OS); D + SoC Compared With SoC Alone OS was defined as the time from the date of randomization until death due to any cause. Any patient not known to have died at the time of analysis was censored based on the last recorded date on which the patient was known to be alive. From baseline until death due to any cause. Assessed until global cohort DCO of 12 March 2021 (maximum of approximately 45 months).

The two primary endpoints are both time-to-event outcomes, but their events are different. PFS counts objective disease progression or death in the absence of progression, whereas OS counts death from any cause. That distinction is fundamental when interpreting the two hazard ratios.

5. Analysis Populations and Covariate Adjustment

The posted primary analyses state that the global-cohort full analysis set (FAS) included all randomized patients. PFS and OS were analyzed as primary outcome measures for the D + SoC versus SoC Alone comparison.

Analysis featureRegistry information
Primary efficacy populationGlobal cohort FAS; all randomized patients
PFS primary comparisonD + SoC vs SoC Alone
OS primary comparisonD + SoC vs SoC Alone
Time-to-event modelStratified Cox proportional-hazards model
Covariate adjustmentPD-L1, histology, and disease stage
Stratification variables in modelPD-L1 (TC ≥50% vs <50%), histology (squamous vs non-squamous), disease stage (IVA vs IVB)
TiesEfron method

This is an important feature of the analysis. The reported hazard ratios are not simple unadjusted ratios of event rates. They come from a stratified Cox model that accounts for specified baseline factors. The resulting HR therefore represents a model-based relative comparison conditional on the structure of that analysis.

6. Statistical Methodology

Log-rank testing for time-to-event outcomes

The registry reports the log-rank test for the primary PFS and OS analyses and for the secondary PFS and OS comparisons involving the T + D + SoC arm. The log-rank test compares survival experience between groups across the observed follow-up rather than comparing a single time point.

Stratified Cox proportional-hazards model

The hazard ratios and confidence intervals were estimated from a stratified Cox proportional-hazards model. The model adjusted for PD-L1 tumor expression, histology, and disease stage. The Efron method was used to handle tied event times.

Conceptual hazard-ratio interpretation
HR = estimated hazard in treatment group relative to comparator

For these analyses, an HR below 1 was specified in the registry as favoring the treatment group being associated with a longer time to the event. The HR is a relative time-to-event measure; it is not an absolute survival probability or a proportion of patients benefiting.

Logistic regression for objective response

Objective Response Rate (ORR) is a binary endpoint, so the registry analysis used logistic regression. The analysis adjusted for PD-L1 tumor expression, histology, and disease stage. The confidence interval was calculated using a profile likelihood approach.

Conceptual odds-ratio interpretation
OR = odds of response in treatment group ÷ odds of response in comparator group

An OR above 1 favors the treatment group in these analyses. The odds ratio should not be read as a risk ratio, and an OR of 1.90 does not mean that the response probability is 90 percentage points higher.

Kaplan-Meier estimation

The OS registry definition states that median OS was calculated using the Kaplan-Meier technique. Kaplan-Meier estimation is designed for right-censored time-to-event data: patients contribute follow-up until the event occurs or until their available observation ends without a documented event.

Kaplan-Meier survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at a given event time and ni represents the number at risk immediately before that time.

7. Primary Results: Progression-Free Survival

The registry reports a formal primary analysis of PFS for D + SoC versus SoC Alone. The analysis used the log-rank test, with the hazard ratio and confidence interval estimated from a stratified Cox proportional-hazards model.

Primary PFS hazard ratio

0.74

95% CI: 0.620–0.885   ·   P = 0.00093

Superiority analysis: D + SoC vs SoC Alone

MeasureReported result
EndpointProgression-Free Survival (PFS)
ComparisonD + SoC vs SoC Alone
MethodLog-rank test
Effect measureHazard ratio
HR0.74
95% CI0.620–0.885
P-value0.00093
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 0.74 means that, under the fitted proportional-hazards model, the estimated instantaneous rate of progression or death was approximately 26% lower with D + SoC than with SoC Alone. This is a relative hazard interpretation, not a statement that 26% of patients were protected from progression or death.

The 95% CI of 0.620–0.885 describes the statistical uncertainty around the estimated hazard ratio under the analysis framework. Because the entire interval is below 1, the interval is consistent with a lower estimated hazard for D + SoC relative to SoC Alone.

The P-value of 0.00093 addresses the statistical evidence against the null hypothesis specified for the comparison; it does not measure the magnitude or clinical importance of the treatment effect. The effect magnitude is conveyed by the HR, while the CI conveys precision.

The analysis was stratified and adjusted for PD-L1 tumor expression, histology, and disease stage. Interpretation also depends on the Cox model framework and its proportional-hazards assumption. The HR should therefore not be treated as a literal constant risk ratio applying identically at every time point.

8. Primary Results: Overall Survival

The second primary endpoint was OS, defined as time from randomization to death from any cause. Patients not known to have died at analysis were censored at the last recorded date they were known to be alive.

Primary OS hazard ratio

0.86

95% CI: 0.724–1.016   ·   P = 0.07581

Superiority analysis: D + SoC vs SoC Alone

MeasureReported result
EndpointOverall Survival (OS)
ComparisonD + SoC vs SoC Alone
MethodLog-rank test
Effect measureHazard ratio
HR0.86
95% CI0.724–1.016
P-value0.07581
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 0.86 corresponds to an estimated instantaneous death rate approximately 14% lower for D + SoC than for SoC Alone under the fitted model. That statement describes the estimated relative hazard; it does not mean that 14% of patients avoided death or that individual patients experienced a 14% reduction in their probability of dying.

The 95% CI of 0.724–1.016 is wider in the sense that it extends across 1. The interval therefore includes values representing no difference in the modeled hazard as well as values favoring D + SoC.

The P-value of 0.07581 is a measure of compatibility with the specified null hypothesis under the statistical testing framework. It is not a measure of effect size, and it should not be interpreted as the probability that the treatment has no effect.

The contrast between the PFS and OS results illustrates why endpoints should be interpreted separately. PFS and OS measure different events, and OS can also be influenced by events and treatments occurring after the initial randomized therapy. The ClinicalTrials.gov record does not provide a basis for attributing the difference between the PFS and OS estimates to any single mechanism.

9. Reading the Two Primary Endpoints Together

Primary endpointHR95% CIP-valueStatistical signal in posted analysis
PFS; D + SoC vs SoC Alone0.740.620–0.8850.00093CI entirely below 1
OS; D + SoC vs SoC Alone0.860.724–1.0160.07581CI crosses 1

The two estimates tell different statistical stories. The PFS estimate is below 1 and its 95% confidence interval remains below 1. The OS estimate is also below 1, but its confidence interval extends above 1. It is therefore important not to collapse the two endpoints into a single overall number.

The HR of 0.74 for PFS and the HR of 0.86 for OS also should not be compared as if they measured the same biological event. One concerns progression or death, while the other concerns death alone. The magnitude of one HR cannot be directly translated into the magnitude of the other.

Interpretive principle: a point estimate below 1 is not, by itself, sufficient to establish a statistically supported treatment difference. The confidence interval and prespecified testing framework provide essential context for interpreting the estimate.

10. Secondary Results: PFS With Tremelimumab

The registry also reports a secondary PFS analysis comparing T + D + SoC with SoC Alone. This comparison used the log-rank test and a stratified Cox proportional-hazards model with the same listed adjustment factors.

Secondary PFS hazard ratio

0.72

95% CI: 0.600–0.860   ·   P = 0.00031

T + D + SoC vs SoC Alone

MeasureReported result
EndpointPFS
ComparisonT + D + SoC vs SoC Alone
MethodLog-rank test
HR0.72
95% CI0.600–0.860
P-value0.00031
Hypothesis typeSuperiority

An HR of 0.72 means the fitted model estimates an approximately 28% lower instantaneous rate of progression or death for T + D + SoC relative to SoC Alone. The 95% CI remains below 1, from 0.600 to 0.860. As with the primary PFS analysis, this is a relative time-to-event interpretation rather than an absolute probability of remaining progression-free.

11. Secondary Results: Overall Survival With Tremelimumab

The registry reports a secondary OS comparison of T + D + SoC versus SoC Alone.

Secondary OS hazard ratio

0.77

95% CI: 0.650–0.916   ·   P = 0.00304

T + D + SoC vs SoC Alone

MeasureReported result
EndpointOS
ComparisonT + D + SoC vs SoC Alone
MethodLog-rank test
HR0.77
95% CI0.650–0.916
P-value0.00304
Hypothesis typeSuperiority

An HR of 0.77 corresponds to an estimated instantaneous death rate approximately 23% lower for T + D + SoC than for SoC Alone under the fitted model. The 95% CI of 0.650–0.916 remains below 1, indicating that the interval is consistent with a lower modeled hazard for the T + D + SoC group.

Multiplicity matters: the existence of a small P-value does not, by itself, establish that every secondary comparison has the same confirmatory status as a primary endpoint. The registry identifies these analyses as secondary and the ClinicalTrials.gov record does not provide a complete alpha-allocation or hierarchical testing scheme. The P-values should therefore be interpreted according to the trial's prespecified multiplicity framework rather than treated as isolated tests.

12. Secondary Results: Objective Response Rate

ORR was analyzed as a binary endpoint. The registry identifies logistic regression as the analysis method and reports odds ratios from adjusted models. The denominator was a subset of the FAS consisting of patients with measurable disease at baseline.

ComparisonOdds ratio95% CIMethod
D + SoC vs SoC Alone1.901.382–2.619Logistic regression
T + D + SoC vs SoC Alone1.721.260–2.367Logistic regression

How to interpret the odds ratios

The reported OR of 1.90 for D + SoC versus SoC Alone means that the estimated odds of objective response were 1.90 times as high in the D + SoC group under the adjusted logistic model. It does not mean that 90% more patients responded, because odds and probabilities are different quantities.

Likewise, the OR of 1.72 for T + D + SoC versus SoC Alone means that the estimated odds of objective response were 1.72 times as high under the adjusted model. The ClinicalTrials.gov record does not provide response percentages, so an absolute response-rate difference cannot be calculated without introducing information outside the permitted dataset.

Clinical Biostats interpretation

The confidence intervals provide the precision information for the two OR estimates. For D + SoC, the 95% CI is 1.382–2.619. For T + D + SoC, it is 1.260–2.367. Neither interval includes 1, so the reported intervals are consistent with odds of response above those in the comparator under the respective models.

Because the analysis population is restricted to patients with measurable disease at baseline, these ORs should not be casually described as if they were calculated from every randomized participant. The analysis population is part of the definition of the statistical result.

13. Secondary Results: Time From Randomization to Second Progression

The registry reports two secondary PFS2 analyses. PFS2 is a time-to-event endpoint analyzed using a stratified Cox proportional-hazards model with adjustment for PD-L1, histology, and disease stage.

ComparisonHR95% CIAnalysis
D + SoC vs SoC Alone0.790.666–0.928Stratified Cox proportional-hazards model
T + D + SoC vs SoC Alone0.750.632–0.883Stratified Cox proportional-hazards model

The D + SoC PFS2 HR of 0.79 corresponds to an approximately 21% lower estimated instantaneous event rate under the model. The T + D + SoC HR of 0.75 corresponds to an approximately 25% lower estimated instantaneous event rate relative to SoC Alone.

Unlike the primary analyses, the ClinicalTrials.gov record does not report P-values for these two PFS2 results. The confidence intervals therefore provide the available precision information in the ClinicalTrials.gov recordset. The registry describes both analyses as secondary.

14. Complete Posted Statistical-Analysis Summary

EndpointComparisonMethodEffect95% CIP-value
Primary PFSD + SoC vs SoC AloneLog-rank; stratified Cox modelHR 0.740.620–0.8850.00093
Primary OSD + SoC vs SoC AloneLog-rank; stratified Cox modelHR 0.860.724–1.0160.07581
Secondary PFST + D + SoC vs SoC AloneLog-rank; stratified Cox modelHR 0.720.600–0.8600.00031
Secondary OST + D + SoC vs SoC AloneLog-rank; stratified Cox modelHR 0.770.650–0.9160.00304
Secondary ORRD + SoC vs SoC AloneLogistic regressionOR 1.901.382–2.619Not reported in registry-reported analysis
Secondary ORRT + D + SoC vs SoC AloneLogistic regressionOR 1.721.260–2.367Not reported in registry-reported analysis
Secondary PFS2D + SoC vs SoC AloneStratified Cox modelHR 0.790.666–0.928Not reported in registry-reported analysis
Secondary PFS2T + D + SoC vs SoC AloneStratified Cox modelHR 0.750.632–0.883Not reported in registry-reported analysis

15. Statistical Methods Explained

Why use a Cox proportional-hazards model for PFS and OS?

PFS and OS are time-to-event outcomes, so the analysis must account not only for whether an event occurs but also for when it occurs and for censoring. A Cox model provides a way to estimate a relative hazard while incorporating covariate adjustment. In POSEIDON, the registry specifies a stratified Cox model for the reported HR estimates.

What does an HR of 0.74 mean?

An HR of 0.74 means that the fitted model estimates the instantaneous event rate in the treatment group at approximately 74% of that in the comparator group. Equivalently, the estimated relative hazard is approximately 26% lower. It does not mean that 26% of patients avoided the event, nor does it provide an absolute difference in survival probability.

Why is the confidence interval important?

A point estimate is only one estimate from the observed data. The 95% confidence interval communicates how precisely the effect has been estimated under the statistical model and sampling framework. For the primary PFS result, the interval is 0.620–0.885. For primary OS, it is 0.724–1.016. Those intervals convey materially different levels of compatibility with a null hazard ratio of 1.

Why was logistic regression used for ORR?

ORR is binary: a patient either meets the prespecified response definition or does not. Logistic regression models the probability of a binary outcome through its odds and allows covariates to be included. POSEIDON's posted analyses adjusted for PD-L1 tumor expression, histology, and disease stage.

What does an OR of 1.90 mean?

An OR of 1.90 means that the modeled odds of response are 1.90 times the comparator odds. Odds are not the same as probability. For example, an odds ratio cannot be converted into a percentage-point increase without knowing the underlying comparator probability.

Why does stratification matter?

The posted Cox analyses adjust for PD-L1 status, histology, and disease stage, and describe these as stratification factors in the model. Adjustment can improve the alignment between the statistical analysis and the randomized design factors and can account for important baseline differences in the model structure. It does not turn a randomized trial into a non-randomized study; rather, it specifies how the randomized comparison is estimated.

Why should PFS and OS not be treated as interchangeable?

PFS counts progression or death, whereas OS counts death from any cause. A therapy can therefore produce a different relative effect on PFS than on OS. The two endpoints also have different susceptibility to events occurring after progression, making it inappropriate to infer one directly from the other.

16. Understanding the Primary PFS Result in Context

Relative effect

HR 0.74 describes the estimated relative hazard of progression or death under the Cox model.

Precision

The 95% CI of 0.620–0.885 describes uncertainty around the estimated HR.

Statistical evidence

The reported P-value is 0.00093. It addresses the specified hypothesis test, not the size of the treatment effect.

Model dependence

The HR comes from a stratified Cox model and should be interpreted within that model's assumptions.

The most useful way to read the primary PFS result is therefore to keep three quantities separate: effect size, precision, and statistical evidence. The HR of 0.74 describes the estimated relative effect. The interval 0.620–0.885 describes uncertainty around that estimate. The P-value of 0.00093 describes the result of the specified hypothesis test.

None of these quantities directly tells a patient how long they will remain progression-free. Nor does the HR describe the absolute difference in the probability of progression or death at a particular time point. Those are different estimands.

17. Understanding the Primary OS Result in Context

The primary OS analysis provides a useful contrast with PFS. The point estimate is below 1 at 0.86, but the 95% CI extends from 0.724 to 1.016. The corresponding P-value is 0.07581.

This does not mean that the HR is "zero effect." The point estimate remains an estimated relative hazard below 1. Instead, the appropriate statistical description is that the estimate is 0.86 and the registry-reported 95% confidence interval includes 1. The interval therefore includes the null value as well as values favoring D + SoC.

Clinical Biostats interpretation

A confidence interval crossing 1 should not be summarized as proving that the treatment and comparator are identical. It means the interval includes the null hazard ratio under the stated analysis. Conversely, a point estimate below 1 should not be summarized as proof of a beneficial effect without considering the interval and the prespecified testing framework.

18. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected patients divided by patients at risk. These figures are reported separately from the efficacy analyses.

Randomized armSerious adverse events affected / at risk
T + D + SoC146/330
D + SoC134/334
SoC Alone117/333

These counts should be read as affected patients divided by patients at risk, exactly as reported in the registry summary. They are not hazard ratios, odds ratios, or formal between-arm tests.

A useful statistical distinction is that efficacy and safety answer different questions. The efficacy analyses described above are anchored to randomized treatment assignment in the global-cohort FAS. The ClinicalTrials.gov record instead describes the number of patients affected by serious adverse events within each arm. A safety count alone does not establish whether a treatment caused a difference in event incidence without a prespecified comparative analysis and appropriate exposure definition.

19. Registry-Specific Design Caveat: China Tail

China-tail caveat: The registry states that the study also incorporates a China tail, comprising additional patients randomized after the end of the global cohort recruitment. Efficacy and safety of patients randomized in China will be reported at a later date once this separate analysis has been completed.

This matters when interpreting the ClinicalTrials.gov record because the posted primary analyses are explicitly described as applying to the global cohort. The China-tail population is therefore a distinct registry-design consideration rather than something that should be silently combined with the global-cohort estimates presented above.

20. Limitations

21. Why This Trial Matters Statistically

POSEIDON is a useful teaching case because it combines randomized three-arm design with multiple types of clinical-trial estimands. Its primary endpoints are time-to-event outcomes, while ORR introduces a binary outcome requiring a different regression framework. The trial therefore demonstrates why statistical methods should be selected according to the endpoint rather than applied uniformly across all outcomes.

ConceptHow it appears in POSEIDON
RandomizationRandomized phase 3 parallel-group design
Three-arm comparisonT + D + SoC, D + SoC, and SoC Alone
Time-to-event endpointsPFS and OS are primary endpoints; PFS2 is secondary
Log-rank testUsed for posted PFS and OS comparisons
Cox proportional-hazards modelUsed for HR estimation in PFS, OS, and PFS2 analyses
Stratified analysisModels adjust for PD-L1, histology, and disease stage
Efron methodUsed to handle tied event times in the Cox model
Logistic regressionUsed for ORR
Odds ratioReported effect measure for ORR
Confidence intervals95% CIs accompany the posted HR and OR estimates
Full analysis setGlobal-cohort FAS included all randomized patients
Endpoint-specific populationORR denominator restricted to patients with measurable disease at baseline
MultiplicityPrimary and secondary comparisons require distinction in interpretation
Safety analysisSerious adverse events reported by randomized arm
Separate cohort considerationRegistry identifies a China tail requiring a separate later analysis

22. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary PFS analysis reports HR 0.74 with 95% CI 0.620–0.885 and P = 0.00093. The primary OS analysis reports HR 0.86 with 95% CI 0.724–1.016 and P = 0.07581.

Effect-size interpretation

The PFS HR corresponds to an approximately 26% lower estimated instantaneous hazard, while the OS HR corresponds to an approximately 14% lower estimated instantaneous hazard, under their respective Cox models.

Precision interpretation

The confidence intervals show that the precision of the two estimates differs and that the OS interval includes the null hazard ratio of 1.

Clinical interpretation

Clinical meaning cannot be reduced to a single HR or P-value. Absolute outcomes, duration of follow-up, subsequent treatment, adverse events, and the characteristics of the analyzed population all contribute to interpretation.

23. A Closer Look at Covariate Adjustment

The Cox analyses were adjusted for three prespecified characteristics described in the registry analysis text: PD-L1 tumor expression, histology, and disease stage. Specifically, the analysis text identifies PD-L1 as TC ≥50% versus <50%, histology as squamous versus non-squamous, and disease stage as IVA versus IVB.

Adjustment factorCategories reported in analysisWhy it matters statistically
PD-L1 tumor expressionTC ≥50% vs <50%Allows the survival model to account for a prespecified tumor-expression factor.
HistologySquamous vs non-squamousAccounts for histologic classification in the fitted comparison.
Disease stageIVA vs IVBAccounts for disease-stage category in the fitted comparison.

Adjustment should not be confused with "controlling away" the treatment effect. The treatment comparison remains the comparison between randomized groups. The model instead specifies how the relative hazard is estimated while incorporating the listed factors.

For readers learning clinical-trial statistics, this is an important distinction between randomization and model adjustment. Randomization establishes the design basis for causal comparison. Regression adjustment specifies the statistical model used to estimate the treatment effect with additional structure.

24. Why the P-Value Does Not Measure Effect Size

The POSEIDON results provide several useful examples of why P-values and effect measures must be kept conceptually separate.

QuantityWhat it tells youWhat it does not tell you
HRRelative time-to-event effect under the Cox modelAbsolute survival probability or proportion benefiting
ORRelative odds of responsePercentage-point difference in response probability
95% CIPrecision/uncertainty around the estimated effect under the analysis frameworkRange of outcomes that individual patients will experience
P-valueEvidence against a specified null hypothesis under the testing frameworkProbability that the treatment works, or magnitude of the effect

For example, the primary PFS P-value is 0.00093, while the primary OS P-value is 0.07581. The appropriate interpretation is not that one P-value represents a "large effect" and the other a "small effect." Effect magnitude is described by the HR and its confidence interval; the P-value answers a different statistical question.

25. Multiplicity and Multiple Comparisons

POSEIDON has multiple efficacy analyses: two primary endpoints, secondary PFS and OS comparisons involving the tremelimumab-containing arm, two ORR comparisons, and two PFS2 comparisons. This creates a statistical distinction between the number of estimates reported and the number of confirmatory hypotheses for which type I error is controlled.

Primary endpoints

PFS and OS are both registered primary endpoints for the D + SoC versus SoC Alone comparison.

Secondary endpoints

PFS, OS, ORR, and PFS2 analyses involving additional comparisons are identified as secondary in the ClinicalTrials.gov record.

Nominal P-values

A P-value is always tied to a particular hypothesis test; it does not automatically encode the effect of testing many hypotheses.

Interpretive caution

The ClinicalTrials.gov record does not provide a complete multiplicity-control procedure, so the secondary results should not be assigned a stronger confirmatory interpretation than the registry supports.

26. Time-to-Event Analysis: What Censoring Means

Both primary endpoints depend on follow-up over time. Some participants may have an observed event, while others may remain event-free when their available follow-up ends. Such observations are handled through censoring rather than being treated as if the event occurred at the end of follow-up.

The OS registry definition explicitly states that patients not known to have died at the time of analysis were censored using the last recorded date on which they were known to be alive. The PFS definition similarly incorporates time from randomization to progression or death.

Why censoring matters
Observed follow-up = information available up to event or censoring

A censored patient still contributes information to the analysis. The observation is not equivalent to a patient who experienced the event at the censoring time.

This is one reason why a simple calculation based only on the number of events divided by the number of patients cannot reproduce a Kaplan-Meier or Cox analysis. The statistical methods use the timing of events and the evolving risk set.

27. Proportional-Hazards Assumption

The Cox model is based on a proportional-hazards framework. In simplified terms, the model treats the relative hazard between treatment groups as having a stable multiplicative structure over time, conditional on the model specification.

Why this matters

An HR such as 0.74 is therefore not simply a percentage reduction in cumulative probability. It is a model-based summary of relative hazard. If the underlying hazards vary substantially in a way that violates proportionality, a single HR can become less descriptive of the full time-varying treatment effect.

The ClinicalTrials.gov record identifies the Cox proportional-hazards model but do not provide a formal proportional-hazards diagnostic. The HR should therefore be presented as the reported model-based effect estimate rather than as a complete description of the survival curves.

28. Trial Timeline

June 1, 2017

Trial start

The registry lists June 1, 2017 as the study start date.

March 12, 2021

Primary completion

The registry lists March 12, 2021 as the primary completion date. The OS primary endpoint time frame also identifies a global cohort data cutoff of March 12, 2021.

Registry status

Active, not recruiting

the ClinicalTrials.gov record lists the current status as ACTIVE_NOT_RECRUITING.

29. Trial Status and Registry Scope

Status
ACTIVE_NOT_RECRUITING
Results
Results posted: Yes
Outcome measures
14 posted outcome measures
Statistical analyses
8 posted statistical analyses

The presence of posted statistical analyses is important because this page can distinguish formal reported comparisons from methodological explanation. For the eight analyses reported, effect estimates and confidence intervals are available for all eight, while P-values are available for the four posted survival comparisons involving PFS and OS and are not posted on ClinicalTrials.gov for the ORR and PFS2 records.

30. Important Statistical Takeaways

Primary PFS

HR 0.74; 95% CI 0.620–0.885; P = 0.00093. The confidence interval remains below 1.

Primary OS

HR 0.86; 95% CI 0.724–1.016; P = 0.07581. The confidence interval includes 1.

Tremelimumab-containing arm

Secondary PFS HR 0.72 and secondary OS HR 0.77 versus SoC Alone, with the registry-reported confidence intervals below 1.

Response

ORR odds ratios were 1.90 for D + SoC and 1.72 for T + D + SoC versus SoC Alone.

These numbers should be interpreted as separate estimates addressing separate endpoints and comparisons. A statistically disciplined summary preserves the distinction between primary and secondary endpoints, time-to-event and binary outcomes, relative effects and absolute effects, and P-values and confidence intervals.

31. Related Tutorials

Learn more about the methods used in this trial:

32. Related Statistical Calculators

33. Sources

Continue through the Clinical Biostats statistical library

Explore the underlying methods used in randomized clinical-trial analysis, from survival models and confidence intervals to logistic regression and odds ratios.

34. Record Summary

POSEIDON provides a compact example of how modern randomized clinical-trial statistics combine multiple analysis frameworks. The primary endpoints, PFS and OS, are time-to-event outcomes analyzed with log-rank testing and stratified Cox proportional-hazards models. The models adjust for PD-L1 tumor expression, histology, and disease stage and use the Efron method for tied event times. ORR is analyzed differently, using logistic regression and odds ratios.

The primary PFS comparison of D + SoC versus SoC Alone reports an HR of 0.74 with a 95% CI of 0.620–0.885 and P = 0.00093. The primary OS comparison reports an HR of 0.86 with a 95% CI of 0.724–1.016 and P = 0.07581. Secondary analyses report HRs of 0.72 for PFS and 0.77 for OS for T + D + SoC versus SoC Alone, along with ORR odds ratios of 1.90 and 1.72 for the two treatment comparisons reported in the registry.

The central statistical lesson is that these numbers should not be reduced to a single "trial result." Hazard ratios describe relative event hazards under a Cox model; odds ratios describe relative odds of a binary response; confidence intervals describe uncertainty around those estimates; and P-values address specified hypothesis tests. The distinction between primary and secondary analyses, the use of the global-cohort FAS, the measurable-disease restriction for ORR, the unmasked three-arm design, and the separate China-tail analysis are all part of the statistical context required to interpret the posted results accurately.

Clinical Biostats methodology: A trial-results page should reconstruct the statistical story rather than merely repeat reported numbers. The objective is to identify the estimand, analysis population, statistical model, effect measure, uncertainty, and testing framework for each result while keeping reported evidence separate from educational interpretation.