This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record for NCT03164616 and the statistical analyses posted there.
1. Trial at a Glance
POSEIDON was a randomized, open-label, parallel-group phase 3 trial evaluating durvalumab plus tremelimumab with chemotherapy, durvalumab with chemotherapy, or chemotherapy alone in patients with non-small cell lung cancer. The registry reports 1,186 enrolled participants across three arms.
| Feature | POSEIDON |
|---|---|
| Phase | Phase 3 |
| Condition | Non Small Cell Lung Cancer NSCLC |
| Design | Randomized, parallel-group, unmasked |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 1,186 |
| Arms | 3 |
| Primary endpoints | Progression-Free Survival (PFS) and Overall Survival (OS) |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 14 |
| Statistical analyses posted | 8 |
| Lead sponsor | AstraZeneca |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT03164616 |
2. Clinical Question
The central statistical question was whether treatment strategies incorporating durvalumab, with or without tremelimumab, and chemotherapy could improve time-to-event outcomes compared with chemotherapy alone. The registry's two primary comparisons were specifically durvalumab plus standard of care (D + SoC) versus SoC alone for PFS and OS.
Population
Patients with non-small cell lung cancer enrolled in the phase 3 POSEIDON trial.
Intervention strategy
Durvalumab plus chemotherapy, with tremelimumab plus durvalumab plus chemotherapy forming a separate randomized arm.
Comparator
Standard-of-care chemotherapy alone, represented in the registry as SoC Alone.
Primary question
Does D + SoC improve PFS and OS compared with SoC alone under the prespecified superiority framework?
3. Trial Design
T + D + SoC
- Tremelimumab
- Durvalumab
- Standard-of-care chemotherapy
D + SoC
- Durvalumab
- Standard-of-care chemotherapy
SoC Alone
- Standard-of-care chemotherapy
The registry identifies the allocation as RANDOMIZED, the design model as PARALLEL, and masking as NONE. Thus, the treatment comparison is based on randomized assignment, but the trial was not described as blinded.
4. Endpoints
| Endpoint | Registry definition | Time frame |
|---|---|---|
| Progression-Free Survival (PFS); D + SoC Compared With SoC Alone | PFS (per RECIST version 1.1 [RECIST 1.1] using Blinded Independent Central Review [BICR] assessments) was defined as time from date of randomization until date of objective disease progression or death (by any cause in the absence of progression), regardless of whether the patient withdrew from randomized therapy or received another anticancer therapy prior to progression. | Tumor scans performed at baseline, Week 6, Week 12 and then every 8 weeks relative to date of randomization until radiological progression. Assessed until global cohort DCO of 24 July 2019 (maximum of approximately 25 months). |
| Overall Survival (OS); D + SoC Compared With SoC Alone | OS was defined as the time from the date of randomization until death due to any cause. Any patient not known to have died at the time of analysis was censored based on the last recorded date on which the patient was known to be alive. | From baseline until death due to any cause. Assessed until global cohort DCO of 12 March 2021 (maximum of approximately 45 months). |
The two primary endpoints are both time-to-event outcomes, but their events are different. PFS counts objective disease progression or death in the absence of progression, whereas OS counts death from any cause. That distinction is fundamental when interpreting the two hazard ratios.
5. Analysis Populations and Covariate Adjustment
The posted primary analyses state that the global-cohort full analysis set (FAS) included all randomized patients. PFS and OS were analyzed as primary outcome measures for the D + SoC versus SoC Alone comparison.
| Analysis feature | Registry information |
|---|---|
| Primary efficacy population | Global cohort FAS; all randomized patients |
| PFS primary comparison | D + SoC vs SoC Alone |
| OS primary comparison | D + SoC vs SoC Alone |
| Time-to-event model | Stratified Cox proportional-hazards model |
| Covariate adjustment | PD-L1, histology, and disease stage |
| Stratification variables in model | PD-L1 (TC ≥50% vs <50%), histology (squamous vs non-squamous), disease stage (IVA vs IVB) |
| Ties | Efron method |
This is an important feature of the analysis. The reported hazard ratios are not simple unadjusted ratios of event rates. They come from a stratified Cox model that accounts for specified baseline factors. The resulting HR therefore represents a model-based relative comparison conditional on the structure of that analysis.
6. Statistical Methodology
Log-rank testing for time-to-event outcomes
The registry reports the log-rank test for the primary PFS and OS analyses and for the secondary PFS and OS comparisons involving the T + D + SoC arm. The log-rank test compares survival experience between groups across the observed follow-up rather than comparing a single time point.
Stratified Cox proportional-hazards model
The hazard ratios and confidence intervals were estimated from a stratified Cox proportional-hazards model. The model adjusted for PD-L1 tumor expression, histology, and disease stage. The Efron method was used to handle tied event times.
For these analyses, an HR below 1 was specified in the registry as favoring the treatment group being associated with a longer time to the event. The HR is a relative time-to-event measure; it is not an absolute survival probability or a proportion of patients benefiting.
Logistic regression for objective response
Objective Response Rate (ORR) is a binary endpoint, so the registry analysis used logistic regression. The analysis adjusted for PD-L1 tumor expression, histology, and disease stage. The confidence interval was calculated using a profile likelihood approach.
An OR above 1 favors the treatment group in these analyses. The odds ratio should not be read as a risk ratio, and an OR of 1.90 does not mean that the response probability is 90 percentage points higher.
Kaplan-Meier estimation
The OS registry definition states that median OS was calculated using the Kaplan-Meier technique. Kaplan-Meier estimation is designed for right-censored time-to-event data: patients contribute follow-up until the event occurs or until their available observation ends without a documented event.
Here, di represents the number of events at a given event time and ni represents the number at risk immediately before that time.
7. Primary Results: Progression-Free Survival
The registry reports a formal primary analysis of PFS for D + SoC versus SoC Alone. The analysis used the log-rank test, with the hazard ratio and confidence interval estimated from a stratified Cox proportional-hazards model.
Primary PFS hazard ratio
95% CI: 0.620–0.885 · P = 0.00093
Superiority analysis: D + SoC vs SoC Alone
| Measure | Reported result |
|---|---|
| Endpoint | Progression-Free Survival (PFS) |
| Comparison | D + SoC vs SoC Alone |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| HR | 0.74 |
| 95% CI | 0.620–0.885 |
| P-value | 0.00093 |
| Hypothesis type | Superiority |
An HR of 0.74 means that, under the fitted proportional-hazards model, the estimated instantaneous rate of progression or death was approximately 26% lower with D + SoC than with SoC Alone. This is a relative hazard interpretation, not a statement that 26% of patients were protected from progression or death.
The 95% CI of 0.620–0.885 describes the statistical uncertainty around the estimated hazard ratio under the analysis framework. Because the entire interval is below 1, the interval is consistent with a lower estimated hazard for D + SoC relative to SoC Alone.
The P-value of 0.00093 addresses the statistical evidence against the null hypothesis specified for the comparison; it does not measure the magnitude or clinical importance of the treatment effect. The effect magnitude is conveyed by the HR, while the CI conveys precision.
The analysis was stratified and adjusted for PD-L1 tumor expression, histology, and disease stage. Interpretation also depends on the Cox model framework and its proportional-hazards assumption. The HR should therefore not be treated as a literal constant risk ratio applying identically at every time point.
8. Primary Results: Overall Survival
The second primary endpoint was OS, defined as time from randomization to death from any cause. Patients not known to have died at analysis were censored at the last recorded date they were known to be alive.
Primary OS hazard ratio
95% CI: 0.724–1.016 · P = 0.07581
Superiority analysis: D + SoC vs SoC Alone
| Measure | Reported result |
|---|---|
| Endpoint | Overall Survival (OS) |
| Comparison | D + SoC vs SoC Alone |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| HR | 0.86 |
| 95% CI | 0.724–1.016 |
| P-value | 0.07581 |
| Hypothesis type | Superiority |
An HR of 0.86 corresponds to an estimated instantaneous death rate approximately 14% lower for D + SoC than for SoC Alone under the fitted model. That statement describes the estimated relative hazard; it does not mean that 14% of patients avoided death or that individual patients experienced a 14% reduction in their probability of dying.
The 95% CI of 0.724–1.016 is wider in the sense that it extends across 1. The interval therefore includes values representing no difference in the modeled hazard as well as values favoring D + SoC.
The P-value of 0.07581 is a measure of compatibility with the specified null hypothesis under the statistical testing framework. It is not a measure of effect size, and it should not be interpreted as the probability that the treatment has no effect.
The contrast between the PFS and OS results illustrates why endpoints should be interpreted separately. PFS and OS measure different events, and OS can also be influenced by events and treatments occurring after the initial randomized therapy. The ClinicalTrials.gov record does not provide a basis for attributing the difference between the PFS and OS estimates to any single mechanism.
9. Reading the Two Primary Endpoints Together
| Primary endpoint | HR | 95% CI | P-value | Statistical signal in posted analysis |
|---|---|---|---|---|
| PFS; D + SoC vs SoC Alone | 0.74 | 0.620–0.885 | 0.00093 | CI entirely below 1 |
| OS; D + SoC vs SoC Alone | 0.86 | 0.724–1.016 | 0.07581 | CI crosses 1 |
The two estimates tell different statistical stories. The PFS estimate is below 1 and its 95% confidence interval remains below 1. The OS estimate is also below 1, but its confidence interval extends above 1. It is therefore important not to collapse the two endpoints into a single overall number.
The HR of 0.74 for PFS and the HR of 0.86 for OS also should not be compared as if they measured the same biological event. One concerns progression or death, while the other concerns death alone. The magnitude of one HR cannot be directly translated into the magnitude of the other.
10. Secondary Results: PFS With Tremelimumab
The registry also reports a secondary PFS analysis comparing T + D + SoC with SoC Alone. This comparison used the log-rank test and a stratified Cox proportional-hazards model with the same listed adjustment factors.
Secondary PFS hazard ratio
95% CI: 0.600–0.860 · P = 0.00031
T + D + SoC vs SoC Alone
| Measure | Reported result |
|---|---|
| Endpoint | PFS |
| Comparison | T + D + SoC vs SoC Alone |
| Method | Log-rank test |
| HR | 0.72 |
| 95% CI | 0.600–0.860 |
| P-value | 0.00031 |
| Hypothesis type | Superiority |
An HR of 0.72 means the fitted model estimates an approximately 28% lower instantaneous rate of progression or death for T + D + SoC relative to SoC Alone. The 95% CI remains below 1, from 0.600 to 0.860. As with the primary PFS analysis, this is a relative time-to-event interpretation rather than an absolute probability of remaining progression-free.
11. Secondary Results: Overall Survival With Tremelimumab
The registry reports a secondary OS comparison of T + D + SoC versus SoC Alone.
Secondary OS hazard ratio
95% CI: 0.650–0.916 · P = 0.00304
T + D + SoC vs SoC Alone
| Measure | Reported result |
|---|---|
| Endpoint | OS |
| Comparison | T + D + SoC vs SoC Alone |
| Method | Log-rank test |
| HR | 0.77 |
| 95% CI | 0.650–0.916 |
| P-value | 0.00304 |
| Hypothesis type | Superiority |
An HR of 0.77 corresponds to an estimated instantaneous death rate approximately 23% lower for T + D + SoC than for SoC Alone under the fitted model. The 95% CI of 0.650–0.916 remains below 1, indicating that the interval is consistent with a lower modeled hazard for the T + D + SoC group.
12. Secondary Results: Objective Response Rate
ORR was analyzed as a binary endpoint. The registry identifies logistic regression as the analysis method and reports odds ratios from adjusted models. The denominator was a subset of the FAS consisting of patients with measurable disease at baseline.
| Comparison | Odds ratio | 95% CI | Method |
|---|---|---|---|
| D + SoC vs SoC Alone | 1.90 | 1.382–2.619 | Logistic regression |
| T + D + SoC vs SoC Alone | 1.72 | 1.260–2.367 | Logistic regression |
How to interpret the odds ratios
The reported OR of 1.90 for D + SoC versus SoC Alone means that the estimated odds of objective response were 1.90 times as high in the D + SoC group under the adjusted logistic model. It does not mean that 90% more patients responded, because odds and probabilities are different quantities.
Likewise, the OR of 1.72 for T + D + SoC versus SoC Alone means that the estimated odds of objective response were 1.72 times as high under the adjusted model. The ClinicalTrials.gov record does not provide response percentages, so an absolute response-rate difference cannot be calculated without introducing information outside the permitted dataset.
The confidence intervals provide the precision information for the two OR estimates. For D + SoC, the 95% CI is 1.382–2.619. For T + D + SoC, it is 1.260–2.367. Neither interval includes 1, so the reported intervals are consistent with odds of response above those in the comparator under the respective models.
Because the analysis population is restricted to patients with measurable disease at baseline, these ORs should not be casually described as if they were calculated from every randomized participant. The analysis population is part of the definition of the statistical result.
13. Secondary Results: Time From Randomization to Second Progression
The registry reports two secondary PFS2 analyses. PFS2 is a time-to-event endpoint analyzed using a stratified Cox proportional-hazards model with adjustment for PD-L1, histology, and disease stage.
| Comparison | HR | 95% CI | Analysis |
|---|---|---|---|
| D + SoC vs SoC Alone | 0.79 | 0.666–0.928 | Stratified Cox proportional-hazards model |
| T + D + SoC vs SoC Alone | 0.75 | 0.632–0.883 | Stratified Cox proportional-hazards model |
The D + SoC PFS2 HR of 0.79 corresponds to an approximately 21% lower estimated instantaneous event rate under the model. The T + D + SoC HR of 0.75 corresponds to an approximately 25% lower estimated instantaneous event rate relative to SoC Alone.
Unlike the primary analyses, the ClinicalTrials.gov record does not report P-values for these two PFS2 results. The confidence intervals therefore provide the available precision information in the ClinicalTrials.gov recordset. The registry describes both analyses as secondary.
14. Complete Posted Statistical-Analysis Summary
| Endpoint | Comparison | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|---|
| Primary PFS | D + SoC vs SoC Alone | Log-rank; stratified Cox model | HR 0.74 | 0.620–0.885 | 0.00093 |
| Primary OS | D + SoC vs SoC Alone | Log-rank; stratified Cox model | HR 0.86 | 0.724–1.016 | 0.07581 |
| Secondary PFS | T + D + SoC vs SoC Alone | Log-rank; stratified Cox model | HR 0.72 | 0.600–0.860 | 0.00031 |
| Secondary OS | T + D + SoC vs SoC Alone | Log-rank; stratified Cox model | HR 0.77 | 0.650–0.916 | 0.00304 |
| Secondary ORR | D + SoC vs SoC Alone | Logistic regression | OR 1.90 | 1.382–2.619 | Not reported in registry-reported analysis |
| Secondary ORR | T + D + SoC vs SoC Alone | Logistic regression | OR 1.72 | 1.260–2.367 | Not reported in registry-reported analysis |
| Secondary PFS2 | D + SoC vs SoC Alone | Stratified Cox model | HR 0.79 | 0.666–0.928 | Not reported in registry-reported analysis |
| Secondary PFS2 | T + D + SoC vs SoC Alone | Stratified Cox model | HR 0.75 | 0.632–0.883 | Not reported in registry-reported analysis |
15. Statistical Methods Explained
Why use a Cox proportional-hazards model for PFS and OS?
PFS and OS are time-to-event outcomes, so the analysis must account not only for whether an event occurs but also for when it occurs and for censoring. A Cox model provides a way to estimate a relative hazard while incorporating covariate adjustment. In POSEIDON, the registry specifies a stratified Cox model for the reported HR estimates.
What does an HR of 0.74 mean?
An HR of 0.74 means that the fitted model estimates the instantaneous event rate in the treatment group at approximately 74% of that in the comparator group. Equivalently, the estimated relative hazard is approximately 26% lower. It does not mean that 26% of patients avoided the event, nor does it provide an absolute difference in survival probability.
Why is the confidence interval important?
A point estimate is only one estimate from the observed data. The 95% confidence interval communicates how precisely the effect has been estimated under the statistical model and sampling framework. For the primary PFS result, the interval is 0.620–0.885. For primary OS, it is 0.724–1.016. Those intervals convey materially different levels of compatibility with a null hazard ratio of 1.
Why was logistic regression used for ORR?
ORR is binary: a patient either meets the prespecified response definition or does not. Logistic regression models the probability of a binary outcome through its odds and allows covariates to be included. POSEIDON's posted analyses adjusted for PD-L1 tumor expression, histology, and disease stage.
What does an OR of 1.90 mean?
An OR of 1.90 means that the modeled odds of response are 1.90 times the comparator odds. Odds are not the same as probability. For example, an odds ratio cannot be converted into a percentage-point increase without knowing the underlying comparator probability.
Why does stratification matter?
The posted Cox analyses adjust for PD-L1 status, histology, and disease stage, and describe these as stratification factors in the model. Adjustment can improve the alignment between the statistical analysis and the randomized design factors and can account for important baseline differences in the model structure. It does not turn a randomized trial into a non-randomized study; rather, it specifies how the randomized comparison is estimated.
Why should PFS and OS not be treated as interchangeable?
PFS counts progression or death, whereas OS counts death from any cause. A therapy can therefore produce a different relative effect on PFS than on OS. The two endpoints also have different susceptibility to events occurring after progression, making it inappropriate to infer one directly from the other.
16. Understanding the Primary PFS Result in Context
Relative effect
HR 0.74 describes the estimated relative hazard of progression or death under the Cox model.
Precision
The 95% CI of 0.620–0.885 describes uncertainty around the estimated HR.
Statistical evidence
The reported P-value is 0.00093. It addresses the specified hypothesis test, not the size of the treatment effect.
Model dependence
The HR comes from a stratified Cox model and should be interpreted within that model's assumptions.
The most useful way to read the primary PFS result is therefore to keep three quantities separate: effect size, precision, and statistical evidence. The HR of 0.74 describes the estimated relative effect. The interval 0.620–0.885 describes uncertainty around that estimate. The P-value of 0.00093 describes the result of the specified hypothesis test.
None of these quantities directly tells a patient how long they will remain progression-free. Nor does the HR describe the absolute difference in the probability of progression or death at a particular time point. Those are different estimands.
17. Understanding the Primary OS Result in Context
The primary OS analysis provides a useful contrast with PFS. The point estimate is below 1 at 0.86, but the 95% CI extends from 0.724 to 1.016. The corresponding P-value is 0.07581.
This does not mean that the HR is "zero effect." The point estimate remains an estimated relative hazard below 1. Instead, the appropriate statistical description is that the estimate is 0.86 and the registry-reported 95% confidence interval includes 1. The interval therefore includes the null value as well as values favoring D + SoC.
A confidence interval crossing 1 should not be summarized as proving that the treatment and comparator are identical. It means the interval includes the null hazard ratio under the stated analysis. Conversely, a point estimate below 1 should not be summarized as proof of a beneficial effect without considering the interval and the prespecified testing framework.
18. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected patients divided by patients at risk. These figures are reported separately from the efficacy analyses.
| Randomized arm | Serious adverse events affected / at risk |
|---|---|
| T + D + SoC | 146/330 |
| D + SoC | 134/334 |
| SoC Alone | 117/333 |
These counts should be read as affected patients divided by patients at risk, exactly as reported in the registry summary. They are not hazard ratios, odds ratios, or formal between-arm tests.
A useful statistical distinction is that efficacy and safety answer different questions. The efficacy analyses described above are anchored to randomized treatment assignment in the global-cohort FAS. The ClinicalTrials.gov record instead describes the number of patients affected by serious adverse events within each arm. A safety count alone does not establish whether a treatment caused a difference in event incidence without a prespecified comparative analysis and appropriate exposure definition.
19. Registry-Specific Design Caveat: China Tail
This matters when interpreting the ClinicalTrials.gov record because the posted primary analyses are explicitly described as applying to the global cohort. The China-tail population is therefore a distinct registry-design consideration rather than something that should be silently combined with the global-cohort estimates presented above.
20. Limitations
- Three-arm design: the trial contains three randomized arms, but the registry-reported primary analyses specifically concern D + SoC versus SoC Alone. Secondary analyses include T + D + SoC comparisons.
- Primary OS precision: the primary OS HR is 0.86 with a 95% CI of 0.724–1.016, so the interval includes 1.
- Hazard-ratio interpretation: Cox HRs are model-based relative measures. Their interpretation depends on the proportional-hazards framework and should not be treated as an absolute risk difference.
- Censoring: time-to-event analyses depend on how follow-up ends for patients without observed events. The OS definition specifies censoring based on the last recorded date on which the patient was known to be alive.
- Analysis populations: the primary efficacy analyses use the global-cohort FAS, while ORR uses a subset of the FAS with measurable disease at baseline.
- Multiplicity: the registry reports multiple primary and secondary comparisons. The ClinicalTrials.gov record does not provide a complete multiplicity hierarchy or alpha-allocation scheme, so secondary P-values should not automatically be interpreted as independently confirmatory.
- Unmasked design: masking is listed as none. This does not invalidate randomized comparisons, but lack of masking can matter for endpoints that involve investigator assessment or other potentially subjective processes.
- China tail: the registry states that additional patients were randomized after global cohort recruitment ended and that their separate efficacy and safety analysis will be reported later.
- Limited absolute-effect information in the ClinicalTrials.gov recordset: the posted statistical-analysis records provide HRs, ORs, confidence intervals, and selected P-values, but the ClinicalTrials.gov record does not provide all corresponding median survival or response-rate values.
21. Why This Trial Matters Statistically
POSEIDON is a useful teaching case because it combines randomized three-arm design with multiple types of clinical-trial estimands. Its primary endpoints are time-to-event outcomes, while ORR introduces a binary outcome requiring a different regression framework. The trial therefore demonstrates why statistical methods should be selected according to the endpoint rather than applied uniformly across all outcomes.
| Concept | How it appears in POSEIDON |
|---|---|
| Randomization | Randomized phase 3 parallel-group design |
| Three-arm comparison | T + D + SoC, D + SoC, and SoC Alone |
| Time-to-event endpoints | PFS and OS are primary endpoints; PFS2 is secondary |
| Log-rank test | Used for posted PFS and OS comparisons |
| Cox proportional-hazards model | Used for HR estimation in PFS, OS, and PFS2 analyses |
| Stratified analysis | Models adjust for PD-L1, histology, and disease stage |
| Efron method | Used to handle tied event times in the Cox model |
| Logistic regression | Used for ORR |
| Odds ratio | Reported effect measure for ORR |
| Confidence intervals | 95% CIs accompany the posted HR and OR estimates |
| Full analysis set | Global-cohort FAS included all randomized patients |
| Endpoint-specific population | ORR denominator restricted to patients with measurable disease at baseline |
| Multiplicity | Primary and secondary comparisons require distinction in interpretation |
| Safety analysis | Serious adverse events reported by randomized arm |
| Separate cohort consideration | Registry identifies a China tail requiring a separate later analysis |
22. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The primary PFS analysis reports HR 0.74 with 95% CI 0.620–0.885 and P = 0.00093. The primary OS analysis reports HR 0.86 with 95% CI 0.724–1.016 and P = 0.07581.
Effect-size interpretation
The PFS HR corresponds to an approximately 26% lower estimated instantaneous hazard, while the OS HR corresponds to an approximately 14% lower estimated instantaneous hazard, under their respective Cox models.
Precision interpretation
The confidence intervals show that the precision of the two estimates differs and that the OS interval includes the null hazard ratio of 1.
Clinical interpretation
Clinical meaning cannot be reduced to a single HR or P-value. Absolute outcomes, duration of follow-up, subsequent treatment, adverse events, and the characteristics of the analyzed population all contribute to interpretation.
23. A Closer Look at Covariate Adjustment
The Cox analyses were adjusted for three prespecified characteristics described in the registry analysis text: PD-L1 tumor expression, histology, and disease stage. Specifically, the analysis text identifies PD-L1 as TC ≥50% versus <50%, histology as squamous versus non-squamous, and disease stage as IVA versus IVB.
| Adjustment factor | Categories reported in analysis | Why it matters statistically |
|---|---|---|
| PD-L1 tumor expression | TC ≥50% vs <50% | Allows the survival model to account for a prespecified tumor-expression factor. |
| Histology | Squamous vs non-squamous | Accounts for histologic classification in the fitted comparison. |
| Disease stage | IVA vs IVB | Accounts for disease-stage category in the fitted comparison. |
Adjustment should not be confused with "controlling away" the treatment effect. The treatment comparison remains the comparison between randomized groups. The model instead specifies how the relative hazard is estimated while incorporating the listed factors.
For readers learning clinical-trial statistics, this is an important distinction between randomization and model adjustment. Randomization establishes the design basis for causal comparison. Regression adjustment specifies the statistical model used to estimate the treatment effect with additional structure.
24. Why the P-Value Does Not Measure Effect Size
The POSEIDON results provide several useful examples of why P-values and effect measures must be kept conceptually separate.
| Quantity | What it tells you | What it does not tell you |
|---|---|---|
| HR | Relative time-to-event effect under the Cox model | Absolute survival probability or proportion benefiting |
| OR | Relative odds of response | Percentage-point difference in response probability |
| 95% CI | Precision/uncertainty around the estimated effect under the analysis framework | Range of outcomes that individual patients will experience |
| P-value | Evidence against a specified null hypothesis under the testing framework | Probability that the treatment works, or magnitude of the effect |
For example, the primary PFS P-value is 0.00093, while the primary OS P-value is 0.07581. The appropriate interpretation is not that one P-value represents a "large effect" and the other a "small effect." Effect magnitude is described by the HR and its confidence interval; the P-value answers a different statistical question.
25. Multiplicity and Multiple Comparisons
POSEIDON has multiple efficacy analyses: two primary endpoints, secondary PFS and OS comparisons involving the tremelimumab-containing arm, two ORR comparisons, and two PFS2 comparisons. This creates a statistical distinction between the number of estimates reported and the number of confirmatory hypotheses for which type I error is controlled.
Primary endpoints
PFS and OS are both registered primary endpoints for the D + SoC versus SoC Alone comparison.
Secondary endpoints
PFS, OS, ORR, and PFS2 analyses involving additional comparisons are identified as secondary in the ClinicalTrials.gov record.
Nominal P-values
A P-value is always tied to a particular hypothesis test; it does not automatically encode the effect of testing many hypotheses.
Interpretive caution
The ClinicalTrials.gov record does not provide a complete multiplicity-control procedure, so the secondary results should not be assigned a stronger confirmatory interpretation than the registry supports.
26. Time-to-Event Analysis: What Censoring Means
Both primary endpoints depend on follow-up over time. Some participants may have an observed event, while others may remain event-free when their available follow-up ends. Such observations are handled through censoring rather than being treated as if the event occurred at the end of follow-up.
The OS registry definition explicitly states that patients not known to have died at the time of analysis were censored using the last recorded date on which they were known to be alive. The PFS definition similarly incorporates time from randomization to progression or death.
A censored patient still contributes information to the analysis. The observation is not equivalent to a patient who experienced the event at the censoring time.
This is one reason why a simple calculation based only on the number of events divided by the number of patients cannot reproduce a Kaplan-Meier or Cox analysis. The statistical methods use the timing of events and the evolving risk set.
27. Proportional-Hazards Assumption
The Cox model is based on a proportional-hazards framework. In simplified terms, the model treats the relative hazard between treatment groups as having a stable multiplicative structure over time, conditional on the model specification.
An HR such as 0.74 is therefore not simply a percentage reduction in cumulative probability. It is a model-based summary of relative hazard. If the underlying hazards vary substantially in a way that violates proportionality, a single HR can become less descriptive of the full time-varying treatment effect.
The ClinicalTrials.gov record identifies the Cox proportional-hazards model but do not provide a formal proportional-hazards diagnostic. The HR should therefore be presented as the reported model-based effect estimate rather than as a complete description of the survival curves.
28. Trial Timeline
Trial start
The registry lists June 1, 2017 as the study start date.
Primary completion
The registry lists March 12, 2021 as the primary completion date. The OS primary endpoint time frame also identifies a global cohort data cutoff of March 12, 2021.
Active, not recruiting
the ClinicalTrials.gov record lists the current status as ACTIVE_NOT_RECRUITING.
29. Trial Status and Registry Scope
The presence of posted statistical analyses is important because this page can distinguish formal reported comparisons from methodological explanation. For the eight analyses reported, effect estimates and confidence intervals are available for all eight, while P-values are available for the four posted survival comparisons involving PFS and OS and are not posted on ClinicalTrials.gov for the ORR and PFS2 records.
30. Important Statistical Takeaways
Primary PFS
HR 0.74; 95% CI 0.620–0.885; P = 0.00093. The confidence interval remains below 1.
Primary OS
HR 0.86; 95% CI 0.724–1.016; P = 0.07581. The confidence interval includes 1.
Tremelimumab-containing arm
Secondary PFS HR 0.72 and secondary OS HR 0.77 versus SoC Alone, with the registry-reported confidence intervals below 1.
Response
ORR odds ratios were 1.90 for D + SoC and 1.72 for T + D + SoC versus SoC Alone.
These numbers should be interpreted as separate estimates addressing separate endpoints and comparisons. A statistically disciplined summary preserves the distinction between primary and secondary endpoints, time-to-event and binary outcomes, relative effects and absolute effects, and P-values and confidence intervals.
31. Related Tutorials
Learn more about the methods used in this trial:
32. Related Statistical Calculators
33. Sources
- ClinicalTrials.gov: POSEIDON, NCT03164616.
- Linked publication: PubMed record for PMID 39243945.
- Linked publication: PubMed record for PMID 37777827.
- Linked publication: PubMed record for PMID 36327426.
Continue through the Clinical Biostats statistical library
Explore the underlying methods used in randomized clinical-trial analysis, from survival models and confidence intervals to logistic regression and odds ratios.
34. Record Summary
POSEIDON provides a compact example of how modern randomized clinical-trial statistics combine multiple analysis frameworks. The primary endpoints, PFS and OS, are time-to-event outcomes analyzed with log-rank testing and stratified Cox proportional-hazards models. The models adjust for PD-L1 tumor expression, histology, and disease stage and use the Efron method for tied event times. ORR is analyzed differently, using logistic regression and odds ratios.
The primary PFS comparison of D + SoC versus SoC Alone reports an HR of 0.74 with a 95% CI of 0.620–0.885 and P = 0.00093. The primary OS comparison reports an HR of 0.86 with a 95% CI of 0.724–1.016 and P = 0.07581. Secondary analyses report HRs of 0.72 for PFS and 0.77 for OS for T + D + SoC versus SoC Alone, along with ORR odds ratios of 1.90 and 1.72 for the two treatment comparisons reported in the registry.
The central statistical lesson is that these numbers should not be reduced to a single "trial result." Hazard ratios describe relative event hazards under a Cox model; odds ratios describe relative odds of a binary response; confidence intervals describe uncertainty around those estimates; and P-values address specified hypothesis tests. The distinction between primary and secondary analyses, the use of the global-cohort FAS, the measurable-disease restriction for ORR, the unmasked three-arm design, and the separate China-tail analysis are all part of the statistical context required to interpret the posted results accurately.