← Clinical Trials
Extensive-Stage SCLC Phase 3 Randomized NCT03043872

CASPIAN: Complete Statistical Analysis of Durvalumab in Extensive-Stage Small Cell Lung Cancer

An independent statistical analysis of the randomized phase 3 CASPIAN trial evaluating durvalumab with platinum-based chemotherapy, with or without tremelimumab, versus platinum-based chemotherapy in untreated extensive-stage small cell lung cancer.

Trial start: 2017-03-27  ·  Primary completion: 2020-01-27  ·  Enrollment: 987
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information posted for NCT03043872 on ClinicalTrials.gov and in the ClinicalTrials.gov record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

CASPIAN is a randomized, open-label, parallel-group phase 3 trial evaluating durvalumab with platinum-based chemotherapy, with or without tremelimumab, against platinum-based chemotherapy in untreated extensive-stage small cell lung cancer.

987
Enrollment
Randomized trial
3
Arms
Parallel-group design
0.75
Global OS HR
D + EP vs EP
0.82
Global OS HR
D + T + EP vs EP
FeatureCASPIAN
Trial nameCASPIAN
PhasePhase 3
PopulationUntreated extensive-stage small cell lung cancer
DesignRandomized, parallel-group, unmasked
AllocationRandomized
Arms3
Primary endpoints4 registered overall-survival analyses across the global and China cohorts and prespecified analysis stages
Outcome measures posted32
Statistical analyses posted18
Primary-endpoint analyses posted6
Lead sponsorAstraZeneca
ClinicalTrials.govNCT03043872

2. Clinical Question

The central statistical question was whether adding durvalumab to platinum-based chemotherapy, with or without tremelimumab, was associated with a different overall-survival experience than platinum-based chemotherapy alone in patients with untreated extensive-stage small cell lung cancer.

Population

Patients with Small Cell Lung Carcinoma Extensive Disease who had not previously received treatment for the trial population's disease setting.

Intervention

Durvalumab plus platinum-based chemotherapy, evaluated both with tremelimumab and without tremelimumab.

Comparator

Platinum-based chemotherapy represented by the EP arm.

Primary question

Does treatment assignment change overall survival, as measured from randomization until death from any cause?

3. Trial Design

01
Randomize987 enrolled
02
3 armsD + T + EP, D + EP, EP
03
AssessOS, PFS, response and safety
04
InterimGlobal OS analysis
05
FinalGlobal and China analyses
ARM A

Durvalumab + tremelimumab + EP

  • Durvalumab
  • Tremelimumab
  • Platinum-based chemotherapy
  • Etoposide
ARM B

Durvalumab + EP

  • Durvalumab
  • Platinum-based chemotherapy
  • Etoposide
ARM C

EP

  • Platinum-based chemotherapy
  • Etoposide
  • Carboplatin or cisplatin were the registered platinum interventions
Allocation
Randomized allocation was used to create the three parallel treatment groups.
Masking
The trial was unmasked; the registry lists masking as none.
Primary purpose
Treatment.
Study status
Active, not recruiting in the ClinicalTrials.gov record.

4. Primary Endpoints

The registry lists four primary overall-survival endpoint analyses. They are distinguished by cohort, analysis stage, and randomized comparison rather than being four different biological endpoints.

Primary endpointRegistered time frameAnalysis
Overall Survival (OS) in the Global Cohort; Global Cohort Interim Analysis; D + EP Compared With EP From baseline until death due to any cause; assessed until global cohort interim analysis DCO Kaplan-Meier estimation with log-rank testing and stratified Cox modeling for HR estimation
OS in the Global Cohort; Global Cohort Final Analysis; D + EP Compared With EP and D + T + EP Compared With EP From baseline until death due to any cause; assessed until global cohort final analysis DCO Log-rank testing and stratified Cox modeling
OS in the China Cohort; China Cohort First Analysis; D + EP Compared With EP From baseline until death due to any cause; assessed until China cohort first analysis DCO Log-rank testing and stratified Cox modeling
OS in the China Cohort; China Cohort Second Analysis; D + EP Compared With EP and D + T + EP Compared With EP From baseline until death due to any cause; assessed until China cohort second analysis DCO Log-rank testing and stratified Cox modeling

For OS, the registry definition is time from the date of randomization until death due to any cause. Patients not known to have died at the time of analysis were censored at the last recorded date on which they were known to be alive. Median OS was calculated using the Kaplan-Meier technique.

5. Analysis Populations and Cohorts

The registry distinguishes a global full analysis set (FAS) from a China FAS. The global FAS included all patients randomized prior to the end of global recruitment. The China FAS included all randomized patients in the China cohort.

PopulationRegistry definition / role
Global FASAll patients randomized prior to the end of global recruitment.
China FASAll randomized patients in the China cohort.
Response denominatorFor global ORR, the denominator was a subset of the FAS population who had measurable disease at baseline.
China cohort interpretation: The registry states that the China cohort was evaluated to assess consistency of efficacy and safety with the global cohort in order to meet regulatory requirements and was not powered for a formal assessment of statistical significance. The China cohort statistical analyses were therefore considered exploratory.

6. Statistical Methodology

Kaplan-Meier estimation

Overall survival and progression-free survival are time-to-event endpoints. The registry states that median OS was calculated using the Kaplan-Meier technique. Kaplan-Meier estimation is appropriate when some patients have not experienced the event by the analysis cutoff because those observations can be right-censored rather than treated as if the event occurred.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

At each event time, the estimated survival probability is updated using the number of events and the number of patients at risk immediately before that time.

Log-rank test

The registry reports the log-rank test as the method for the OS and PFS time-to-event comparisons. Conceptually, the log-rank test compares the observed and expected numbers of events between treatment groups over follow-up, while accounting for censoring.

Stratified Cox proportional-hazards model

The registry reports that the HRs and confidence intervals were calculated using a stratified Cox proportional-hazards model. The model adjusted for planned platinum therapy in Cycle 1, specifically carboplatin or cisplatin, and the registry notes that ties were handled using the Efron approach for the reported global analyses.

Hazard-ratio interpretation
HR < 1  →  lower estimated hazard in the first-named treatment group

For example, the reported global final-analysis HR of 0.75 for D + EP versus EP corresponds to an estimated hazard approximately 25% lower under the fitted model. This is a relative time-to-event measure, not a statement that 25% of patients avoided death or that every patient experienced the same reduction.

Logistic regression

Objective response rate was analyzed using logistic regression. The registry reports odds ratios as the effect measure. For ORR, an odds ratio greater than 1 favors the first-named treatment group in the corresponding comparison.

Odds-ratio interpretation
OR = odds(response in treatment group) / odds(response in comparator group)

An OR of 1.61 does not mean that response was 61 percentage points higher. It means the estimated odds of response were 1.61 times those in the comparator under the logistic model.

Stratification and covariate adjustment

The registry identifies covariate adjustment and stratified analysis as concepts used in the time-to-event analyses. Planned platinum therapy in Cycle 1 was specifically identified as an adjustment factor in the Cox analyses. Stratification helps account for prespecified design factors while estimating the treatment effect.

7. Global Cohort Overall Survival — Interim Analysis

The first reported primary OS analysis in the registry-reported statistical data compared D + EP with EP in the global FAS. The registry states that this interim OS analysis was performed after approximately 318 OS events had occurred in the prespecified design.

Global interim OS hazard ratio

0.73

95% CI: 0.591–0.909   ·   P = 0.0047

D + EP vs EP

The analysis used a log-rank framework, with the HR and confidence interval derived from the corresponding stratified Cox model. The global FAS included all patients randomized prior to the end of global recruitment.

Clinical Biostats interpretation

The HR of 0.73 means that, under the fitted time-to-event model, the estimated instantaneous hazard of death was approximately 27% lower for D + EP than for EP during the analyzed follow-up.

The HR does not mean that 27% of patients survived because of treatment, that 27% fewer patients died overall, or that each patient experienced a 27% reduction in their individual probability of death.

The 95% CI of 0.591–0.909 describes uncertainty around the estimated relative hazard. It does not describe the range of individual treatment effects. Because this was an interim analysis, the confidence interval and P-value must also be understood in the context of the prespecified sequential-testing framework.

The reported P-value of 0.0047 measures evidence against the relevant null hypothesis under the specified analysis; it is not a measure of the size or clinical importance of the treatment effect.

Interim alpha spending

The registry states that the global interim OS analysis used a Lan-DeMets alpha-spending function with an O'Brien-Fleming-type boundary, using the actual number of events observed as a proportion of the planned total. The boundary for declaring statistical significance at this interim analysis was 0.0178 for a 4% overall alpha.

Why an interim boundary is needed

Looking at accumulating outcome data creates multiple opportunities to declare a treatment effect. An alpha-spending approach allocates the permitted type I error across the sequential looks rather than treating every interim P-value as if it came from a single final analysis.

Why the nominal P-value is not enough

The relevant comparison at an interim look is between the observed evidence and the boundary specified for that information time. The registry's boundary was more stringent than a conventional unadjusted threshold.

8. Global Cohort Overall Survival — Final Analysis

D + EP versus EP

Global final OS hazard ratio

0.75

95% CI: 0.625–0.910   ·   P = 0.0032

D + EP vs EP

The final global analysis used the global FAS and compared D + EP with EP. The registry reports a stratified Cox proportional-hazards model for the HR and confidence interval, adjusting for planned platinum therapy in Cycle 1 and handling tied event times using the Efron approach.

Clinical Biostats interpretation

An HR of 0.75 corresponds to an estimated hazard approximately 25% lower for D + EP than EP under the fitted model. It is a relative hazard measure over the analyzed follow-up, not an absolute risk reduction and not a statement about the proportion of patients benefiting.

The 95% CI of 0.625–0.910 gives the statistical uncertainty around the estimated HR. Its width reflects the precision of the estimate; it does not describe the distribution of outcomes among individual patients.

The P-value of 0.0032 quantifies evidence against the null hypothesis under the specified statistical framework. It should not be interpreted as the probability that the treatment effect is real, nor as a measure of effect magnitude.

The Cox interpretation also depends on the proportional-hazards framework. If the hazards are not approximately proportional over time, a single HR may compress important differences in the shapes of the survival curves.

D + T + EP versus EP

Global final OS hazard ratio

0.82

95% CI: 0.682–0.995   ·   P = 0.0451

D + T + EP vs EP

This comparison was also based on the global FAS. The registry reports a log-rank analysis with the HR and confidence interval calculated from a stratified Cox proportional-hazards model. The final-analysis alpha level was adjusted using a generalized Haybittle-Peto method to account for alpha spent at the interim analysis and maintain control of overall type I error.

Clinical Biostats interpretation

The HR of 0.82 corresponds to an estimated hazard approximately 18% lower for D + T + EP than EP under the fitted model.

The 95% CI of 0.682–0.995 is relatively close to 1 at its upper boundary. That makes the precision of the estimate particularly important when interpreting the size of the relative effect.

The P-value of 0.0451 should not be interpreted without the trial's prespecified alpha adjustment. The registry reports a final-analysis significance boundary of 0.0418 for a 5% overall alpha. Thus, the numerical P-value and the prespecified decision boundary are distinct pieces of the statistical framework.

As with the other OS analyses, the HR is not an absolute survival difference and does not imply that every individual experienced the same relative change in hazard.

9. Secondary Overall Survival: D + T + EP versus D + EP

Global OS hazard ratio

1.08

95% CI: 0.890–1.309   ·   P = 0.4352

D + T + EP vs D + EP

This secondary analysis directly compared the two durvalumab-containing strategies. The registry reports a log-rank analysis with a stratified Cox model for the HR and confidence interval, adjusting for planned platinum therapy in Cycle 1 and using the Efron approach for ties.

Clinical Biostats interpretation

An HR of 1.08 is above 1, corresponding to an estimated hazard approximately 8% higher for D + T + EP than D + EP under the fitted model. Because the confidence interval includes 1, the estimate is compatible with no difference in the relative hazard under the model.

The 95% CI of 0.890–1.309 is also important because it spans effects in both directions. It therefore provides substantially more information than the P-value alone about the uncertainty surrounding the estimated treatment contrast.

The P-value of 0.4352 is not a probability that the two treatments are equivalent. It indicates that the observed data are not unusual under the null hypothesis used for this comparison. It does not establish clinical equivalence or prove that the two treatment strategies have identical effects.

10. Progression-Free Survival

PFS was a secondary time-to-event outcome in the registry analyses. Tumor scans were performed at baseline, Week 6, Week 12, then every 8 weeks relative to the date of randomization until RECIST 1.1-defined progression. Assessed until China cohort second analysis DCO (maximum of approximately 29 months)..

Global comparisonHR95% CIP-value
D + EP vs EP0.800.665–0.9590.0157
D + T + EP vs EP0.840.696–1.0050.0568
D + T + EP vs D + EP1.030.857–1.2350.7540

D + EP versus EP

Global PFS hazard ratio

0.80

95% CI: 0.665–0.959   ·   P = 0.0157

Clinical Biostats interpretation

An HR of 0.80 corresponds to an estimated hazard approximately 20% lower for D + EP than EP under the fitted model. The endpoint is progression-free survival, so the event process differs from OS: the relevant event is progression or death according to the registered PFS definition.

The 95% CI of 0.665–0.959 indicates uncertainty around the estimated relative hazard. It does not give a range of median PFS values or individual patient outcomes.

The P-value of 0.0157 is evidence against the null hypothesis within the specified analysis. It does not measure effect size and should not be used alone to judge whether the observed difference is clinically meaningful.

D + T + EP versus EP

Global PFS hazard ratio

0.84

95% CI: 0.696–1.005   ·   P = 0.0568

Clinical Biostats interpretation

The HR of 0.84 corresponds to an estimated hazard approximately 16% lower for D + T + EP than EP under the fitted model.

The confidence interval of 0.696–1.005 crosses 1. The estimate therefore has appreciable uncertainty about the direction of the relative hazard, even though the point estimate is below 1.

The P-value of 0.0568 should be interpreted as a measure of evidence under the specified hypothesis test, not as a measure of treatment effect size and not as evidence that the treatment strategies are equivalent.

D + T + EP versus D + EP

Global PFS hazard ratio

1.03

95% CI: 0.857–1.235   ·   P = 0.7540

Clinical Biostats interpretation

The HR of 1.03 is close to 1, corresponding to an estimated hazard approximately 3% higher for D + T + EP than D + EP under the fitted model.

The 95% CI of 0.857–1.235 includes 1 and allows for both a lower and a higher hazard. The interval therefore communicates uncertainty that a point estimate alone cannot show.

The P-value of 0.7540 does not establish equivalence between the regimens. It indicates limited evidence against the null hypothesis for this particular comparison.

11. Objective Response Rate

ORR was analyzed as a binary endpoint using logistic regression. For the global cohort, the denominator was a subset of the FAS population who had measurable disease at baseline.

Global comparisonOdds ratio95% CIP-value
D + EP vs EP1.611.086–2.4010.0177
D + T + EP vs EP1.190.817–1.7460.3611

D + EP versus EP

Global ORR odds ratio

1.61

95% CI: 1.086–2.401   ·   P = 0.0177

Clinical Biostats interpretation

An OR of 1.61 means that the estimated odds of objective response were 1.61 times the odds in the EP group under the logistic regression model.

This does not mean that 61% more patients responded, nor does it mean that the response probability was 61 percentage points higher. Odds and probabilities are related but are not interchangeable.

The 95% CI of 1.086–2.401 describes uncertainty around the odds ratio. The P-value of 0.0177 measures evidence against the specified null hypothesis; it is not a measure of the magnitude of response improvement.

D + T + EP versus EP

Global ORR odds ratio

1.19

95% CI: 0.817–1.746   ·   P = 0.3611

Clinical Biostats interpretation

The OR of 1.19 indicates estimated response odds 1.19 times those in EP under the logistic model.

The confidence interval of 0.817–1.746 includes 1, so the estimate is compatible with both lower and higher response odds. The interval also illustrates why the point estimate should not be interpreted without its uncertainty.

The P-value of 0.3611 does not establish that the treatments have the same response rate. It indicates limited evidence against the null hypothesis in this analysis.

12. China Cohort Overall Survival

The China cohort analyses were exploratory. The registry explicitly states that the study was not designed or powered to show statistical significance for efficacy endpoints in the China cohort.

AnalysisComparisonHR95% CIP-value
First analysisD + EP vs EP0.650.414–1.0290.0664
Second analysisD + EP vs EP0.750.504–1.1060.1455
Second analysisD + T + EP vs EP0.650.439–0.9640.0314
Second analysisD + T + EP vs D + EP0.860.574–1.2770.4470

D + EP versus EP — first China analysis

China OS hazard ratio

0.65

95% CI: 0.414–1.029   ·   P = 0.0664

Clinical Biostats interpretation

The HR of 0.65 corresponds to an estimated hazard approximately 35% lower for D + EP than EP under the fitted model.

However, the 95% CI of 0.414–1.029 includes 1. More importantly, the registry states that the China cohort was not powered for formal statistical significance and that its analyses were exploratory. Therefore, this estimate should be interpreted primarily as an exploratory assessment of consistency rather than as an independently powered confirmatory test.

The P-value of 0.0664 is not a measure of the effect size and should not be interpreted as the probability that the treatment effect is absent.

D + EP versus EP — second China analysis

China OS hazard ratio

0.75

95% CI: 0.504–1.106   ·   P = 0.1455

Clinical Biostats interpretation

The point estimate corresponds to an estimated hazard approximately 25% lower for D + EP than EP, but the 95% CI of 0.504–1.106 includes 1 and is therefore compatible with a range of relative effects.

Because the China cohort was not powered for formal statistical significance, the P-value of 0.1455 should not be used to turn this exploratory analysis into a separate confirmatory conclusion.

D + T + EP versus EP — second China analysis

China OS hazard ratio

0.65

95% CI: 0.439–0.964   ·   P = 0.0314

Clinical Biostats interpretation

The HR of 0.65 corresponds to an estimated hazard approximately 35% lower for D + T + EP than EP under the fitted model.

The 95% CI of 0.439–0.964 is below 1 at its upper boundary, but the registry's explicit statement that the China cohort was not powered for formal statistical significance remains essential. The analysis was exploratory and intended to evaluate consistency with the global cohort.

The P-value of 0.0314 is therefore not appropriately interpreted as if it came from a separately powered confirmatory trial with the same inferential role as the global primary analysis.

D + T + EP versus D + EP — second China analysis

China OS hazard ratio

0.86

95% CI: 0.574–1.277   ·   P = 0.4470

Clinical Biostats interpretation

The HR of 0.86 corresponds to an estimated hazard approximately 14% lower for D + T + EP than D + EP under the fitted model.

The 95% CI of 0.574–1.277 includes 1 and is wide enough to allow for both a lower and a higher hazard. The P-value of 0.4470 does not establish equivalence, and the exploratory status of the China cohort further limits confirmatory interpretation.

13. China Cohort Progression-Free Survival

ComparisonHR95% CIP-value
D + EP vs EP0.970.661–1.4370.8934
D + T + EP vs EP0.720.487–1.0680.1035
D + T + EP vs D + EP0.760.522–1.1160.1673

All three China PFS analyses were exploratory under the registry's stated framework. Each used log-rank testing with a stratified Cox model for HR estimation and adjustment for planned platinum therapy.

D + EP vs EP

HR 0.97; 95% CI 0.661–1.437; P = 0.8934. The estimate is close to 1, while the confidence interval permits materially different relative hazards.

D + T + EP vs EP

HR 0.72; 95% CI 0.487–1.068; P = 0.1035. The point estimate is below 1, but the interval includes 1.

D + T + EP vs D + EP

HR 0.76; 95% CI 0.522–1.116; P = 0.1673. The point estimate favors the first-named group directionally, but the interval includes 1.

14. China Cohort Objective Response Rate

ComparisonOdds ratio95% CIP-value
D + EP vs EP1.390.610–3.2440.4320
D + T + EP vs EP2.070.874–5.1180.0986

These ORR analyses used logistic regression. As with the China OS and PFS analyses, the registry states that the China cohort was not designed or powered for formal statistical significance and that these analyses were exploratory.

Clinical Biostats interpretation

The OR of 1.39 for D + EP versus EP corresponds to estimated response odds 1.39 times those of EP. Its 95% CI of 0.610–3.244 is wide and includes 1.

The OR of 2.07 for D + T + EP versus EP corresponds to estimated response odds 2.07 times those of EP. Its 95% CI of 0.874–5.118 includes 1 and is also wide.

The P-values of 0.4320 and 0.0986 are hypothesis-test results, not measures of effect magnitude. Given the exploratory design and limited power of the China cohort, the estimates are most appropriately considered in the context of consistency with the global results rather than as independent confirmatory evidence.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized cohort and treatment arm. The data are presented as affected patients divided by the number at risk.

CohortArmSerious adverse events
GlobalD + T + EP126 / 266
GlobalD + EP86 / 265
GlobalEP97 / 266
ChinaD + T + EP31 / 65
ChinaD + EP26 / 61
ChinaEP22 / 62

The reported figures are counts of affected patients relative to the corresponding number at risk. They should not be converted into a different safety metric without the underlying definitions, follow-up periods, and exposure information.

Safety interpretation: Serious adverse-event frequencies should be interpreted separately from efficacy estimates. A randomized efficacy HR and a safety event count answer different statistical questions and generally require different analysis populations and exposure considerations.

16. Interim Analysis, Alpha Spending, and Multiplicity

CASPIAN provides a useful example of why interim analyses cannot simply be treated as an additional opportunity to inspect a conventional P-value. The global interim OS analysis used a Lan-DeMets alpha-spending function with an O'Brien-Fleming-type boundary.

Design elementRegistry-supported detail
Interim endpointOverall survival in the global cohort
Interim comparisonD + EP vs EP
Interim event targetApproximately 318 OS events
Alpha-spending frameworkLan-DeMets alpha spending
Boundary typeO'Brien-Fleming type
Interim significance boundary0.0178
Overall alpha referenced4%
Final alpha adjustmentGeneralized Haybittle-Peto method
Final significance boundary0.0418
Final overall alpha referenced5%

The final analysis for D + T + EP versus EP used an adjusted alpha level to account for actual alpha spent at the interim analysis and the actual final number of events, with the stated objective of maintaining overall type I error control.

Why this matters

If investigators could repeatedly examine an accumulating trial and declare success whenever a conventional threshold was crossed, the probability of a false-positive conclusion would increase. Alpha spending addresses this by defining how much type I error may be used as information accumulates.

The distinction is important when reading the reported P-values. The P-value is a property of the observed data under the specified test, whereas the prespecified boundary determines how that evidence is interpreted within the sequential design.

17. Stratified Analysis and Covariate Adjustment

The registry repeatedly identifies stratified analysis and covariate adjustment as components of the time-to-event methodology. The reported global Cox analyses specifically adjusted for planned platinum therapy in Cycle 1: carboplatin or cisplatin.

Why stratify?

Stratification allows the analysis to account for prespecified factors that can influence the event process without requiring the same baseline hazard function to apply across every stratum.

Why adjust?

Adjustment can improve the precision or preserve the intended treatment comparison when the adjusted variable is related to outcome and was specified as part of the analysis plan.

The registry's reported model also used the Efron approach for tied event times. Ties arise when multiple patients share the same recorded event time; the Efron method provides one way to approximate the partial likelihood contribution of tied events.

18. Statistical Methods Explained

Why was a log-rank test used?

OS and PFS are time-to-event endpoints, so the analysis must account for both the timing of events and right censoring. The log-rank test compares the event experience between groups across follow-up rather than reducing each patient to a simple binary event/no-event outcome.

What does an HR of 0.75 mean?

For the global final OS comparison of D + EP versus EP, an HR of 0.75 means the fitted model estimated approximately 25% lower instantaneous hazard for the D + EP group. It does not mean 25% fewer deaths in absolute terms and does not mean that every patient experienced a 25% reduction in their personal risk.

Why is the confidence interval as important as the HR?

A point estimate alone does not show how precisely the treatment effect has been estimated. The 95% CI provides a range describing statistical uncertainty around the estimated effect under the model and sampling framework. For example, the global final D + EP versus EP OS HR of 0.75 has a 95% CI of 0.625–0.910.

Why doesn't a P-value measure effect size?

A P-value quantifies how compatible the observed data are with a specified null hypothesis under the statistical model. It depends on both the magnitude of the observed effect and the amount of information in the analysis. A small P-value can therefore accompany a modest estimate in a large dataset, while a larger P-value can occur with a substantial point estimate when uncertainty is high.

Why was logistic regression used for ORR?

ORR is a binary outcome: a patient either met the prespecified response criterion or did not. Logistic regression models the probability of the binary outcome and naturally produces an odds ratio. The CASPIAN registry reports this method for the global and China ORR analyses.

Why does the China cohort require special interpretation?

The registry explicitly states that the China cohort was not designed or powered for formal statistical significance and that its analyses were exploratory. Therefore, the China estimates can describe the observed data and contribute to an assessment of consistency, but their P-values should not be interpreted as if they represented an independently powered confirmatory trial.

Why does the interim analysis change interpretation of the P-value?

An interim look creates a sequential-testing problem. CASPIAN used Lan-DeMets alpha spending with an O'Brien-Fleming-type boundary. The reported interim P-value therefore needs to be interpreted against the prespecified interim boundary rather than against an arbitrary conventional threshold.

19. Limitations and Interpretation Issues

  • Exploratory China cohort: the registry states that the China cohort was not powered for formal assessment of statistical significance. Its efficacy and safety analyses were exploratory.
  • Interim analysis: the global interim OS result was generated under a group-sequential framework. The appropriate inferential threshold was determined by alpha spending rather than by treating the interim P-value as an unadjusted final analysis.
  • Multiple comparisons: CASPIAN contains several treatment contrasts, analysis stages, cohorts, and endpoints. The role of each comparison must therefore be distinguished rather than treating every reported P-value as equivalent evidence.
  • Hazard-ratio assumptions: Cox HRs are model-based. A single HR is most naturally interpreted under a proportional-hazards framework; if hazards vary substantially over time, the HR can summarize rather than fully describe the treatment difference.
  • Censoring: Kaplan-Meier and Cox methods depend on appropriate handling of censored observations. A patient censored at the last known date alive contributes follow-up information up to that point but is not treated as having experienced death.
  • ORR denominator: the global ORR analysis used a subset of the FAS with measurable disease at baseline. Therefore, its analysis population differs conceptually from the full randomized population used for OS.
  • Relative versus absolute effects: HRs and ORs are relative measures. They do not directly communicate absolute survival probabilities or absolute differences in response probability.
  • Three-arm structure: the presence of two active strategies means that the D + T + EP versus D + EP comparison answers a different question from either comparison against EP.

20. Why This Trial Matters Statistically

CASPIAN is a useful teaching case because it combines randomized treatment comparisons, multiple treatment arms, time-to-event endpoints, binary response outcomes, interim monitoring, stratified Cox modeling, logistic regression, and an explicitly exploratory regional cohort.

ConceptHow it appears in CASPIAN
RandomizationRandomized phase 3 parallel-group design with 987 enrolled patients
Three-arm comparisonD + T + EP, D + EP, and EP
Kaplan-Meier estimationUsed for median OS and relevant time-to-event estimation
Log-rank testReported for OS and PFS comparisons
Hazard ratioPrimary relative effect measure for OS and secondary measure for PFS
Stratified Cox modelUsed to calculate HRs and confidence intervals
Covariate adjustmentPlatinum therapy in Cycle 1 was included in the reported Cox analyses
Logistic regressionUsed for global and China ORR analyses
Odds ratioEffect measure for the binary ORR endpoint
Interim analysisGlobal OS interim analysis based on approximately 318 OS events
Alpha spendingLan-DeMets function with an O'Brien-Fleming-type boundary
Multiplicity / sequential testingFinal alpha was adjusted after the interim analysis
Exploratory analysisChina cohort analyses were explicitly not powered for formal significance

21. Overall Statistical Picture

The ClinicalTrials.gov record shows a consistent distinction between the global confirmatory framework and the exploratory China cohort. In the global cohort, the D + EP versus EP comparison produced OS HRs of 0.73 at interim analysis and 0.75 at final analysis, while the corresponding PFS HR was 0.80 and the ORR OR was 1.61.

The D + T + EP versus EP comparison produced a global final OS HR of 0.82 and a PFS HR of 0.84. The direct D + T + EP versus D + EP comparisons produced an OS HR of 1.08 and a PFS HR of 1.03. These comparisons illustrate why the reference group matters: the same treatment can have a different statistical interpretation depending on which randomized group serves as the comparator.

The China analyses are more uncertain and explicitly exploratory. Their confidence intervals are generally wider, and the registry cautions that the cohort was not powered for formal statistical significance. The appropriate statistical reading is therefore to examine the point estimates, confidence intervals, and direction of effects while retaining the stated limitation on inferential strength.

22. A Practical Guide to Reading the CASPIAN Results

Start with the estimand

Ask which patients, treatment contrast, endpoint, and analysis time point are being compared. D + EP versus EP is not the same question as D + T + EP versus D + EP.

Then read the effect measure

OS and PFS use hazard ratios, whereas ORR uses odds ratios. The numerical interpretation of 0.75 is therefore fundamentally different from the interpretation of 1.61.

Then read the confidence interval

The interval indicates precision and possible effect sizes under the model. It should be read alongside the point estimate rather than after it as an afterthought.

Finally read the design

Interim monitoring, alpha spending, multiple comparisons, stratification, and exploratory regional analyses determine how the numerical results should be interpreted.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats knowledge graph

Connect the endpoints and methods used in CASPIAN to deeper statistical tutorials and analysis tools.

26. Record Summary

CASPIAN provides a compact teaching example of several central clinical-trial statistical methods. The trial used randomized parallel-group allocation with three treatment arms and evaluated overall survival and progression-free survival as time-to-event outcomes, with logistic regression used for objective response rate. The reported global OS analyses incorporated log-rank testing, stratified Cox modeling, covariate adjustment for planned platinum therapy, and sequential alpha control through Lan-DeMets alpha spending with an O'Brien-Fleming-type boundary.

The numerical results need to be read in the context of their analysis stage and comparison. The global D + EP versus EP OS HR was 0.73 at interim analysis and 0.75 at final analysis. The global D + T + EP versus EP final OS HR was 0.82, while the direct D + T + EP versus D + EP comparison produced an HR of 1.08. PFS and ORR produced corresponding but distinct effect measures.

The China cohort demonstrates another important statistical principle: a numerical estimate and a formal confirmatory conclusion are not synonymous. The registry explicitly states that the China cohort was not powered for formal statistical significance and that its analyses were exploratory. Confidence intervals, analysis populations, study design, and prespecified inferential procedures therefore matter as much as the individual point estimates.

Clinical Biostats methodology: A trial-results page should not merely repeat reported numbers. The goal is to reconstruct the statistical story of the trial while clearly distinguishing the registered analysis, the reported estimates, and the educational interpretation of those estimates.