This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information posted for NCT03043872 on ClinicalTrials.gov and in the ClinicalTrials.gov record.
1. Trial at a Glance
CASPIAN is a randomized, open-label, parallel-group phase 3 trial evaluating durvalumab with platinum-based chemotherapy, with or without tremelimumab, against platinum-based chemotherapy in untreated extensive-stage small cell lung cancer.
| Feature | CASPIAN |
|---|---|
| Trial name | CASPIAN |
| Phase | Phase 3 |
| Population | Untreated extensive-stage small cell lung cancer |
| Design | Randomized, parallel-group, unmasked |
| Allocation | Randomized |
| Arms | 3 |
| Primary endpoints | 4 registered overall-survival analyses across the global and China cohorts and prespecified analysis stages |
| Outcome measures posted | 32 |
| Statistical analyses posted | 18 |
| Primary-endpoint analyses posted | 6 |
| Lead sponsor | AstraZeneca |
| ClinicalTrials.gov | NCT03043872 |
2. Clinical Question
The central statistical question was whether adding durvalumab to platinum-based chemotherapy, with or without tremelimumab, was associated with a different overall-survival experience than platinum-based chemotherapy alone in patients with untreated extensive-stage small cell lung cancer.
Population
Patients with Small Cell Lung Carcinoma Extensive Disease who had not previously received treatment for the trial population's disease setting.
Intervention
Durvalumab plus platinum-based chemotherapy, evaluated both with tremelimumab and without tremelimumab.
Comparator
Platinum-based chemotherapy represented by the EP arm.
Primary question
Does treatment assignment change overall survival, as measured from randomization until death from any cause?
3. Trial Design
Durvalumab + tremelimumab + EP
- Durvalumab
- Tremelimumab
- Platinum-based chemotherapy
- Etoposide
Durvalumab + EP
- Durvalumab
- Platinum-based chemotherapy
- Etoposide
EP
- Platinum-based chemotherapy
- Etoposide
- Carboplatin or cisplatin were the registered platinum interventions
4. Primary Endpoints
The registry lists four primary overall-survival endpoint analyses. They are distinguished by cohort, analysis stage, and randomized comparison rather than being four different biological endpoints.
| Primary endpoint | Registered time frame | Analysis |
|---|---|---|
| Overall Survival (OS) in the Global Cohort; Global Cohort Interim Analysis; D + EP Compared With EP | From baseline until death due to any cause; assessed until global cohort interim analysis DCO | Kaplan-Meier estimation with log-rank testing and stratified Cox modeling for HR estimation |
| OS in the Global Cohort; Global Cohort Final Analysis; D + EP Compared With EP and D + T + EP Compared With EP | From baseline until death due to any cause; assessed until global cohort final analysis DCO | Log-rank testing and stratified Cox modeling |
| OS in the China Cohort; China Cohort First Analysis; D + EP Compared With EP | From baseline until death due to any cause; assessed until China cohort first analysis DCO | Log-rank testing and stratified Cox modeling |
| OS in the China Cohort; China Cohort Second Analysis; D + EP Compared With EP and D + T + EP Compared With EP | From baseline until death due to any cause; assessed until China cohort second analysis DCO | Log-rank testing and stratified Cox modeling |
For OS, the registry definition is time from the date of randomization until death due to any cause. Patients not known to have died at the time of analysis were censored at the last recorded date on which they were known to be alive. Median OS was calculated using the Kaplan-Meier technique.
5. Analysis Populations and Cohorts
The registry distinguishes a global full analysis set (FAS) from a China FAS. The global FAS included all patients randomized prior to the end of global recruitment. The China FAS included all randomized patients in the China cohort.
| Population | Registry definition / role |
|---|---|
| Global FAS | All patients randomized prior to the end of global recruitment. |
| China FAS | All randomized patients in the China cohort. |
| Response denominator | For global ORR, the denominator was a subset of the FAS population who had measurable disease at baseline. |
6. Statistical Methodology
Kaplan-Meier estimation
Overall survival and progression-free survival are time-to-event endpoints. The registry states that median OS was calculated using the Kaplan-Meier technique. Kaplan-Meier estimation is appropriate when some patients have not experienced the event by the analysis cutoff because those observations can be right-censored rather than treated as if the event occurred.
At each event time, the estimated survival probability is updated using the number of events and the number of patients at risk immediately before that time.
Log-rank test
The registry reports the log-rank test as the method for the OS and PFS time-to-event comparisons. Conceptually, the log-rank test compares the observed and expected numbers of events between treatment groups over follow-up, while accounting for censoring.
Stratified Cox proportional-hazards model
The registry reports that the HRs and confidence intervals were calculated using a stratified Cox proportional-hazards model. The model adjusted for planned platinum therapy in Cycle 1, specifically carboplatin or cisplatin, and the registry notes that ties were handled using the Efron approach for the reported global analyses.
For example, the reported global final-analysis HR of 0.75 for D + EP versus EP corresponds to an estimated hazard approximately 25% lower under the fitted model. This is a relative time-to-event measure, not a statement that 25% of patients avoided death or that every patient experienced the same reduction.
Logistic regression
Objective response rate was analyzed using logistic regression. The registry reports odds ratios as the effect measure. For ORR, an odds ratio greater than 1 favors the first-named treatment group in the corresponding comparison.
An OR of 1.61 does not mean that response was 61 percentage points higher. It means the estimated odds of response were 1.61 times those in the comparator under the logistic model.
Stratification and covariate adjustment
The registry identifies covariate adjustment and stratified analysis as concepts used in the time-to-event analyses. Planned platinum therapy in Cycle 1 was specifically identified as an adjustment factor in the Cox analyses. Stratification helps account for prespecified design factors while estimating the treatment effect.
7. Global Cohort Overall Survival — Interim Analysis
The first reported primary OS analysis in the registry-reported statistical data compared D + EP with EP in the global FAS. The registry states that this interim OS analysis was performed after approximately 318 OS events had occurred in the prespecified design.
Global interim OS hazard ratio
95% CI: 0.591–0.909 · P = 0.0047
D + EP vs EP
The analysis used a log-rank framework, with the HR and confidence interval derived from the corresponding stratified Cox model. The global FAS included all patients randomized prior to the end of global recruitment.
The HR of 0.73 means that, under the fitted time-to-event model, the estimated instantaneous hazard of death was approximately 27% lower for D + EP than for EP during the analyzed follow-up.
The HR does not mean that 27% of patients survived because of treatment, that 27% fewer patients died overall, or that each patient experienced a 27% reduction in their individual probability of death.
The 95% CI of 0.591–0.909 describes uncertainty around the estimated relative hazard. It does not describe the range of individual treatment effects. Because this was an interim analysis, the confidence interval and P-value must also be understood in the context of the prespecified sequential-testing framework.
The reported P-value of 0.0047 measures evidence against the relevant null hypothesis under the specified analysis; it is not a measure of the size or clinical importance of the treatment effect.
Interim alpha spending
The registry states that the global interim OS analysis used a Lan-DeMets alpha-spending function with an O'Brien-Fleming-type boundary, using the actual number of events observed as a proportion of the planned total. The boundary for declaring statistical significance at this interim analysis was 0.0178 for a 4% overall alpha.
Why an interim boundary is needed
Looking at accumulating outcome data creates multiple opportunities to declare a treatment effect. An alpha-spending approach allocates the permitted type I error across the sequential looks rather than treating every interim P-value as if it came from a single final analysis.
Why the nominal P-value is not enough
The relevant comparison at an interim look is between the observed evidence and the boundary specified for that information time. The registry's boundary was more stringent than a conventional unadjusted threshold.
8. Global Cohort Overall Survival — Final Analysis
D + EP versus EP
Global final OS hazard ratio
95% CI: 0.625–0.910 · P = 0.0032
D + EP vs EP
The final global analysis used the global FAS and compared D + EP with EP. The registry reports a stratified Cox proportional-hazards model for the HR and confidence interval, adjusting for planned platinum therapy in Cycle 1 and handling tied event times using the Efron approach.
An HR of 0.75 corresponds to an estimated hazard approximately 25% lower for D + EP than EP under the fitted model. It is a relative hazard measure over the analyzed follow-up, not an absolute risk reduction and not a statement about the proportion of patients benefiting.
The 95% CI of 0.625–0.910 gives the statistical uncertainty around the estimated HR. Its width reflects the precision of the estimate; it does not describe the distribution of outcomes among individual patients.
The P-value of 0.0032 quantifies evidence against the null hypothesis under the specified statistical framework. It should not be interpreted as the probability that the treatment effect is real, nor as a measure of effect magnitude.
The Cox interpretation also depends on the proportional-hazards framework. If the hazards are not approximately proportional over time, a single HR may compress important differences in the shapes of the survival curves.
D + T + EP versus EP
Global final OS hazard ratio
95% CI: 0.682–0.995 · P = 0.0451
D + T + EP vs EP
This comparison was also based on the global FAS. The registry reports a log-rank analysis with the HR and confidence interval calculated from a stratified Cox proportional-hazards model. The final-analysis alpha level was adjusted using a generalized Haybittle-Peto method to account for alpha spent at the interim analysis and maintain control of overall type I error.
The HR of 0.82 corresponds to an estimated hazard approximately 18% lower for D + T + EP than EP under the fitted model.
The 95% CI of 0.682–0.995 is relatively close to 1 at its upper boundary. That makes the precision of the estimate particularly important when interpreting the size of the relative effect.
The P-value of 0.0451 should not be interpreted without the trial's prespecified alpha adjustment. The registry reports a final-analysis significance boundary of 0.0418 for a 5% overall alpha. Thus, the numerical P-value and the prespecified decision boundary are distinct pieces of the statistical framework.
As with the other OS analyses, the HR is not an absolute survival difference and does not imply that every individual experienced the same relative change in hazard.
9. Secondary Overall Survival: D + T + EP versus D + EP
Global OS hazard ratio
95% CI: 0.890–1.309 · P = 0.4352
D + T + EP vs D + EP
This secondary analysis directly compared the two durvalumab-containing strategies. The registry reports a log-rank analysis with a stratified Cox model for the HR and confidence interval, adjusting for planned platinum therapy in Cycle 1 and using the Efron approach for ties.
An HR of 1.08 is above 1, corresponding to an estimated hazard approximately 8% higher for D + T + EP than D + EP under the fitted model. Because the confidence interval includes 1, the estimate is compatible with no difference in the relative hazard under the model.
The 95% CI of 0.890–1.309 is also important because it spans effects in both directions. It therefore provides substantially more information than the P-value alone about the uncertainty surrounding the estimated treatment contrast.
The P-value of 0.4352 is not a probability that the two treatments are equivalent. It indicates that the observed data are not unusual under the null hypothesis used for this comparison. It does not establish clinical equivalence or prove that the two treatment strategies have identical effects.
10. Progression-Free Survival
PFS was a secondary time-to-event outcome in the registry analyses. Tumor scans were performed at baseline, Week 6, Week 12, then every 8 weeks relative to the date of randomization until RECIST 1.1-defined progression. Assessed until China cohort second analysis DCO (maximum of approximately 29 months)..
| Global comparison | HR | 95% CI | P-value |
|---|---|---|---|
| D + EP vs EP | 0.80 | 0.665–0.959 | 0.0157 |
| D + T + EP vs EP | 0.84 | 0.696–1.005 | 0.0568 |
| D + T + EP vs D + EP | 1.03 | 0.857–1.235 | 0.7540 |
D + EP versus EP
Global PFS hazard ratio
95% CI: 0.665–0.959 · P = 0.0157
An HR of 0.80 corresponds to an estimated hazard approximately 20% lower for D + EP than EP under the fitted model. The endpoint is progression-free survival, so the event process differs from OS: the relevant event is progression or death according to the registered PFS definition.
The 95% CI of 0.665–0.959 indicates uncertainty around the estimated relative hazard. It does not give a range of median PFS values or individual patient outcomes.
The P-value of 0.0157 is evidence against the null hypothesis within the specified analysis. It does not measure effect size and should not be used alone to judge whether the observed difference is clinically meaningful.
D + T + EP versus EP
Global PFS hazard ratio
95% CI: 0.696–1.005 · P = 0.0568
The HR of 0.84 corresponds to an estimated hazard approximately 16% lower for D + T + EP than EP under the fitted model.
The confidence interval of 0.696–1.005 crosses 1. The estimate therefore has appreciable uncertainty about the direction of the relative hazard, even though the point estimate is below 1.
The P-value of 0.0568 should be interpreted as a measure of evidence under the specified hypothesis test, not as a measure of treatment effect size and not as evidence that the treatment strategies are equivalent.
D + T + EP versus D + EP
Global PFS hazard ratio
95% CI: 0.857–1.235 · P = 0.7540
The HR of 1.03 is close to 1, corresponding to an estimated hazard approximately 3% higher for D + T + EP than D + EP under the fitted model.
The 95% CI of 0.857–1.235 includes 1 and allows for both a lower and a higher hazard. The interval therefore communicates uncertainty that a point estimate alone cannot show.
The P-value of 0.7540 does not establish equivalence between the regimens. It indicates limited evidence against the null hypothesis for this particular comparison.
11. Objective Response Rate
ORR was analyzed as a binary endpoint using logistic regression. For the global cohort, the denominator was a subset of the FAS population who had measurable disease at baseline.
| Global comparison | Odds ratio | 95% CI | P-value |
|---|---|---|---|
| D + EP vs EP | 1.61 | 1.086–2.401 | 0.0177 |
| D + T + EP vs EP | 1.19 | 0.817–1.746 | 0.3611 |
D + EP versus EP
Global ORR odds ratio
95% CI: 1.086–2.401 · P = 0.0177
An OR of 1.61 means that the estimated odds of objective response were 1.61 times the odds in the EP group under the logistic regression model.
This does not mean that 61% more patients responded, nor does it mean that the response probability was 61 percentage points higher. Odds and probabilities are related but are not interchangeable.
The 95% CI of 1.086–2.401 describes uncertainty around the odds ratio. The P-value of 0.0177 measures evidence against the specified null hypothesis; it is not a measure of the magnitude of response improvement.
D + T + EP versus EP
Global ORR odds ratio
95% CI: 0.817–1.746 · P = 0.3611
The OR of 1.19 indicates estimated response odds 1.19 times those in EP under the logistic model.
The confidence interval of 0.817–1.746 includes 1, so the estimate is compatible with both lower and higher response odds. The interval also illustrates why the point estimate should not be interpreted without its uncertainty.
The P-value of 0.3611 does not establish that the treatments have the same response rate. It indicates limited evidence against the null hypothesis in this analysis.
12. China Cohort Overall Survival
The China cohort analyses were exploratory. The registry explicitly states that the study was not designed or powered to show statistical significance for efficacy endpoints in the China cohort.
| Analysis | Comparison | HR | 95% CI | P-value |
|---|---|---|---|---|
| First analysis | D + EP vs EP | 0.65 | 0.414–1.029 | 0.0664 |
| Second analysis | D + EP vs EP | 0.75 | 0.504–1.106 | 0.1455 |
| Second analysis | D + T + EP vs EP | 0.65 | 0.439–0.964 | 0.0314 |
| Second analysis | D + T + EP vs D + EP | 0.86 | 0.574–1.277 | 0.4470 |
D + EP versus EP — first China analysis
China OS hazard ratio
95% CI: 0.414–1.029 · P = 0.0664
The HR of 0.65 corresponds to an estimated hazard approximately 35% lower for D + EP than EP under the fitted model.
However, the 95% CI of 0.414–1.029 includes 1. More importantly, the registry states that the China cohort was not powered for formal statistical significance and that its analyses were exploratory. Therefore, this estimate should be interpreted primarily as an exploratory assessment of consistency rather than as an independently powered confirmatory test.
The P-value of 0.0664 is not a measure of the effect size and should not be interpreted as the probability that the treatment effect is absent.
D + EP versus EP — second China analysis
China OS hazard ratio
95% CI: 0.504–1.106 · P = 0.1455
The point estimate corresponds to an estimated hazard approximately 25% lower for D + EP than EP, but the 95% CI of 0.504–1.106 includes 1 and is therefore compatible with a range of relative effects.
Because the China cohort was not powered for formal statistical significance, the P-value of 0.1455 should not be used to turn this exploratory analysis into a separate confirmatory conclusion.
D + T + EP versus EP — second China analysis
China OS hazard ratio
95% CI: 0.439–0.964 · P = 0.0314
The HR of 0.65 corresponds to an estimated hazard approximately 35% lower for D + T + EP than EP under the fitted model.
The 95% CI of 0.439–0.964 is below 1 at its upper boundary, but the registry's explicit statement that the China cohort was not powered for formal statistical significance remains essential. The analysis was exploratory and intended to evaluate consistency with the global cohort.
The P-value of 0.0314 is therefore not appropriately interpreted as if it came from a separately powered confirmatory trial with the same inferential role as the global primary analysis.
D + T + EP versus D + EP — second China analysis
China OS hazard ratio
95% CI: 0.574–1.277 · P = 0.4470
The HR of 0.86 corresponds to an estimated hazard approximately 14% lower for D + T + EP than D + EP under the fitted model.
The 95% CI of 0.574–1.277 includes 1 and is wide enough to allow for both a lower and a higher hazard. The P-value of 0.4470 does not establish equivalence, and the exploratory status of the China cohort further limits confirmatory interpretation.
13. China Cohort Progression-Free Survival
| Comparison | HR | 95% CI | P-value |
|---|---|---|---|
| D + EP vs EP | 0.97 | 0.661–1.437 | 0.8934 |
| D + T + EP vs EP | 0.72 | 0.487–1.068 | 0.1035 |
| D + T + EP vs D + EP | 0.76 | 0.522–1.116 | 0.1673 |
All three China PFS analyses were exploratory under the registry's stated framework. Each used log-rank testing with a stratified Cox model for HR estimation and adjustment for planned platinum therapy.
D + EP vs EP
HR 0.97; 95% CI 0.661–1.437; P = 0.8934. The estimate is close to 1, while the confidence interval permits materially different relative hazards.
D + T + EP vs EP
HR 0.72; 95% CI 0.487–1.068; P = 0.1035. The point estimate is below 1, but the interval includes 1.
D + T + EP vs D + EP
HR 0.76; 95% CI 0.522–1.116; P = 0.1673. The point estimate favors the first-named group directionally, but the interval includes 1.
14. China Cohort Objective Response Rate
| Comparison | Odds ratio | 95% CI | P-value |
|---|---|---|---|
| D + EP vs EP | 1.39 | 0.610–3.244 | 0.4320 |
| D + T + EP vs EP | 2.07 | 0.874–5.118 | 0.0986 |
These ORR analyses used logistic regression. As with the China OS and PFS analyses, the registry states that the China cohort was not designed or powered for formal statistical significance and that these analyses were exploratory.
The OR of 1.39 for D + EP versus EP corresponds to estimated response odds 1.39 times those of EP. Its 95% CI of 0.610–3.244 is wide and includes 1.
The OR of 2.07 for D + T + EP versus EP corresponds to estimated response odds 2.07 times those of EP. Its 95% CI of 0.874–5.118 includes 1 and is also wide.
The P-values of 0.4320 and 0.0986 are hypothesis-test results, not measures of effect magnitude. Given the exploratory design and limited power of the China cohort, the estimates are most appropriately considered in the context of consistency with the global results rather than as independent confirmatory evidence.
15. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized cohort and treatment arm. The data are presented as affected patients divided by the number at risk.
| Cohort | Arm | Serious adverse events |
|---|---|---|
| Global | D + T + EP | 126 / 266 |
| Global | D + EP | 86 / 265 |
| Global | EP | 97 / 266 |
| China | D + T + EP | 31 / 65 |
| China | D + EP | 26 / 61 |
| China | EP | 22 / 62 |
The reported figures are counts of affected patients relative to the corresponding number at risk. They should not be converted into a different safety metric without the underlying definitions, follow-up periods, and exposure information.
16. Interim Analysis, Alpha Spending, and Multiplicity
CASPIAN provides a useful example of why interim analyses cannot simply be treated as an additional opportunity to inspect a conventional P-value. The global interim OS analysis used a Lan-DeMets alpha-spending function with an O'Brien-Fleming-type boundary.
| Design element | Registry-supported detail |
|---|---|
| Interim endpoint | Overall survival in the global cohort |
| Interim comparison | D + EP vs EP |
| Interim event target | Approximately 318 OS events |
| Alpha-spending framework | Lan-DeMets alpha spending |
| Boundary type | O'Brien-Fleming type |
| Interim significance boundary | 0.0178 |
| Overall alpha referenced | 4% |
| Final alpha adjustment | Generalized Haybittle-Peto method |
| Final significance boundary | 0.0418 |
| Final overall alpha referenced | 5% |
The final analysis for D + T + EP versus EP used an adjusted alpha level to account for actual alpha spent at the interim analysis and the actual final number of events, with the stated objective of maintaining overall type I error control.
Why this matters
If investigators could repeatedly examine an accumulating trial and declare success whenever a conventional threshold was crossed, the probability of a false-positive conclusion would increase. Alpha spending addresses this by defining how much type I error may be used as information accumulates.
The distinction is important when reading the reported P-values. The P-value is a property of the observed data under the specified test, whereas the prespecified boundary determines how that evidence is interpreted within the sequential design.
17. Stratified Analysis and Covariate Adjustment
The registry repeatedly identifies stratified analysis and covariate adjustment as components of the time-to-event methodology. The reported global Cox analyses specifically adjusted for planned platinum therapy in Cycle 1: carboplatin or cisplatin.
Why stratify?
Stratification allows the analysis to account for prespecified factors that can influence the event process without requiring the same baseline hazard function to apply across every stratum.
Why adjust?
Adjustment can improve the precision or preserve the intended treatment comparison when the adjusted variable is related to outcome and was specified as part of the analysis plan.
The registry's reported model also used the Efron approach for tied event times. Ties arise when multiple patients share the same recorded event time; the Efron method provides one way to approximate the partial likelihood contribution of tied events.
18. Statistical Methods Explained
Why was a log-rank test used?
OS and PFS are time-to-event endpoints, so the analysis must account for both the timing of events and right censoring. The log-rank test compares the event experience between groups across follow-up rather than reducing each patient to a simple binary event/no-event outcome.
What does an HR of 0.75 mean?
For the global final OS comparison of D + EP versus EP, an HR of 0.75 means the fitted model estimated approximately 25% lower instantaneous hazard for the D + EP group. It does not mean 25% fewer deaths in absolute terms and does not mean that every patient experienced a 25% reduction in their personal risk.
Why is the confidence interval as important as the HR?
A point estimate alone does not show how precisely the treatment effect has been estimated. The 95% CI provides a range describing statistical uncertainty around the estimated effect under the model and sampling framework. For example, the global final D + EP versus EP OS HR of 0.75 has a 95% CI of 0.625–0.910.
Why doesn't a P-value measure effect size?
A P-value quantifies how compatible the observed data are with a specified null hypothesis under the statistical model. It depends on both the magnitude of the observed effect and the amount of information in the analysis. A small P-value can therefore accompany a modest estimate in a large dataset, while a larger P-value can occur with a substantial point estimate when uncertainty is high.
Why was logistic regression used for ORR?
ORR is a binary outcome: a patient either met the prespecified response criterion or did not. Logistic regression models the probability of the binary outcome and naturally produces an odds ratio. The CASPIAN registry reports this method for the global and China ORR analyses.
Why does the China cohort require special interpretation?
The registry explicitly states that the China cohort was not designed or powered for formal statistical significance and that its analyses were exploratory. Therefore, the China estimates can describe the observed data and contribute to an assessment of consistency, but their P-values should not be interpreted as if they represented an independently powered confirmatory trial.
Why does the interim analysis change interpretation of the P-value?
An interim look creates a sequential-testing problem. CASPIAN used Lan-DeMets alpha spending with an O'Brien-Fleming-type boundary. The reported interim P-value therefore needs to be interpreted against the prespecified interim boundary rather than against an arbitrary conventional threshold.
19. Limitations and Interpretation Issues
- Exploratory China cohort: the registry states that the China cohort was not powered for formal assessment of statistical significance. Its efficacy and safety analyses were exploratory.
- Interim analysis: the global interim OS result was generated under a group-sequential framework. The appropriate inferential threshold was determined by alpha spending rather than by treating the interim P-value as an unadjusted final analysis.
- Multiple comparisons: CASPIAN contains several treatment contrasts, analysis stages, cohorts, and endpoints. The role of each comparison must therefore be distinguished rather than treating every reported P-value as equivalent evidence.
- Hazard-ratio assumptions: Cox HRs are model-based. A single HR is most naturally interpreted under a proportional-hazards framework; if hazards vary substantially over time, the HR can summarize rather than fully describe the treatment difference.
- Censoring: Kaplan-Meier and Cox methods depend on appropriate handling of censored observations. A patient censored at the last known date alive contributes follow-up information up to that point but is not treated as having experienced death.
- ORR denominator: the global ORR analysis used a subset of the FAS with measurable disease at baseline. Therefore, its analysis population differs conceptually from the full randomized population used for OS.
- Relative versus absolute effects: HRs and ORs are relative measures. They do not directly communicate absolute survival probabilities or absolute differences in response probability.
- Three-arm structure: the presence of two active strategies means that the D + T + EP versus D + EP comparison answers a different question from either comparison against EP.
20. Why This Trial Matters Statistically
CASPIAN is a useful teaching case because it combines randomized treatment comparisons, multiple treatment arms, time-to-event endpoints, binary response outcomes, interim monitoring, stratified Cox modeling, logistic regression, and an explicitly exploratory regional cohort.
| Concept | How it appears in CASPIAN |
|---|---|
| Randomization | Randomized phase 3 parallel-group design with 987 enrolled patients |
| Three-arm comparison | D + T + EP, D + EP, and EP |
| Kaplan-Meier estimation | Used for median OS and relevant time-to-event estimation |
| Log-rank test | Reported for OS and PFS comparisons |
| Hazard ratio | Primary relative effect measure for OS and secondary measure for PFS |
| Stratified Cox model | Used to calculate HRs and confidence intervals |
| Covariate adjustment | Platinum therapy in Cycle 1 was included in the reported Cox analyses |
| Logistic regression | Used for global and China ORR analyses |
| Odds ratio | Effect measure for the binary ORR endpoint |
| Interim analysis | Global OS interim analysis based on approximately 318 OS events |
| Alpha spending | Lan-DeMets function with an O'Brien-Fleming-type boundary |
| Multiplicity / sequential testing | Final alpha was adjusted after the interim analysis |
| Exploratory analysis | China cohort analyses were explicitly not powered for formal significance |
21. Overall Statistical Picture
The ClinicalTrials.gov record shows a consistent distinction between the global confirmatory framework and the exploratory China cohort. In the global cohort, the D + EP versus EP comparison produced OS HRs of 0.73 at interim analysis and 0.75 at final analysis, while the corresponding PFS HR was 0.80 and the ORR OR was 1.61.
The D + T + EP versus EP comparison produced a global final OS HR of 0.82 and a PFS HR of 0.84. The direct D + T + EP versus D + EP comparisons produced an OS HR of 1.08 and a PFS HR of 1.03. These comparisons illustrate why the reference group matters: the same treatment can have a different statistical interpretation depending on which randomized group serves as the comparator.
The China analyses are more uncertain and explicitly exploratory. Their confidence intervals are generally wider, and the registry cautions that the cohort was not powered for formal statistical significance. The appropriate statistical reading is therefore to examine the point estimates, confidence intervals, and direction of effects while retaining the stated limitation on inferential strength.
22. A Practical Guide to Reading the CASPIAN Results
Start with the estimand
Ask which patients, treatment contrast, endpoint, and analysis time point are being compared. D + EP versus EP is not the same question as D + T + EP versus D + EP.
Then read the effect measure
OS and PFS use hazard ratios, whereas ORR uses odds ratios. The numerical interpretation of 0.75 is therefore fundamentally different from the interpretation of 1.61.
Then read the confidence interval
The interval indicates precision and possible effect sizes under the model. It should be read alongside the point estimate rather than after it as an afterthought.
Finally read the design
Interim monitoring, alpha spending, multiple comparisons, stratification, and exploratory regional analyses determine how the numerical results should be interpreted.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NCT03043872 — CASPIAN.
- Linked publication: PubMed record — PMID 40331625.
- Linked publication: PubMed record — PMID 38811992.
- Linked publication: PubMed record — PMID 33285097.
- Linked publication: PubMed record — PMID 31590988.
Continue through the Clinical Biostats knowledge graph
Connect the endpoints and methods used in CASPIAN to deeper statistical tutorials and analysis tools.
26. Record Summary
CASPIAN provides a compact teaching example of several central clinical-trial statistical methods. The trial used randomized parallel-group allocation with three treatment arms and evaluated overall survival and progression-free survival as time-to-event outcomes, with logistic regression used for objective response rate. The reported global OS analyses incorporated log-rank testing, stratified Cox modeling, covariate adjustment for planned platinum therapy, and sequential alpha control through Lan-DeMets alpha spending with an O'Brien-Fleming-type boundary.
The numerical results need to be read in the context of their analysis stage and comparison. The global D + EP versus EP OS HR was 0.73 at interim analysis and 0.75 at final analysis. The global D + T + EP versus EP final OS HR was 0.82, while the direct D + T + EP versus D + EP comparison produced an HR of 1.08. PFS and ORR produced corresponding but distinct effect measures.
The China cohort demonstrates another important statistical principle: a numerical estimate and a formal confirmatory conclusion are not synonymous. The registry explicitly states that the China cohort was not powered for formal statistical significance and that its analyses were exploratory. Confidence intervals, analysis populations, study design, and prespecified inferential procedures therefore matter as much as the individual point estimates.