This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
TOPAZ-1 is a randomized, parallel-group, quadruple-masked phase 3 treatment trial in biliary tract neoplasms. The registry reports 810 enrolled participants, two intervention groups, three registered primary endpoints, and posted statistical analyses for overall survival, progression-free survival, and objective response rate.
| Feature | TOPAZ-1 |
|---|---|
| Trial name | TOPAZ-1 |
| ClinicalTrials.gov identifier | NCT03875235 |
| Phase | Phase 3 |
| Condition | Biliary Tract Neoplasms |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 810 |
| Interventions | Durvalumab; placebo |
| Lead sponsor | AstraZeneca |
| Sponsor type | Industry |
| Status | Active, not recruiting |
2. Clinical Question
The central statistical question is whether treatment with durvalumab plus gemcitabine/cisplatin differs from placebo plus gemcitabine/cisplatin with respect to the registered time-to-event and response outcomes.
Population
Patients represented by the registry condition of biliary tract neoplasms enrolled in the TOPAZ-1 phase 3 treatment trial.
Intervention
Durvalumab in combination with gemcitabine and cisplatin.
Comparator
Placebo in combination with gemcitabine and cisplatin.
Primary question
Does the durvalumab-containing regimen produce a superior treatment effect on the registered efficacy endpoints compared with the placebo-containing regimen?
The registry identifies superiority as the hypothesis type for the posted statistical analyses. This is important: the analysis is framed as testing whether the treatment groups differ in favor of the durvalumab-containing regimen, rather than testing whether two regimens are sufficiently similar for a non-inferiority claim.
3. Trial Design
Durvalumab combination
- Durvalumab
- Gemcitabine
- Cisplatin
Placebo combination
- Placebo
- Gemcitabine
- Cisplatin
Randomization is the foundation of the causal comparison. Because treatment assignment is randomized, the principal efficacy comparison can be made between treatment groups rather than treating differences in observed outcomes as evidence that the groups were self-selected. The quadruple-masked design additionally reduces the opportunity for knowledge of treatment assignment to influence trial conduct or assessment.
4. Endpoints
The registry identifies three primary endpoints, all involving overall survival. The first is overall survival itself; the other two are fixed-time overall-survival rates calculated from the Kaplan-Meier estimator.
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Overall Survival (OS) | From date of randomization until death due to any cause. Assessed up to maximum of approximately 27 months (from date of randomization to primary analysis data cut-off)). OS is defined as time from randomization until death due to any cause; patients not known to have died at analysis are censored at the last recorded date known alive. | Time-to-event |
| Overall Survival (OS) Rate at 18 Months | From date of randomization until death due to any cause. Calculated at 18 months using the Kaplan-Meier technique. | Time-to-event |
| Overall Survival (OS) Rate at 24 Months | From date of randomization until death due to any cause. Calculated at 24 months using the Kaplan-Meier technique. | Time-to-event |
Although the registry lists three primary endpoints, the posted formal statistical analysis is for Overall Survival (OS). The registry also posts secondary analyses for progression-free survival and objective response rate. Those secondary analyses provide complementary information about disease progression and tumor response, but they should not be relabeled as primary endpoints.
5. Analysis Populations and Statistical Framework
| Outcome | Analysis population | Comparison | Method | Effect measure |
|---|---|---|---|---|
| Overall Survival | Full analysis set | Durvalumab + gemcitabine + cisplatin vs placebo + gemcitabine + cisplatin | Log-rank test; stratified Cox model for HR | Hazard ratio |
| Progression-free Survival | Full analysis set | Durvalumab + gemcitabine + cisplatin vs placebo + gemcitabine + cisplatin | Log-rank test; stratified Cox model for HR | Hazard ratio |
| Objective Response Rate | Full analysis set — subjects with measurable disease at baseline | Durvalumab + gemcitabine + cisplatin vs placebo + gemcitabine + cisplatin | Cochran-Mantel-Haenszel test | Odds ratio |
The statistical framework combines two important classes of methods. Log-rank testing and Cox regression are appropriate for time-to-event outcomes, while the Cochran-Mantel-Haenszel (CMH) test is used for the categorical response endpoint. The registry also identifies covariate adjustment, intention-to-treat analysis, interim analysis/alpha spending, and stratified analysis as concepts in the time-to-event analyses.
6. Statistical Methodology
Kaplan-Meier estimation
The registry explicitly specifies the Kaplan-Meier technique for OS and for calculating OS rates at 18 and 24 months. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not have experienced the event by the time of analysis.
Here, di represents the number of events at an event time and ni represents the number at risk immediately before that time. The estimator therefore updates the estimated survival probability as events occur while accounting for the information contributed before censoring.
For TOPAZ-1, this matters particularly for the fixed-time OS endpoints. An OS rate at a specified time is a survival probability estimated from the entire observed time-to-event experience up to that point; it is not simply the proportion of enrolled participants who happen to be alive in a crude count.
Log-rank test
The registry reports the log-rank test for the OS and PFS comparisons. The log-rank test compares the observed and expected pattern of events between randomized groups across follow-up, rather than comparing only a single time point.
This is well matched to the structure of a time-to-event endpoint because participants can enter the risk set, experience an event, or become censored at different times. A conventional comparison of means would not use that information appropriately.
Stratified Cox proportional-hazards model
The registry states that the HR and confidence interval for OS and PFS were estimated from a stratified Cox proportional-hazards model adjusting for disease status and primary tumor location.
A hazard ratio is a relative, model-based measure of the event rate over follow-up. It is not a probability, a median, an absolute risk difference, or the percentage of participants who benefit.
The stratification and covariate adjustment are statistically important because the model is not simply comparing two unadjusted survival curves. The registry specifically identifies disease status and primary tumor location as adjustment variables for the reported HR and confidence interval.
Cochran-Mantel-Haenszel test
The ORR analysis uses a Cochran-Mantel-Haenszel test. The registry states that the odds ratio and confidence interval were estimated from a stratified CMH test adjusting for disease status and primary tumor location.
The CMH framework is useful when a binary outcome is compared across treatment groups while accounting for a stratification factor or set of strata. In TOPAZ-1, the resulting effect measure is an odds ratio rather than a hazard ratio.
Full analysis set
The posted efficacy analyses use the full analysis set. For ORR, the registry specifies the full analysis set restricted to subjects with measurable disease at baseline. This distinction matters because the population contributing to an objective response analysis is not necessarily identical to the population conceptually eligible for a time-to-event endpoint.
7. Primary Result: Overall Survival
The registry reports a formal primary statistical analysis of overall survival using the full analysis set. The groups compared were durvalumab plus gemcitabine/cisplatin versus placebo plus gemcitabine/cisplatin.
Hazard ratio for overall survival
97% two-sided CI: 0.64–0.99 · P = 0.021
Log-rank test; superiority hypothesis
| Primary endpoint | Analysis | Estimate | Confidence interval | P-value |
|---|---|---|---|---|
| Overall Survival (OS) | Log-rank test; stratified Cox proportional-hazards model | HR 0.80 | 97% two-sided CI 0.64–0.99 | 0.021 |
The registry notes that the 2-sided significance level for OS at the second interim analysis was 3%. This is an important design feature: the reported P-value should be interpreted in the context of the prespecified interim-analysis framework rather than as if it came from an unplanned single final look at the data.
An HR of 0.80 means that, under the fitted time-to-event model, the estimated instantaneous rate of death in the durvalumab-containing group was 0.80 times that in the placebo-containing group over the analyzed follow-up. Equivalently, 0.80 corresponds to a 20% lower estimated hazard under the model.
That interpretation does not mean that 20% of patients avoided death, that each patient experienced exactly a 20% reduction in risk, or that the absolute probability of death was reduced by 20 percentage points. A hazard ratio is a relative time-to-event measure.
The 97% two-sided confidence interval of 0.64–0.99 describes statistical uncertainty around the estimated hazard ratio under the model and analysis framework. It does not describe the range of effects experienced by individual patients.
The P-value of 0.021 is evidence against the null hypothesis in the prespecified superiority framework, but it is not a measure of effect size. Effect size is conveyed by the HR, while the confidence interval describes precision. The P-value also does not say that there is a particular probability that the null hypothesis is true.
Because the analysis uses a Cox proportional-hazards model, interpretation of a single HR should also be considered in light of the proportional-hazards assumption. The registry does not provide a proportional-hazards diagnostic in the ClinicalTrials.gov record, so the HR should not be interpreted as a guarantee that the relative hazard was constant at every time point.
Why the 97% confidence interval matters
The unusual confidence level is directly relevant to interpretation. The registry reports a 97% two-sided confidence interval for the primary OS HR, rather than the more familiar 95% interval. That choice should be retained when describing the formal primary analysis because the inferential framework includes interim monitoring and alpha allocation.
The interval of 0.64–0.99 is relatively close to the null value of 1.00 at its upper boundary. Thus, the numerical precision of the estimate matters: the observed HR is 0.80, but the interval communicates that the compatible values under the stated statistical framework extend considerably below 1 and approach 1.
Why the interim-analysis alpha matters
The registry states that the second interim OS analysis used a 2-sided significance level of 3%. Without such prespecification, repeatedly inspecting accumulating data can increase the probability of a false-positive conclusion. An interim boundary is intended to preserve the overall type I error while permitting a trial to be evaluated before all planned information has accumulated.
8. Secondary Result: Progression-Free Survival
Progression-free survival was analyzed in the full analysis set using the log-rank test. The registry describes tumor assessments every 6 weeks after randomization for the first 24 weeks and then every 8 weeks thereafter until the date specified in the registry outcome record.
Hazard ratio for progression-free survival
95.19% two-sided CI: 0.63–0.89 · P = 0.001
Log-rank test; superiority hypothesis
| Secondary endpoint | Analysis | Estimate | Confidence interval | P-value |
|---|---|---|---|---|
| Progression-free Survival (PFS) | Log-rank test; stratified Cox proportional-hazards model | HR 0.75 | 95.19% two-sided CI 0.63–0.89 | 0.001 |
The registry states that the 2-sided significance level for PFS at the second interim analysis was 4.81%. The HR and confidence interval were estimated from the stratified Cox proportional-hazards model adjusting for disease status and primary tumor location.
An HR of 0.75 corresponds to a 25% lower estimated hazard of the PFS event in the durvalumab-containing group under the fitted model. The relevant event is defined by the trial's PFS endpoint framework; the HR should therefore be understood as a time-to-event comparison rather than a simple percentage of patients whose tumors progressed.
The 95.19% two-sided confidence interval of 0.63–0.89 quantifies uncertainty around the estimated HR. It remains below 1.00 throughout the reported interval, but the interval is not a prediction interval for individual patients and does not indicate that every patient has the same relative treatment effect.
The P-value of 0.001 provides evidence against the superiority null hypothesis under the prespecified interim-analysis framework. It does not indicate that the effect is "0.001 in size," nor does a smaller P-value by itself establish a larger or more clinically important effect.
As with OS, the HR comes from a Cox proportional-hazards model. The proportional-hazards assumption is therefore relevant when reducing the entire follow-up experience to a single HR. The ClinicalTrials.gov record does not provide a formal diagnostic of that assumption.
Why PFS and OS should not be conflated
PFS and OS are both time-to-event outcomes, but they answer different questions. PFS focuses on the time until progression or death under the trial's endpoint definition, whereas OS measures time from randomization until death from any cause. A treatment can therefore have different numerical effects on the two endpoints.
The TOPAZ-1 registry results illustrate this distinction directly: the posted HR is 0.75 for PFS and 0.80 for OS. Those numbers should be interpreted separately rather than treated as interchangeable measures of the same outcome.
9. Secondary Result: Objective Response Rate
Objective response rate was analyzed as a binary endpoint in the full analysis set among subjects with measurable disease at baseline. Tumor assessments were performed per RECIST 1.1 every 6 weeks for the first 24 weeks relative to randomization and then at the subsequent schedule specified in the registry outcome record.
Odds ratio for objective response
95% two-sided CI: 1.11–2.31 · P = 0.011
Cochran-Mantel-Haenszel test; superiority hypothesis
| Secondary endpoint | Analysis | Estimate | Confidence interval | P-value |
|---|---|---|---|---|
| Objective Response Rate (ORR) | Cochran-Mantel-Haenszel test | OR 1.60 | 95% two-sided CI 1.11–2.31 | 0.011 |
The registry states that the OR and confidence interval were estimated from a stratified CMH test adjusting for disease status and primary tumor location.
An odds ratio of 1.60 means that the estimated odds of objective response were 1.60 times as high in the durvalumab-containing group as in the placebo-containing group under the stratified analysis. This is an odds comparison, not a risk ratio.
The distinction matters because an odds ratio of 1.60 does not mean that 60% of patients responded, nor does it mean that the probability of response increased by 60%. The underlying response probabilities would be required to translate the odds ratio into an absolute risk difference.
The 95% confidence interval of 1.11–2.31 quantifies uncertainty around the estimated odds ratio. The entire reported interval is above 1.00, but the interval does not describe the range of individual patient responses.
The P-value of 0.011 addresses evidence against the null hypothesis within the specified statistical framework. It does not measure the magnitude or clinical importance of the response effect. The odds ratio and confidence interval are the more direct measures of the estimated effect and its precision.
Why ORR uses a different statistical method
Unlike OS and PFS, ORR is a binary outcome: a participant either meets the prespecified response definition or does not. There is no event-time structure requiring Kaplan-Meier estimation for the primary ORR comparison. The registry therefore uses the CMH framework and reports an odds ratio rather than a hazard ratio.
10. Bringing the Efficacy Results Together
| Endpoint | Role | Effect measure | Estimate | CI | P-value |
|---|---|---|---|---|---|
| Overall Survival | Primary | Hazard ratio | 0.80 | 97% two-sided: 0.64–0.99 | 0.021 |
| Progression-free Survival | Secondary | Hazard ratio | 0.75 | 95.19% two-sided: 0.63–0.89 | 0.001 |
| Objective Response Rate | Secondary | Odds ratio | 1.60 | 95% two-sided: 1.11–2.31 | 0.011 |
These three estimates describe different dimensions of the same randomized comparison. The OS and PFS HRs are both below 1.00, indicating lower estimated event hazards for the durvalumab-containing regimen under their respective Cox models. The ORR odds ratio is above 1.00, indicating higher estimated odds of response under the stratified categorical analysis.
The estimates should not be collapsed into one overall numerical "benefit." Each endpoint has a different definition, statistical scale, censoring structure, and clinical interpretation. In particular, the ORR odds ratio cannot be compared numerically with the OS or PFS hazard ratios as if all three were measuring the same quantity.
11. Interim Analysis and Alpha Spending
Interim analysis is an explicit part of the TOPAZ-1 statistical framework reported by the registry. The OS analysis notes that the second interim analysis was pre-specified after approximately 397 OS events occurred.
OS interim threshold
The 2-sided significance level for OS at the second interim analysis was 3%.
PFS interim threshold
The 2-sided significance level for PFS at the second interim analysis was 4.81%.
These thresholds illustrate why a P-value cannot be interpreted without knowing the design under which it was generated. A trial that looks at accumulating data repeatedly needs a prespecified error-control strategy. Otherwise, the probability of obtaining at least one apparently positive result by chance can exceed the intended type I error rate.
The registry data identify interim analysis / alpha spending as a concept in the posted analyses. This page therefore treats the reported P-values in the context of the interim design rather than applying a generic single-analysis interpretation.
12. Covariate Adjustment and Stratification
The posted OS and PFS analyses use a stratified Cox proportional-hazards model adjusting for disease status and primary tumor location. The ORR analysis similarly uses a stratified CMH test adjusting for these variables.
Adjustment does not change the basic randomized comparison into an observational analysis. Instead, the model incorporates prespecified covariate information into estimation of the treatment effect. For a stratified model, the treatment effect is estimated while allowing the baseline event process to differ across the specified strata.
The fact that the same two variables appear in the posted analyses is also useful pedagogically: it demonstrates that statistical adjustment is not synonymous with adding arbitrary variables to a regression model. The variables used here are explicitly identified in the registry's analysis notes.
13. Missing Data, Censoring, and Time-to-Event Interpretation
The OS definition explicitly addresses participants who have not died by the time of analysis. Such patients are censored based on the last recorded date on which they were known to be alive. This is a standard right-censoring structure for survival analysis.
Censoring does not mean that the participant is treated as having survived forever. Instead, the participant contributes observed survival information up to the censoring time. The validity of the resulting survival analysis depends on assumptions about the relationship between censoring and the underlying event process.
A participant contributes information until the first observed event or censoring time. Participants who remain alive at analysis therefore still contribute information even though their ultimate survival time is not yet observed.
The registry's the ClinicalTrials.gov record do not describe a separate imputation procedure for missing OS outcomes. That is not surprising for a standard time-to-event endpoint: censoring is incorporated directly into Kaplan-Meier and Cox methods rather than replacing every censored observation with an imputed survival time.
14. Statistical Methods Explained
Why was a log-rank test used for OS and PFS?
OS and PFS are time-to-event endpoints, so participants can experience the event at different times or be censored before the event is observed. The log-rank test compares the event experience between treatment groups across the follow-up period and therefore uses the ordering of event times rather than reducing the outcome to a single binary status.
What does an OS hazard ratio of 0.80 mean?
Under the reported Cox model, an HR of 0.80 means the estimated instantaneous death rate in the durvalumab-containing group was 0.80 times that of the placebo-containing group. It corresponds to a 20% lower estimated hazard. It does not mean a 20-percentage-point improvement in survival or that 20% of participants avoided death.
Why is the OS confidence interval 97% rather than 95%?
The registry reports a 97% two-sided confidence interval for the primary OS HR and states that the second interim analysis used a 2-sided significance level of 3%. The confidence level therefore reflects the prespecified interim-analysis inferential framework rather than a generic preference for a particular confidence level.
Why does the PFS analysis use a different confidence level?
The registry reports a 95.19% two-sided confidence interval for PFS and states that the 2-sided significance level for PFS at the second interim analysis was 4.81%. These values sum to the complementary relationship expected between the stated two-sided alpha level and confidence level.
Why is ORR reported with an odds ratio instead of a hazard ratio?
ORR is a binary outcome, whereas OS and PFS are time-to-event outcomes. The registry therefore uses a stratified Cochran-Mantel-Haenszel analysis for ORR and reports an odds ratio. The statistical scale changes because the underlying endpoint structure changes.
What does an OR of 1.60 mean?
An OR of 1.60 means that the estimated odds of objective response were 1.60 times as high in the durvalumab-containing group as in the placebo-containing group under the stratified analysis. It does not mean that the response probability increased by exactly 60%, because odds and probabilities are different quantities.
Why does interim analysis change how a P-value should be read?
If a trial is examined more than once while data accumulate, an unadjusted testing strategy can increase the chance of a false-positive conclusion. TOPAZ-1's registry-reported analysis notes explicitly identify an interim-analysis framework and endpoint-specific significance levels. Therefore, the reported P-values should be interpreted against those prespecified thresholds rather than against an automatically assumed single-look threshold.
15. Safety Results
The ClinicalTrials.gov record includes serious adverse events by treatment arm. The reported counts are presented as affected participants over participants at risk.
| Serious adverse events | Affected / at risk |
|---|---|
| Durvalumab + Gemcitabine + Cisplatin | 168 / 338 |
| Placebo + Gemcitabine + Cisplatin | 152 / 342 |
These figures should be read as serious adverse-event counts relative to the reported at-risk denominators. They are not the same endpoint as overall adverse events, grade-specific adverse events, or treatment discontinuations.
The safety analysis is also conceptually different from the efficacy analysis. Efficacy is anchored to the randomized comparison and the full analysis set specified in the posted analyses. Safety is inherently related to treatment exposure, so the denominator and analysis population used for a particular safety table must be respected rather than inferred from the efficacy population.
16. Trial Timeline
Trial start
The registry lists April 16, 2019 as the study start date.
Primary completion
The registry lists August 11, 2021 as the primary completion date.
Active, not recruiting
The ClinicalTrials.gov record identifies the trial status as ACTIVE_NOT_RECRUITING.
The registry also identifies a second interim analysis for OS after approximately 397 OS events. This event-driven structure is typical of time-to-event trials: information accumulates according to the number of observed events rather than simply according to elapsed calendar time.
17. Primary Endpoint Interpretation in Detail
The reported HR of 0.80 summarizes the relative death hazard estimated from the stratified Cox model. Its direction is below 1.00, corresponding to a lower estimated hazard in the durvalumab-containing group.
The HR does not directly state the median survival, the absolute survival probability at a specific time, the number of deaths in each group, or the proportion of patients who personally benefited. None of those additional numerical quantities are reported in the ClinicalTrials.gov record used for this page.
The 97% two-sided interval of 0.64–0.99 communicates the uncertainty surrounding the HR estimate under the specified analysis. The upper endpoint approaching 1.00 is relevant to precision even though the entire reported interval remains below 1.00.
The P-value of 0.021 quantifies the compatibility of the observed test statistic with the null hypothesis under the statistical model and prespecified testing framework. It is not the probability that the null hypothesis is true and it is not a measure of clinical effect size.
18. Interpreting the Secondary Endpoints Without Overstating Them
The secondary results provide a coherent set of statistical signals: the PFS HR is below 1.00 and the ORR odds ratio is above 1.00. But statistical interpretation should preserve the distinction between endpoint scales.
Time-to-event evidence
The PFS HR of 0.75 describes a relative difference in the estimated event hazard under the Cox model.
Response evidence
The OR of 1.60 describes a relative difference in the odds of objective response under the stratified CMH analysis.
Different estimands
An HR and an OR are not directly comparable numerical measures. Each answers a different statistical question.
Analysis populations
ORR is restricted to the full analysis set among subjects with measurable disease at baseline, while OS and PFS use the full analysis set.
This is a useful example of why a trial should not be summarized by a single "positive" or "negative" statistic. The statistical story consists of endpoint definitions, analysis populations, effect measures, confidence intervals, hypothesis-testing thresholds, and the design that generated the data.
19. Multiplicity and Endpoint Hierarchy
TOPAZ-1 has three registered primary endpoints: OS, OS rate at 18 months, and OS rate at 24 months. The registry-reported formal statistical analyses include one primary endpoint analysis, specifically the OS analysis with HR 0.80, 97% two-sided CI 0.64–0.99, and P = 0.021.
| Endpoint / analysis | Role in the ClinicalTrials.gov record | Posted statistical result |
|---|---|---|
| Overall Survival | Primary endpoint | Formal analysis posted |
| OS Rate at 18 Months | Primary endpoint | Registry states it is calculated using Kaplan-Meier; no separate statistical analysis is reported in the ClinicalTrials.gov record used here |
| OS Rate at 24 Months | Primary endpoint | Registry states it is calculated using Kaplan-Meier; no separate statistical analysis is reported in the ClinicalTrials.gov record used here |
| Progression-free Survival | Secondary endpoint | HR 0.75; 95.19% two-sided CI 0.63–0.89; P = 0.001 |
| Objective Response Rate | Secondary endpoint | OR 1.60; 95% two-sided CI 1.11–2.31; P = 0.011 |
The distinction between "registered primary endpoint" and "posted formal statistical analysis" is important. A registry can contain several registered outcome measures while only some of them have an accompanying formal statistical-analysis record. The absence of a separate posted analysis in the ClinicalTrials.gov record should not be filled with an inferred test or a reconstructed P-value.
20. Limitations
- Registry-level evidence: This analysis is deliberately limited to the ClinicalTrials.gov record and the statistical information contained there. It does not import numerical results from external publications.
- Incomplete endpoint time-frame text: The registry-reported OS registry time frame ends with the phrase "from date of" in the maximum approximately 27-month description. The missing text is not reconstructed.
- Limited primary-endpoint analysis detail: Three primary endpoints are registered, but the ClinicalTrials.gov record contains one formal primary endpoint analysis. The fixed-time OS endpoints are therefore described according to their registry definitions without inventing additional inferential results.
- Hazard-ratio assumptions: Cox proportional-hazards models rely on assumptions about the relationship of hazards over time. The ClinicalTrials.gov record does not provide a formal diagnostic for proportional hazards.
- Censoring: OS uses right-censoring for patients not known to have died at the time of analysis. As with any survival analysis, interpretation depends on the assumptions underlying the censoring mechanism.
- Analysis population: The efficacy results are based on the full analysis set, while ORR additionally requires measurable disease at baseline. Results should not be generalized to an unspecified population.
- Interim testing: The reported OS and PFS P-values were generated within an interim-analysis framework with endpoint-specific significance levels. They should not be interpreted as generic single-look P-values.
- Safety scope: The ClinicalTrials.gov record provides serious adverse-event counts by arm but not a complete safety dataset. Additional adverse-event categories cannot be inferred.
- No reconstructed results: Median survival, subgroup results, event counts beyond those explicitly reported, absolute response rates, or additional follow-up estimates are not added when they are absent from the trial data.
21. Why This Trial Matters Statistically
TOPAZ-1 is a useful statistical teaching case because it combines randomized treatment comparison, quadruple masking, multiple time-to-event endpoints, a binary response endpoint, stratified analysis, covariate adjustment, interim monitoring, and different effect measures within the same trial.
| Concept | How it appears in TOPAZ-1 |
|---|---|
| Randomization | The trial is randomized, supporting comparison of the assigned treatment groups. |
| Blinding | The registry identifies quadruple masking. |
| Kaplan-Meier estimation | Used for OS and specifically for OS rates at 18 and 24 months. |
| Log-rank test | Used for the posted OS and PFS comparisons. |
| Hazard ratio | Reported for OS and PFS from stratified Cox proportional-hazards models. |
| Confidence intervals | Reported alongside the OS, PFS, and ORR effect estimates. |
| Interim analysis | The second interim analysis was prespecified after approximately 397 OS events. |
| Alpha spending | The registry-reported analysis notes specify endpoint-specific interim significance levels. |
| Covariate adjustment | Disease status and primary tumor location were used in the reported adjusted analyses. |
| Cochran-Mantel-Haenszel test | Used for the binary ORR analysis. |
| Odds ratio | Used as the effect measure for ORR. |
| Time-to-event endpoints | OS and PFS incorporate event times and censoring rather than simple binary follow-up status. |
The most instructive feature is the contrast between the hazard ratio and the odds ratio. A reader who understands why those measures are different is already equipped to interpret much of the statistical structure of the trial correctly.
22. Statistical Methods Explained: A Deeper Walkthrough
Why is randomization important before any statistical test is performed?
Randomization establishes the treatment assignment mechanism before outcomes are observed. Statistical testing then evaluates the observed treatment-group differences under that randomized framework. The statistical model does not create the randomization; it analyzes the information generated by it.
Why use a stratified Cox model rather than only a raw Kaplan-Meier comparison?
Kaplan-Meier estimation describes survival over time, while the Cox model provides a quantitative relative-effect estimate through the hazard ratio. Stratification and adjustment also allow the reported analysis to incorporate disease status and primary tumor location as specified in the registry.
Why can a confidence interval be more informative than a P-value?
The P-value primarily addresses evidence against a null hypothesis. The confidence interval also shows the range of effect estimates compatible with the statistical model at the stated confidence level. For TOPAZ-1, the OS interval of 0.64–0.99 communicates both the direction of the estimated effect and its uncertainty.
Why does the OS analysis have a 97% CI?
The confidence level is connected to the prespecified interim-testing framework. The registry reports a 2-sided OS significance level of 3% at the second interim analysis, corresponding to a 97% two-sided confidence level for the reported HR.
Why is the PFS confidence level 95.19%?
The registry reports a 2-sided PFS significance level of 4.81% at the second interim analysis. The complementary confidence level is therefore 95.19%. This is another demonstration that confidence intervals and hypothesis tests are two expressions of the same underlying inferential framework.
Why is the ORR analysis restricted to measurable disease at baseline?
An objective response requires an evaluable baseline disease burden against which tumor response can be assessed. The registry explicitly defines the ORR analysis population as the full analysis set among subjects with measurable disease at baseline.
Why should OS and PFS not be treated as independent pieces of evidence without considering the trial design?
Both endpoints arise from the same randomized participants and are evaluated within the same clinical-trial framework. Their statistical interpretation therefore depends on the prespecified endpoint structure, interim monitoring, and any multiplicity strategy. Simply counting the number of P-values below a conventional threshold would ignore that design.
23. What the Trial's Effect Measures Mean
HR 0.80 for OS
The estimated instantaneous death hazard was 0.80 times that of the comparator under the reported stratified Cox model.
HR 0.75 for PFS
The estimated PFS event hazard was 0.75 times that of the comparator under the reported stratified Cox model.
OR 1.60 for ORR
The estimated odds of objective response were 1.60 times those of the comparator under the stratified CMH analysis.
P-values
The P-values quantify evidence against the respective null hypotheses within the stated inferential framework; they are not effect-size measures.
These measures also have different null values. The null value for a hazard ratio is 1.00, and the null value for an odds ratio is also 1.00. Values below 1 for the HR indicate lower estimated event hazard in the treatment group, whereas values above 1 for the OR indicate higher estimated odds of response in the treatment group.
24. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The reported OS HR is 0.80 with a 97% two-sided CI of 0.64–0.99 and P = 0.021. The PFS HR is 0.75 with a 95.19% two-sided CI of 0.63–0.89 and P = 0.001. ORR has an OR of 1.60 with a 95% two-sided CI of 1.11–2.31 and P = 0.011.
Clinical interpretation
The statistical results indicate differences between the randomized treatment groups on the reported efficacy endpoints. Clinical meaning requires consideration of the endpoint definitions, magnitude and precision of effects, treatment context, and safety rather than relying on P-values alone.
The distinction is deliberate. Statistical significance and clinical importance are related but not identical concepts. A statistical analysis can establish evidence against a null hypothesis without, by itself, defining how meaningful the treatment difference is for an individual patient.
25. Related Tutorials
Learn more about the methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: NCT03875235 — TOPAZ-1. Official registry record for the trial design, endpoints, posted results, and statistical analyses used on this page.
- PubMed record: PMID 42424063.
- PubMed record: PMID 40622010.
- PubMed record: PMID 40381735.
- PubMed record: PMID 38823398.
- PubMed record: PMID 38697156.
Continue through the Clinical Biostats statistical tutorials
Explore the survival-analysis, categorical-data, confidence-interval, and clinical-trial methods that appear in TOPAZ-1 and other randomized studies.
28. Record Summary
TOPAZ-1 provides a compact example of how modern randomized-trial statistics combine several complementary methods. The trial is randomized, parallel, quadruple-masked, and designed for treatment comparison. Its registered primary endpoints are time-to-event outcomes based on overall survival, including fixed-time OS rates calculated by Kaplan-Meier estimation. The posted formal OS analysis uses a log-rank test and a stratified Cox proportional-hazards model, with an HR of 0.80, a 97% two-sided confidence interval of 0.64–0.99, and P = 0.021.
The secondary analyses add two different statistical perspectives. PFS uses the same broad survival-analysis framework and reports an HR of 0.75 with a 95.19% two-sided confidence interval of 0.63–0.89 and P = 0.001. ORR is analyzed as a binary endpoint using a stratified Cochran-Mantel-Haenszel test and reports an OR of 1.60 with a 95% two-sided confidence interval of 1.11–2.31 and P = 0.011.
The statistical interpretation depends on more than the point estimates. The analysis populations, censoring rules, covariate adjustment, interim significance levels, confidence levels, and distinction between hazard ratios and odds ratios all affect what the reported numbers mean. The safety data in the ClinicalTrials.gov record further show why efficacy and safety must be examined as separate statistical dimensions rather than compressed into a single measure.