← Clinical Trials
Biliary Tract Cancer Phase 3 Survival Analysis NCT03875235

TOPAZ-1: Complete Statistical Analysis of Durvalumab in Advanced Biliary Tract Cancer

An independent statistical review of the randomized phase 3 TOPAZ-1 trial evaluating durvalumab plus gemcitabine/cisplatin versus placebo plus gemcitabine/cisplatin in patients with advanced biliary tract cancer.

Trial start: April 16, 2019  ·  Primary completion: August 11, 2021  ·  Enrollment: 810
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

TOPAZ-1 is a randomized, parallel-group, quadruple-masked phase 3 treatment trial in biliary tract neoplasms. The registry reports 810 enrolled participants, two intervention groups, three registered primary endpoints, and posted statistical analyses for overall survival, progression-free survival, and objective response rate.

810
Enrollment
ClinicalTrials.gov record
2
Arms
Parallel design
0.80
OS HR
97% CI 0.64–0.99
0.75
PFS HR
95.19% CI 0.63–0.89
FeatureTOPAZ-1
Trial nameTOPAZ-1
ClinicalTrials.gov identifierNCT03875235
PhasePhase 3
ConditionBiliary Tract Neoplasms
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment810
InterventionsDurvalumab; placebo
Lead sponsorAstraZeneca
Sponsor typeIndustry
StatusActive, not recruiting

2. Clinical Question

The central statistical question is whether treatment with durvalumab plus gemcitabine/cisplatin differs from placebo plus gemcitabine/cisplatin with respect to the registered time-to-event and response outcomes.

Population

Patients represented by the registry condition of biliary tract neoplasms enrolled in the TOPAZ-1 phase 3 treatment trial.

Intervention

Durvalumab in combination with gemcitabine and cisplatin.

Comparator

Placebo in combination with gemcitabine and cisplatin.

Primary question

Does the durvalumab-containing regimen produce a superior treatment effect on the registered efficacy endpoints compared with the placebo-containing regimen?

The registry identifies superiority as the hypothesis type for the posted statistical analyses. This is important: the analysis is framed as testing whether the treatment groups differ in favor of the durvalumab-containing regimen, rather than testing whether two regimens are sufficiently similar for a non-inferiority claim.

3. Trial Design

01
Randomize 810 enrolled
02
Parallel groups Two treatment groups
03
Quadruple masked Registry design
04
Assess outcomes Survival and response
05
Analyze FAS efficacy set
ARM A

Durvalumab combination

  • Durvalumab
  • Gemcitabine
  • Cisplatin
ARM B

Placebo combination

  • Placebo
  • Gemcitabine
  • Cisplatin
Allocation
Randomized
Design model
Parallel
Masking
Quadruple
Primary purpose
Treatment

Randomization is the foundation of the causal comparison. Because treatment assignment is randomized, the principal efficacy comparison can be made between treatment groups rather than treating differences in observed outcomes as evidence that the groups were self-selected. The quadruple-masked design additionally reduces the opportunity for knowledge of treatment assignment to influence trial conduct or assessment.

4. Endpoints

The registry identifies three primary endpoints, all involving overall survival. The first is overall survival itself; the other two are fixed-time overall-survival rates calculated from the Kaplan-Meier estimator.

EndpointRegistry definition / time frameEndpoint type
Overall Survival (OS) From date of randomization until death due to any cause. Assessed up to maximum of approximately 27 months (from date of randomization to primary analysis data cut-off)). OS is defined as time from randomization until death due to any cause; patients not known to have died at analysis are censored at the last recorded date known alive. Time-to-event
Overall Survival (OS) Rate at 18 Months From date of randomization until death due to any cause. Calculated at 18 months using the Kaplan-Meier technique. Time-to-event
Overall Survival (OS) Rate at 24 Months From date of randomization until death due to any cause. Calculated at 24 months using the Kaplan-Meier technique. Time-to-event

Although the registry lists three primary endpoints, the posted formal statistical analysis is for Overall Survival (OS). The registry also posts secondary analyses for progression-free survival and objective response rate. Those secondary analyses provide complementary information about disease progression and tumor response, but they should not be relabeled as primary endpoints.

5. Analysis Populations and Statistical Framework

OutcomeAnalysis populationComparisonMethodEffect measure
Overall Survival Full analysis set Durvalumab + gemcitabine + cisplatin vs placebo + gemcitabine + cisplatin Log-rank test; stratified Cox model for HR Hazard ratio
Progression-free Survival Full analysis set Durvalumab + gemcitabine + cisplatin vs placebo + gemcitabine + cisplatin Log-rank test; stratified Cox model for HR Hazard ratio
Objective Response Rate Full analysis set — subjects with measurable disease at baseline Durvalumab + gemcitabine + cisplatin vs placebo + gemcitabine + cisplatin Cochran-Mantel-Haenszel test Odds ratio

The statistical framework combines two important classes of methods. Log-rank testing and Cox regression are appropriate for time-to-event outcomes, while the Cochran-Mantel-Haenszel (CMH) test is used for the categorical response endpoint. The registry also identifies covariate adjustment, intention-to-treat analysis, interim analysis/alpha spending, and stratified analysis as concepts in the time-to-event analyses.

6. Statistical Methodology

Kaplan-Meier estimation

The registry explicitly specifies the Kaplan-Meier technique for OS and for calculating OS rates at 18 and 24 months. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not have experienced the event by the time of analysis.

Conceptual survival estimator
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at an event time and ni represents the number at risk immediately before that time. The estimator therefore updates the estimated survival probability as events occur while accounting for the information contributed before censoring.

For TOPAZ-1, this matters particularly for the fixed-time OS endpoints. An OS rate at a specified time is a survival probability estimated from the entire observed time-to-event experience up to that point; it is not simply the proportion of enrolled participants who happen to be alive in a crude count.

Log-rank test

The registry reports the log-rank test for the OS and PFS comparisons. The log-rank test compares the observed and expected pattern of events between randomized groups across follow-up, rather than comparing only a single time point.

This is well matched to the structure of a time-to-event endpoint because participants can enter the risk set, experience an event, or become censored at different times. A conventional comparison of means would not use that information appropriately.

Stratified Cox proportional-hazards model

The registry states that the HR and confidence interval for OS and PFS were estimated from a stratified Cox proportional-hazards model adjusting for disease status and primary tumor location.

Hazard-ratio interpretation
HR < 1  →  lower estimated instantaneous event rate in the durvalumab group

A hazard ratio is a relative, model-based measure of the event rate over follow-up. It is not a probability, a median, an absolute risk difference, or the percentage of participants who benefit.

The stratification and covariate adjustment are statistically important because the model is not simply comparing two unadjusted survival curves. The registry specifically identifies disease status and primary tumor location as adjustment variables for the reported HR and confidence interval.

Cochran-Mantel-Haenszel test

The ORR analysis uses a Cochran-Mantel-Haenszel test. The registry states that the odds ratio and confidence interval were estimated from a stratified CMH test adjusting for disease status and primary tumor location.

The CMH framework is useful when a binary outcome is compared across treatment groups while accounting for a stratification factor or set of strata. In TOPAZ-1, the resulting effect measure is an odds ratio rather than a hazard ratio.

Full analysis set

The posted efficacy analyses use the full analysis set. For ORR, the registry specifies the full analysis set restricted to subjects with measurable disease at baseline. This distinction matters because the population contributing to an objective response analysis is not necessarily identical to the population conceptually eligible for a time-to-event endpoint.

7. Primary Result: Overall Survival

The registry reports a formal primary statistical analysis of overall survival using the full analysis set. The groups compared were durvalumab plus gemcitabine/cisplatin versus placebo plus gemcitabine/cisplatin.

Hazard ratio for overall survival

0.80

97% two-sided CI: 0.64–0.99   ·   P = 0.021

Log-rank test; superiority hypothesis

Primary endpointAnalysisEstimateConfidence intervalP-value
Overall Survival (OS) Log-rank test; stratified Cox proportional-hazards model HR 0.80 97% two-sided CI 0.64–0.99 0.021

The registry notes that the 2-sided significance level for OS at the second interim analysis was 3%. This is an important design feature: the reported P-value should be interpreted in the context of the prespecified interim-analysis framework rather than as if it came from an unplanned single final look at the data.

Clinical Biostats interpretation

An HR of 0.80 means that, under the fitted time-to-event model, the estimated instantaneous rate of death in the durvalumab-containing group was 0.80 times that in the placebo-containing group over the analyzed follow-up. Equivalently, 0.80 corresponds to a 20% lower estimated hazard under the model.

That interpretation does not mean that 20% of patients avoided death, that each patient experienced exactly a 20% reduction in risk, or that the absolute probability of death was reduced by 20 percentage points. A hazard ratio is a relative time-to-event measure.

The 97% two-sided confidence interval of 0.64–0.99 describes statistical uncertainty around the estimated hazard ratio under the model and analysis framework. It does not describe the range of effects experienced by individual patients.

The P-value of 0.021 is evidence against the null hypothesis in the prespecified superiority framework, but it is not a measure of effect size. Effect size is conveyed by the HR, while the confidence interval describes precision. The P-value also does not say that there is a particular probability that the null hypothesis is true.

Because the analysis uses a Cox proportional-hazards model, interpretation of a single HR should also be considered in light of the proportional-hazards assumption. The registry does not provide a proportional-hazards diagnostic in the ClinicalTrials.gov record, so the HR should not be interpreted as a guarantee that the relative hazard was constant at every time point.

Why the 97% confidence interval matters

The unusual confidence level is directly relevant to interpretation. The registry reports a 97% two-sided confidence interval for the primary OS HR, rather than the more familiar 95% interval. That choice should be retained when describing the formal primary analysis because the inferential framework includes interim monitoring and alpha allocation.

The interval of 0.64–0.99 is relatively close to the null value of 1.00 at its upper boundary. Thus, the numerical precision of the estimate matters: the observed HR is 0.80, but the interval communicates that the compatible values under the stated statistical framework extend considerably below 1 and approach 1.

Why the interim-analysis alpha matters

The registry states that the second interim OS analysis used a 2-sided significance level of 3%. Without such prespecification, repeatedly inspecting accumulating data can increase the probability of a false-positive conclusion. An interim boundary is intended to preserve the overall type I error while permitting a trial to be evaluated before all planned information has accumulated.

8. Secondary Result: Progression-Free Survival

Progression-free survival was analyzed in the full analysis set using the log-rank test. The registry describes tumor assessments every 6 weeks after randomization for the first 24 weeks and then every 8 weeks thereafter until the date specified in the registry outcome record.

Hazard ratio for progression-free survival

0.75

95.19% two-sided CI: 0.63–0.89   ·   P = 0.001

Log-rank test; superiority hypothesis

Secondary endpointAnalysisEstimateConfidence intervalP-value
Progression-free Survival (PFS) Log-rank test; stratified Cox proportional-hazards model HR 0.75 95.19% two-sided CI 0.63–0.89 0.001

The registry states that the 2-sided significance level for PFS at the second interim analysis was 4.81%. The HR and confidence interval were estimated from the stratified Cox proportional-hazards model adjusting for disease status and primary tumor location.

Clinical Biostats interpretation

An HR of 0.75 corresponds to a 25% lower estimated hazard of the PFS event in the durvalumab-containing group under the fitted model. The relevant event is defined by the trial's PFS endpoint framework; the HR should therefore be understood as a time-to-event comparison rather than a simple percentage of patients whose tumors progressed.

The 95.19% two-sided confidence interval of 0.63–0.89 quantifies uncertainty around the estimated HR. It remains below 1.00 throughout the reported interval, but the interval is not a prediction interval for individual patients and does not indicate that every patient has the same relative treatment effect.

The P-value of 0.001 provides evidence against the superiority null hypothesis under the prespecified interim-analysis framework. It does not indicate that the effect is "0.001 in size," nor does a smaller P-value by itself establish a larger or more clinically important effect.

As with OS, the HR comes from a Cox proportional-hazards model. The proportional-hazards assumption is therefore relevant when reducing the entire follow-up experience to a single HR. The ClinicalTrials.gov record does not provide a formal diagnostic of that assumption.

Why PFS and OS should not be conflated

PFS and OS are both time-to-event outcomes, but they answer different questions. PFS focuses on the time until progression or death under the trial's endpoint definition, whereas OS measures time from randomization until death from any cause. A treatment can therefore have different numerical effects on the two endpoints.

The TOPAZ-1 registry results illustrate this distinction directly: the posted HR is 0.75 for PFS and 0.80 for OS. Those numbers should be interpreted separately rather than treated as interchangeable measures of the same outcome.

9. Secondary Result: Objective Response Rate

Objective response rate was analyzed as a binary endpoint in the full analysis set among subjects with measurable disease at baseline. Tumor assessments were performed per RECIST 1.1 every 6 weeks for the first 24 weeks relative to randomization and then at the subsequent schedule specified in the registry outcome record.

Odds ratio for objective response

1.60

95% two-sided CI: 1.11–2.31   ·   P = 0.011

Cochran-Mantel-Haenszel test; superiority hypothesis

Secondary endpointAnalysisEstimateConfidence intervalP-value
Objective Response Rate (ORR) Cochran-Mantel-Haenszel test OR 1.60 95% two-sided CI 1.11–2.31 0.011

The registry states that the OR and confidence interval were estimated from a stratified CMH test adjusting for disease status and primary tumor location.

Clinical Biostats interpretation

An odds ratio of 1.60 means that the estimated odds of objective response were 1.60 times as high in the durvalumab-containing group as in the placebo-containing group under the stratified analysis. This is an odds comparison, not a risk ratio.

The distinction matters because an odds ratio of 1.60 does not mean that 60% of patients responded, nor does it mean that the probability of response increased by 60%. The underlying response probabilities would be required to translate the odds ratio into an absolute risk difference.

The 95% confidence interval of 1.11–2.31 quantifies uncertainty around the estimated odds ratio. The entire reported interval is above 1.00, but the interval does not describe the range of individual patient responses.

The P-value of 0.011 addresses evidence against the null hypothesis within the specified statistical framework. It does not measure the magnitude or clinical importance of the response effect. The odds ratio and confidence interval are the more direct measures of the estimated effect and its precision.

Why ORR uses a different statistical method

Unlike OS and PFS, ORR is a binary outcome: a participant either meets the prespecified response definition or does not. There is no event-time structure requiring Kaplan-Meier estimation for the primary ORR comparison. The registry therefore uses the CMH framework and reports an odds ratio rather than a hazard ratio.

10. Bringing the Efficacy Results Together

EndpointRoleEffect measureEstimateCIP-value
Overall Survival Primary Hazard ratio 0.80 97% two-sided: 0.64–0.99 0.021
Progression-free Survival Secondary Hazard ratio 0.75 95.19% two-sided: 0.63–0.89 0.001
Objective Response Rate Secondary Odds ratio 1.60 95% two-sided: 1.11–2.31 0.011

These three estimates describe different dimensions of the same randomized comparison. The OS and PFS HRs are both below 1.00, indicating lower estimated event hazards for the durvalumab-containing regimen under their respective Cox models. The ORR odds ratio is above 1.00, indicating higher estimated odds of response under the stratified categorical analysis.

The estimates should not be collapsed into one overall numerical "benefit." Each endpoint has a different definition, statistical scale, censoring structure, and clinical interpretation. In particular, the ORR odds ratio cannot be compared numerically with the OS or PFS hazard ratios as if all three were measuring the same quantity.

11. Interim Analysis and Alpha Spending

Interim analysis is an explicit part of the TOPAZ-1 statistical framework reported by the registry. The OS analysis notes that the second interim analysis was pre-specified after approximately 397 OS events occurred.

OS interim threshold

The 2-sided significance level for OS at the second interim analysis was 3%.

PFS interim threshold

The 2-sided significance level for PFS at the second interim analysis was 4.81%.

These thresholds illustrate why a P-value cannot be interpreted without knowing the design under which it was generated. A trial that looks at accumulating data repeatedly needs a prespecified error-control strategy. Otherwise, the probability of obtaining at least one apparently positive result by chance can exceed the intended type I error rate.

The registry data identify interim analysis / alpha spending as a concept in the posted analyses. This page therefore treats the reported P-values in the context of the interim design rather than applying a generic single-analysis interpretation.

What the ClinicalTrials.gov record establishes: the registry documents the interim-analysis timing and the endpoint-specific significance levels. The ClinicalTrials.gov record does not provide enough detail to reconstruct the complete alpha-spending function or the full sequence of information times, so this page does not infer those details.

12. Covariate Adjustment and Stratification

The posted OS and PFS analyses use a stratified Cox proportional-hazards model adjusting for disease status and primary tumor location. The ORR analysis similarly uses a stratified CMH test adjusting for these variables.

Disease status
Included as an adjustment variable in the reported OS, PFS, and ORR analyses.
Primary tumor location
Included as an adjustment variable in the reported OS, PFS, and ORR analyses.
OS / PFS
Stratified Cox proportional-hazards model for HR and confidence interval.
ORR
Stratified Cochran-Mantel-Haenszel analysis for OR and confidence interval.

Adjustment does not change the basic randomized comparison into an observational analysis. Instead, the model incorporates prespecified covariate information into estimation of the treatment effect. For a stratified model, the treatment effect is estimated while allowing the baseline event process to differ across the specified strata.

The fact that the same two variables appear in the posted analyses is also useful pedagogically: it demonstrates that statistical adjustment is not synonymous with adding arbitrary variables to a regression model. The variables used here are explicitly identified in the registry's analysis notes.

13. Missing Data, Censoring, and Time-to-Event Interpretation

The OS definition explicitly addresses participants who have not died by the time of analysis. Such patients are censored based on the last recorded date on which they were known to be alive. This is a standard right-censoring structure for survival analysis.

Censoring does not mean that the participant is treated as having survived forever. Instead, the participant contributes observed survival information up to the censoring time. The validity of the resulting survival analysis depends on assumptions about the relationship between censoring and the underlying event process.

Time-to-event contribution
Observed time = min(event time, censoring time)

A participant contributes information until the first observed event or censoring time. Participants who remain alive at analysis therefore still contribute information even though their ultimate survival time is not yet observed.

The registry's the ClinicalTrials.gov record do not describe a separate imputation procedure for missing OS outcomes. That is not surprising for a standard time-to-event endpoint: censoring is incorporated directly into Kaplan-Meier and Cox methods rather than replacing every censored observation with an imputed survival time.

Important distinction: censoring and missing-data imputation are not interchangeable concepts. A censored survival observation still contributes known follow-up time. The registry-reported TOPAZ-1 data do not document a separate imputation method for the OS analysis, so none is inferred here.

14. Statistical Methods Explained

Why was a log-rank test used for OS and PFS?

OS and PFS are time-to-event endpoints, so participants can experience the event at different times or be censored before the event is observed. The log-rank test compares the event experience between treatment groups across the follow-up period and therefore uses the ordering of event times rather than reducing the outcome to a single binary status.

What does an OS hazard ratio of 0.80 mean?

Under the reported Cox model, an HR of 0.80 means the estimated instantaneous death rate in the durvalumab-containing group was 0.80 times that of the placebo-containing group. It corresponds to a 20% lower estimated hazard. It does not mean a 20-percentage-point improvement in survival or that 20% of participants avoided death.

Why is the OS confidence interval 97% rather than 95%?

The registry reports a 97% two-sided confidence interval for the primary OS HR and states that the second interim analysis used a 2-sided significance level of 3%. The confidence level therefore reflects the prespecified interim-analysis inferential framework rather than a generic preference for a particular confidence level.

Why does the PFS analysis use a different confidence level?

The registry reports a 95.19% two-sided confidence interval for PFS and states that the 2-sided significance level for PFS at the second interim analysis was 4.81%. These values sum to the complementary relationship expected between the stated two-sided alpha level and confidence level.

Why is ORR reported with an odds ratio instead of a hazard ratio?

ORR is a binary outcome, whereas OS and PFS are time-to-event outcomes. The registry therefore uses a stratified Cochran-Mantel-Haenszel analysis for ORR and reports an odds ratio. The statistical scale changes because the underlying endpoint structure changes.

What does an OR of 1.60 mean?

An OR of 1.60 means that the estimated odds of objective response were 1.60 times as high in the durvalumab-containing group as in the placebo-containing group under the stratified analysis. It does not mean that the response probability increased by exactly 60%, because odds and probabilities are different quantities.

Why does interim analysis change how a P-value should be read?

If a trial is examined more than once while data accumulate, an unadjusted testing strategy can increase the chance of a false-positive conclusion. TOPAZ-1's registry-reported analysis notes explicitly identify an interim-analysis framework and endpoint-specific significance levels. Therefore, the reported P-values should be interpreted against those prespecified thresholds rather than against an automatically assumed single-look threshold.

15. Safety Results

The ClinicalTrials.gov record includes serious adverse events by treatment arm. The reported counts are presented as affected participants over participants at risk.

Serious adverse eventsAffected / at risk
Durvalumab + Gemcitabine + Cisplatin 168 / 338
Placebo + Gemcitabine + Cisplatin 152 / 342

These figures should be read as serious adverse-event counts relative to the reported at-risk denominators. They are not the same endpoint as overall adverse events, grade-specific adverse events, or treatment discontinuations.

The safety analysis is also conceptually different from the efficacy analysis. Efficacy is anchored to the randomized comparison and the full analysis set specified in the posted analyses. Safety is inherently related to treatment exposure, so the denominator and analysis population used for a particular safety table must be respected rather than inferred from the efficacy population.

Safety interpretation: The ClinicalTrials.gov record supports reporting serious adverse events as 168/338 versus 152/342. They do not provide enough information to construct a complete adverse-event profile, calculate additional safety categories, or attribute causality to individual events.

16. Trial Timeline

April 16, 2019

Trial start

The registry lists April 16, 2019 as the study start date.

August 11, 2021

Primary completion

The registry lists August 11, 2021 as the primary completion date.

Current registry status

Active, not recruiting

The ClinicalTrials.gov record identifies the trial status as ACTIVE_NOT_RECRUITING.

The registry also identifies a second interim analysis for OS after approximately 397 OS events. This event-driven structure is typical of time-to-event trials: information accumulates according to the number of observed events rather than simply according to elapsed calendar time.

17. Primary Endpoint Interpretation in Detail

What the OS estimate tells us

The reported HR of 0.80 summarizes the relative death hazard estimated from the stratified Cox model. Its direction is below 1.00, corresponding to a lower estimated hazard in the durvalumab-containing group.

What the OS estimate does not tell us

The HR does not directly state the median survival, the absolute survival probability at a specific time, the number of deaths in each group, or the proportion of patients who personally benefited. None of those additional numerical quantities are reported in the ClinicalTrials.gov record used for this page.

What the confidence interval tells us

The 97% two-sided interval of 0.64–0.99 communicates the uncertainty surrounding the HR estimate under the specified analysis. The upper endpoint approaching 1.00 is relevant to precision even though the entire reported interval remains below 1.00.

What the P-value tells us

The P-value of 0.021 quantifies the compatibility of the observed test statistic with the null hypothesis under the statistical model and prespecified testing framework. It is not the probability that the null hypothesis is true and it is not a measure of clinical effect size.

18. Interpreting the Secondary Endpoints Without Overstating Them

The secondary results provide a coherent set of statistical signals: the PFS HR is below 1.00 and the ORR odds ratio is above 1.00. But statistical interpretation should preserve the distinction between endpoint scales.

Time-to-event evidence

The PFS HR of 0.75 describes a relative difference in the estimated event hazard under the Cox model.

Response evidence

The OR of 1.60 describes a relative difference in the odds of objective response under the stratified CMH analysis.

Different estimands

An HR and an OR are not directly comparable numerical measures. Each answers a different statistical question.

Analysis populations

ORR is restricted to the full analysis set among subjects with measurable disease at baseline, while OS and PFS use the full analysis set.

This is a useful example of why a trial should not be summarized by a single "positive" or "negative" statistic. The statistical story consists of endpoint definitions, analysis populations, effect measures, confidence intervals, hypothesis-testing thresholds, and the design that generated the data.

19. Multiplicity and Endpoint Hierarchy

TOPAZ-1 has three registered primary endpoints: OS, OS rate at 18 months, and OS rate at 24 months. The registry-reported formal statistical analyses include one primary endpoint analysis, specifically the OS analysis with HR 0.80, 97% two-sided CI 0.64–0.99, and P = 0.021.

Endpoint / analysisRole in the ClinicalTrials.gov recordPosted statistical result
Overall Survival Primary endpoint Formal analysis posted
OS Rate at 18 Months Primary endpoint Registry states it is calculated using Kaplan-Meier; no separate statistical analysis is reported in the ClinicalTrials.gov record used here
OS Rate at 24 Months Primary endpoint Registry states it is calculated using Kaplan-Meier; no separate statistical analysis is reported in the ClinicalTrials.gov record used here
Progression-free Survival Secondary endpoint HR 0.75; 95.19% two-sided CI 0.63–0.89; P = 0.001
Objective Response Rate Secondary endpoint OR 1.60; 95% two-sided CI 1.11–2.31; P = 0.011

The distinction between "registered primary endpoint" and "posted formal statistical analysis" is important. A registry can contain several registered outcome measures while only some of them have an accompanying formal statistical-analysis record. The absence of a separate posted analysis in the ClinicalTrials.gov record should not be filled with an inferred test or a reconstructed P-value.

20. Limitations

21. Why This Trial Matters Statistically

TOPAZ-1 is a useful statistical teaching case because it combines randomized treatment comparison, quadruple masking, multiple time-to-event endpoints, a binary response endpoint, stratified analysis, covariate adjustment, interim monitoring, and different effect measures within the same trial.

ConceptHow it appears in TOPAZ-1
Randomization The trial is randomized, supporting comparison of the assigned treatment groups.
Blinding The registry identifies quadruple masking.
Kaplan-Meier estimation Used for OS and specifically for OS rates at 18 and 24 months.
Log-rank test Used for the posted OS and PFS comparisons.
Hazard ratio Reported for OS and PFS from stratified Cox proportional-hazards models.
Confidence intervals Reported alongside the OS, PFS, and ORR effect estimates.
Interim analysis The second interim analysis was prespecified after approximately 397 OS events.
Alpha spending The registry-reported analysis notes specify endpoint-specific interim significance levels.
Covariate adjustment Disease status and primary tumor location were used in the reported adjusted analyses.
Cochran-Mantel-Haenszel test Used for the binary ORR analysis.
Odds ratio Used as the effect measure for ORR.
Time-to-event endpoints OS and PFS incorporate event times and censoring rather than simple binary follow-up status.

The most instructive feature is the contrast between the hazard ratio and the odds ratio. A reader who understands why those measures are different is already equipped to interpret much of the statistical structure of the trial correctly.

22. Statistical Methods Explained: A Deeper Walkthrough

Why is randomization important before any statistical test is performed?

Randomization establishes the treatment assignment mechanism before outcomes are observed. Statistical testing then evaluates the observed treatment-group differences under that randomized framework. The statistical model does not create the randomization; it analyzes the information generated by it.

Why use a stratified Cox model rather than only a raw Kaplan-Meier comparison?

Kaplan-Meier estimation describes survival over time, while the Cox model provides a quantitative relative-effect estimate through the hazard ratio. Stratification and adjustment also allow the reported analysis to incorporate disease status and primary tumor location as specified in the registry.

Why can a confidence interval be more informative than a P-value?

The P-value primarily addresses evidence against a null hypothesis. The confidence interval also shows the range of effect estimates compatible with the statistical model at the stated confidence level. For TOPAZ-1, the OS interval of 0.64–0.99 communicates both the direction of the estimated effect and its uncertainty.

Why does the OS analysis have a 97% CI?

The confidence level is connected to the prespecified interim-testing framework. The registry reports a 2-sided OS significance level of 3% at the second interim analysis, corresponding to a 97% two-sided confidence level for the reported HR.

Why is the PFS confidence level 95.19%?

The registry reports a 2-sided PFS significance level of 4.81% at the second interim analysis. The complementary confidence level is therefore 95.19%. This is another demonstration that confidence intervals and hypothesis tests are two expressions of the same underlying inferential framework.

Why is the ORR analysis restricted to measurable disease at baseline?

An objective response requires an evaluable baseline disease burden against which tumor response can be assessed. The registry explicitly defines the ORR analysis population as the full analysis set among subjects with measurable disease at baseline.

Why should OS and PFS not be treated as independent pieces of evidence without considering the trial design?

Both endpoints arise from the same randomized participants and are evaluated within the same clinical-trial framework. Their statistical interpretation therefore depends on the prespecified endpoint structure, interim monitoring, and any multiplicity strategy. Simply counting the number of P-values below a conventional threshold would ignore that design.

23. What the Trial's Effect Measures Mean

HR 0.80 for OS

The estimated instantaneous death hazard was 0.80 times that of the comparator under the reported stratified Cox model.

HR 0.75 for PFS

The estimated PFS event hazard was 0.75 times that of the comparator under the reported stratified Cox model.

OR 1.60 for ORR

The estimated odds of objective response were 1.60 times those of the comparator under the stratified CMH analysis.

P-values

The P-values quantify evidence against the respective null hypotheses within the stated inferential framework; they are not effect-size measures.

These measures also have different null values. The null value for a hazard ratio is 1.00, and the null value for an odds ratio is also 1.00. Values below 1 for the HR indicate lower estimated event hazard in the treatment group, whereas values above 1 for the OR indicate higher estimated odds of response in the treatment group.

24. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The reported OS HR is 0.80 with a 97% two-sided CI of 0.64–0.99 and P = 0.021. The PFS HR is 0.75 with a 95.19% two-sided CI of 0.63–0.89 and P = 0.001. ORR has an OR of 1.60 with a 95% two-sided CI of 1.11–2.31 and P = 0.011.

Clinical interpretation

The statistical results indicate differences between the randomized treatment groups on the reported efficacy endpoints. Clinical meaning requires consideration of the endpoint definitions, magnitude and precision of effects, treatment context, and safety rather than relying on P-values alone.

The distinction is deliberate. Statistical significance and clinical importance are related but not identical concepts. A statistical analysis can establish evidence against a null hypothesis without, by itself, defining how meaningful the treatment difference is for an individual patient.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through the Clinical Biostats statistical tutorials

Explore the survival-analysis, categorical-data, confidence-interval, and clinical-trial methods that appear in TOPAZ-1 and other randomized studies.

28. Record Summary

TOPAZ-1 provides a compact example of how modern randomized-trial statistics combine several complementary methods. The trial is randomized, parallel, quadruple-masked, and designed for treatment comparison. Its registered primary endpoints are time-to-event outcomes based on overall survival, including fixed-time OS rates calculated by Kaplan-Meier estimation. The posted formal OS analysis uses a log-rank test and a stratified Cox proportional-hazards model, with an HR of 0.80, a 97% two-sided confidence interval of 0.64–0.99, and P = 0.021.

The secondary analyses add two different statistical perspectives. PFS uses the same broad survival-analysis framework and reports an HR of 0.75 with a 95.19% two-sided confidence interval of 0.63–0.89 and P = 0.001. ORR is analyzed as a binary endpoint using a stratified Cochran-Mantel-Haenszel test and reports an OR of 1.60 with a 95% two-sided confidence interval of 1.11–2.31 and P = 0.011.

The statistical interpretation depends on more than the point estimates. The analysis populations, censoring rules, covariate adjustment, interim significance levels, confidence levels, and distinction between hazard ratios and odds ratios all affect what the reported numbers mean. The safety data in the ClinicalTrials.gov record further show why efficacy and safety must be examined as separate statistical dimensions rather than compressed into a single measure.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. Numerical estimates, confidence intervals, P-values, endpoint definitions, and analysis populations on this page are restricted to the registry-reported TOPAZ-1 trial data; where the registry does not provide a result or sufficient detail, no value has been inferred.