← Clinical Trials
Chronic Lymphocytic Leukemia Phase 3 Completed NCT02242942

CLL14: Complete Statistical Analysis of Obinutuzumab + Venetoclax in Chronic Lymphocytic Leukemia

An independent statistical analysis of the randomized phase 3 CLL14 trial comparing obinutuzumab plus venetoclax with obinutuzumab plus chlorambucil, focusing on progression-free survival, response, minimal residual disease, and the categorical and time-to-event methods reported in the registry.

Trial start: 2014-12-31  ·  Primary completion: 2018-08-17  ·  Sponsor: Hoffmann-La Roche
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

CLL14 was a randomized, parallel-group, open-label phase 3 treatment trial in chronic lymphocytic leukemia. The registry reports a comparison of obinutuzumab plus venetoclax with obinutuzumab plus chlorambucil, with progression-free survival based on investigator assessment according to IWCLL criteria as the registered primary endpoint.

445
Enrollment
ClinicalTrials.gov record
3
Arms
Parallel design
0.35
Primary PFS HR
95% CI 0.23–0.53
<0.0001
Primary P-value
Superiority analysis
FeatureCLL14
PhasePhase 3
ConditionLymphocytic Leukemia, Chronic
DesignRandomized, parallel
MaskingNone
Primary purposeTreatment
Enrollment445
Primary endpointProgression Free Survival (PFS) Based on Investigator Assessment According to IWCLL Criteria
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted23
Statistical analyses posted10
ClinicalTrials.govNCT02242942

2. Clinical Question

The principal statistical question was whether obinutuzumab plus venetoclax produced a different time-to-event outcome from obinutuzumab plus chlorambucil, as measured by investigator-assessed progression-free survival under the IWCLL criteria. The registered hypothesis type for the formal analyses was superiority.

Population

Participants with chronic lymphocytic leukemia enrolled in the CLL14 trial.

Intervention

Obinutuzumab plus venetoclax.

Comparator

Obinutuzumab plus chlorambucil.

Primary question

Does obinutuzumab plus venetoclax improve investigator-assessed progression-free survival relative to obinutuzumab plus chlorambucil?

3. Trial Design

01
Enroll445 participants
02
RandomizeRandomized allocation
03
Parallel armsThree registered arms
04
AssessPFS, response, MRD
05
AnalyzeITT and reported methods
RANDOMIZED COMPARISON

Obinutuzumab + Chlorambucil

  • Comparator treatment group in the posted statistical analyses.
  • Compared directly with obinutuzumab + venetoclax for the reported efficacy endpoints.
  • Included in the reported ITT efficacy analyses.
RANDOMIZED COMPARISON

Obinutuzumab + Venetoclax

  • Intervention treatment group in the posted statistical analyses.
  • Compared directly with obinutuzumab + chlorambucil for the reported efficacy endpoints.
  • Included in the reported ITT efficacy analyses.
Third arm: the ClinicalTrials.gov record lists 3 arms, and the safety data separately identify a Safety Run-in Obinutuzumab + Venetoclax group with 13 participants at risk. The formal statistical analyses in the ClinicalTrials.gov record compare obinutuzumab + chlorambucil with obinutuzumab + venetoclax.

4. Trial Timing and Registry Status

2014-12-31

Trial start

The registry lists December 31, 2014 as the study start date.

2018-08-17

Primary completion

The registry lists August 17, 2018 as the primary completion date.

Completed

Registry status

The trial is listed as completed, with results posted on ClinicalTrials.gov.

5. Primary Endpoint

EndpointRegistry definition / time frameAnalysis
Progression Free Survival (PFS) Based on Investigator Assessment According to IWCLL Criteria Baseline until disease progression or death up to approximately 3.75 years Log-rank test; hazard ratio estimated by Cox regression model

6. Secondary Endpoints

The posted statistical analyses cover several categories of secondary efficacy outcomes. Time-to-event PFS was analyzed using a log-rank framework, while binary response and minimal-residual-disease outcomes were analyzed using either the Cochran-Mantel-Haenszel test or chi-squared test.

Secondary endpointTime frameMethodEffect measure
PFS based on Institutional Review Committee assessments according to IWCLL criteria Baseline until disease progression or death up to approximately 3.75 years Log-rank test Hazard ratio
Overall response at completion of treatment 3 months after treatment completion, approximately 15 months) Cochran-Mantel-Haenszel test Difference in response rates
Complete response rate at completion of treatment 3 months after treatment completion, approximately 15 months) Cochran-Mantel-Haenszel test Difference in response rates
MRD negativity in peripheral blood at completion of treatment 3 months after treatment completion, approximately 15 months) Cochran-Mantel-Haenszel test Difference in MRD-negative rates
MRD negativity in bone marrow at completion of treatment 3 months after treatment completion, approximately 15 months) Cochran-Mantel-Haenszel test Difference in MRD-negative rates
MRD negativity in peripheral blood at completion of combination treatment Day 1 Cycle 9 or 3 months after last IV infusion, approximately 9 months Chi-squared test Difference in MRD-negative rates
MRD negativity in bone marrow at completion of combination treatment Day 1 Cycle 9 or 3 months after last IV infusion, approximately 9 months Chi-squared test Difference in MRD-negative rates
Overall response at completion of combination treatment Day 1 Cycle 7 or 28 days after last IV infusion, approximately 6 months Cochran-Mantel-Haenszel test Difference in response rates
Best response achieved Baseline up to completion of treatment assessment, up to approximately 15 months) Cochran-Mantel-Haenszel test Difference in response rates

7. Primary Result: Investigator-Assessed Progression-Free Survival

The primary analysis compared obinutuzumab plus chlorambucil with obinutuzumab plus venetoclax in the intention-to-treat population, defined in the registry as all randomized participants. The reported method was the log-rank test, with hazard ratios estimated using a Cox regression model.

Hazard ratio for progression or death

0.35

95% CI: 0.23–0.53   ·   P < 0.0001

Two-sided 95% confidence interval; superiority hypothesis.

Primary endpointComparisonEstimate95% CIP-value
Investigator-assessed PFS according to IWCLL criteria Obinutuzumab + Chlorambucil vs Obinutuzumab + Venetoclax HR 0.35 0.23–0.53 <0.0001
Analysis details: the registry states that hazard ratios were estimated by a Cox regression model with Binet and Geographic region as stratification factors.
Clinical Biostats interpretation

An HR of 0.35 means that, under the fitted Cox model, the estimated instantaneous rate of progression or death for the obinutuzumab-plus-venetoclax group was approximately 35% of the corresponding rate in the obinutuzumab-plus-chlorambucil group. Equivalently, the estimated hazard was approximately 65% lower in relative terms.

The HR is not a statement that 65% of patients were protected from progression or death, nor does it mean that every patient experienced a 65% reduction in individual risk. It is a relative time-to-event measure derived from the statistical model.

The two-sided 95% CI of 0.23–0.53 describes uncertainty around the estimated hazard ratio. It does not describe the range of treatment effects across individual patients. Because the entire interval is below 1, the reported data are consistent with a lower estimated hazard in the obinutuzumab-plus-venetoclax group under this analysis.

The P-value <0.0001 addresses evidence against the null hypothesis specified for the statistical comparison; it does not measure the magnitude or clinical importance of the treatment effect. The magnitude is described by the HR and its confidence interval.

As with other Cox-model hazard ratios, interpretation should account for the proportional-hazards structure implicit in summarizing the treatment contrast with a single HR. The ClinicalTrials.gov record does not report a separate assessment of that assumption.

8. Secondary Result: IRC-Assessed Progression-Free Survival

The second reported PFS analysis used assessments from an Institutional Review Committee according to IWCLL criteria. Like the primary PFS analysis, it used the ITT population and a log-rank test, with the hazard ratio estimated from a Cox regression model stratified by Binet and Geographic region.

Hazard ratio for progression or death

0.33

95% CI: 0.22–0.51   ·   P < 0.0001

Two-sided 95% confidence interval; superiority hypothesis.

Clinical Biostats interpretation

The IRC-assessed HR of 0.33 corresponds to an estimated instantaneous progression-or-death hazard approximately 33% as large in the obinutuzumab-plus-venetoclax group as in the obinutuzumab-plus-chlorambucil group, or approximately a 67% lower estimated hazard under the fitted model.

The 95% CI of 0.22–0.51 gives the statistical uncertainty around that relative estimate. It should not be interpreted as a prediction interval for individual patients. The P-value of <0.0001 provides evidence against the null hypothesis used for the comparison; it is not a measure of the size of the observed treatment effect.

The similarity in direction between investigator-assessed and IRC-assessed PFS is relevant descriptively because the two analyses use different assessment sources. It does not turn the secondary endpoint into an independent replication of the primary hypothesis, and the ClinicalTrials.gov record does not provide an adjusted multiplicity framework for interpreting the entire set of posted secondary analyses.

9. Secondary Results: Overall Response

Overall response was analyzed as a binary outcome using the Cochran-Mantel-Haenszel test. The analysis population was the ITT population. The reported effect measure was the difference in response rates, with two-sided 95% confidence intervals constructed using the Anderson-Hauck method.

Overall response at approximately 15 months)

13.43 percentage points

95% CI: 5.47–21.38   ·   P = 0.0007

Assessment at treatment completion, 3 months after treatment completion.

Clinical Biostats interpretation

The reported difference in response rates of 13.43 percentage points describes the contrast between the two treatment groups using the registry's stated direction of comparison: obinutuzumab plus chlorambucil versus obinutuzumab plus venetoclax. A positive reported difference therefore represents a higher response rate for the latter group under the registry's comparison.

The 95% CI of 5.47–21.38 quantifies uncertainty around that difference. Unlike a hazard ratio, a difference in response rates is an absolute percentage-point contrast and is directly interpretable on the probability scale.

The P-value of 0.0007 assesses the evidence against the null comparison used by the CMH analysis. It does not tell us that the effect is "0.0007 large," nor does it establish that every individual patient has the same response benefit.

10. Secondary Results: Complete Response Rate

Complete response rate was assessed at the completion-of-treatment assessment, 3 months after treatment completion at approximately 15 months. The registry reports a Cochran-Mantel-Haenszel analysis in the ITT population.

Complete response rate difference

26.39 percentage points

95% CI: 17.41–35.36   ·   P < 0.0001

Two-sided 95% confidence interval using the Anderson-Hauck method.

Clinical Biostats interpretation

The estimated difference of 26.39 percentage points indicates a substantially different observed complete-response proportion between the two randomized treatment groups at the specified assessment time. Because the comparison is expressed as an absolute difference, it is not directly interchangeable with the PFS hazard ratio.

The 95% CI of 17.41–35.36 indicates that the estimated difference is subject to sampling uncertainty, while remaining entirely above zero. The P-value of <0.0001 supplies evidence against the null hypothesis of no difference in the reported CMH comparison but does not itself describe the magnitude of the difference.

The endpoint is also time-specific: it concerns the completion-of-treatment assessment approximately 15 months into the specified time frame, rather than a time-to-event measure over the entire follow-up period.

11. Secondary Results: Minimal Residual Disease

Several MRD-negativity endpoints were posted. The ClinicalTrials.gov record distinguishes peripheral blood and bone marrow measurements and also distinguish the approximately 9-month completion-of-combination-treatment assessment from the approximately 15-month completion-of-treatment assessment.

MRD endpointAssessmentDifference95% CIP-value
Peripheral blood MRD negativity Completion of treatment, approximately 15 months) 40.28 31.45–49.10 <0.0001
Bone marrow MRD negativity Completion of treatment, approximately 15 months) 39.81 31.27–48.36 <0.0001
Peripheral blood MRD negativity Completion of combination treatment, approximately 9 months 32.87 23.76–41.98 <0.0001
Bone marrow MRD negativity Completion of combination treatment, approximately 9 months 38.43 30.15–46.71 <0.0001

The approximately 15-month peripheral-blood and bone-marrow analyses used the Cochran-Mantel-Haenszel test. The approximately 9-month analyses used the chi-squared test. The registry states that the 95% confidence intervals for the differences in rates were constructed using the Anderson-Hauck method.

Clinical Biostats interpretation

These estimates are differences in MRD-negative rates, not hazard ratios. For example, the reported peripheral-blood estimate of 40.28 at approximately 15 months represents a 40.28-percentage-point difference under the registry's specified comparison.

The confidence intervals provide uncertainty around the percentage-point differences. They do not indicate the range of MRD values among individual participants. The consistently positive intervals reported for these four analyses indicate that the corresponding observed rate differences were separated from zero under their respective statistical comparisons.

The different assessment times also matter. The approximately 9-month endpoints and approximately 15-month endpoints answer related but distinct questions about MRD negativity. They should not be treated as repeated measurements of exactly the same endpoint at the same time.

12. Secondary Result: Overall Response at Approximately 6 Months

The registry also reports overall response at the completion-of-combination-treatment assessment, defined as Day 1 Cycle 7 or 28 days after the last IV infusion, approximately 6 months.

Overall response difference

1.85 percentage points

95% CI: -4.63–8.33   ·   P = 0.5612

Cochran-Mantel-Haenszel analysis in the ITT population.

Clinical Biostats interpretation

The point estimate of 1.85 percentage points is close to zero relative to the uncertainty represented by its 95% CI. The interval spans both negative and positive values, from -4.63 to 8.33.

The P-value of 0.5612 does not provide evidence against the null hypothesis for this particular comparison. Importantly, this should not be converted into a claim that the treatments are equivalent or identical. A nonsignificant superiority test means that the analysis did not demonstrate a statistically detectable difference under its specified framework; it is not a formal non-inferiority or equivalence analysis.

This result also illustrates why the assessment time matters. A response comparison at approximately 6 months can differ from a response comparison at approximately 15 months without the two analyses being contradictory.

13. Secondary Result: Best Response Achieved

The registry defines this outcome as the percentage of participants by best response achieved, including complete response, complete response with incomplete marrow recovery, partial response, stable disease, or progressive disease, from baseline through the completion-of-treatment assessment up to approximately 15 months.

Best-response difference

0.93 percentage points

95% CI: -4.66–6.51   ·   P = 0.7169

Cochran-Mantel-Haenszel analysis stratified by the IvRS randomization stratification factors.

Clinical Biostats interpretation

The reported difference of 0.93 percentage points is close to zero, while the 95% CI of -4.66–6.51 includes zero. The P-value of 0.7169 indicates that this analysis did not demonstrate a statistically detectable difference under the reported CMH framework.

The important statistical distinction is between failure to demonstrate a difference and demonstration of equivalence. The ClinicalTrials.gov record identifies the hypothesis as superiority, not equivalence or non-inferiority. Therefore, the result should not be reinterpreted as proof that the two treatments have identical best-response distributions.

14. Statistical Methodology

Intention-to-treat analysis

The registry defines the ITT population as all randomized participants. The primary PFS analysis and the posted secondary efficacy analyses described here use that population. An ITT analysis preserves treatment assignment as the basis of comparison rather than redefining groups according to treatment received or subsequent treatment experience.

Core principle
Analyze participants according to the treatment to which they were randomized.

The principal statistical advantage is preservation of the randomized comparison. This is particularly important for causal interpretation because randomization is the mechanism intended to balance prognostic factors between treatment groups.

Log-rank test

The primary PFS analysis and the IRC-assessed PFS analysis used the log-rank test. This is a time-to-event comparison that uses information from the entire observed follow-up rather than comparing only one fixed time point.

For each event time, the analysis considers the participants still at risk immediately before that time and compares the observed number of events with what would be expected under the null hypothesis of equal survival distributions. Repeating that comparison across event times produces the overall test statistic.

Cox regression and hazard ratios

The registry states that hazard ratios were estimated using a Cox regression model. The primary and IRC-assessed PFS analyses used Binet and Geographic region as stratification factors.

Hazard-ratio interpretation
HR = estimated instantaneous event rate in treatment group ÷ estimated instantaneous event rate in comparator group

An HR below 1 indicates a lower estimated event hazard for the numerator group under the fitted model. It is not an absolute risk difference and is not the same as a ratio of cumulative event probabilities.

Cochran-Mantel-Haenszel testing

The CMH test was used for several binary response and MRD outcomes. This method can compare categorical outcomes while accounting for stratification factors. The registry specifically notes stratification by IvRS randomization stratification factors for the best-response analysis.

Chi-squared testing

The approximately 9-month peripheral-blood and bone-marrow MRD-negativity analyses used the chi-squared test. For a binary outcome, the test evaluates whether the observed distribution of the outcome differs between treatment groups under the specified null hypothesis.

Anderson-Hauck confidence intervals

The registry states that the 95% confidence intervals for differences in rates were constructed using the Anderson-Hauck method. This is important because the interval construction is part of the statistical specification of the reported percentage-point effects; the confidence intervals should not be treated as if they had necessarily been generated by a generic Wald approximation.

15. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint. Some participants can experience progression or death early, while others remain event-free at the end of their observed follow-up and are therefore censored. A log-rank analysis is designed for this structure because it uses event timing and accounts for censoring rather than treating every participant as if everyone had been observed for exactly the same duration.

What does an HR of 0.35 mean?

An HR of 0.35 means that the fitted Cox model estimates the instantaneous progression-or-death hazard in the obinutuzumab-plus-venetoclax group at about 35% of the comparator hazard. The complementary interpretation is an approximately 65% lower estimated hazard. It does not mean a 65% absolute reduction in the probability of progression or death.

Why is the confidence interval important?

The point estimate alone does not describe statistical precision. For the primary PFS analysis, the HR is 0.35, while the 95% CI is 0.23–0.53. The interval communicates uncertainty around the estimated relative effect. It is therefore more informative to report the HR together with its confidence interval than to report the HR alone.

Why does the P-value not measure effect size?

A P-value summarizes the compatibility of the observed data with the null hypothesis under the statistical model. It is affected by both the magnitude of an effect and the amount of information available. A very small P-value therefore does not mean that the effect is necessarily large, clinically important, or precisely estimated to a particular degree.

Why use both investigator and IRC assessments?

The trial reports PFS using two assessment sources. The primary endpoint uses investigator assessment, while a secondary endpoint uses Institutional Review Committee assessment. Comparing results from these approaches can provide useful information about the consistency of the observed time-to-event comparison, but the secondary analysis remains a separate endpoint rather than a second primary endpoint.

What is the difference between a response-rate difference and a hazard ratio?

A response-rate difference compares proportions at a defined assessment. A hazard ratio compares instantaneous event rates over time. A difference such as 26.39 percentage points for complete response is therefore answering a different statistical question from a PFS HR of 0.35. Neither measure substitutes for the other.

Does a nonsignificant superiority test establish equivalence?

No. For the approximately 6-month overall-response analysis, the P-value is 0.5612 and the 95% CI includes zero. This means the superiority analysis did not demonstrate a statistically detectable difference. It does not establish that the treatments are equivalent, because equivalence would require a separately specified equivalence framework and margins.

16. Reading the Primary PFS Result Correctly

Relative effect

HR 0.35 summarizes the estimated relative time-to-event effect under the Cox model.

Precision

The 95% CI of 0.23–0.53 quantifies uncertainty around the estimated hazard ratio.

Statistical evidence

P < 0.0001 indicates strong evidence against the null hypothesis used for the superiority comparison.

What is not reported

The ClinicalTrials.gov record does not provide median PFS, Kaplan-Meier survival probabilities, event counts, or a survival curve for this page.

This distinction is important. A hazard ratio is a compact summary of a time-to-event comparison, but it does not contain every clinically relevant feature of the survival distributions. Median PFS, fixed-time PFS probabilities, numbers of events, and the shape of the Kaplan-Meier curves would answer additional questions. Because those values are not contained in the ClinicalTrials.gov record, they are not added here.

Educational note: a Kaplan-Meier curve cannot be reconstructed reliably from a hazard ratio and confidence interval alone. Valid reconstruction requires event and censoring information or an appropriately digitized source curve.

17. Stratification and the Role of Randomization

The registry identifies Binet and Geographic region as stratification factors for the Cox regression analyses of PFS. Stratification allows the model to account for differences in baseline hazard across the specified strata without requiring a single common baseline hazard for all strata.

The best-response analysis also states that its P-value was assessed using a CMH test stratified by the IvRS randomization stratification factors. Thus, the ClinicalTrials.gov record shows stratification being used in both time-to-event and categorical analyses, although the exact full list of IvRS factors is not provided in the registry extract.

Why stratification matters
Treatment effect estimate + prespecified stratification structure → analysis aligned with the randomized design

Stratification does not create balance after the fact. Rather, it incorporates the trial's prespecified randomization structure into the statistical comparison.

18. Multiplicity and the Set of Reported Analyses

The registry reports 23 outcome measures and 10 statistical analyses, including one posted primary-endpoint analysis and nine posted secondary analyses in the ClinicalTrials.gov record.

FeatureRegistry data reported
Registered primary endpoints1
Primary endpoint analyses posted1
Primary analyses with estimate + CI1
Statistical analyses posted10
Outcome measures posted23
Hypothesis typeSuperiority

The ClinicalTrials.gov record identifies superiority as the hypothesis type for the posted analyses but do not provide an explicit multiplicity-adjustment procedure or alpha-allocation scheme for the entire family of secondary endpoints. Consequently, each secondary P-value should be understood as the reported result of its individual analysis rather than automatically interpreted as an independently confirmatory test of a new hypothesis family.

Statistical caution: when many endpoints are evaluated, the collection of nominal P-values cannot automatically be interpreted as though each test had its own fully protected type I error rate. A formal multiplicity interpretation requires the prespecified testing hierarchy, alpha allocation, or adjustment procedure. Those details are not contained in the ClinicalTrials.gov record.

19. Missing Data, Censoring, and Analysis Assumptions

The primary endpoint is a time-to-event outcome, so censoring is intrinsically part of the analysis. Participants who have not experienced progression or death by the end of their observed follow-up can contribute information up to the point at which their event status is no longer observed.

For binary response and MRD outcomes, the ClinicalTrials.gov record identifies the analysis population and statistical tests but do not provide a missing-data or imputation specification. Accordingly, this page does not assume a particular imputation method.

IssueWhat the ClinicalTrials.gov record supportsWhat is not specified
Censoring PFS is a time-to-event endpoint analyzed with log-rank and Cox methods. A detailed censoring algorithm is not provided in the registry-reported extract.
Missing binary outcomes Binary endpoints were analyzed in the ITT population using CMH or chi-squared methods. No imputation method is provided.
Proportional hazards Hazard ratios were estimated using Cox regression. No formal proportional-hazards diagnostic is provided.
Interim analysis No interim-analysis method is identified in the ClinicalTrials.gov record. Timing, stopping boundaries, and alpha spending are not reported here.

20. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm using affected participants divided by participants at risk. The available safety figures include the two principal randomized treatment groups and the separate safety run-in group.

Treatment groupSerious adverse eventsAt risk
Obinutuzumab + Chlorambucil 90 214
Obinutuzumab + Venetoclax 104 212
Safety Run-in Obinutuzumab + Venetoclax 10 13

The ClinicalTrials.gov record does not provide a formal between-group hypothesis test for serious adverse events, confidence intervals, exposure-adjusted incidence rates, or detailed adverse-event categories. Therefore, the safety information is presented descriptively rather than converted into an unsupported comparative statistical conclusion.

Clinical Biostats interpretation

The denominators matter. The serious-adverse-event counts are accompanied by the number at risk for each group, and the safety run-in has a much smaller denominator of 13. A raw count alone would therefore be misleading as a comparison.

Safety and efficacy also have different statistical structures. The primary PFS analysis is anchored to randomization and uses time-to-event methods, whereas the ClinicalTrials.gov record is a descriptive count of participants affected by serious adverse events. The two should not be collapsed into a single numerical "benefit-risk" statistic.

21. What the Posted Secondary Results Show Statistically

EndpointEstimate95% CIP-valueMethod
IRC-assessed PFSHR 0.330.22–0.51<0.0001Log-rank / Cox
Overall response, approximately 15 months)13.435.47–21.380.0007CMH
Complete response rate, approximately 15 months)26.3917.41–35.36<0.0001CMH
Peripheral blood MRD negativity, approximately 15 months)40.2831.45–49.10<0.0001CMH
Bone marrow MRD negativity, approximately 15 months)39.8131.27–48.36<0.0001CMH
Peripheral blood MRD negativity, approximately 9 months32.8723.76–41.98<0.0001Chi-squared
Bone marrow MRD negativity, approximately 9 months38.4330.15–46.71<0.0001Chi-squared
Overall response, approximately 6 months1.85-4.63–8.330.5612CMH
Best response, up to approximately 15 months)0.93-4.66–6.510.7169CMH

Several statistical patterns are worth noticing. The time-to-event analyses report hazard ratios substantially below 1. The later response and MRD analyses report positive percentage-point differences with confidence intervals that do not cross zero. In contrast, the approximately 6-month overall-response and best-response analyses have confidence intervals that cross zero and larger P-values.

These results should not be treated as contradictory simply because some secondary endpoints are more strongly separated than others. Different endpoints measure different biological and clinical phenomena, use different assessment times, and have different statistical information content.

22. Why the Assessment Time Matters

CLL14's posted secondary endpoints illustrate an important principle in longitudinal clinical-trial analysis: an endpoint defined at one assessment time is not interchangeable with the same type of endpoint assessed at another time.

Approximately 6 months

Overall response at Day 1 Cycle 7 or 28 days after the last IV infusion: difference 1.85, 95% CI -4.63–8.33.

Approximately 9 months

Peripheral-blood and bone-marrow MRD negativity were analyzed at the completion of combination treatment.

Approximately 15 months)

Response, complete response, and MRD-negativity outcomes were assessed at completion of treatment, 3 months after treatment completion.

Up to approximately 3.75 years

The PFS endpoints followed participants from baseline until disease progression or death over the registered time frame.

A time-specific binary endpoint asks, in effect, "what proportion had the specified status at this assessment?" A PFS endpoint instead asks "how long until progression or death?" Those questions require different statistical methods and should be interpreted on their own scales.

23. Limitations

24. Why This Trial Matters Statistically

CLL14 is a useful statistical teaching case because the registry data bring together two major classes of clinical-trial endpoints: time-to-event outcomes and binary categorical outcomes. The trial therefore shows why a single clinical study can require several statistical tools rather than one universal test.

ConceptHow it appears in CLL14
RandomizationThe trial is randomized with a parallel-group design.
Intention-to-treat analysisThe ITT population is defined as all randomized participants.
Time-to-event endpointPrimary PFS is measured from baseline until disease progression or death.
Log-rank testUsed for the primary and IRC-assessed PFS comparisons.
Hazard ratioUsed to quantify the relative PFS treatment effect.
Cox regressionUsed to estimate PFS hazard ratios.
Stratified analysisBinet and Geographic region were used as Cox-model stratification factors.
CMH testUsed for multiple response and MRD categorical endpoints.
Chi-squared testUsed for approximately 9-month MRD-negativity comparisons.
Confidence intervalsTwo-sided 95% intervals accompany the posted effect estimates.
Multiple endpointsThe registry contains 23 posted outcome measures and 10 statistical analyses.
Safety analysisSerious adverse events are reported by treatment group using affected/at-risk counts.

The educational value is particularly clear when the PFS HR of 0.35 is placed beside the response-rate difference of 13.43 percentage points. These are not competing ways of expressing exactly the same statistic. The first summarizes a longitudinal event process; the second summarizes a binary outcome at a defined assessment.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Calculators

27. Sources

Continue through the Clinical Biostats statistical pathway

Use the methods illustrated by CLL14 to explore time-to-event analysis, categorical-data methods, confidence intervals, and clinical-trial analysis.

28. Record Summary

CLL14 provides a compact example of how modern clinical-trial statistics combine randomized treatment comparisons, a time-to-event primary endpoint, stratified survival analysis, hazard ratios, and several categorical secondary endpoints. The primary investigator-assessed PFS analysis reported an HR of 0.35 with a two-sided 95% CI of 0.23–0.53 and P < 0.0001. The registry also reports an IRC-assessed PFS HR of 0.33, together with differences in response and MRD-negativity rates analyzed using CMH or chi-squared methods.

The most important statistical lesson is that these results should be read according to the scale and structure of each endpoint. Hazard ratios describe relative time-to-event effects; percentage-point differences describe contrasts in binary outcomes at defined assessment times; confidence intervals describe statistical precision; and P-values address evidence against specified null hypotheses rather than measuring effect size.

Clinical Biostats methodology: A trial-results page should distinguish the numerical evidence reported by the registry from the statistical interpretation applied to that evidence. For CLL14, this means preserving the registry's endpoint definitions, analysis populations, methods, effect measures, confidence intervals, and P-values while avoiding unsupported survival summaries, subgroup results, baseline characteristics, or design assumptions not contained in the ClinicalTrials.gov record.