← Clinical Trials
CLL / SLL Phase 2 Randomized NCT02910583

CAPTIVATE: Complete Statistical Analysis of Ibrutinib Plus Venetoclax in CLL/SLL

An independent statistical analysis of the randomized phase 2 CAPTIVATE trial evaluating ibrutinib plus venetoclax in treatment-naive chronic lymphocytic leukemia and small lymphocytic lymphoma, with particular attention to the registered disease-free survival and complete response endpoints.

Completed  ·  September 28, 2016 – November 12, 2020  ·  Enrollment: 323
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information contained in the ClinicalTrials.gov record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

CAPTIVATE was a randomized, parallel, triple-masked phase 2 clinical trial of ibrutinib plus venetoclax in subjects with treatment-naive chronic lymphocytic leukemia or small lymphocytic lymphoma. The registry reports 323 enrolled participants, 5 arms, 2 registered primary endpoints, 26 posted outcome measures, and 2 posted statistical analyses.

323
Enrollment
Participants
5
Trial arms
Parallel design
2
Primary endpoints
Binary + time-to-event
2
Statistical analyses
Both primary
FeatureCAPTIVATE
PhasePhase 2
PopulationSubjects with treatment-naive chronic lymphocytic leukemia / small lymphocytic lymphoma
DesignRandomized, parallel, triple-masked
AllocationRandomized
Primary purposeTreatment
Enrollment323
Trial arms5
InterventionsIbrutinib; venetoclax; placebo
StatusCompleted
Lead sponsorPharmacyclics LLC.
Sponsor typeIndustry
ClinicalTrials.govNCT02910583

2. Clinical Question

The clinical setting was treatment-naive chronic lymphocytic leukemia or small lymphocytic lymphoma. The registered trial intervention set included ibrutinib, venetoclax, and placebo. The principal statistical questions represented in the posted primary analyses concern disease-free survival after randomization among participants with a confirmed undetectable minimal residual disease clinical response and complete response / complete response with incomplete blood count recovery in the FD cohort.

Population

Subjects with treatment-naive chronic lymphocytic leukemia / small lymphocytic lymphoma (CLL/SLL).

Intervention

The trial included ibrutinib and venetoclax as drug interventions. The ClinicalTrials.gov record identifies the intervention combination as ibrutinib plus venetoclax.

Comparator

For the primary disease-free survival comparison, confirmed uMRD randomized participants were assigned to blinded ibrutinib or blinded placebo.

Primary questions

Among the relevant analysis populations, what is the 1-year disease-free survival rate difference between the randomized blinded groups, and what is the complete response / CRi rate in the FD cohort?

3. Trial Design

01
Enroll323 participants
02
RandomizeRandomized allocation
03
AssessClinical and MRD response
04
Primary analysesDFS and CR/CRi rate
05
Follow-upRegistered endpoint time frames
Design model
Parallel-group randomized phase 2 design.
Masking
Triple masking.
Primary endpoint types
One binary endpoint and one time-to-event endpoint.
Trial period
Start: September 28, 2016. Primary completion: November 12, 2020.
Important design boundary: the ClinicalTrials.gov record reports 5 arms and the interventions ibrutinib, venetoclax, and placebo, but they do not provide a complete arm-by-arm treatment schedule or allocation ratio. This analysis therefore does not reconstruct dosing schedules, arm sizes, or treatment sequences beyond what is explicitly present in the posted statistical analyses.

4. Analysis Populations

The two posted primary analyses use different analysis populations. This distinction is central to interpreting the results because the estimand is not simply a comparison among all 323 enrolled participants.

Analysis populationRole in the registry analysis
Confirmed uMRD Randomized PopulationPrimary population for the 1-year DFS analysis. Includes participants who achieved confirmed MRD-negative clinical response at the end of the pre-randomization phase and were randomized to either the blinded placebo arm or blinded ibrutinib arm.
FD Cohort, Non-Del 17p PopulationPer-protocol population used for the primary analysis of the CRR endpoint in the FD cohort.

The difference between these populations illustrates an important statistical principle: a trial's enrollment count and a particular endpoint's analysis population can represent different sets of participants. An effect estimate must always be interpreted in the population to which it applies.

5. Primary Endpoints

EndpointRegistry definition / time frameEndpoint typePosted analysis
MRD Cohort: 1-Year Disease-Free Survival (DFS) Rate in Confirmed uMRD Randomized Participants DFS is defined as time from randomization date to MRD-positive relapse, or disease progression per investigator assessment (per 2008 International Workshop for Chronic Lymphocytic Leukemia [IWCLL] criteria [Halleck et al.]) or death from any cause, whichever occurred first. 1-year DFS estimated using Kaplan-Meier method at 12 months landmark time. Time frame: 1 year after randomization. Time-to-event Yes
FD Cohort: Complete Response Rate (CRR; Complete Response/Complete Response With Incomplete Blood Count Recovery [CR/CRi]) Rate CR/CRi rate is defined as the percentage of participants achieving a best overall response of complete response (CR), CR with incomplete blood count recovery (CRi) per 2008 IWCLL criteria (Halleck et al.) on or prior to initiation of subsequent antineoplastic therapy or, if applicable, reintroduction of study treatment, whichever occurred earlier. Time frame: from the first dose of ibrutinib to the first confirmed PD, for a median follow-up of 69.0 months. Binary / count-rate Yes

6. Statistical Methodology

Kaplan-Meier estimation for disease-free survival

The registry specifies the Kaplan-Meier method for estimating the 1-year DFS rate at the 12-month landmark. DFS is a time-to-event endpoint because each participant can be followed from randomization until one of several qualifying events occurs, while some participants may remain event-free at the end of their observed follow-up.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at event time ti, while ni represents participants at risk immediately before that time. The method allows participants without an observed event to contribute information until censoring.

Z test / Wald approach for the DFS rate difference

The posted analysis compares the 1-year DFS rates using a Z test, normalized here as a Wald / z-test for categorical data. The effect measure is the difference in rates, equivalently a risk difference when the outcome is expressed as a proportion at the specified landmark.

Rate-difference framework
Difference in rates = Rateibrutinib − Rateplacebo

A positive value means the estimated 1-year DFS rate was higher in the blinded ibrutinib group than in the blinded placebo group. The magnitude is expressed in percentage-point units because the registry reports the estimate as a difference in rates.

Exact / Clopper-Pearson binomial method for CRR

The CRR analysis is a binary-response problem: participants either achieved CR/CRi according to the registered definition or they did not. The registry identifies an asymptotic test for binomial proportion in the reported method field and normalizes the method as an exact / Clopper-Pearson binomial approach.

The analysis was explicitly described as per protocol and was based on the FD Cohort, Non-Del 17p Population. That distinction matters because a per-protocol analysis can answer a different question from an analysis that retains everyone according to randomized assignment.

Superiority hypothesis

Both posted primary analyses are classified in the registry as superiority hypotheses. The statistical question is therefore whether the observed treatment-group comparison provides evidence of a difference in the prespecified direction rather than whether a treatment meets a non-inferiority margin.

7. Primary Results: 1-Year Disease-Free Survival

The first posted primary analysis evaluates the MRD Cohort: 1-Year Disease-Free Survival (DFS) Rate in Confirmed uMRD Randomized Participants, with DFS measured 1 year after randomization.

Difference in 1-year DFS rates

4.7

95% CI: −1.6 to 10.9   ·   P = 0.1475

Comparison: blinded ibrutinib vs blinded placebo in the confirmed uMRD randomized population.

Analysis elementReported result
EndpointMRD Cohort: 1-Year Disease-Free Survival (DFS) Rate in Confirmed uMRD Randomized Participants
Time frame1 year after randomization
PopulationConfirmed uMRD Randomized Population
ComparisonRandomized to Ibrutinib (Blinded) vs randomized to Placebo (Blinded)
MethodZ test; normalized as Wald / z-test
Effect measureDifference in Rates / risk difference
Estimate4.7
95% CI−1.6 to 10.9
P-value0.1475
HypothesisSuperiority
Clinical Biostats interpretation

The estimated difference in 1-year DFS rates was 4.7 percentage points, comparing the randomized blinded ibrutinib group with the randomized blinded placebo group in participants who had achieved confirmed uMRD before randomization. In other words, the point estimate favored the blinded ibrutinib group for the specified 1-year DFS rate.

The estimate does not mean that 4.7% of all enrolled participants were prevented from experiencing an event, nor does it represent a hazard ratio. It is a difference between two time-specific DFS rates in a particular analysis population.

The 95% confidence interval extends from −1.6 to 10.9. Because the interval crosses zero, the data are compatible with a small negative difference as well as a larger positive difference under the statistical framework used for the analysis. That interval is therefore important for understanding the uncertainty around the point estimate.

The P-value of 0.1475 is a measure of the statistical evidence against the null comparison under the specified test; it is not a measure of the size of the treatment effect. A P-value does not tell us that the treatment effect is 14.75%, nor does it quantify the probability that the null hypothesis is true.

Because DFS is estimated using Kaplan-Meier methodology, censoring and the timing of events are part of the analysis. The endpoint also combines MRD-positive relapse, investigator-assessed disease progression according to the registered IWCLL framework, and death from any cause. The result should therefore be interpreted as the registered composite DFS endpoint rather than as a result for any one component in isolation.

8. Primary Results: Complete Response Rate

The second posted primary analysis evaluates the FD Cohort: Complete Response Rate (CRR; CR/CRi) Rate. The registry defines this as the percentage of participants achieving a best overall response of CR or CRi according to the registered 2008 IWCLL criteria, within the specified treatment-response window.

Statistical test for CRR

P < 0.0001

Superiority hypothesis · asymptotic test for binomial proportion / normalized exact binomial framework

Analysis population: FD Cohort, Non-Del 17p Population, per protocol.

Analysis elementReported result
EndpointFD Cohort: Complete Response Rate (CRR; CR/CRi) Rate
Time frameFrom the first dose of ibrutinib to the first confirmed PD, for a median follow-up of 69.0 months.
Analysis populationFD Cohort, Non-Del 17p Population
Analysis frameworkPer protocol
Method as reportedAsymptotic test for binomial proportion
Method normalizedExact / Clopper-Pearson binomial
HypothesisSuperiority
P-value< 0.0001
Clinical Biostats interpretation

The registry reports a P-value of < 0.0001 for the superiority analysis of CR/CRi rate in the FD Cohort, Non-Del 17p Population. This indicates strong statistical evidence against the null comparison under the reported binomial testing framework.

Importantly, the ClinicalTrials.gov record does not provide a CRR estimate or a confidence interval. The P-value therefore cannot be converted into a response-rate difference from the ClinicalTrials.gov record. A P-value alone does not establish the magnitude of the difference.

The analysis population is also explicitly per protocol. That means the result should not be presented as though it were an all-randomized-participant estimate. Per-protocol analyses can be useful for evaluating outcomes among participants meeting protocol-defined analysis criteria, but their interpretation differs from an intention-to-treat comparison.

The exact numerical size of the CR/CRi treatment effect cannot be determined from the posted statistical-analysis fields reported here without introducing an estimate from another source, which is outside the scope of this page.

Why the missing estimate matters: statistical significance and effect size are different quantities. Even a very small P-value does not tell the reader whether the absolute response-rate difference is small, moderate, or large. For an educational analysis, the estimate and its uncertainty should be reported whenever they are available; here, only the reported P-value is available in the ClinicalTrials.gov record.

9. Secondary Endpoint Results

The registry data state that 26 outcome measures were posted, but the ClinicalTrials.gov recordset identifies only 2 formal statistical analyses, both corresponding to the registered primary endpoints. No additional secondary-endpoint estimates, confidence intervals, or P-values are provided in the ClinicalTrials.gov record.

Accordingly, this page does not assign numerical results to secondary endpoints. For a binary secondary endpoint, a binomial proportion or a comparison of proportions would generally be appropriate depending on the estimand and design. For a time-to-event secondary endpoint, Kaplan-Meier estimation with an appropriate between-group comparison would generally be considered. Those general methods do not constitute reported CAPTIVATE results and are therefore not presented as such.

10. Safety Results

The ClinicalTrials.gov record includes serious adverse-event counts by cohort and, for part of the MRD cohort, by response/randomization group. These figures are reported as affected participants divided by participants at risk.

Cohort / groupSerious adverse events affected / at risk
FD Cohort: Pre-Dose1 / 159
MRD Cohort: Pre-Dose2 / 164
FD Cohort: All Participants37 / 159
MRD Cohort: All Participants63 / 164
MRD Cohort: Confirmed uMRD — IbrVen → Ibr15 / 43
MRD Cohort: Confirmed uMRD — IbrVen → Pbo14 / 43
MRD Cohort: uMRD Not Confirmed — IbrVen →13 / 31
Safety-data boundary: the registry-reported serious-adverse-event field ends with an incomplete entry after the uMRD-not-confirmed group. The table therefore reports only the complete affected/at-risk counts explicitly registry-reported and does not infer or reconstruct the omitted portion.

These are descriptive serious-adverse-event counts rather than an efficacy comparison. In particular, the ClinicalTrials.gov record does not provide a formal statistical test, confidence interval, exposure-adjusted incidence rate, or complete adverse-event profile for these safety categories. The counts should therefore be interpreted as reported safety summaries rather than as a quantitative treatment-effect analysis.

11. Randomization and Blinding

CAPTIVATE is registered as randomized, with a parallel design and triple masking. Randomization is important because it creates the basis for comparing outcomes between assigned groups without relying solely on adjustment for observed baseline characteristics.

Why randomization matters

Random assignment helps balance measured and unmeasured prognostic factors in expectation, supporting causal interpretation of treatment-group differences.

Why blinding matters

Triple masking can reduce the potential for knowledge of treatment assignment to influence treatment administration, assessment, reporting, or other trial processes, depending on who was masked.

Five arms

The registry identifies 5 arms, but the ClinicalTrials.gov record does not specify the complete arm-level allocation structure.

Analysis populations

The primary analyses are restricted to endpoint-specific populations, illustrating that randomization at enrollment does not mean every endpoint uses every enrolled participant.

12. Time-to-Event Endpoint: What DFS Means Here

The CAPTIVATE DFS endpoint is more specific than a generic "time to progression" outcome. The registry defines DFS as the time from randomization to the first of MRD-positive relapse, disease progression per investigator assessment under the registered 2008 IWCLL criteria, or death from any cause.

Registered DFS event rule
DFS event = MRD-positive relapse OR disease progression OR death, whichever occurs first

The 1-year DFS rate is then estimated using the Kaplan-Meier method at the 12-month landmark.

This structure has two statistical consequences. First, participants can have different follow-up durations and may be censored if an event has not been observed by the end of available follow-up. Second, because the endpoint is a composite time-to-event outcome, the reported DFS rate does not isolate the contribution of MRD-positive relapse, progression, and death separately.

13. Statistical Methods Explained

Why was Kaplan-Meier estimation used for 1-year DFS?

DFS is defined as a time from randomization until an event, and not every participant necessarily has an event observed during the available follow-up. Kaplan-Meier estimation is designed for this type of right-censored time-to-event data. It uses the sequence of observed event times and the number of participants at risk at each event time to estimate the probability of remaining event-free over time.

What does a risk difference of 4.7 mean?

The reported 4.7 is a difference in the estimated 1-year DFS rates between the randomized blinded ibrutinib and placebo groups in the confirmed uMRD population. It is an absolute difference, not a ratio. A positive value indicates a higher estimated DFS rate in the first group under the direction of the comparison.

Why is the confidence interval important?

The 95% confidence interval of −1.6 to 10.9 describes uncertainty around the estimated 4.7-point difference under the statistical framework used. It crosses zero, so the ClinicalTrials.gov record does not exclude a small difference in the opposite direction or a larger positive difference at the conventional confidence level represented by the interval.

Why does P = 0.1475 not measure treatment effect size?

A P-value measures how unusual the observed data would be under the specified null hypothesis and test assumptions. It is not an effect-size metric. The effect size here is the reported difference in rates, while the confidence interval communicates uncertainty around that effect estimate.

Why is the CRR analysis described as per protocol?

The registry explicitly states that the primary analysis of the CRR endpoint for the FD cohort was based on the FD Cohort, Non-Del 17p Population only and identifies the analysis as per protocol. This means the result applies to that protocol-defined analysis population rather than automatically to all randomized or enrolled participants.

What is the role of the exact / Clopper-Pearson binomial method?

CR/CRi is a binary response outcome, so the underlying quantity is a binomial proportion. The exact / Clopper-Pearson framework is a method for constructing confidence intervals for a binomial proportion without relying on the same large-sample approximation as a simple Wald interval. The ClinicalTrials.gov record identifies this method as the normalized statistical-method classification for the CRR analysis.

Why should the two primary analyses not be treated as identical?

They answer different statistical questions in different populations. The DFS analysis is a time-to-event comparison among confirmed uMRD randomized participants and reports a rate difference with a confidence interval and P-value. The CRR analysis is a binary-response analysis in the FD Cohort, Non-Del 17p Population and reports a P-value without an estimate or confidence interval in the ClinicalTrials.gov record.

14. Confidence Intervals and Statistical Precision

The DFS result is particularly useful for teaching the distinction between a point estimate and its uncertainty. The estimated rate difference is 4.7, but the associated 95% confidence interval ranges from −1.6 to 10.9.

Point estimate

The point estimate is the single value produced by the analysis: 4.7 percentage points.

Interval estimate

The 95% CI of −1.6 to 10.9 displays the uncertainty surrounding that estimate under the stated statistical framework.

Zero is important

For a difference measure, zero represents no difference between the compared rates. The reported interval crosses zero.

Precision is not significance

A confidence interval communicates precision and plausible values under the model; a P-value addresses evidence against a null hypothesis. Neither alone describes the full clinical meaning of an endpoint.

15. Missing Data, Censoring, and What the Registry Does Not Establish

The DFS endpoint is explicitly analyzed with Kaplan-Meier methodology, so censoring is inherently relevant. Participants who have not experienced the registered DFS event by the end of observed follow-up may contribute information up to their censoring time.

The ClinicalTrials.gov record does not specify a separate missing-data imputation procedure for the DFS endpoint, nor do they describe a particular imputation model for CRR. Accordingly, this page does not attribute an imputation strategy to CAPTIVATE.

Statistical distinction: censoring in a time-to-event analysis is not the same as ordinary missing binary observations. Kaplan-Meier methods explicitly incorporate time at risk and censoring, whereas a binary response analysis requires a rule for determining which participants are included in the response denominator.

16. Multiplicity and Other Design Topics

The ClinicalTrials.gov record identifies 2 registered primary endpoints and classify both formal statistical analyses as superiority tests. However, they do not provide an alpha-allocation scheme, multiplicity-adjustment procedure, interim-analysis plan, or hierarchical testing strategy.

Design topicWhat the ClinicalTrials.gov record establishes
MultiplicityTwo registered primary endpoints are present. No specific multiplicity-adjustment procedure is reported.
Interim analysisNo interim-analysis method is provided in the ClinicalTrials.gov record.
Alpha spendingNo alpha-spending procedure is provided.
Non-inferiority marginNot applicable to the registry-reported primary hypotheses; both are classified as superiority.
CrossoverNo crossover procedure is specified in the ClinicalTrials.gov record.
Factorial designThe design is registered as parallel, not factorial.
Bayesian methodsNo Bayesian statistical method is reported.
StratificationNo stratification factors are reported in the ClinicalTrials.gov record.

The absence of a registry-reported procedure should not be interpreted as evidence that no such procedure existed in the full protocol or statistical analysis plan. It means only that the procedure is not part of the ClinicalTrials.gov record.

17. Understanding the CRR Analysis

The CRR endpoint is defined as the percentage of participants achieving a best overall response of CR or CRi according to the registered 2008 IWCLL criteria, before the specified subsequent-therapy or reintroduction boundary.

Binary endpoint framework
CRR = participants achieving CR or CRi ÷ participants in the specified analysis population

The numerator and denominator must be defined according to the protocol's response and analysis rules. The ClinicalTrials.gov record provides the analysis population and P-value but not the numerical response count or estimated CRR.

This is a useful example of why a P-value should not be reported in isolation. The P-value < 0.0001 communicates strong evidence under the specified test, but without the corresponding response-rate estimate, the magnitude of the difference cannot be evaluated from this dataset alone.

18. Longitudinal Trial History

September 28, 2016

Trial start

CAPTIVATE began according to the registry profile.

November 12, 2020

Primary completion

The registry profile lists November 12, 2020 as the primary completion date.

Completed

Results posted

The registry record is classified as completed and reports 26 outcome measures and 2 statistical analyses.

19. Primary Results in Statistical Context

EndpointEffect / result95% CIP-valueKey interpretation
1-Year DFS rate in confirmed uMRD randomized participants Difference in rates = 4.7 −1.6 to 10.9 0.1475 Point estimate favors blinded ibrutinib, but the reported CI crosses zero.
CR/CRi rate in FD Cohort, Non-Del 17p Population Estimate not reported Not reported < 0.0001 Strong statistical evidence under the reported superiority test, but effect magnitude cannot be determined from the registry-reported analysis fields.

The two results therefore require different styles of interpretation. The DFS analysis gives an effect estimate and a confidence interval, allowing direct discussion of both magnitude and uncertainty. The CRR analysis provides a P-value but no corresponding numerical estimate in the ClinicalTrials.gov record, so its statistical evidence can be described without manufacturing an effect size.

20. Important Limitations and Interpretation Issues

21. Why This Trial Matters Statistically

CAPTIVATE is a useful teaching case because its primary analyses illustrate several distinct statistical ideas within the same randomized trial. The study combines a landmark time-to-event endpoint with a binary response endpoint, and the two analyses use different populations and statistical frameworks.

ConceptHow it appears in CAPTIVATE
RandomizationRandomized allocation in a parallel phase 2 design.
BlindingTriple masking is registered.
Kaplan-Meier estimationUsed to estimate 1-year DFS at the 12-month landmark.
Time-to-event endpointDFS is measured from randomization until MRD-positive relapse, progression, or death.
Risk differenceThe DFS analysis reports a difference in rates of 4.7.
Confidence intervalThe DFS estimate has a 95% CI of −1.6 to 10.9.
Wald / z-testThe DFS comparison uses a reported Z test, normalized as a Wald / z-test.
Binary endpointCRR is based on whether participants achieved CR or CRi.
Exact binomial methodsThe CRR statistical method is normalized as exact / Clopper-Pearson binomial.
Per-protocol analysisThe primary CRR analysis is based on the FD Cohort, Non-Del 17p Population and is described as per protocol.
Superiority testingBoth posted primary analyses are classified as superiority hypotheses.
Analysis populationsThe two primary endpoints use different endpoint-specific populations.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue with the statistical methods

Explore the broader Clinical Biostats tutorials and statistical calculators related to randomized trials, binary endpoints, confidence intervals, and time-to-event analysis.

25. Record Summary

CAPTIVATE provides a compact example of how a randomized clinical trial can require different statistical frameworks for different primary endpoints. Its registered 1-year DFS endpoint is a time-to-event outcome estimated by Kaplan-Meier methodology and compared using a Z test, with a reported difference in rates of 4.7, a 95% confidence interval of −1.6 to 10.9, and P = 0.1475. Its CRR endpoint is a binary response measure analyzed in the FD Cohort, Non-Del 17p Population on a per-protocol basis, with a reported P-value < 0.0001 but no numerical effect estimate or confidence interval in the ClinicalTrials.gov record.

The most important statistical lesson is that these results should be read together with their analysis populations, endpoint definitions, effect measures, and uncertainty. A P-value without an effect estimate does not describe magnitude, while a point estimate without its confidence interval does not adequately communicate precision. CAPTIVATE also demonstrates why the distinction between time-to-event and binary endpoints, and between randomized and per-protocol populations, is essential when interpreting clinical-trial evidence.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical explanation. Where the ClinicalTrials.gov record provides an estimate and confidence interval, both are interpreted directly. Where only a P-value is available, the analysis does not manufacture an effect size from external sources.