This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information contained in the ClinicalTrials.gov record.
1. Trial at a Glance
CAPTIVATE was a randomized, parallel, triple-masked phase 2 clinical trial of ibrutinib plus venetoclax in subjects with treatment-naive chronic lymphocytic leukemia or small lymphocytic lymphoma. The registry reports 323 enrolled participants, 5 arms, 2 registered primary endpoints, 26 posted outcome measures, and 2 posted statistical analyses.
| Feature | CAPTIVATE |
|---|---|
| Phase | Phase 2 |
| Population | Subjects with treatment-naive chronic lymphocytic leukemia / small lymphocytic lymphoma |
| Design | Randomized, parallel, triple-masked |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 323 |
| Trial arms | 5 |
| Interventions | Ibrutinib; venetoclax; placebo |
| Status | Completed |
| Lead sponsor | Pharmacyclics LLC. |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT02910583 |
2. Clinical Question
The clinical setting was treatment-naive chronic lymphocytic leukemia or small lymphocytic lymphoma. The registered trial intervention set included ibrutinib, venetoclax, and placebo. The principal statistical questions represented in the posted primary analyses concern disease-free survival after randomization among participants with a confirmed undetectable minimal residual disease clinical response and complete response / complete response with incomplete blood count recovery in the FD cohort.
Population
Subjects with treatment-naive chronic lymphocytic leukemia / small lymphocytic lymphoma (CLL/SLL).
Intervention
The trial included ibrutinib and venetoclax as drug interventions. The ClinicalTrials.gov record identifies the intervention combination as ibrutinib plus venetoclax.
Comparator
For the primary disease-free survival comparison, confirmed uMRD randomized participants were assigned to blinded ibrutinib or blinded placebo.
Primary questions
Among the relevant analysis populations, what is the 1-year disease-free survival rate difference between the randomized blinded groups, and what is the complete response / CRi rate in the FD cohort?
3. Trial Design
4. Analysis Populations
The two posted primary analyses use different analysis populations. This distinction is central to interpreting the results because the estimand is not simply a comparison among all 323 enrolled participants.
| Analysis population | Role in the registry analysis |
|---|---|
| Confirmed uMRD Randomized Population | Primary population for the 1-year DFS analysis. Includes participants who achieved confirmed MRD-negative clinical response at the end of the pre-randomization phase and were randomized to either the blinded placebo arm or blinded ibrutinib arm. |
| FD Cohort, Non-Del 17p Population | Per-protocol population used for the primary analysis of the CRR endpoint in the FD cohort. |
The difference between these populations illustrates an important statistical principle: a trial's enrollment count and a particular endpoint's analysis population can represent different sets of participants. An effect estimate must always be interpreted in the population to which it applies.
5. Primary Endpoints
| Endpoint | Registry definition / time frame | Endpoint type | Posted analysis |
|---|---|---|---|
| MRD Cohort: 1-Year Disease-Free Survival (DFS) Rate in Confirmed uMRD Randomized Participants | DFS is defined as time from randomization date to MRD-positive relapse, or disease progression per investigator assessment (per 2008 International Workshop for Chronic Lymphocytic Leukemia [IWCLL] criteria [Halleck et al.]) or death from any cause, whichever occurred first. 1-year DFS estimated using Kaplan-Meier method at 12 months landmark time. Time frame: 1 year after randomization. | Time-to-event | Yes |
| FD Cohort: Complete Response Rate (CRR; Complete Response/Complete Response With Incomplete Blood Count Recovery [CR/CRi]) Rate | CR/CRi rate is defined as the percentage of participants achieving a best overall response of complete response (CR), CR with incomplete blood count recovery (CRi) per 2008 IWCLL criteria (Halleck et al.) on or prior to initiation of subsequent antineoplastic therapy or, if applicable, reintroduction of study treatment, whichever occurred earlier. Time frame: from the first dose of ibrutinib to the first confirmed PD, for a median follow-up of 69.0 months. | Binary / count-rate | Yes |
6. Statistical Methodology
Kaplan-Meier estimation for disease-free survival
The registry specifies the Kaplan-Meier method for estimating the 1-year DFS rate at the 12-month landmark. DFS is a time-to-event endpoint because each participant can be followed from randomization until one of several qualifying events occurs, while some participants may remain event-free at the end of their observed follow-up.
Here, di represents events at event time ti, while ni represents participants at risk immediately before that time. The method allows participants without an observed event to contribute information until censoring.
Z test / Wald approach for the DFS rate difference
The posted analysis compares the 1-year DFS rates using a Z test, normalized here as a Wald / z-test for categorical data. The effect measure is the difference in rates, equivalently a risk difference when the outcome is expressed as a proportion at the specified landmark.
A positive value means the estimated 1-year DFS rate was higher in the blinded ibrutinib group than in the blinded placebo group. The magnitude is expressed in percentage-point units because the registry reports the estimate as a difference in rates.
Exact / Clopper-Pearson binomial method for CRR
The CRR analysis is a binary-response problem: participants either achieved CR/CRi according to the registered definition or they did not. The registry identifies an asymptotic test for binomial proportion in the reported method field and normalizes the method as an exact / Clopper-Pearson binomial approach.
The analysis was explicitly described as per protocol and was based on the FD Cohort, Non-Del 17p Population. That distinction matters because a per-protocol analysis can answer a different question from an analysis that retains everyone according to randomized assignment.
Superiority hypothesis
Both posted primary analyses are classified in the registry as superiority hypotheses. The statistical question is therefore whether the observed treatment-group comparison provides evidence of a difference in the prespecified direction rather than whether a treatment meets a non-inferiority margin.
7. Primary Results: 1-Year Disease-Free Survival
The first posted primary analysis evaluates the MRD Cohort: 1-Year Disease-Free Survival (DFS) Rate in Confirmed uMRD Randomized Participants, with DFS measured 1 year after randomization.
Difference in 1-year DFS rates
95% CI: −1.6 to 10.9 · P = 0.1475
Comparison: blinded ibrutinib vs blinded placebo in the confirmed uMRD randomized population.
| Analysis element | Reported result |
|---|---|
| Endpoint | MRD Cohort: 1-Year Disease-Free Survival (DFS) Rate in Confirmed uMRD Randomized Participants |
| Time frame | 1 year after randomization |
| Population | Confirmed uMRD Randomized Population |
| Comparison | Randomized to Ibrutinib (Blinded) vs randomized to Placebo (Blinded) |
| Method | Z test; normalized as Wald / z-test |
| Effect measure | Difference in Rates / risk difference |
| Estimate | 4.7 |
| 95% CI | −1.6 to 10.9 |
| P-value | 0.1475 |
| Hypothesis | Superiority |
The estimated difference in 1-year DFS rates was 4.7 percentage points, comparing the randomized blinded ibrutinib group with the randomized blinded placebo group in participants who had achieved confirmed uMRD before randomization. In other words, the point estimate favored the blinded ibrutinib group for the specified 1-year DFS rate.
The estimate does not mean that 4.7% of all enrolled participants were prevented from experiencing an event, nor does it represent a hazard ratio. It is a difference between two time-specific DFS rates in a particular analysis population.
The 95% confidence interval extends from −1.6 to 10.9. Because the interval crosses zero, the data are compatible with a small negative difference as well as a larger positive difference under the statistical framework used for the analysis. That interval is therefore important for understanding the uncertainty around the point estimate.
The P-value of 0.1475 is a measure of the statistical evidence against the null comparison under the specified test; it is not a measure of the size of the treatment effect. A P-value does not tell us that the treatment effect is 14.75%, nor does it quantify the probability that the null hypothesis is true.
Because DFS is estimated using Kaplan-Meier methodology, censoring and the timing of events are part of the analysis. The endpoint also combines MRD-positive relapse, investigator-assessed disease progression according to the registered IWCLL framework, and death from any cause. The result should therefore be interpreted as the registered composite DFS endpoint rather than as a result for any one component in isolation.
8. Primary Results: Complete Response Rate
The second posted primary analysis evaluates the FD Cohort: Complete Response Rate (CRR; CR/CRi) Rate. The registry defines this as the percentage of participants achieving a best overall response of CR or CRi according to the registered 2008 IWCLL criteria, within the specified treatment-response window.
Statistical test for CRR
Superiority hypothesis · asymptotic test for binomial proportion / normalized exact binomial framework
Analysis population: FD Cohort, Non-Del 17p Population, per protocol.
| Analysis element | Reported result |
|---|---|
| Endpoint | FD Cohort: Complete Response Rate (CRR; CR/CRi) Rate |
| Time frame | From the first dose of ibrutinib to the first confirmed PD, for a median follow-up of 69.0 months. |
| Analysis population | FD Cohort, Non-Del 17p Population |
| Analysis framework | Per protocol |
| Method as reported | Asymptotic test for binomial proportion |
| Method normalized | Exact / Clopper-Pearson binomial |
| Hypothesis | Superiority |
| P-value | < 0.0001 |
The registry reports a P-value of < 0.0001 for the superiority analysis of CR/CRi rate in the FD Cohort, Non-Del 17p Population. This indicates strong statistical evidence against the null comparison under the reported binomial testing framework.
Importantly, the ClinicalTrials.gov record does not provide a CRR estimate or a confidence interval. The P-value therefore cannot be converted into a response-rate difference from the ClinicalTrials.gov record. A P-value alone does not establish the magnitude of the difference.
The analysis population is also explicitly per protocol. That means the result should not be presented as though it were an all-randomized-participant estimate. Per-protocol analyses can be useful for evaluating outcomes among participants meeting protocol-defined analysis criteria, but their interpretation differs from an intention-to-treat comparison.
The exact numerical size of the CR/CRi treatment effect cannot be determined from the posted statistical-analysis fields reported here without introducing an estimate from another source, which is outside the scope of this page.
9. Secondary Endpoint Results
The registry data state that 26 outcome measures were posted, but the ClinicalTrials.gov recordset identifies only 2 formal statistical analyses, both corresponding to the registered primary endpoints. No additional secondary-endpoint estimates, confidence intervals, or P-values are provided in the ClinicalTrials.gov record.
Accordingly, this page does not assign numerical results to secondary endpoints. For a binary secondary endpoint, a binomial proportion or a comparison of proportions would generally be appropriate depending on the estimand and design. For a time-to-event secondary endpoint, Kaplan-Meier estimation with an appropriate between-group comparison would generally be considered. Those general methods do not constitute reported CAPTIVATE results and are therefore not presented as such.
10. Safety Results
The ClinicalTrials.gov record includes serious adverse-event counts by cohort and, for part of the MRD cohort, by response/randomization group. These figures are reported as affected participants divided by participants at risk.
| Cohort / group | Serious adverse events affected / at risk |
|---|---|
| FD Cohort: Pre-Dose | 1 / 159 |
| MRD Cohort: Pre-Dose | 2 / 164 |
| FD Cohort: All Participants | 37 / 159 |
| MRD Cohort: All Participants | 63 / 164 |
| MRD Cohort: Confirmed uMRD — IbrVen → Ibr | 15 / 43 |
| MRD Cohort: Confirmed uMRD — IbrVen → Pbo | 14 / 43 |
| MRD Cohort: uMRD Not Confirmed — IbrVen → | 13 / 31 |
These are descriptive serious-adverse-event counts rather than an efficacy comparison. In particular, the ClinicalTrials.gov record does not provide a formal statistical test, confidence interval, exposure-adjusted incidence rate, or complete adverse-event profile for these safety categories. The counts should therefore be interpreted as reported safety summaries rather than as a quantitative treatment-effect analysis.
11. Randomization and Blinding
CAPTIVATE is registered as randomized, with a parallel design and triple masking. Randomization is important because it creates the basis for comparing outcomes between assigned groups without relying solely on adjustment for observed baseline characteristics.
Why randomization matters
Random assignment helps balance measured and unmeasured prognostic factors in expectation, supporting causal interpretation of treatment-group differences.
Why blinding matters
Triple masking can reduce the potential for knowledge of treatment assignment to influence treatment administration, assessment, reporting, or other trial processes, depending on who was masked.
Five arms
The registry identifies 5 arms, but the ClinicalTrials.gov record does not specify the complete arm-level allocation structure.
Analysis populations
The primary analyses are restricted to endpoint-specific populations, illustrating that randomization at enrollment does not mean every endpoint uses every enrolled participant.
12. Time-to-Event Endpoint: What DFS Means Here
The CAPTIVATE DFS endpoint is more specific than a generic "time to progression" outcome. The registry defines DFS as the time from randomization to the first of MRD-positive relapse, disease progression per investigator assessment under the registered 2008 IWCLL criteria, or death from any cause.
The 1-year DFS rate is then estimated using the Kaplan-Meier method at the 12-month landmark.
This structure has two statistical consequences. First, participants can have different follow-up durations and may be censored if an event has not been observed by the end of available follow-up. Second, because the endpoint is a composite time-to-event outcome, the reported DFS rate does not isolate the contribution of MRD-positive relapse, progression, and death separately.
13. Statistical Methods Explained
Why was Kaplan-Meier estimation used for 1-year DFS?
DFS is defined as a time from randomization until an event, and not every participant necessarily has an event observed during the available follow-up. Kaplan-Meier estimation is designed for this type of right-censored time-to-event data. It uses the sequence of observed event times and the number of participants at risk at each event time to estimate the probability of remaining event-free over time.
What does a risk difference of 4.7 mean?
The reported 4.7 is a difference in the estimated 1-year DFS rates between the randomized blinded ibrutinib and placebo groups in the confirmed uMRD population. It is an absolute difference, not a ratio. A positive value indicates a higher estimated DFS rate in the first group under the direction of the comparison.
Why is the confidence interval important?
The 95% confidence interval of −1.6 to 10.9 describes uncertainty around the estimated 4.7-point difference under the statistical framework used. It crosses zero, so the ClinicalTrials.gov record does not exclude a small difference in the opposite direction or a larger positive difference at the conventional confidence level represented by the interval.
Why does P = 0.1475 not measure treatment effect size?
A P-value measures how unusual the observed data would be under the specified null hypothesis and test assumptions. It is not an effect-size metric. The effect size here is the reported difference in rates, while the confidence interval communicates uncertainty around that effect estimate.
Why is the CRR analysis described as per protocol?
The registry explicitly states that the primary analysis of the CRR endpoint for the FD cohort was based on the FD Cohort, Non-Del 17p Population only and identifies the analysis as per protocol. This means the result applies to that protocol-defined analysis population rather than automatically to all randomized or enrolled participants.
What is the role of the exact / Clopper-Pearson binomial method?
CR/CRi is a binary response outcome, so the underlying quantity is a binomial proportion. The exact / Clopper-Pearson framework is a method for constructing confidence intervals for a binomial proportion without relying on the same large-sample approximation as a simple Wald interval. The ClinicalTrials.gov record identifies this method as the normalized statistical-method classification for the CRR analysis.
Why should the two primary analyses not be treated as identical?
They answer different statistical questions in different populations. The DFS analysis is a time-to-event comparison among confirmed uMRD randomized participants and reports a rate difference with a confidence interval and P-value. The CRR analysis is a binary-response analysis in the FD Cohort, Non-Del 17p Population and reports a P-value without an estimate or confidence interval in the ClinicalTrials.gov record.
14. Confidence Intervals and Statistical Precision
The DFS result is particularly useful for teaching the distinction between a point estimate and its uncertainty. The estimated rate difference is 4.7, but the associated 95% confidence interval ranges from −1.6 to 10.9.
Point estimate
The point estimate is the single value produced by the analysis: 4.7 percentage points.
Interval estimate
The 95% CI of −1.6 to 10.9 displays the uncertainty surrounding that estimate under the stated statistical framework.
Zero is important
For a difference measure, zero represents no difference between the compared rates. The reported interval crosses zero.
Precision is not significance
A confidence interval communicates precision and plausible values under the model; a P-value addresses evidence against a null hypothesis. Neither alone describes the full clinical meaning of an endpoint.
15. Missing Data, Censoring, and What the Registry Does Not Establish
The DFS endpoint is explicitly analyzed with Kaplan-Meier methodology, so censoring is inherently relevant. Participants who have not experienced the registered DFS event by the end of observed follow-up may contribute information up to their censoring time.
The ClinicalTrials.gov record does not specify a separate missing-data imputation procedure for the DFS endpoint, nor do they describe a particular imputation model for CRR. Accordingly, this page does not attribute an imputation strategy to CAPTIVATE.
16. Multiplicity and Other Design Topics
The ClinicalTrials.gov record identifies 2 registered primary endpoints and classify both formal statistical analyses as superiority tests. However, they do not provide an alpha-allocation scheme, multiplicity-adjustment procedure, interim-analysis plan, or hierarchical testing strategy.
| Design topic | What the ClinicalTrials.gov record establishes |
|---|---|
| Multiplicity | Two registered primary endpoints are present. No specific multiplicity-adjustment procedure is reported. |
| Interim analysis | No interim-analysis method is provided in the ClinicalTrials.gov record. |
| Alpha spending | No alpha-spending procedure is provided. |
| Non-inferiority margin | Not applicable to the registry-reported primary hypotheses; both are classified as superiority. |
| Crossover | No crossover procedure is specified in the ClinicalTrials.gov record. |
| Factorial design | The design is registered as parallel, not factorial. |
| Bayesian methods | No Bayesian statistical method is reported. |
| Stratification | No stratification factors are reported in the ClinicalTrials.gov record. |
The absence of a registry-reported procedure should not be interpreted as evidence that no such procedure existed in the full protocol or statistical analysis plan. It means only that the procedure is not part of the ClinicalTrials.gov record.
17. Understanding the CRR Analysis
The CRR endpoint is defined as the percentage of participants achieving a best overall response of CR or CRi according to the registered 2008 IWCLL criteria, before the specified subsequent-therapy or reintroduction boundary.
The numerator and denominator must be defined according to the protocol's response and analysis rules. The ClinicalTrials.gov record provides the analysis population and P-value but not the numerical response count or estimated CRR.
This is a useful example of why a P-value should not be reported in isolation. The P-value < 0.0001 communicates strong evidence under the specified test, but without the corresponding response-rate estimate, the magnitude of the difference cannot be evaluated from this dataset alone.
18. Longitudinal Trial History
Trial start
CAPTIVATE began according to the registry profile.
Primary completion
The registry profile lists November 12, 2020 as the primary completion date.
Results posted
The registry record is classified as completed and reports 26 outcome measures and 2 statistical analyses.
19. Primary Results in Statistical Context
| Endpoint | Effect / result | 95% CI | P-value | Key interpretation |
|---|---|---|---|---|
| 1-Year DFS rate in confirmed uMRD randomized participants | Difference in rates = 4.7 | −1.6 to 10.9 | 0.1475 | Point estimate favors blinded ibrutinib, but the reported CI crosses zero. |
| CR/CRi rate in FD Cohort, Non-Del 17p Population | Estimate not reported | Not reported | < 0.0001 | Strong statistical evidence under the reported superiority test, but effect magnitude cannot be determined from the registry-reported analysis fields. |
The two results therefore require different styles of interpretation. The DFS analysis gives an effect estimate and a confidence interval, allowing direct discussion of both magnitude and uncertainty. The CRR analysis provides a P-value but no corresponding numerical estimate in the ClinicalTrials.gov record, so its statistical evidence can be described without manufacturing an effect size.
20. Important Limitations and Interpretation Issues
- Endpoint-specific populations: the primary DFS and CRR analyses use different analysis populations, so their results should not be interpreted as if they describe the same participants.
- Per-protocol CRR analysis: the CRR analysis is explicitly described as per protocol and restricted to the FD Cohort, Non-Del 17p Population.
- DFS uncertainty: the 95% CI for the 1-year DFS rate difference extends from −1.6 to 10.9 and therefore includes zero.
- CRR effect size unavailable in the registry-reported analysis: the CRR result provides P < 0.0001 but no estimate or confidence interval in the ClinicalTrials.gov record.
- Composite DFS endpoint: MRD-positive relapse, investigator-assessed progression, and death are combined into one time-to-event endpoint.
- Kaplan-Meier assumptions: the interpretation of a Kaplan-Meier estimate depends on appropriate handling of censoring and follow-up. The ClinicalTrials.gov record does not provide the detailed censoring rules beyond the registered endpoint definition.
- Multiple primary endpoints: two registered primary endpoints are present, but the ClinicalTrials.gov record does not describe how multiplicity was controlled.
- Limited design detail: the ClinicalTrials.gov recordset does not specify the complete 5-arm allocation structure, stratification factors, interim-analysis plan, or crossover rules.
- Registry scope: this page is limited to the ClinicalTrials.gov record and does not import numerical findings from the linked publications.
21. Why This Trial Matters Statistically
CAPTIVATE is a useful teaching case because its primary analyses illustrate several distinct statistical ideas within the same randomized trial. The study combines a landmark time-to-event endpoint with a binary response endpoint, and the two analyses use different populations and statistical frameworks.
| Concept | How it appears in CAPTIVATE |
|---|---|
| Randomization | Randomized allocation in a parallel phase 2 design. |
| Blinding | Triple masking is registered. |
| Kaplan-Meier estimation | Used to estimate 1-year DFS at the 12-month landmark. |
| Time-to-event endpoint | DFS is measured from randomization until MRD-positive relapse, progression, or death. |
| Risk difference | The DFS analysis reports a difference in rates of 4.7. |
| Confidence interval | The DFS estimate has a 95% CI of −1.6 to 10.9. |
| Wald / z-test | The DFS comparison uses a reported Z test, normalized as a Wald / z-test. |
| Binary endpoint | CRR is based on whether participants achieved CR or CRi. |
| Exact binomial methods | The CRR statistical method is normalized as exact / Clopper-Pearson binomial. |
| Per-protocol analysis | The primary CRR analysis is based on the FD Cohort, Non-Del 17p Population and is described as per protocol. |
| Superiority testing | Both posted primary analyses are classified as superiority hypotheses. |
| Analysis populations | The two primary endpoints use different endpoint-specific populations. |
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: CAPTIVATE, NCT02910583.
- PubMed: PubMed record, PMID 41843779.
- PubMed: PubMed record, PMID 37315225.
- PubMed: PubMed record, PMID 34618601.
Continue with the statistical methods
Explore the broader Clinical Biostats tutorials and statistical calculators related to randomized trials, binary endpoints, confidence intervals, and time-to-event analysis.
25. Record Summary
CAPTIVATE provides a compact example of how a randomized clinical trial can require different statistical frameworks for different primary endpoints. Its registered 1-year DFS endpoint is a time-to-event outcome estimated by Kaplan-Meier methodology and compared using a Z test, with a reported difference in rates of 4.7, a 95% confidence interval of −1.6 to 10.9, and P = 0.1475. Its CRR endpoint is a binary response measure analyzed in the FD Cohort, Non-Del 17p Population on a per-protocol basis, with a reported P-value < 0.0001 but no numerical effect estimate or confidence interval in the ClinicalTrials.gov record.
The most important statistical lesson is that these results should be read together with their analysis populations, endpoint definitions, effect measures, and uncertainty. A P-value without an effect estimate does not describe magnitude, while a point estimate without its confidence interval does not adequately communicate precision. CAPTIVATE also demonstrates why the distinction between time-to-event and binary endpoints, and between randomized and per-protocol populations, is essential when interpreting clinical-trial evidence.