← Clinical Trials
Advanced Hepatocellular Carcinoma Phase 3 Completed NCT03713593

LEAP-002: Complete Statistical Analysis of Lenvatinib Plus Pembrolizumab in Advanced Hepatocellular Carcinoma

An independent statistical review of the randomized phase 3 LEAP-002 trial comparing lenvatinib plus pembrolizumab with lenvatinib plus placebo as first-line therapy in participants with advanced hepatocellular carcinoma.

Trial start: 2018-12-31  ·  Primary completion: 2022-06-21  ·  Enrollment: 794
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

LEAP-002 was a randomized, double-masked, parallel phase 3 trial evaluating lenvatinib plus pembrolizumab versus lenvatinib plus saline placebo as first-line therapy in participants with advanced hepatocellular carcinoma. The trial enrolled 794 participants and had two primary time-to-event endpoints: progression-free survival and overall survival.

794
Enrolled
2 treatment arms
3
Phase
Randomized, parallel
0.834
PFS HR
95% CI 0.712–0.978
0.840
OS HR
95% CI 0.708–0.997
FeatureLEAP-002
Trial nameLEAP-002
ClinicalTrials.gov identifierNCT03713593
PhasePhase 3
StatusCompleted
ConditionCarcinoma, Hepatocellular
Enrollment794
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Primary endpointsProgression-free survival and overall survival
Primary endpoint typeTime-to-event
Results postedYes
Statistical analyses posted7

2. Clinical Question

The clinical question was whether adding pembrolizumab to lenvatinib, compared with lenvatinib plus placebo, produced a difference in the primary time-to-event endpoints of progression-free survival and overall survival when used as first-line therapy in participants with advanced hepatocellular carcinoma.

Population

Participants with advanced hepatocellular carcinoma receiving first-line therapy.

Intervention

Lenvatinib in combination with pembrolizumab.

Comparator

Lenvatinib in combination with saline placebo.

Primary question

How does lenvatinib plus pembrolizumab compare with lenvatinib plus placebo for PFS and OS?

3. Trial Design

01
Randomize 794 participants
02
Two arms Lenvatinib + pembrolizumab or placebo
03
Double mask Double-masked trial design
04
Assess PFS, OS, response and safety
05
Analyze Time-to-event and categorical methods
ARM A

Lenvatinib + Pembrolizumab

  • Lenvatinib
  • Pembrolizumab
ARM B

Lenvatinib + Placebo

  • Lenvatinib
  • Saline placebo

The ClinicalTrials.gov record identifies the trial as randomized, parallel, double-masked, and treatment-focused. The enrollment was 794 participants. The ClinicalTrials.gov record does not provide an allocation ratio, dosing schedule, treatment-cycle schedule, or crossover information, so those details are not inferred here.

Why the masking matters statistically. The double-masked design reduces the opportunity for knowledge of assigned treatment to influence aspects of trial conduct and assessment. For PFS in particular, the registered definition specifies disease progression as determined by blinded independent central review, which provides an additional layer of separation between treatment assignment and progression assessment.

4. Randomization, Stratification, and Analysis Population

The primary efficacy analyses used all randomized participants. Participants were analyzed in the treatment group to which they were randomized. This is an intention-to-treat-style analysis principle and is important because it preserves the treatment comparison created by randomization rather than reassigning participants according to treatment actually received.

Analysis featureRegistry-supported description
Primary efficacy populationAll randomized participants
Analysis assignmentParticipants analyzed in the treatment group to which they were randomized
PFS modelCox proportional-hazards model
OS comparisonLog-rank test
Stratification / covariatesGeographic region, MPVI or extrahepatic spread or both, AFP and ECOG PS

The PFS and OS analysis notes both identify stratification by geographic region, portal-vein invasion/extrahepatic spread status, AFP status, and ECOG performance status. The registry-reported PFS analysis describes these factors as covariates or stratification variables in the Cox model. This adjustment attempts to account for important prognostic structure while preserving the randomized treatment comparison.

5. Primary Endpoints

EndpointRegistered definitionTime framePrimary analysis
Progression-free Survival (PFS) Per RECIST 1.1 PFS was defined as the time from the date of the first documentation of disease progression, as determined by blinded independent central review (BICR) per RECIST 1.1, or death due to any cause (whichever occurred first). Up to approximately 41 months Cox proportional-hazards model
Overall Survival (OS) OS was defined as the time from randomization until death from any cause. Up to approximately 41 months Log-rank test; Cox model used in the analysis description
Registry-definition note: The ClinicalTrials.gov definition for PFS contains a longer RECIST 1.1 progression description, but the provided trial-data extract ends mid-sentence after the lesion-diameter criterion. This page therefore does not reconstruct the omitted wording from outside sources.

6. Statistical Methodology

Time-to-event analysis

Both primary endpoints are time-to-event outcomes. The analysis is therefore concerned not only with whether an event occurred, but also with the time until that event. Participants who have not experienced the event during available follow-up contribute information up to their censoring time.

Cox proportional-hazards model

The primary PFS analysis used a Cox proportional-hazards model. The analysis note specifies Efron's method for handling tied event times and treatment as a covariate, with stratification by geographic region, MPVI or extrahepatic spread or both, AFP and ECOG PS.

Hazard-ratio framework
HR = estimated hazard in the lenvatinib + pembrolizumab group ÷ estimated hazard in the lenvatinib + placebo group

An HR below 1 indicates a lower estimated instantaneous event rate in the pembrolizumab-containing group under the fitted model. It is a relative time-to-event measure, not an absolute probability of experiencing an event.

Log-rank test

The OS analysis used a log-rank test. The registry analysis note describes a one-sided p-value based on a log-rank test stratified by geographic region, portal vein invasion/extrahepatic spread/both, AFP status, and ECOG status. The same analysis description also identifies a Cox model with Efron's method for tied event times and treatment as a covariate.

Score-based confidence intervals for proportions

The secondary ORR analyses used a score-based confidence-interval approach described in the normalized registry methods as the Miettinen-Nurminen, Newcombe, Wilson family. The analysis notes specifically identify the Miettinen-Nurminen method stratified by geographic region, MPVI or extrahepatic spread or both, AFP and ECOG PS.

Risk-difference framework
Risk difference = response proportion in lenvatinib + pembrolizumab − response proportion in lenvatinib + placebo

A positive risk difference means the observed response proportion was higher in the pembrolizumab-containing group. The confidence interval describes statistical uncertainty around that between-group difference.

Covariate adjustment and stratification

Stratification is important because the treatment comparison is not interpreted as though all participants had identical baseline prognostic characteristics. The reported analyses account for geographic region and disease or clinical factors including portal-vein invasion or extrahepatic spread, AFP and ECOG performance status.

7. Primary Results: Progression-Free Survival

The primary PFS comparison used all randomized participants, analyzed according to randomized treatment assignment. The registry reports a hazard ratio comparing lenvatinib plus pembrolizumab with lenvatinib plus placebo of 0.834, with a two-sided 95% confidence interval of 0.712 to 0.978.

Progression-free survival

HR 0.834

95% CI: 0.712–0.978   ·   Two-sided 95% CI

Cox proportional-hazards model with Efron's method of tie handling and stratification/covariate adjustment.

Clinical Biostats interpretation

An HR of 0.834 means that, under the fitted Cox model, the estimated instantaneous rate of progression or death was lower in the lenvatinib-plus-pembrolizumab group than in the lenvatinib-plus-placebo group. Expressed as a simple relative interpretation, 0.834 corresponds to an estimated hazard that is about 16.6% lower in the pembrolizumab-containing group.

The HR does not mean that 16.6% of participants avoided progression, that individual patients experienced exactly a 16.6% reduction in risk, or that median PFS differed by 16.6%. It is a model-based relative comparison of event hazards over time.

The 95% CI of 0.712–0.978 describes uncertainty around the estimated hazard ratio under the statistical model. Because the interval is relatively close to 1 at its upper boundary, the estimate should not be treated as a precise statement of a large effect.

The ClinicalTrials.gov record does not report a PFS p-value. Therefore, no p-value is inferred from the confidence interval. A confidence interval and a p-value answer related but distinct questions, and the p-value is not a measure of effect size.

Interpretation of a Cox HR also depends on the proportional-hazards framework. The registry extract does not provide a diagnostic assessment of the proportional-hazards assumption, so this page does not claim that the assumption was formally verified.

8. Primary Results: Overall Survival

The primary OS comparison used the randomized population and compared lenvatinib plus pembrolizumab with lenvatinib plus placebo. The registry reports an OS hazard ratio of 0.840, with a two-sided 95% confidence interval of 0.708 to 0.997 and a reported p-value of 0.0227.

Overall survival

HR 0.840

95% CI: 0.708–0.997   ·   P = 0.0227

Log-rank comparison with the reported stratification; Cox model with Efron's method of tie handling in the analysis description.

Clinical Biostats interpretation

An HR of 0.840 corresponds to an estimated instantaneous hazard of death about 16.0% lower in the lenvatinib-plus-pembrolizumab group under the fitted time-to-event model.

The HR does not mean that 16.0% of participants survived because of pembrolizumab, nor does it provide an absolute survival probability. It is a relative measure of the death hazard between the randomized groups.

The 95% CI of 0.708–0.997 quantifies uncertainty around the HR. The upper confidence limit is close to 1, so the interval indicates a comparatively narrow separation from the null value in the direction of the observed effect.

The reported P = 0.0227 is evidence from the specified statistical test against its null hypothesis. It does not measure the size or clinical importance of the effect. The registry describes this as a one-sided p-value based on a stratified log-rank test.

Because the analysis involved a time-to-event endpoint, censoring and the proportional-hazards model are important considerations. The ClinicalTrials.gov record does not provide a formal proportional-hazards diagnostic, so no conclusion about that assumption is made here.

Primary endpointComparisonEstimate95% CIP-valueMethod
PFS Lenvatinib + pembrolizumab vs lenvatinib + placebo HR 0.834 0.712–0.978 Not reported in the ClinicalTrials.gov record Cox proportional-hazards model
OS Lenvatinib + pembrolizumab vs lenvatinib + placebo HR 0.840 0.708–0.997 0.0227 Log-rank test; Cox model in analysis description

9. Secondary Results: Objective Response Rate

The registry reports ORR as a secondary binary endpoint, analyzed using a score-based confidence-interval approach and expressed as a difference in percentage between the randomized groups.

ORR per RECIST 1.1

Risk difference 8.5 percentage points

95% CI: 2.8–14.2 percentage points

Score-based CI using the Miettinen-Nurminen method, stratified by geographic region, MPVI or extrahepatic spread or both, AFP and ECOG PS.

The positive risk difference indicates a higher reported objective response proportion in the lenvatinib-plus-pembrolizumab group. The confidence interval quantifies uncertainty around the between-group difference. The registry-reported analysis does not report the separate ORR percentages, so they are not reconstructed from the risk difference.

ORR per modified RECIST

Objective response rate per mRECIST

Risk difference 6.7 percentage points

95% CI: 0.0–13.4 percentage points

Score-based CI using the Miettinen-Nurminen method, with the same reported stratification factors.

The lower confidence limit of 0.0 means that the registry-reported interval reaches the no-difference value. This illustrates why the confidence interval should be presented alongside the point estimate rather than relying only on the estimated difference.

10. Secondary Results: Time to Disease Progression

Time to disease progression was analyzed separately under both RECIST 1.1 and modified RECIST definitions. These endpoints are time-to-event outcomes and were analyzed with Cox proportional-hazards models using the reported covariate and stratification framework.

Secondary endpointEffect estimate95% CIMethod
Time to Disease Progression, RECIST 1.1 HR 0.79 0.66–0.93 Cox proportional-hazards model
Time to Disease Progression, mRECIST HR 0.73 0.61–0.88 Cox proportional-hazards model
Clinical Biostats interpretation

The TTP HR of 0.79 under RECIST 1.1 corresponds to an estimated instantaneous rate of documented disease progression about 21% lower in the pembrolizumab-containing group under the fitted model. The mRECIST HR of 0.73 corresponds to an estimated rate about 27% lower.

These are relative hazard interpretations, not absolute reductions in the probability of progression. They also should not be interpreted as interchangeable with PFS: PFS includes death as an event, whereas the registry labels these secondary endpoints specifically as time to disease progression.

The confidence intervals provide the relevant measure of statistical precision: 0.66–0.93 for RECIST 1.1 TTP and 0.61–0.88 for mRECIST TTP.

11. Secondary Results: Progression-Free Survival by mRECIST

PFS per modified RECIST

HR 0.80

95% CI: 0.68–0.94

Cox proportional-hazards model with treatment as a covariate and reported stratification by geographic region, MPVI or extrahepatic spread or both, AFP and ECOG PS.

Clinical Biostats interpretation

An HR of 0.80 corresponds to an estimated instantaneous hazard about 20% lower in the lenvatinib-plus-pembrolizumab group under the fitted model.

The 95% CI of 0.68–0.94 expresses uncertainty around that relative estimate. As with the primary PFS result, this does not establish an absolute 20% reduction in the proportion of participants who progress or die.

The ClinicalTrials.gov record does not report a p-value for this secondary endpoint, so no formal significance claim is added.

12. Statistical Results at a Glance

EndpointRoleEffect95% CIInterpretive scale
PFS per RECIST 1.1 Primary HR 0.834 0.712–0.978 Lower hazard of progression or death
OS Primary HR 0.840 0.708–0.997 Lower hazard of death
ORR per RECIST 1.1 Secondary Risk difference 8.5 2.8–14.2 Percentage-point difference
TTP per RECIST 1.1 Secondary HR 0.79 0.66–0.93 Lower hazard of disease progression
PFS per mRECIST Secondary HR 0.80 0.68–0.94 Lower hazard of progression or death
ORR per mRECIST Secondary Risk difference 6.7 0.0–13.4 Percentage-point difference
TTP per mRECIST Secondary HR 0.73 0.61–0.88 Lower hazard of disease progression

13. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm using affected participants over participants at risk. The reported counts were 185/395 for lenvatinib plus pembrolizumab and 159/395 for lenvatinib plus placebo.

Serious adverse events
Lenvatinib + pembrolizumab
185/395
Lenvatinib + placebo
159/395
Safety measureLenvatinib + PembrolizumabLenvatinib + Placebo
Serious adverse events185/395159/395

The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or p-value for these serious-adverse-event counts. Accordingly, the counts are presented descriptively rather than converted into a comparative inference.

14. Statistical Methods Explained

Why was a Cox proportional-hazards model used for PFS?

PFS is a time-to-event endpoint. A Cox model uses information about the timing of events while allowing participants without an observed event during follow-up to contribute censored information. It produces a hazard ratio that summarizes the relative event hazard between treatment groups while accommodating the reported stratification and covariate structure.

What does an HR of 0.834 mean?

An HR of 0.834 means that the fitted model estimates the instantaneous event hazard in the lenvatinib-plus-pembrolizumab group at 0.834 times that of the lenvatinib-plus-placebo group. The corresponding relative interpretation is approximately a 16.6% lower estimated hazard. It does not mean that 16.6% fewer participants necessarily experienced the event.

Why is the confidence interval important?

The point estimate is only one estimate from the observed trial data. The 95% CI of 0.712–0.978 for PFS describes the uncertainty around the estimated HR under the specified statistical framework. The width and location of the interval help communicate precision and how close the estimated effect remains to the null value of 1.

Why is the OS p-value not an effect size?

The reported OS p-value of 0.0227 addresses evidence against the null hypothesis under the specified test. It does not tell us that the treatment effect is "2.27%" or that the treatment reduces mortality by a particular percentage. The HR and its confidence interval provide the effect-size information.

Why were stratification factors used?

The reported analyses account for geographic region, MPVI or extrahepatic spread or both, AFP and ECOG performance status. Stratification allows the time-to-event comparison to be made while respecting clinically relevant structure in the randomized population. It can also improve the efficiency and interpretability of the treatment comparison when the stratification variables are prognostically relevant.

What does a risk difference of 8.5 mean for ORR?

A risk difference of 8.5 means that the objective response proportion in the lenvatinib-plus-pembrolizumab group was estimated to be 8.5 percentage points higher than in the lenvatinib-plus-placebo group. It is different from a hazard ratio because ORR is a binary endpoint rather than a time-to-event endpoint.

Why use the Miettinen-Nurminen method for ORR?

The registry identifies a score-based confidence-interval approach, specifically the Miettinen-Nurminen method, for the ORR comparisons. This provides an interval estimate for the difference between proportions while incorporating the structure of the two-group comparison. In LEAP-002, the analysis was additionally stratified by the reported geographic and clinical factors.

15. One-Sided Versus Two-Sided Inference

The ClinicalTrials.gov record contains an important distinction between confidence intervals and the reported OS hypothesis test. The OS result is accompanied by a two-sided 95% confidence interval, while the analysis note describes the p-value as one-sided.

QuantityReported specificationHow to interpret it
OS confidence interval95%, two-sidedInterval estimate around the hazard ratio
OS p-value0.0227, one-sided in analysis noteHypothesis-test result under the reported one-sided log-rank framework
PFS confidence interval95%, two-sidedInterval estimate around the hazard ratio
ORR confidence intervals95%, two-sidedInterval estimates around risk differences

This distinction matters because a p-value and a confidence interval are not interchangeable. The confidence interval communicates uncertainty about the effect estimate, whereas the p-value addresses a specified null hypothesis under a specified testing framework.

16. Multiplicity and Multiple Endpoints

LEAP-002 has two registered primary endpoints, PFS and OS, and the ClinicalTrials.gov record also contain five secondary statistical analyses. The trial therefore illustrates an important principle in confirmatory clinical-trial statistics: the number and hierarchy of endpoints affect how individual statistical results should be interpreted.

Endpoint groupRegistry roleStatistical method
PFS per RECIST 1.1PrimaryCox proportional-hazards model
OSPrimaryLog-rank test; Cox model in analysis description
ORR per RECIST 1.1SecondaryMiettinen-Nurminen score-based CI
TTP per RECIST 1.1SecondaryCox proportional-hazards model
PFS per mRECISTSecondaryCox proportional-hazards model
ORR per mRECISTSecondaryMiettinen-Nurminen score-based CI
TTP per mRECISTSecondaryCox proportional-hazards model
Multiplicity caution: The ClinicalTrials.gov record does not provide an alpha-allocation scheme, hierarchical testing procedure, or multiplicity-adjustment rule for all seven posted statistical analyses. Therefore, individual secondary estimates should not be presented as though they automatically carry the same confirmatory interpretation as a primary endpoint.

17. Non-Inferiority, Bayesian Methods, and Crossover

The ClinicalTrials.gov record does not identify a non-inferiority margin, a Bayesian analysis, or a crossover analysis. The primary PFS analysis is labeled a descriptive assessment in the registry data rather than a non-inferiority analysis, and the OS analysis is explicitly associated with a superiority hypothesis.

No non-inferiority margin reported

The ClinicalTrials.gov record does not provide a non-inferiority margin, so the results should not be interpreted through a non-inferiority framework.

No Bayesian method reported

The normalized methods identify Cox, log-rank, and score-based proportion methods; no Bayesian method is listed.

No crossover analysis reported

The ClinicalTrials.gov record does not identify a crossover design or a crossover-adjusted analysis.

Superiority hypothesis for OS

The OS statistical analysis is explicitly classified as a superiority analysis in the ClinicalTrials.gov record.

18. Missing Data and Imputation

The registry analysis records do not specify a missing-data or imputation method for the primary or secondary analyses. For time-to-event endpoints, censoring is intrinsic to the analysis because participants may reach the end of available follow-up without experiencing the event. However, the registry extract reported here does not provide detailed censoring rules or sensitivity analyses.

Why this matters: A time-to-event analysis does not simply discard every participant who has not experienced the endpoint. Their observed follow-up contributes information until censoring. The exact handling of censoring can influence estimates, so detailed protocol and statistical-analysis-plan rules would be needed for a complete assessment beyond the ClinicalTrials.gov record.

19. Limitations of the Reported Statistical Evidence

20. Why This Trial Matters Statistically

LEAP-002 is a useful teaching case because its registry results bring together randomized treatment comparison, two primary time-to-event endpoints, blinded assessment, stratified Cox modeling, log-rank testing, and score-based confidence intervals for binary response endpoints.

ConceptHow it appears in LEAP-002
RandomizationRandomized phase 3 parallel-group design with 794 enrolled participants.
BlindingDouble-masked trial design.
Time-to-event endpointsPFS and OS are the two registered primary endpoints.
Hazard ratioUsed for PFS, OS, TTP, and mRECIST PFS analyses.
Cox proportional-hazards modelPrimary PFS method and also reported in the OS and secondary time-to-event analysis descriptions.
Log-rank testReported method for the primary OS comparison.
Stratified analysisGeographic region, MPVI or extrahepatic spread or both, AFP and ECOG PS are included in the reported analysis framework.
Risk differenceUsed for ORR comparisons under RECIST 1.1 and mRECIST.
Score-based CIMiettinen-Nurminen method used for ORR difference estimates.
Analysis populationAll randomized participants analyzed according to randomized treatment group for the primary efficacy analyses.
MultiplicityTwo primary endpoints plus multiple secondary statistical analyses make endpoint hierarchy important.

21. Statistical Methods Explained: A Deeper Walkthrough

Why does PFS use the time of progression or death?

The registered PFS endpoint combines two possible events: documented disease progression or death from any cause, whichever occurs first. This means a participant who dies before a documented progression still contributes a PFS event. The endpoint therefore captures both disease-control failure and death within one time-to-event measure.

Why is the OS endpoint simpler to define?

The registered OS definition is the time from randomization until death from any cause. Unlike PFS, it does not require radiologic determination of progression. The event definition is therefore directly tied to survival status.

Why does blinded independent central review matter for PFS?

Progression can require interpretation of imaging against prespecified criteria. The registry specifically states that progression for PFS was determined by blinded independent central review according to RECIST 1.1. Blinding helps reduce the opportunity for knowledge of randomized treatment to influence the classification of progression.

Why is the analysis population important?

The primary analyses use all randomized participants and analyze them according to randomized treatment assignment. This means the treatment comparison remains anchored to the randomization process. If participants discontinue treatment or receive other therapy, simply removing them from the primary efficacy population could compromise that randomized comparison.

Why are ORR and PFS not interchangeable?

ORR is a binary endpoint: a participant either meets the response definition or does not. PFS is a time-to-event endpoint: the timing of progression or death is part of the outcome. Consequently, ORR is naturally summarized with proportions and risk differences, while PFS is analyzed using survival methods such as Cox models.

Why can two hazard ratios not be directly treated as percentage differences in response?

A hazard ratio compares event hazards over time. A risk difference compares probabilities or proportions. Although an HR of 0.80 can be described as an estimated 20% lower hazard, it cannot be translated into a 20-percentage-point difference in event-free patients without additional information.

22. What the Primary Hazard Ratios Do — and Do Not — Mean

PFS

The PFS HR of 0.834 indicates an estimated lower instantaneous hazard of progression or death for lenvatinib plus pembrolizumab relative to lenvatinib plus placebo under the reported Cox model. The corresponding relative interpretation is approximately a 16.6% lower estimated hazard.

It does not mean that 16.6% of patients were protected from progression, nor does it provide an absolute probability of remaining progression-free at any particular time.

OS

The OS HR of 0.840 indicates an estimated lower instantaneous hazard of death for lenvatinib plus pembrolizumab relative to lenvatinib plus placebo under the reported time-to-event analysis. The corresponding relative interpretation is approximately a 16.0% lower estimated hazard.

It does not mean that 16.0% more patients survived, nor does it imply a particular difference in median survival or survival probability at a particular time point.

Confidence intervals

The 95% CI for PFS is 0.712–0.978, while the 95% CI for OS is 0.708–0.997. These intervals describe uncertainty around the corresponding hazard-ratio estimates under the reported statistical framework. They do not describe the range of individual patient outcomes.

23. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The registry-reported analyses produced hazard ratios below 1 for both primary time-to-event endpoints. The OS analysis reports P = 0.0227, while the registry-reported PFS record reports a 95% confidence interval without a p-value.

Clinical interpretation

The ClinicalTrials.gov record describes differences between randomized groups in PFS, OS, disease progression, and objective response. The ClinicalTrials.gov record does not include median survival or time-specific survival estimates, so those outcomes are not inferred.

24. Overall Statistical Reading of the Trial

The statistical story of LEAP-002 is internally consistent across several endpoint types in the ClinicalTrials.gov record. The primary PFS HR is 0.834, and the primary OS HR is 0.840. The secondary time-to-progression analyses also have HRs below 1: 0.79 under RECIST 1.1 and 0.73 under mRECIST. Secondary PFS under mRECIST has an HR of 0.80.

The response endpoints use a different effect measure. The reported ORR difference is 8.5 percentage points under RECIST 1.1 and 6.7 percentage points under mRECIST. Their confidence intervals are 2.8–14.2 and 0.0–13.4, respectively.

These results should not be collapsed into one statistic. Time-to-event endpoints and binary response endpoints answer different questions, and the choice of effect measure follows from the structure of the endpoint. The primary OS result additionally includes a reported one-sided p-value of 0.0227, while the PFS result is reported with a two-sided 95% confidence interval of 0.712–0.978 but no p-value in the ClinicalTrials.gov record.

Statistical bottom line: The registry-reported LEAP-002 results show hazard-ratio estimates below 1 for both primary endpoints and all four reported secondary time-to-event analyses, together with positive reported risk differences for both ORR endpoints. The appropriate interpretation remains endpoint-specific: hazard ratios describe relative event hazards, risk differences describe absolute percentage-point differences in response, confidence intervals describe precision, and p-values address prespecified hypothesis tests rather than effect magnitude.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through Clinical Biostats

Connect this trial's endpoints and statistical methods to deeper tutorials and practical statistical calculators.

28. Record Summary

LEAP-002 provides a useful example of how a randomized phase 3 oncology trial can combine multiple statistical frameworks within a single evidence package. The two primary endpoints are time-to-event outcomes, with PFS analyzed using a stratified Cox proportional-hazards model and OS evaluated using a stratified log-rank framework with a Cox-model description. The secondary endpoints add both additional time-to-event analyses and binary response analyses using score-based confidence intervals.

The most important statistical distinction is between effect size, precision, and hypothesis testing. The hazard ratios describe relative event hazards; the risk differences describe percentage-point differences in response; the confidence intervals describe uncertainty around those estimates; and the reported OS p-value addresses the specified superiority test. None of these quantities alone provides a complete description of treatment effect.

Clinical Biostats methodology: A trial-results page should not merely repeat numerical registry fields. The goal is to explain what each estimate measures, how the statistical method produces it, what the confidence interval contributes, and which conclusions are supported by the available data without extending beyond the reported evidence.