← Clinical Trials
Advanced / Metastatic NSCLC Phase 3 Time-to-Event Analysis NCT04538664

PAPILLON: Complete Statistical Analysis of Amivantamab in Advanced or Metastatic NSCLC

An independent statistical review of the randomized phase 3 PAPILLON trial comparing amivantamab plus chemotherapy with chemotherapy alone in participants with advanced or metastatic non-small cell lung cancer characterized by epidermal growth factor receptor (EGFR) exon 20 insertions.

Trial start: 2020-10-13  ·  Primary completion: 2023-05-03  ·  Status: Active, not recruiting
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information posted on ClinicalTrials.gov for PAPILLON.

1. Trial at a Glance

PAPILLON was a randomized, parallel-group, open-label phase 3 trial evaluating amivantamab plus chemotherapy versus chemotherapy alone in participants with advanced or metastatic non-small cell lung cancer characterized by EGFR exon 20 insertions. The registry reports one primary endpoint, progression-free survival assessed by blinded independent central review.

308
Enrolled
2 treatment arms
3
Phase
Randomized phase 3
0.395
PFS HR
95% CI 0.296–0.528
<0.0001
P-value
Superiority analysis
FeaturePAPILLON
Trial namePAPILLON
NCT identifierNCT04538664
PhasePhase 3
Therapeutic areaOncology
ConditionCarcinoma, Non-Small-Cell Lung
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment308.0
Arms2
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted16
Statistical analyses posted1
Lead sponsorJanssen Research & Development, LLC
Sponsor typeIndustry

2. Clinical Question

The central statistical question was whether adding amivantamab to chemotherapy would improve progression-free survival compared with chemotherapy alone in participants with advanced or metastatic non-small cell lung cancer characterized by EGFR exon 20 insertions.

Population

Participants with advanced or metastatic non-small cell lung cancer characterized by epidermal growth factor receptor (EGFR) exon 20 insertions.

Intervention

Amivantamab plus chemotherapy, with amivantamab, pemetrexed, and carboplatin listed among the interventions.

Comparator

Chemotherapy alone, with pemetrexed and carboplatin listed among the interventions.

Primary question

Does amivantamab plus chemotherapy improve progression-free survival compared with chemotherapy alone?

3. Trial Design

Allocation
Randomized
Design model
Parallel
Masking
None
Primary purpose
Treatment
Enrollment
308.0 participants
Primary endpoint
Progression-free survival, a time-to-event endpoint
01
Randomization 308.0 participants
02
Two arms Combination vs chemotherapy alone
03
Follow-up Time-to-event observation
04
BICR RECIST version 1.1 assessment
05
PFS analysis Log-rank test and hazard ratio
ARM A

Amivantamab + Chemotherapy

  • Amivantamab
  • Pemetrexed
  • Carboplatin
ARM B

Chemotherapy Alone

  • Pemetrexed
  • Carboplatin

The trial was open-label because the registry specifies no masking. The primary efficacy assessment nevertheless used blinded independent central review (BICR) of tumor assessments according to RECIST version 1.1. This distinction is statistically important: lack of masking at the treatment level does not mean that the primary radiologic endpoint was necessarily assessed without independent review.

4. Endpoints

EndpointRegistered definition / time frameEndpoint type
Progression-Free Survival (PFS) According to Response Evaluation Criteria in Solid Tumors (RECIST) Version 1.1 as Assessed by Blinded Independent Central Review (BICR) From randomization to either disease progression or death whichever occurs first (up to 29 months) Time-to-event

The registry definition states that PFS is the time from randomization until the date of objective disease progression based on BICR using RECIST version 1.1 or death by any cause in the absence of progression, whichever came first. Participants who had not progressed or died at the time of analysis were censored at the time of the latest date of their last evaluable RECIST version 1.1 assessment.

Why the endpoint definition matters: PFS combines two possible events—objective disease progression and death. A participant who experiences either event reaches the endpoint, while a participant who remains without progression or death can contribute follow-up until the prespecified censoring rule is reached.

5. Statistical Analysis Population

The posted primary analysis used the full analysis set. This included all randomized participants, classified according to their assigned treatment arm regardless of the actual treatment received.

PopulationDefinition / statistical role
Full analysis set All randomized participants, classified according to assigned treatment arm regardless of actual treatment received; used for the posted primary PFS analysis.
Safety population The ClinicalTrials.gov record separately reports serious adverse events by treatment arm.

This treatment-assignment principle is closely related to the intention-to-treat concept. It keeps the comparison anchored to the randomized groups rather than allowing treatment received after randomization to redefine the primary comparison.

6. Statistical Methodology

Log-rank test

The registry reports a log-rank test as the method used for the primary PFS comparison. The log-rank test is designed for comparing time-to-event distributions between treatment groups while accounting for the timing of events and censored observations.

What the log-rank test evaluates
H0: the treatment groups have the same event-time distribution

For PAPILLON, the event is defined by the registered PFS endpoint: objective disease progression based on BICR using RECIST version 1.1 or death, whichever occurs first.

Hazard ratio

The reported effect measure was a hazard ratio (HR). The primary comparison was Arm B, chemotherapy alone, versus Arm A, amivantamab plus chemotherapy.

Reported comparison
HR = 0.395    (Arm B: Chemotherapy Alone vs Arm A: Amivantamab + Chemotherapy)

Because the treatment comparison is reported as Arm B versus Arm A, an HR below 1 corresponds to a lower estimated hazard in Arm B relative to Arm A under that comparison direction. Equivalently, reversing the conceptual direction, the estimate indicates a substantially lower estimated PFS event hazard for Arm A relative to Arm B.

Confidence interval

The analysis reports a two-sided 95% confidence interval from 0.296 to 0.528. A confidence interval communicates statistical precision around the estimated hazard ratio; it is not a range containing the outcomes that individual participants will experience.

Superiority hypothesis

The registry identifies the hypothesis type as superiority. The objective was therefore to evaluate whether the randomized treatment groups differed in the direction specified by the trial's superiority framework, rather than to establish that two treatments were sufficiently similar within a prespecified non-inferiority margin.

Intention-to-treat analysis

The analysis population included all randomized participants according to assigned treatment arm regardless of actual treatment received. This is important because randomization creates the foundation for a fair treatment comparison. Departing from assignment after randomization does not cause a participant to be reclassified into the other randomized group for the primary efficacy comparison.

7. Results: Progression-Free Survival

The registry contains a formal statistical analysis for the primary endpoint. The analysis compared Arm B, chemotherapy alone, with Arm A, amivantamab plus chemotherapy, using a log-rank test and a hazard ratio.

Primary PFS hazard ratio

0.395

95% CI: 0.296–0.528   ·   P < 0.0001

Hypothesis type: superiority  ·  Analysis population: full analysis set

Primary endpointArm B: Chemotherapy AloneArm A: Amivantamab + ChemotherapyReported analysis
Progression-Free Survival Comparator Intervention HR 0.395; 95% CI 0.296–0.528; P < 0.0001; log-rank test
Clinical Biostats interpretation

What the estimate means: The reported HR of 0.395 compares chemotherapy alone with amivantamab plus chemotherapy in the direction specified by the registry. Because the estimate is below 1, the estimated hazard of a PFS event is lower in the amivantamab-plus-chemotherapy group when the comparison is expressed in the reverse direction. A simple transformation, 1 − 0.395, gives approximately 60.5% lower estimated hazard for amivantamab plus chemotherapy relative to chemotherapy alone.

What it does not mean: An HR of 0.395 does not mean that 39.5% of participants avoided progression, that 60.5% of participants benefited, or that each individual participant had exactly a 60.5% reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute probability.

What the confidence interval says: The two-sided 95% CI of 0.296–0.528 describes uncertainty around the estimated hazard ratio under the analysis framework. It indicates that the observed estimate is not being presented as an exact population value. The interval remains below 1, which is consistent with the reported superiority result.

Why the p-value is different: P < 0.0001 addresses the statistical evidence against the null hypothesis in the log-rank comparison. It does not measure the magnitude of the treatment effect. Effect size is described by the hazard ratio and its confidence interval.

Censoring matters: PFS includes participants who have not yet progressed or died at analysis, with censoring according to the registered rule. The hazard ratio therefore uses information from both observed events and censored follow-up rather than simply comparing the percentage of participants who progressed.

Proportional-hazards caution: A single hazard ratio summarizes a relative hazard over the analyzed time period. Its interpretation is most straightforward when the relative hazards are reasonably stable over time. The ClinicalTrials.gov record does not provide a separate assessment of the proportional-hazards assumption, so the HR should not be interpreted as an absolute risk reduction at every time point.

8. Reading the PFS Result in Context

The PAPILLON result illustrates why a time-to-event analysis is different from a simple binary endpoint. Two participants can both remain progression-free at a particular calendar time while having contributed very different amounts of follow-up information. Survival methods preserve the timing of progression, death, and censoring rather than collapsing the entire study into one yes/no outcome.

Relative effect

The HR of 0.395 summarizes the relative difference in the instantaneous PFS event hazard under the reported comparison.

Statistical uncertainty

The 95% CI of 0.296–0.528 shows the uncertainty around the estimated hazard ratio.

Hypothesis evidence

The reported P-value of <0.0001 comes from the log-rank comparison and addresses the null hypothesis rather than effect magnitude.

Endpoint construction

PFS ends at objective disease progression or death, whichever occurs first, with censoring for participants without either event at analysis.

The distinction between these four pieces of information is central to interpreting clinical-trial results. The hazard ratio describes relative effect, the confidence interval describes precision, the p-value addresses evidence against a null hypothesis, and the endpoint definition determines exactly what event is being analyzed.

9. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These figures are presented as affected participants divided by participants at risk.

Safety measureAffected / at risk
Arm A: Amivantamab + Chemotherapy 56/151
Arm B: Chemotherapy Alone 48/155

These safety counts should be interpreted separately from the primary PFS analysis. The PFS comparison uses randomized treatment assignment and a time-to-event endpoint, whereas a serious-adverse-event summary describes participants affected among those at risk for the reported safety measure.

Do not convert these counts into an efficacy conclusion. The serious-adverse-event figures do not measure progression-free survival, and the PFS hazard ratio does not summarize safety. Efficacy and safety answer different statistical questions and should remain separate when interpreting the evidence.

10. Statistical Methods Explained

Why was a log-rank test used?

PFS is a time-to-event endpoint, so the timing of progression or death matters. The log-rank test compares the event-time experience of the randomized groups while accommodating right-censored participants. It is therefore suited to the registered PFS structure rather than treating PFS as a simple binary outcome.

What does a hazard ratio of 0.395 mean?

The HR of 0.395 is the reported relative effect measure for the primary PFS comparison, with chemotherapy alone as Arm B and amivantamab plus chemotherapy as Arm A in the registry's comparison. Expressed in the clinically intuitive reverse direction, 1 − 0.395 = 0.605, so the estimate corresponds to approximately a 60.5% lower estimated hazard for the amivantamab-plus-chemotherapy group relative to chemotherapy alone.

Why does the confidence interval matter?

A point estimate alone can make an effect look more precise than it actually is. The 95% CI of 0.296–0.528 provides an interval representation of statistical uncertainty around the HR. It should not be read as the range of treatment effects experienced by individual participants.

Why does P < 0.0001 not mean the treatment effect is "99.99% certain"?

A p-value is a measure of statistical evidence against a specified null hypothesis under the statistical model. It is not the probability that the treatment works, the probability that the null hypothesis is false, or a direct measure of effect size. The magnitude of the PFS association is described by the HR and its confidence interval.

Why is the full analysis set important?

The posted analysis included all randomized participants and classified them according to assigned treatment regardless of actual treatment received. This preserves the randomized comparison and avoids allowing post-randomization treatment receipt to redefine the primary efficacy groups.

Why does censoring matter for PFS?

Not every participant necessarily experiences progression or death before the analysis. Participants who had not progressed or died were censored at the latest date of their last evaluable RECIST version 1.1 assessment. Censoring allows their available follow-up information to contribute without pretending that their eventual event time is known.

Why does BICR matter?

The primary endpoint was assessed according to RECIST version 1.1 by blinded independent central review. Independent central assessment provides a standardized framework for determining objective progression, which is particularly relevant when the trial itself was not masked.

11. Why the Analysis Is Based on Time-to-Event Methods

A useful way to understand the PAPILLON analysis is to separate three related but different quantities: when an event occurs, whether an event has occurred by a given time, and the relative hazard between treatment groups.

Statistical conceptRole in PAPILLON
Time-to-event endpoint PFS measures time from randomization to disease progression or death, whichever occurs first.
Censoring Participants without progression or death at analysis are censored according to the registered assessment rule.
Log-rank test Formal comparison of the treatment-group time-to-event experience.
Hazard ratio Relative effect measure for the primary PFS comparison.
Confidence interval Quantifies uncertainty around the estimated hazard ratio.
P-value Quantifies evidence against the null hypothesis in the reported comparison.

These methods work together rather than serving interchangeable purposes. The log-rank test supplies the formal comparison, the hazard ratio expresses relative magnitude, and the confidence interval provides a precision statement around that magnitude.

12. Randomization and Causal Interpretation

Randomization is one of the most important design features of PAPILLON. By assigning participants to treatment groups before the efficacy outcome is observed, a randomized comparison creates a framework for attributing differences between groups to treatment assignment rather than simply to observed differences between people who chose or received different therapies.

Before outcome measurement

Treatment assignment is randomized rather than determined by the eventual PFS outcome.

Primary analysis

The full analysis set retains participants according to their randomized assignment.

Endpoint review

Progression is assessed according to RECIST version 1.1 by blinded independent central review.

Statistical comparison

The randomized groups are compared using a time-to-event framework and log-rank testing.

Randomization does not make every numerical characteristic identical between groups, nor does it eliminate statistical uncertainty. Its principal value is the validity of the treatment comparison created by the assignment mechanism.

13. Open-Label Design and Independent Assessment

The registry specifies none for masking. That makes the distinction between treatment assignment and endpoint assessment especially relevant. PAPILLON's primary PFS endpoint was not simply a subjective clinical impression; it was defined using RECIST version 1.1 and assessed by blinded independent central review.

Two different design questions
Treatment masking ≠ independent endpoint review

A trial can be open-label while still using an independent, blinded central process for an imaging-based efficacy assessment. The registry specifically identifies BICR for the primary PFS endpoint.

For statistical interpretation, this distinction matters because open-label treatment can affect behaviors and clinical decisions, while independent central review provides a separate mechanism for standardizing the primary radiologic endpoint.

14. Primary Analysis: What the Numbers Tell Us Together

ComponentPAPILLON primary PFS resultInterpretive role
Effect measure Hazard ratio Relative magnitude of the time-to-event treatment effect.
Estimate 0.395 Point estimate of the reported relative hazard.
95% CI 0.296–0.528 Statistical uncertainty and precision around the estimate.
P-value <0.0001 Evidence against the null hypothesis in the reported log-rank comparison.
Hypothesis Superiority The trial's stated hypothesis framework.
Population Full analysis set All randomized participants classified by assigned treatment.
Method Log-rank test Formal comparison of the time-to-event distributions.

The strongest statistical reading comes from considering these elements jointly. The HR alone does not establish statistical evidence; the p-value alone does not describe effect magnitude; and neither provides the endpoint definition. The complete result is the combination of the endpoint, analysis population, comparison direction, statistical test, effect estimate, confidence interval, and hypothesis framework.

15. What the Hazard Ratio Does — and Does Not — Mean

Statistical interpretation

The reported PFS HR of 0.395 means that, for the comparison as reported by the registry, the estimated instantaneous rate of a PFS event in Arm B relative to Arm A was 0.395. Re-expressing the comparison in the direction of amivantamab plus chemotherapy versus chemotherapy alone gives the reciprocal relationship conceptually; the simpler treatment-effect interpretation of the registry-reported estimate is that the amivantamab-plus-chemotherapy group had an approximately 60.5% lower estimated hazard of the PFS event.

What it does not mean

The HR does not mean that 60.5% of participants avoided progression, that 60.5% of participants were cured, or that every participant experienced the same proportional reduction in risk. Hazard is an instantaneous event-rate concept within a time-to-event model, not an individual probability.

Why the confidence interval matters

The two-sided 95% CI of 0.296–0.528 indicates uncertainty around the estimated HR. A confidence interval is about the statistical estimation procedure and the underlying population parameter; it is not a prediction interval for individual patients.

Why the p-value does not measure effect size

The reported P < 0.0001 indicates strong statistical evidence against the null hypothesis under the log-rank analysis. It does not tell us how large the treatment effect is. The HR supplies the effect-size estimate, while the confidence interval supplies its precision.

16. Limitations

17. Why This Trial Matters Statistically

PAPILLON is a useful teaching example because it connects several core principles of randomized clinical-trial statistics in one primary analysis. The endpoint is time-to-event, the primary comparison uses a log-rank test, the effect is expressed as a hazard ratio, the analysis uses the full randomized analysis set, and progression is determined through an independent central review framework.

ConceptHow it appears in PAPILLON
RandomizationThe trial uses randomized allocation to two parallel treatment arms.
Time-to-event analysisThe primary endpoint is PFS from randomization to progression or death, whichever occurs first.
RECIST assessmentProgression is defined using RECIST version 1.1.
Blinded independent reviewThe primary endpoint is assessed by BICR.
Log-rank testThe reported formal comparison uses the log-rank test.
Hazard ratioThe reported effect measure is HR 0.395.
Confidence intervalThe 95% CI is 0.296–0.528.
P-valueThe reported p-value is <0.0001.
Superiority testingThe registered hypothesis type is superiority.
Full analysis setAll randomized participants are classified according to assigned treatment regardless of actual treatment received.
CensoringParticipants without progression or death at analysis are censored at the latest date of their last evaluable RECIST version 1.1 assessment.
Safety analysisSerious adverse events are reported separately by treatment arm.

18. A Practical Framework for Reading PAPILLON

When evaluating the primary result, it is useful to proceed in a fixed order rather than starting with the p-value.

1. Define the endpoint

First establish exactly what counts as a PFS event and how participants without an event are censored.

2. Check the population

Confirm that the primary analysis uses the full analysis set and preserves randomized treatment assignment.

3. Identify the comparison

Read the treatment direction carefully: the reported comparison is Arm B versus Arm A.

4. Read the effect

Interpret HR 0.395 as a relative time-to-event measure rather than an absolute probability.

5. Read the CI

Use 0.296–0.528 to understand the precision of the estimated hazard ratio.

6. Read the p-value last

Use P < 0.0001 to understand the evidence against the null hypothesis, not the magnitude of benefit.

This sequence prevents one of the most common statistical mistakes in trial interpretation: allowing a very small p-value to substitute for an actual description of the treatment effect.

19. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The randomized comparison of PFS produced a hazard ratio of 0.395, with a two-sided 95% confidence interval of 0.296–0.528 and a log-rank p-value of <0.0001 under a superiority framework.

Clinical interpretation

The numerical result describes a substantially lower estimated hazard of progression or death for amivantamab plus chemotherapy relative to chemotherapy alone. The result should be interpreted together with the endpoint definition, censoring rules, analysis population, and separately reported safety experience.

The statistical result is precise about the comparison being made, but it should not be expanded into claims that are not represented by the reported endpoint. In particular, the HR does not provide a direct estimate of an individual's probability of progression-free survival at a specific time point.

20. Trial Timeline

2020-10-13

Trial start

The PAPILLON trial began on 2020-10-13.

2023-05-03

Primary completion

The trial's primary completion date was 2023-05-03.

Current registry status

Active, not recruiting

The ClinicalTrials.gov record identifies the study status as ACTIVE_NOT_RECRUITING.

21. What Is and Is Not Established by the Posted Analysis

QuestionWhat the ClinicalTrials.gov record supports
Was there a formal primary analysis? Yes. One statistical analysis is posted for the primary PFS endpoint.
What endpoint was analyzed? PFS according to RECIST version 1.1 assessed by BICR.
What statistical test was used? Log-rank test.
What was the effect measure? Hazard ratio.
What was the estimate? 0.395.
What was the 95% CI? 0.296–0.528, two-sided.
What was the p-value? <0.0001.
What hypothesis type was registered? Superiority.
Are median PFS values available in the ClinicalTrials.gov record? No numerical median PFS estimate is included in the ClinicalTrials.gov record.
Are subgroup treatment effects available? No subgroup estimates are included in the registry-reported statistical analysis.
Can a Kaplan-Meier curve be reconstructed? Not from the registry-reported summary statistics alone.

This distinction is important for a complete statistical record. A high-quality analysis should describe what has actually been estimated without filling missing quantities with values derived from memory, unrelated publications, or assumptions about the underlying patient-level data.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Statistical Calculators

24. Sources

Continue through Clinical Biostats

Use the related tutorials and statistical calculators to explore the survival-analysis concepts illustrated by PAPILLON.

25. Record Summary

PAPILLON is a randomized phase 3 trial with a parallel design and no masking, enrolling 308.0 participants across two treatment arms. Its registered primary endpoint is progression-free survival according to RECIST version 1.1 as assessed by blinded independent central review, measured from randomization to disease progression or death, whichever occurs first, with a time frame of up to 29 months.

The posted formal analysis uses the full analysis set, classifies participants according to randomized treatment assignment regardless of actual treatment received, and compares Arm B, chemotherapy alone, with Arm A, amivantamab plus chemotherapy. The reported log-rank analysis gives a hazard ratio of 0.395, a two-sided 95% confidence interval of 0.296–0.528, and P < 0.0001 under a superiority hypothesis.

The most informative way to read this result is not to treat the p-value as a measure of treatment magnitude. Instead, the statistical story is built from the endpoint definition, randomized analysis population, log-rank comparison, hazard ratio, confidence interval, and censoring framework. Together, these describe a substantially lower estimated PFS event hazard for the amivantamab-plus-chemotherapy group relative to chemotherapy alone while preserving the distinction between relative effect, statistical precision, and individual patient outcomes.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. Where the ClinicalTrials.gov record provides a formal estimate, the estimate is reported exactly; where the record does not provide additional numerical results, no unsupported value is reconstructed.