This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information posted on ClinicalTrials.gov for PAPILLON.
1. Trial at a Glance
PAPILLON was a randomized, parallel-group, open-label phase 3 trial evaluating amivantamab plus chemotherapy versus chemotherapy alone in participants with advanced or metastatic non-small cell lung cancer characterized by EGFR exon 20 insertions. The registry reports one primary endpoint, progression-free survival assessed by blinded independent central review.
| Feature | PAPILLON |
|---|---|
| Trial name | PAPILLON |
| NCT identifier | NCT04538664 |
| Phase | Phase 3 |
| Therapeutic area | Oncology |
| Condition | Carcinoma, Non-Small-Cell Lung |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 308.0 |
| Arms | 2 |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 16 |
| Statistical analyses posted | 1 |
| Lead sponsor | Janssen Research & Development, LLC |
| Sponsor type | Industry |
2. Clinical Question
The central statistical question was whether adding amivantamab to chemotherapy would improve progression-free survival compared with chemotherapy alone in participants with advanced or metastatic non-small cell lung cancer characterized by EGFR exon 20 insertions.
Population
Participants with advanced or metastatic non-small cell lung cancer characterized by epidermal growth factor receptor (EGFR) exon 20 insertions.
Intervention
Amivantamab plus chemotherapy, with amivantamab, pemetrexed, and carboplatin listed among the interventions.
Comparator
Chemotherapy alone, with pemetrexed and carboplatin listed among the interventions.
Primary question
Does amivantamab plus chemotherapy improve progression-free survival compared with chemotherapy alone?
3. Trial Design
Amivantamab + Chemotherapy
- Amivantamab
- Pemetrexed
- Carboplatin
Chemotherapy Alone
- Pemetrexed
- Carboplatin
The trial was open-label because the registry specifies no masking. The primary efficacy assessment nevertheless used blinded independent central review (BICR) of tumor assessments according to RECIST version 1.1. This distinction is statistically important: lack of masking at the treatment level does not mean that the primary radiologic endpoint was necessarily assessed without independent review.
4. Endpoints
| Endpoint | Registered definition / time frame | Endpoint type |
|---|---|---|
| Progression-Free Survival (PFS) According to Response Evaluation Criteria in Solid Tumors (RECIST) Version 1.1 as Assessed by Blinded Independent Central Review (BICR) | From randomization to either disease progression or death whichever occurs first (up to 29 months) | Time-to-event |
The registry definition states that PFS is the time from randomization until the date of objective disease progression based on BICR using RECIST version 1.1 or death by any cause in the absence of progression, whichever came first. Participants who had not progressed or died at the time of analysis were censored at the time of the latest date of their last evaluable RECIST version 1.1 assessment.
5. Statistical Analysis Population
The posted primary analysis used the full analysis set. This included all randomized participants, classified according to their assigned treatment arm regardless of the actual treatment received.
| Population | Definition / statistical role |
|---|---|
| Full analysis set | All randomized participants, classified according to assigned treatment arm regardless of actual treatment received; used for the posted primary PFS analysis. |
| Safety population | The ClinicalTrials.gov record separately reports serious adverse events by treatment arm. |
This treatment-assignment principle is closely related to the intention-to-treat concept. It keeps the comparison anchored to the randomized groups rather than allowing treatment received after randomization to redefine the primary comparison.
6. Statistical Methodology
Log-rank test
The registry reports a log-rank test as the method used for the primary PFS comparison. The log-rank test is designed for comparing time-to-event distributions between treatment groups while accounting for the timing of events and censored observations.
For PAPILLON, the event is defined by the registered PFS endpoint: objective disease progression based on BICR using RECIST version 1.1 or death, whichever occurs first.
Hazard ratio
The reported effect measure was a hazard ratio (HR). The primary comparison was Arm B, chemotherapy alone, versus Arm A, amivantamab plus chemotherapy.
Because the treatment comparison is reported as Arm B versus Arm A, an HR below 1 corresponds to a lower estimated hazard in Arm B relative to Arm A under that comparison direction. Equivalently, reversing the conceptual direction, the estimate indicates a substantially lower estimated PFS event hazard for Arm A relative to Arm B.
Confidence interval
The analysis reports a two-sided 95% confidence interval from 0.296 to 0.528. A confidence interval communicates statistical precision around the estimated hazard ratio; it is not a range containing the outcomes that individual participants will experience.
Superiority hypothesis
The registry identifies the hypothesis type as superiority. The objective was therefore to evaluate whether the randomized treatment groups differed in the direction specified by the trial's superiority framework, rather than to establish that two treatments were sufficiently similar within a prespecified non-inferiority margin.
Intention-to-treat analysis
The analysis population included all randomized participants according to assigned treatment arm regardless of actual treatment received. This is important because randomization creates the foundation for a fair treatment comparison. Departing from assignment after randomization does not cause a participant to be reclassified into the other randomized group for the primary efficacy comparison.
7. Results: Progression-Free Survival
The registry contains a formal statistical analysis for the primary endpoint. The analysis compared Arm B, chemotherapy alone, with Arm A, amivantamab plus chemotherapy, using a log-rank test and a hazard ratio.
Primary PFS hazard ratio
95% CI: 0.296–0.528 · P < 0.0001
Hypothesis type: superiority · Analysis population: full analysis set
| Primary endpoint | Arm B: Chemotherapy Alone | Arm A: Amivantamab + Chemotherapy | Reported analysis |
|---|---|---|---|
| Progression-Free Survival | Comparator | Intervention | HR 0.395; 95% CI 0.296–0.528; P < 0.0001; log-rank test |
What the estimate means: The reported HR of 0.395 compares chemotherapy alone with amivantamab plus chemotherapy in the direction specified by the registry. Because the estimate is below 1, the estimated hazard of a PFS event is lower in the amivantamab-plus-chemotherapy group when the comparison is expressed in the reverse direction. A simple transformation, 1 − 0.395, gives approximately 60.5% lower estimated hazard for amivantamab plus chemotherapy relative to chemotherapy alone.
What it does not mean: An HR of 0.395 does not mean that 39.5% of participants avoided progression, that 60.5% of participants benefited, or that each individual participant had exactly a 60.5% reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute probability.
What the confidence interval says: The two-sided 95% CI of 0.296–0.528 describes uncertainty around the estimated hazard ratio under the analysis framework. It indicates that the observed estimate is not being presented as an exact population value. The interval remains below 1, which is consistent with the reported superiority result.
Why the p-value is different: P < 0.0001 addresses the statistical evidence against the null hypothesis in the log-rank comparison. It does not measure the magnitude of the treatment effect. Effect size is described by the hazard ratio and its confidence interval.
Censoring matters: PFS includes participants who have not yet progressed or died at analysis, with censoring according to the registered rule. The hazard ratio therefore uses information from both observed events and censored follow-up rather than simply comparing the percentage of participants who progressed.
Proportional-hazards caution: A single hazard ratio summarizes a relative hazard over the analyzed time period. Its interpretation is most straightforward when the relative hazards are reasonably stable over time. The ClinicalTrials.gov record does not provide a separate assessment of the proportional-hazards assumption, so the HR should not be interpreted as an absolute risk reduction at every time point.
8. Reading the PFS Result in Context
The PAPILLON result illustrates why a time-to-event analysis is different from a simple binary endpoint. Two participants can both remain progression-free at a particular calendar time while having contributed very different amounts of follow-up information. Survival methods preserve the timing of progression, death, and censoring rather than collapsing the entire study into one yes/no outcome.
Relative effect
The HR of 0.395 summarizes the relative difference in the instantaneous PFS event hazard under the reported comparison.
Statistical uncertainty
The 95% CI of 0.296–0.528 shows the uncertainty around the estimated hazard ratio.
Hypothesis evidence
The reported P-value of <0.0001 comes from the log-rank comparison and addresses the null hypothesis rather than effect magnitude.
Endpoint construction
PFS ends at objective disease progression or death, whichever occurs first, with censoring for participants without either event at analysis.
The distinction between these four pieces of information is central to interpreting clinical-trial results. The hazard ratio describes relative effect, the confidence interval describes precision, the p-value addresses evidence against a null hypothesis, and the endpoint definition determines exactly what event is being analyzed.
9. Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by treatment arm. These figures are presented as affected participants divided by participants at risk.
| Safety measure | Affected / at risk |
|---|---|
| Arm A: Amivantamab + Chemotherapy | 56/151 |
| Arm B: Chemotherapy Alone | 48/155 |
These safety counts should be interpreted separately from the primary PFS analysis. The PFS comparison uses randomized treatment assignment and a time-to-event endpoint, whereas a serious-adverse-event summary describes participants affected among those at risk for the reported safety measure.
10. Statistical Methods Explained
Why was a log-rank test used?
PFS is a time-to-event endpoint, so the timing of progression or death matters. The log-rank test compares the event-time experience of the randomized groups while accommodating right-censored participants. It is therefore suited to the registered PFS structure rather than treating PFS as a simple binary outcome.
What does a hazard ratio of 0.395 mean?
The HR of 0.395 is the reported relative effect measure for the primary PFS comparison, with chemotherapy alone as Arm B and amivantamab plus chemotherapy as Arm A in the registry's comparison. Expressed in the clinically intuitive reverse direction, 1 − 0.395 = 0.605, so the estimate corresponds to approximately a 60.5% lower estimated hazard for the amivantamab-plus-chemotherapy group relative to chemotherapy alone.
Why does the confidence interval matter?
A point estimate alone can make an effect look more precise than it actually is. The 95% CI of 0.296–0.528 provides an interval representation of statistical uncertainty around the HR. It should not be read as the range of treatment effects experienced by individual participants.
Why does P < 0.0001 not mean the treatment effect is "99.99% certain"?
A p-value is a measure of statistical evidence against a specified null hypothesis under the statistical model. It is not the probability that the treatment works, the probability that the null hypothesis is false, or a direct measure of effect size. The magnitude of the PFS association is described by the HR and its confidence interval.
Why is the full analysis set important?
The posted analysis included all randomized participants and classified them according to assigned treatment regardless of actual treatment received. This preserves the randomized comparison and avoids allowing post-randomization treatment receipt to redefine the primary efficacy groups.
Why does censoring matter for PFS?
Not every participant necessarily experiences progression or death before the analysis. Participants who had not progressed or died were censored at the latest date of their last evaluable RECIST version 1.1 assessment. Censoring allows their available follow-up information to contribute without pretending that their eventual event time is known.
Why does BICR matter?
The primary endpoint was assessed according to RECIST version 1.1 by blinded independent central review. Independent central assessment provides a standardized framework for determining objective progression, which is particularly relevant when the trial itself was not masked.
11. Why the Analysis Is Based on Time-to-Event Methods
A useful way to understand the PAPILLON analysis is to separate three related but different quantities: when an event occurs, whether an event has occurred by a given time, and the relative hazard between treatment groups.
| Statistical concept | Role in PAPILLON |
|---|---|
| Time-to-event endpoint | PFS measures time from randomization to disease progression or death, whichever occurs first. |
| Censoring | Participants without progression or death at analysis are censored according to the registered assessment rule. |
| Log-rank test | Formal comparison of the treatment-group time-to-event experience. |
| Hazard ratio | Relative effect measure for the primary PFS comparison. |
| Confidence interval | Quantifies uncertainty around the estimated hazard ratio. |
| P-value | Quantifies evidence against the null hypothesis in the reported comparison. |
These methods work together rather than serving interchangeable purposes. The log-rank test supplies the formal comparison, the hazard ratio expresses relative magnitude, and the confidence interval provides a precision statement around that magnitude.
12. Randomization and Causal Interpretation
Randomization is one of the most important design features of PAPILLON. By assigning participants to treatment groups before the efficacy outcome is observed, a randomized comparison creates a framework for attributing differences between groups to treatment assignment rather than simply to observed differences between people who chose or received different therapies.
Before outcome measurement
Treatment assignment is randomized rather than determined by the eventual PFS outcome.
Primary analysis
The full analysis set retains participants according to their randomized assignment.
Endpoint review
Progression is assessed according to RECIST version 1.1 by blinded independent central review.
Statistical comparison
The randomized groups are compared using a time-to-event framework and log-rank testing.
Randomization does not make every numerical characteristic identical between groups, nor does it eliminate statistical uncertainty. Its principal value is the validity of the treatment comparison created by the assignment mechanism.
13. Open-Label Design and Independent Assessment
The registry specifies none for masking. That makes the distinction between treatment assignment and endpoint assessment especially relevant. PAPILLON's primary PFS endpoint was not simply a subjective clinical impression; it was defined using RECIST version 1.1 and assessed by blinded independent central review.
A trial can be open-label while still using an independent, blinded central process for an imaging-based efficacy assessment. The registry specifically identifies BICR for the primary PFS endpoint.
For statistical interpretation, this distinction matters because open-label treatment can affect behaviors and clinical decisions, while independent central review provides a separate mechanism for standardizing the primary radiologic endpoint.
14. Primary Analysis: What the Numbers Tell Us Together
| Component | PAPILLON primary PFS result | Interpretive role |
|---|---|---|
| Effect measure | Hazard ratio | Relative magnitude of the time-to-event treatment effect. |
| Estimate | 0.395 | Point estimate of the reported relative hazard. |
| 95% CI | 0.296–0.528 | Statistical uncertainty and precision around the estimate. |
| P-value | <0.0001 | Evidence against the null hypothesis in the reported log-rank comparison. |
| Hypothesis | Superiority | The trial's stated hypothesis framework. |
| Population | Full analysis set | All randomized participants classified by assigned treatment. |
| Method | Log-rank test | Formal comparison of the time-to-event distributions. |
The strongest statistical reading comes from considering these elements jointly. The HR alone does not establish statistical evidence; the p-value alone does not describe effect magnitude; and neither provides the endpoint definition. The complete result is the combination of the endpoint, analysis population, comparison direction, statistical test, effect estimate, confidence interval, and hypothesis framework.
15. What the Hazard Ratio Does — and Does Not — Mean
The reported PFS HR of 0.395 means that, for the comparison as reported by the registry, the estimated instantaneous rate of a PFS event in Arm B relative to Arm A was 0.395. Re-expressing the comparison in the direction of amivantamab plus chemotherapy versus chemotherapy alone gives the reciprocal relationship conceptually; the simpler treatment-effect interpretation of the registry-reported estimate is that the amivantamab-plus-chemotherapy group had an approximately 60.5% lower estimated hazard of the PFS event.
The HR does not mean that 60.5% of participants avoided progression, that 60.5% of participants were cured, or that every participant experienced the same proportional reduction in risk. Hazard is an instantaneous event-rate concept within a time-to-event model, not an individual probability.
The two-sided 95% CI of 0.296–0.528 indicates uncertainty around the estimated HR. A confidence interval is about the statistical estimation procedure and the underlying population parameter; it is not a prediction interval for individual patients.
The reported P < 0.0001 indicates strong statistical evidence against the null hypothesis under the log-rank analysis. It does not tell us how large the treatment effect is. The HR supplies the effect-size estimate, while the confidence interval supplies its precision.
16. Limitations
- Single posted formal analysis: The ClinicalTrials.gov record contains one statistical analysis, corresponding to the primary PFS endpoint. The page therefore does not infer additional efficacy estimates that are not included in the ClinicalTrials.gov record.
- No median PFS estimate reported: The available data provides the hazard ratio, confidence interval, and p-value but does not provide a median PFS estimate. A median should not be reconstructed from the HR.
- No subgroup estimates reported: The ClinicalTrials.gov record does not provide subgroup-specific hazard ratios or interaction tests. Treatment-effect heterogeneity therefore cannot be assessed from the available numerical results.
- No Kaplan-Meier coordinates reported: The available information is insufficient to reconstruct a valid Kaplan-Meier curve. A curve should not be generated from the HR and p-value alone.
- Hazard-ratio interpretation: A single HR summarizes a relative time-to-event effect and should not be treated as an absolute risk difference or as a constant individual-level risk reduction.
- Censoring: PFS analysis necessarily includes censoring for participants without progression or death at analysis. Interpretation depends on the validity of the underlying censoring framework.
- Open-label treatment: The registry specifies no masking. Although the primary endpoint was assessed by blinded independent central review, other aspects of an open-label trial can differ from those of a fully masked study.
- Safety and efficacy are distinct: Serious adverse-event counts cannot be incorporated into the PFS hazard ratio and should be evaluated as separate components of the evidence.
- Registry scope: The analysis presented here is limited to the numerical and methodological information posted on ClinicalTrials.gov for the ClinicalTrials.gov record and does not infer additional results from publications or other sources.
17. Why This Trial Matters Statistically
PAPILLON is a useful teaching example because it connects several core principles of randomized clinical-trial statistics in one primary analysis. The endpoint is time-to-event, the primary comparison uses a log-rank test, the effect is expressed as a hazard ratio, the analysis uses the full randomized analysis set, and progression is determined through an independent central review framework.
| Concept | How it appears in PAPILLON |
|---|---|
| Randomization | The trial uses randomized allocation to two parallel treatment arms. |
| Time-to-event analysis | The primary endpoint is PFS from randomization to progression or death, whichever occurs first. |
| RECIST assessment | Progression is defined using RECIST version 1.1. |
| Blinded independent review | The primary endpoint is assessed by BICR. |
| Log-rank test | The reported formal comparison uses the log-rank test. |
| Hazard ratio | The reported effect measure is HR 0.395. |
| Confidence interval | The 95% CI is 0.296–0.528. |
| P-value | The reported p-value is <0.0001. |
| Superiority testing | The registered hypothesis type is superiority. |
| Full analysis set | All randomized participants are classified according to assigned treatment regardless of actual treatment received. |
| Censoring | Participants without progression or death at analysis are censored at the latest date of their last evaluable RECIST version 1.1 assessment. |
| Safety analysis | Serious adverse events are reported separately by treatment arm. |
18. A Practical Framework for Reading PAPILLON
When evaluating the primary result, it is useful to proceed in a fixed order rather than starting with the p-value.
1. Define the endpoint
First establish exactly what counts as a PFS event and how participants without an event are censored.
2. Check the population
Confirm that the primary analysis uses the full analysis set and preserves randomized treatment assignment.
3. Identify the comparison
Read the treatment direction carefully: the reported comparison is Arm B versus Arm A.
4. Read the effect
Interpret HR 0.395 as a relative time-to-event measure rather than an absolute probability.
5. Read the CI
Use 0.296–0.528 to understand the precision of the estimated hazard ratio.
6. Read the p-value last
Use P < 0.0001 to understand the evidence against the null hypothesis, not the magnitude of benefit.
This sequence prevents one of the most common statistical mistakes in trial interpretation: allowing a very small p-value to substitute for an actual description of the treatment effect.
19. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The randomized comparison of PFS produced a hazard ratio of 0.395, with a two-sided 95% confidence interval of 0.296–0.528 and a log-rank p-value of <0.0001 under a superiority framework.
Clinical interpretation
The numerical result describes a substantially lower estimated hazard of progression or death for amivantamab plus chemotherapy relative to chemotherapy alone. The result should be interpreted together with the endpoint definition, censoring rules, analysis population, and separately reported safety experience.
The statistical result is precise about the comparison being made, but it should not be expanded into claims that are not represented by the reported endpoint. In particular, the HR does not provide a direct estimate of an individual's probability of progression-free survival at a specific time point.
20. Trial Timeline
Trial start
The PAPILLON trial began on 2020-10-13.
Primary completion
The trial's primary completion date was 2023-05-03.
Active, not recruiting
The ClinicalTrials.gov record identifies the study status as ACTIVE_NOT_RECRUITING.
21. What Is and Is Not Established by the Posted Analysis
| Question | What the ClinicalTrials.gov record supports |
|---|---|
| Was there a formal primary analysis? | Yes. One statistical analysis is posted for the primary PFS endpoint. |
| What endpoint was analyzed? | PFS according to RECIST version 1.1 assessed by BICR. |
| What statistical test was used? | Log-rank test. |
| What was the effect measure? | Hazard ratio. |
| What was the estimate? | 0.395. |
| What was the 95% CI? | 0.296–0.528, two-sided. |
| What was the p-value? | <0.0001. |
| What hypothesis type was registered? | Superiority. |
| Are median PFS values available in the ClinicalTrials.gov record? | No numerical median PFS estimate is included in the ClinicalTrials.gov record. |
| Are subgroup treatment effects available? | No subgroup estimates are included in the registry-reported statistical analysis. |
| Can a Kaplan-Meier curve be reconstructed? | Not from the registry-reported summary statistics alone. |
This distinction is important for a complete statistical record. A high-quality analysis should describe what has actually been estimated without filling missing quantities with values derived from memory, unrelated publications, or assumptions about the underlying patient-level data.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: NCT04538664 — PAPILLON.
- PubMed: PMID 42192224.
- PubMed: PMID 41671628.
- PubMed: PMID 41184595.
- PubMed: PMID 40922529.
- PubMed: PMID 37870976.
Continue through Clinical Biostats
Use the related tutorials and statistical calculators to explore the survival-analysis concepts illustrated by PAPILLON.
25. Record Summary
PAPILLON is a randomized phase 3 trial with a parallel design and no masking, enrolling 308.0 participants across two treatment arms. Its registered primary endpoint is progression-free survival according to RECIST version 1.1 as assessed by blinded independent central review, measured from randomization to disease progression or death, whichever occurs first, with a time frame of up to 29 months.
The posted formal analysis uses the full analysis set, classifies participants according to randomized treatment assignment regardless of actual treatment received, and compares Arm B, chemotherapy alone, with Arm A, amivantamab plus chemotherapy. The reported log-rank analysis gives a hazard ratio of 0.395, a two-sided 95% confidence interval of 0.296–0.528, and P < 0.0001 under a superiority hypothesis.
The most informative way to read this result is not to treat the p-value as a measure of treatment magnitude. Instead, the statistical story is built from the endpoint definition, randomized analysis population, log-rank comparison, hazard ratio, confidence interval, and censoring framework. Together, these describe a substantially lower estimated PFS event hazard for the amivantamab-plus-chemotherapy group relative to chemotherapy alone while preserving the distinction between relative effect, statistical precision, and individual patient outcomes.