← Clinical Trials
Solid Tumors Phase 1 Randomized NCT01295827

KEYNOTE-001: Complete Statistical Analysis of Pembrolizumab in Advanced Carcinoma, Melanoma, and NSCLC

An independent statistical review of the randomized phase 1 KEYNOTE-001 study evaluating pembrolizumab in participants with progressive locally advanced or metastatic carcinoma, melanoma, or non-small cell lung carcinoma, with registered primary endpoints covering dose-limiting toxicity, adverse events, and overall response rate.

Trial period: 04-Mar-2011 to 05-Nov-2018  ·  Enrollment: 1260  ·  15 arms
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. The numerical trial information on this page is restricted to the ClinicalTrials.gov record. Because the registry extract contains no formal statistical analyses, this page distinguishes registered endpoints from the statistical methods that would ordinarily be used to analyze them.

1. Trial at a Glance

KEYNOTE-001 was a completed randomized, parallel, unmasked phase 1 study of pembrolizumab in participants with progressive locally advanced or metastatic carcinoma, melanoma, or non-small cell lung carcinoma. The ClinicalTrials.gov record reports 1260 participants across 15 arms and four registered primary endpoints, all classified as binary endpoints.

1260
Enrollment
Participants
15
Arms
Parallel design
4
Primary endpoints
All binary
1
Intervention
Pembrolizumab
FeatureKEYNOTE-001
Trial nameKEYNOTE-001
NCT identifierNCT01295827
PhasePhase 1
StatusCompleted
Therapeutic areaOncology
ConditionCancer, Solid Tumor
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment1260.0
Arms15
InterventionPembrolizumab (biological)
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry
Start2011-03-04
Primary completion2018-11-05
Results postedYes
Registry-results distinction: the ClinicalTrials.gov extract says that results are posted and lists four primary endpoints, but no formal statistical analyses were posted.

2. Clinical Question

The study's clinical question can be framed around whether pembrolizumab could be administered and evaluated in participants with progressive locally advanced or metastatic carcinoma, melanoma, or non-small cell lung carcinoma, with particular attention to dose-limiting toxicity, adverse events, and tumor response.

Population

Participants with progressive locally advanced or metastatic carcinoma, melanoma, or non-small cell lung carcinoma, as described in the brief trial title.

Intervention

Pembrolizumab, identified in the registry data as a biological intervention.

Comparator

The registry extract identifies 15 study arms but does not identify a single comparator arm. The study is therefore not represented here as a simple two-arm treatment-versus-control comparison.

Primary questions

How many participants experienced dose-limiting toxicities, adverse events, or objective responses under the registered endpoint definitions?

3. Trial Design

01
Enroll1260 participants
02
Randomize15 study arms
03
TreatPembrolizumab
04
AssessToxicity, AEs, response
05
AnalyzeBinary endpoints
Allocation
Randomized.
Model
Parallel.
Masking
None.
Purpose
Treatment.

The statistical structure is therefore different from a conventional single-comparator confirmatory trial. There are multiple study arms, a randomized allocation, and binary primary endpoints that correspond to different clinical questions. The ClinicalTrials.gov record does not provide an allocation ratio, arm-specific enrollment totals, or a formal hierarchy of comparisons, so those features are not reconstructed here.

4. Endpoints

The registry lists four primary endpoints. All four are classified in the ClinicalTrials.gov record as binary endpoints. Their clinical meanings and observation windows differ substantially, which is important when interpreting any eventual comparison of proportions.

Registered primary endpointTime frameType
Number of Participants Experiencing Dose-Limiting Toxicities (DLTs) According to National Cancer Institute Common Terminology Criteria for Adverse Events Version 4.0 (NCI-CTCAE v.4.0) in Participants With Solid Tumors (Parts A and A1) Up to 28 days in Cycle 1 Binary
Number of Participants Who Experienced an Adverse Event (AE) Up to approximately 91 months (through Final Database cut-off date of 05-Nov-2018) Binary
Overall Response Rate (ORR) According to Response Evaluation Criteria In Solid Tumors Version 1.1 (RECIST 1.1) as Assessed by Integrated Radiology and Oncology (IRO): Melanoma Participants (Parts B Plus D) Up to approximately 53 months (through Interim Database cut-off date of 18-Sep-2015) Binary
ORR According to RECIST 1.1 as Assessed by Independent Review Committee (IRC): Non-Small Cell Lung Cancer (NSCLC) Participants (Parts C Plus F) Up to approximately 53 months (through Interim Database cut-off date of 18-Sep-2015) Binary

Endpoint 1: Dose-limiting toxicity

DLTs were assessed according to NCI-CTCAE v.4.0 during the first cycle, defined as 28 days. The registry-reported definition states that DLTs were toxicities judged by the investigator to be possibly, probably, or definitely related to study drug administration and included specified severe toxicities such as Grade 4 nonhematologic toxicity, Grade 4 hematologic toxicity lasting at least 14 days, and Grade 3 nonhematologic toxicity lasting more than 3 days despite optimal supportive care.

Endpoint 2: Adverse events

An adverse event was defined as any unfavorable and unintended change in the structure, function, or chemistry of the body temporally associated with use of the sponsor's product, whether or not considered related to product use. Worsening of a pre-existing condition could also qualify when temporally associated with product use.

Endpoint 3: Melanoma ORR

For melanoma participants in Parts B plus D, ORR was defined as the percentage of participants in the analysis population with a confirmed complete response or partial response according to RECIST 1.1. Complete response meant disappearance of all lesions. Partial response required at least a 30% decrease in the sum of diameters of target lesions, using baseline sum of diameters as the reference. The study modified RECIST 1.1 to follow a maximum of 10 target lesions and a maximum of 5 target lesions per organ.

Endpoint 4: NSCLC ORR

For NSCLC participants in Parts C plus F, ORR used RECIST 1.1 and was assessed by an independent review committee. The registry-reported definition again specifies confirmed complete or partial response, with partial response requiring at least a 30% decrease in the sum of diameters of target lesions relative to baseline, using the study's modified target-lesion framework.

5. Planned Analysis

No formal statistical comparisons are contained in the ClinicalTrials.gov record. The registry extract states that results are posted, but no formal statistical analyses were posted. The following describes the statistical analysis that would ordinarily be appropriate for the registered binary endpoints; it is not presented as a reconstruction of an unreported ClinicalTrials.gov analysis.

DLT endpoint

The DLT endpoint is binary at the participant level: each participant either experienced a qualifying DLT during the first 28-day cycle or did not. A natural descriptive analysis would report the number and percentage experiencing a DLT in each relevant study arm, together with a confidence interval for each proportion. If prespecified treatment-arm comparisons were required, the comparison could use a two-sample proportion method, exact methods, or a regression model appropriate to the number and structure of arms.

Adverse-event endpoint

The AE endpoint is also binary at the participant level: whether a participant experienced an adverse event during the approximately 91-month observation window through the final database cutoff. Descriptive arm-specific incidence proportions would ordinarily be central to the analysis. Because the observation period is much longer than the DLT window, the two binary endpoints should not be treated as interchangeable measures of toxicity.

Melanoma and NSCLC ORR

ORR is a binary response endpoint. A standard analysis would calculate the proportion of participants with confirmed complete or partial response in the specified analysis population. Confidence intervals for the response proportion would quantify precision. If a formal comparison between study arms were prespecified, the appropriate comparison would depend on the arm structure and analysis plan; a simple two-group test should not be assumed for a 15-arm randomized study.

Binary endpoint framework
p = number of participants with the event or response / number of participants in the analysis population

For DLT, AE, and ORR, the core estimand is a participant-level proportion. The clinical interpretation depends on the endpoint definition, observation window, analysis population, and study arm structure.

6. Statistical Methodology

Binary endpoint analysis

All four registered primary endpoints are classified as binary in the ClinicalTrials.gov record. This means that the primary outcome for each participant can be represented as an indicator such as 1 = event/response and 0 = no event/no response.

For a single arm, the most direct descriptive quantity is the observed proportion:

Observed proportion
p̂ = x / n

Here, x is the number of participants meeting the endpoint definition and n is the number of participants in the relevant analysis population.

A confidence interval around the proportion is useful because the observed proportion is a sample estimate rather than a population constant. The interval reflects statistical uncertainty arising from the finite number of participants.

Why analysis populations matter

The denominator is not merely a bookkeeping detail. A response proportion based on all randomized participants answers a different question from a response proportion based on participants with evaluable imaging. Likewise, a DLT proportion in participants exposed to treatment is different from a proportion based on all randomized participants. The ClinicalTrials.gov record does not specify the detailed analysis populations or denominators for the posted results, so they are not reconstructed here.

Multiple arms require deliberate comparisons

With 15 arms, there are many possible pairwise comparisons. If every possible comparison were tested independently at the usual significance level, the probability of at least one false-positive finding would increase. A formal analysis would therefore need a prespecified comparison strategy and, where appropriate, multiplicity control.

Endpoint-specific observation windows

The DLT endpoint has a short, explicitly defined first-cycle window, whereas the AE endpoint extends through approximately 91 months. ORR has its own approximately 53-month assessment window and is restricted to disease-specific participant groups. These windows define different estimands and prevent the four endpoints from being interpreted as if they were measurements on the same time scale.

RECIST response classification

ORR is not simply a subjective statement that a tumor "improved." The registered definition uses RECIST 1.1 criteria, including confirmed complete or partial response, with a study-specific modification to the number of target lesions followed. The binary endpoint therefore compresses a structured radiologic assessment into a yes/no participant-level response indicator.

Independent review

The NSCLC ORR endpoint was assessed by an Independent Review Committee, while the melanoma ORR endpoint was assessed by Integrated Radiology and Oncology. Independent or centrally structured review can reduce the influence of treatment knowledge or local assessment practices, but the ClinicalTrials.gov record does not provide enough detail to characterize the operational blinding or adjudication procedures further.

7. Statistical Methods Explained

Why is randomization important in a 15-arm study?

Randomization provides the design mechanism for distributing measured and unmeasured participant characteristics across study arms according to the randomization process. In a multi-arm study, this is particularly important because there are many possible treatment-arm comparisons. Randomization does not guarantee identical baseline characteristics in every arm; its role is to support valid comparison under the trial's allocation mechanism.

What does a binary endpoint mean?

A binary endpoint reduces each participant's outcome to two categories. For DLT, the categories are whether the participant experienced a qualifying DLT during the specified first-cycle window. For ORR, the categories are whether the participant achieved a confirmed complete or partial response under the registered RECIST definition.

Why does the DLT time frame matter?

The DLT endpoint is explicitly limited to up to 28 days in Cycle 1. A toxicity occurring later does not automatically become a DLT merely because it is clinically important. The DLT definition is tied to both the severity criteria and the specified observation period.

Why is ORR different from tumor shrinkage as a continuous measurement?

RECIST converts radiologic measurements into categorical response states. ORR then counts participants who reach confirmed complete or partial response. This is useful for a clear clinical endpoint, but it discards information about the precise amount of shrinkage among responders and about participants whose tumors change without meeting the response threshold.

Why should a 15-arm trial not automatically use one pairwise p-value?

A 15-arm design creates multiple possible comparisons. The more independent statistical tests that are performed, the greater the chance of obtaining at least one apparently positive result by chance alone. A valid confirmatory analysis therefore needs a prespecified comparison structure and appropriate handling of multiplicity when multiple hypotheses are tested.

Why can the AE endpoint not be interpreted like the DLT endpoint?

Although both are binary endpoints, they have different definitions and observation windows. DLT is a narrowly defined first-cycle toxicity endpoint, whereas the AE endpoint captures adverse events over approximately 91 months through the final database cutoff. The same statistical form does not make two endpoints clinically equivalent.

Why does an empty statistical-analysis field matter?

It means that the registry extract does not provide the formal treatment-effect estimates, confidence intervals, or p-values that would be needed for a quantitative comparative results section. It would be statistically misleading to create those values from the sample size, endpoint descriptions, or general knowledge of the trial.

8. Safety Results

The ClinicalTrials.gov record includes a partial arm-level listing of serious adverse events, expressed as affected participants over participants at risk. The following entries are reproduced exactly from the registry extract.

Study armSerious adverse events
Solid Tumors: Pembrolizumab 1 mg/kg Q2W1/4
Solid Tumors: Pembrolizumab 3 mg/kg Q2W1/3
Solid Tumors: Pembrolizumab 10 mg/kg Q2W3/10
Solid Tumors: Pembrolizumab Titration Co0/4
Solid Tumors: Pembrolizumab Titration Co2/3
Solid Tumors: Pembrolizumab Titration Co2/6

The registry-reported arm-level safety field continues beyond these entries, but the ClinicalTrials.gov record terminate partway through a subsequent melanoma entry. The incomplete portion is therefore not reconstructed or interpreted as a complete arm-level safety table.

Statistical interpretation

The affected/at-risk format is a direct way to express the observed frequency of serious adverse events. For example, 1/4 means one affected participant among four participants at risk in that listed arm. It does not by itself provide a comparative treatment effect, a confidence interval, or evidence that one arm differs from another.

9. Primary Endpoint 1: Dose-Limiting Toxicities

Registered endpoint

DLT within 28 days

NCI-CTCAE v.4.0  ·  Solid Tumors, Parts A and A1

Binary endpoint; results posted in the ClinicalTrials.gov record.

The endpoint counts participants experiencing dose-limiting toxicities according to the registered NCI-CTCAE v.4.0 definition during the first 28-day cycle. The ClinicalTrials.gov record contains no formal estimate, confidence interval, or p-value for this endpoint.

Clinical Biostats interpretation

The clinically meaningful quantity would ordinarily be the proportion of participants experiencing a qualifying DLT in each relevant arm. A higher proportion would indicate more frequent occurrence of the narrowly defined first-cycle toxicity endpoint. It would not mean that all adverse events were DLTs, because the DLT definition is restricted by severity, duration, attribution, and the first-cycle observation window.

A confidence interval around a DLT proportion would describe uncertainty around the estimated frequency. It would not describe the probability that an individual participant will experience toxicity. A p-value, if a formal comparison were prespecified, would address evidence against a null comparison; it would not measure the size or clinical importance of the toxicity difference.

10. Primary Endpoint 2: Adverse Events

Registered endpoint

Any AE

Up to approximately 91 months

Final database cutoff: 05-Nov-2018

The registered endpoint counts participants who experienced an adverse event through approximately 91 months, using the final database cutoff of 05-Nov-2018. The registry-reported statistical-analysis array contains no formal comparison or effect estimate for this endpoint.

Clinical Biostats interpretation

An AE proportion answers a broad safety question: how many participants experienced at least one event under the registry's AE definition during the specified observation period. Because this is a participant-level binary endpoint, it does not capture the number of separate events per participant, their severity distribution, or their timing unless those features are analyzed separately.

The very long observation window also matters. A participant followed for a long period has more opportunity to experience an AE than a participant observed for a shorter period. A simple binary incidence proportion therefore has to be interpreted in the context of the trial's defined follow-up and analysis population.

11. Primary Endpoint 3: Melanoma Overall Response Rate

Registered endpoint

ORR by RECIST 1.1

Melanoma participants · Parts B plus D

Integrated Radiology and Oncology assessment · approximately 53 months

The melanoma ORR endpoint uses confirmed complete response or partial response under RECIST 1.1, with the study-specific modification to target-lesion tracking. The ClinicalTrials.gov record does not provide a formal statistical estimate, confidence interval, or p-value.

Clinical Biostats interpretation

The natural estimand is the proportion of melanoma participants in the relevant analysis population who achieved a confirmed complete or partial response. ORR therefore focuses on the occurrence of a predefined tumor-response state rather than survival time.

A response proportion does not tell us how long a response lasted, how long participants survived, or whether every participant experienced some degree of tumor shrinkage. Those are different clinical questions requiring different endpoints.

12. Primary Endpoint 4: NSCLC Overall Response Rate

Registered endpoint

ORR by RECIST 1.1

NSCLC participants · Parts C plus F

Independent Review Committee assessment · approximately 53 months

The NSCLC ORR endpoint is based on confirmed complete or partial response according to RECIST 1.1, with assessment by an Independent Review Committee. The ClinicalTrials.gov record contains no formal comparison, confidence interval, or p-value.

Clinical Biostats interpretation

The response endpoint is binary, so the core result would normally be an observed response proportion for the prespecified analysis population. The confidence interval would quantify uncertainty around that proportion. If multiple arms were compared, the statistical interpretation would additionally depend on which comparisons were prespecified and whether multiplicity was controlled.

Independent review is relevant to the measurement process, but it does not itself turn ORR into a time-to-event endpoint or establish a treatment effect. The endpoint remains a categorical response measure.

13. Results Interpretation: What Can and Cannot Be Concluded From the Supplied Data

What is documented

The study was randomized, parallel, unmasked, phase 1, had 1260 participants and 15 arms, and registered four binary primary endpoints.

What is not reported

The statistical-analysis field contains no formal estimates, confidence intervals, p-values, treatment-effect models, or multiplicity procedures.

Why that distinction matters

Knowing the endpoint definition does not reveal the observed event rate, and knowing the enrollment does not permit reconstruction of an unreported treatment comparison.

Educational implication

This trial is useful for studying how a randomized multi-arm phase 1 study can combine toxicity and response endpoints with very different definitions and observation windows.

14. Limitations

15. Why This Trial Matters Statistically

KEYNOTE-001 provides a useful statistical teaching case because it combines randomization, multiple parallel study arms, binary toxicity endpoints, radiologic response endpoints, different assessment populations, and markedly different observation windows.

ConceptHow it appears in KEYNOTE-001
RandomizationThe study allocation is identified as randomized.
Multi-arm designThe study contains 15 arms.
Binary endpointsAll four registered primary endpoints are classified as binary.
Dose-limiting toxicityDLTs are evaluated during the first 28 days in Cycle 1 using NCI-CTCAE v.4.0.
Adverse-event analysisAny AE is followed through approximately 91 months.
Objective responseORR is based on confirmed CR or PR under RECIST 1.1.
Independent reviewNSCLC ORR is assessed by an Independent Review Committee.
Multiple comparisonsFifteen arms create a potentially large set of pairwise comparisons, making prespecified comparison strategy important.
Different estimandsDLT, AE, melanoma ORR, and NSCLC ORR answer different clinical questions.

16. Endpoint Definitions: A Statistical Reading

DLT is an early safety estimand

The DLT endpoint is deliberately narrow. Its 28-day window focuses attention on early treatment-related toxicity. Statistically, the endpoint can be viewed as an incidence proportion over a fixed short interval, provided the analysis population and exposure rules are defined.

AE is a broad cumulative safety estimand

The AE endpoint has a much longer observation window. Its definition is also broad, covering unfavorable and unintended changes temporally associated with use of the sponsor's product whether or not considered related. A participant can therefore meet the endpoint for an event that is not judged treatment-related.

ORR is a categorical efficacy estimand

ORR asks whether a participant achieves a confirmed CR or PR. The endpoint therefore requires both a measurement system and response-classification rules. RECIST 1.1 supplies those rules, while the study's lesion-count modification defines how tumor burden was followed.

Different endpoints should not be collapsed into one statistic

A DLT rate, an AE rate, and an ORR are all proportions, but their numerators, denominators, observation windows, and clinical meanings differ. Statistical similarity in mathematical form does not imply equivalence in interpretation.

17. Randomization and Multi-Arm Inference

Randomization is the principal design feature identified by the registry that supports causal comparison among the randomized study arms. In a parallel randomized study, each participant's assigned arm is determined by the trial's allocation mechanism rather than by an investigator selecting treatment after observing the participant's outcome.

With 15 arms, however, the phrase "the treatment effect" is incomplete without specifying the comparison. There can be multiple pairwise contrasts, combinations of arms, or dose-related contrasts. A statistical analysis must therefore identify the estimand and comparison of interest before the p-value or confidence interval can be interpreted.

Pairwise comparison count
Number of unique pairwise comparisons = k(k − 1) / 2

For a study with k arms, this formula illustrates why multiplicity becomes important as the number of arms increases. It does not imply that every possible comparison was actually performed in KEYNOTE-001.

The ClinicalTrials.gov record does not state that all possible pairwise comparisons were performed. Therefore, the formula is presented only as a statistical illustration of the multi-arm problem, not as a description of the trial's actual testing program.

18. Binary Confidence Intervals

For a binary endpoint, a confidence interval complements the observed proportion by showing how precisely the study estimates the underlying event or response probability. The exact interval method should be selected according to the prespecified analysis plan and sample size rather than assumed from the endpoint label alone.

Estimate

The observed proportion describes the frequency of participants meeting the endpoint definition in the analyzed population.

Precision

The confidence interval describes statistical uncertainty around the estimated proportion.

Comparison

A treatment comparison requires a defined contrast between study arms, not merely two separate proportions.

P-value

If used, a p-value addresses a prespecified null hypothesis. It is not a measure of effect magnitude or clinical importance.

19. Statistical Interpretation of Safety Counts

The serious-adverse-event entries posted on ClinicalTrials.gov for several solid-tumor arms illustrate why denominators must remain visible. A count such as 3/10 contains more information than the numerator alone because it establishes the population at risk in that listed arm.

Listed armAffectedAt riskStatistical reading
Pembrolizumab 1 mg/kg Q2W14One serious-AE-affected participant among four at risk.
Pembrolizumab 3 mg/kg Q2W13One serious-AE-affected participant among three at risk.
Pembrolizumab 10 mg/kg Q2W310Three serious-AE-affected participants among ten at risk.
Pembrolizumab Titration Co04No affected participants among four at risk.
Pembrolizumab Titration Co23Two affected participants among three at risk.
Pembrolizumab Titration Co26Two affected participants among six at risk.

20. Clinical Interpretation vs Statistical Interpretation

Clinical interpretation

The registered endpoints focus on early dose-limiting toxicity, cumulative adverse events, and objective tumor response in defined disease-specific participant groups.

Statistical interpretation

Each primary endpoint is binary, so the core analysis concerns event or response proportions. Formal comparative inference requires the relevant arm denominators, prespecified contrasts, uncertainty estimates, and testing strategy.

21. What the Available Data Do Not Establish

Methodological principle: absence of a the ClinicalTrials.gov record is not evidence that the original investigators performed no statistical analysis. It means only that the ClinicalTrials.gov record does not contain those details. This page therefore avoids attributing unreported methods or results to the trial.

22. Related Tutorials

Learn more about the methods and statistical ideas used to understand this trial:

23. Related Statistical Calculators

24. Sources

Continue through Clinical Biostats

Connect this trial's statistical concepts to deeper tutorials, statistical calculators, and other clinical-trial analyses.

25. Record Summary

KEYNOTE-001 is a useful example of how statistical analysis must follow the structure of the clinical question rather than simply the size of the trial. The registry describes a randomized, parallel phase 1 study with 15 arms and 1260 participants, four binary primary endpoints, and results posted. Those endpoints span early dose-limiting toxicity, long-term adverse-event occurrence, and disease-specific objective response assessed using RECIST 1.1.

The most important statistical lesson is that a binary endpoint is only the starting point. Its interpretation depends on the event definition, observation window, analysis population, study-arm structure, and prespecified comparison. DLT and AE are both safety endpoints but operate on very different time frames; melanoma and NSCLC ORR are both response endpoints but apply to different participant groups and use different assessment organizations.

Because the registry-reported statistical-analysis field contains no formal estimates, confidence intervals, or p-values, a statistically responsible analysis must stop at the information actually available rather than reconstructing numerical treatment effects from incomplete inputs. That distinction between registered endpoint information, reported numerical results, and statistical interpretation is itself an important part of rigorous clinical-trial analysis.

Clinical Biostats methodology: A trial-results page should separate what the registry explicitly reports from what statistical theory says would ordinarily be done with that endpoint. This page therefore uses the registry definitions and available safety counts while clearly labeling standard statistical methods as methodological explanation rather than reported trial results.