This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. The numerical trial information on this page is restricted to the ClinicalTrials.gov record. Because the registry extract contains no formal statistical analyses, this page distinguishes registered endpoints from the statistical methods that would ordinarily be used to analyze them.
1. Trial at a Glance
KEYNOTE-001 was a completed randomized, parallel, unmasked phase 1 study of pembrolizumab in participants with progressive locally advanced or metastatic carcinoma, melanoma, or non-small cell lung carcinoma. The ClinicalTrials.gov record reports 1260 participants across 15 arms and four registered primary endpoints, all classified as binary endpoints.
| Feature | KEYNOTE-001 |
|---|---|
| Trial name | KEYNOTE-001 |
| NCT identifier | NCT01295827 |
| Phase | Phase 1 |
| Status | Completed |
| Therapeutic area | Oncology |
| Condition | Cancer, Solid Tumor |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 1260.0 |
| Arms | 15 |
| Intervention | Pembrolizumab (biological) |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| Start | 2011-03-04 |
| Primary completion | 2018-11-05 |
| Results posted | Yes |
2. Clinical Question
The study's clinical question can be framed around whether pembrolizumab could be administered and evaluated in participants with progressive locally advanced or metastatic carcinoma, melanoma, or non-small cell lung carcinoma, with particular attention to dose-limiting toxicity, adverse events, and tumor response.
Population
Participants with progressive locally advanced or metastatic carcinoma, melanoma, or non-small cell lung carcinoma, as described in the brief trial title.
Intervention
Pembrolizumab, identified in the registry data as a biological intervention.
Comparator
The registry extract identifies 15 study arms but does not identify a single comparator arm. The study is therefore not represented here as a simple two-arm treatment-versus-control comparison.
Primary questions
How many participants experienced dose-limiting toxicities, adverse events, or objective responses under the registered endpoint definitions?
3. Trial Design
The statistical structure is therefore different from a conventional single-comparator confirmatory trial. There are multiple study arms, a randomized allocation, and binary primary endpoints that correspond to different clinical questions. The ClinicalTrials.gov record does not provide an allocation ratio, arm-specific enrollment totals, or a formal hierarchy of comparisons, so those features are not reconstructed here.
4. Endpoints
The registry lists four primary endpoints. All four are classified in the ClinicalTrials.gov record as binary endpoints. Their clinical meanings and observation windows differ substantially, which is important when interpreting any eventual comparison of proportions.
| Registered primary endpoint | Time frame | Type |
|---|---|---|
| Number of Participants Experiencing Dose-Limiting Toxicities (DLTs) According to National Cancer Institute Common Terminology Criteria for Adverse Events Version 4.0 (NCI-CTCAE v.4.0) in Participants With Solid Tumors (Parts A and A1) | Up to 28 days in Cycle 1 | Binary |
| Number of Participants Who Experienced an Adverse Event (AE) | Up to approximately 91 months (through Final Database cut-off date of 05-Nov-2018) | Binary |
| Overall Response Rate (ORR) According to Response Evaluation Criteria In Solid Tumors Version 1.1 (RECIST 1.1) as Assessed by Integrated Radiology and Oncology (IRO): Melanoma Participants (Parts B Plus D) | Up to approximately 53 months (through Interim Database cut-off date of 18-Sep-2015) | Binary |
| ORR According to RECIST 1.1 as Assessed by Independent Review Committee (IRC): Non-Small Cell Lung Cancer (NSCLC) Participants (Parts C Plus F) | Up to approximately 53 months (through Interim Database cut-off date of 18-Sep-2015) | Binary |
Endpoint 1: Dose-limiting toxicity
DLTs were assessed according to NCI-CTCAE v.4.0 during the first cycle, defined as 28 days. The registry-reported definition states that DLTs were toxicities judged by the investigator to be possibly, probably, or definitely related to study drug administration and included specified severe toxicities such as Grade 4 nonhematologic toxicity, Grade 4 hematologic toxicity lasting at least 14 days, and Grade 3 nonhematologic toxicity lasting more than 3 days despite optimal supportive care.
Endpoint 2: Adverse events
An adverse event was defined as any unfavorable and unintended change in the structure, function, or chemistry of the body temporally associated with use of the sponsor's product, whether or not considered related to product use. Worsening of a pre-existing condition could also qualify when temporally associated with product use.
Endpoint 3: Melanoma ORR
For melanoma participants in Parts B plus D, ORR was defined as the percentage of participants in the analysis population with a confirmed complete response or partial response according to RECIST 1.1. Complete response meant disappearance of all lesions. Partial response required at least a 30% decrease in the sum of diameters of target lesions, using baseline sum of diameters as the reference. The study modified RECIST 1.1 to follow a maximum of 10 target lesions and a maximum of 5 target lesions per organ.
Endpoint 4: NSCLC ORR
For NSCLC participants in Parts C plus F, ORR used RECIST 1.1 and was assessed by an independent review committee. The registry-reported definition again specifies confirmed complete or partial response, with partial response requiring at least a 30% decrease in the sum of diameters of target lesions relative to baseline, using the study's modified target-lesion framework.
5. Planned Analysis
DLT endpoint
The DLT endpoint is binary at the participant level: each participant either experienced a qualifying DLT during the first 28-day cycle or did not. A natural descriptive analysis would report the number and percentage experiencing a DLT in each relevant study arm, together with a confidence interval for each proportion. If prespecified treatment-arm comparisons were required, the comparison could use a two-sample proportion method, exact methods, or a regression model appropriate to the number and structure of arms.
Adverse-event endpoint
The AE endpoint is also binary at the participant level: whether a participant experienced an adverse event during the approximately 91-month observation window through the final database cutoff. Descriptive arm-specific incidence proportions would ordinarily be central to the analysis. Because the observation period is much longer than the DLT window, the two binary endpoints should not be treated as interchangeable measures of toxicity.
Melanoma and NSCLC ORR
ORR is a binary response endpoint. A standard analysis would calculate the proportion of participants with confirmed complete or partial response in the specified analysis population. Confidence intervals for the response proportion would quantify precision. If a formal comparison between study arms were prespecified, the appropriate comparison would depend on the arm structure and analysis plan; a simple two-group test should not be assumed for a 15-arm randomized study.
For DLT, AE, and ORR, the core estimand is a participant-level proportion. The clinical interpretation depends on the endpoint definition, observation window, analysis population, and study arm structure.
6. Statistical Methodology
Binary endpoint analysis
All four registered primary endpoints are classified as binary in the ClinicalTrials.gov record. This means that the primary outcome for each participant can be represented as an indicator such as 1 = event/response and 0 = no event/no response.
For a single arm, the most direct descriptive quantity is the observed proportion:
Here, x is the number of participants meeting the endpoint definition and n is the number of participants in the relevant analysis population.
A confidence interval around the proportion is useful because the observed proportion is a sample estimate rather than a population constant. The interval reflects statistical uncertainty arising from the finite number of participants.
Why analysis populations matter
The denominator is not merely a bookkeeping detail. A response proportion based on all randomized participants answers a different question from a response proportion based on participants with evaluable imaging. Likewise, a DLT proportion in participants exposed to treatment is different from a proportion based on all randomized participants. The ClinicalTrials.gov record does not specify the detailed analysis populations or denominators for the posted results, so they are not reconstructed here.
Multiple arms require deliberate comparisons
With 15 arms, there are many possible pairwise comparisons. If every possible comparison were tested independently at the usual significance level, the probability of at least one false-positive finding would increase. A formal analysis would therefore need a prespecified comparison strategy and, where appropriate, multiplicity control.
Endpoint-specific observation windows
The DLT endpoint has a short, explicitly defined first-cycle window, whereas the AE endpoint extends through approximately 91 months. ORR has its own approximately 53-month assessment window and is restricted to disease-specific participant groups. These windows define different estimands and prevent the four endpoints from being interpreted as if they were measurements on the same time scale.
RECIST response classification
ORR is not simply a subjective statement that a tumor "improved." The registered definition uses RECIST 1.1 criteria, including confirmed complete or partial response, with a study-specific modification to the number of target lesions followed. The binary endpoint therefore compresses a structured radiologic assessment into a yes/no participant-level response indicator.
Independent review
The NSCLC ORR endpoint was assessed by an Independent Review Committee, while the melanoma ORR endpoint was assessed by Integrated Radiology and Oncology. Independent or centrally structured review can reduce the influence of treatment knowledge or local assessment practices, but the ClinicalTrials.gov record does not provide enough detail to characterize the operational blinding or adjudication procedures further.
7. Statistical Methods Explained
Why is randomization important in a 15-arm study?
Randomization provides the design mechanism for distributing measured and unmeasured participant characteristics across study arms according to the randomization process. In a multi-arm study, this is particularly important because there are many possible treatment-arm comparisons. Randomization does not guarantee identical baseline characteristics in every arm; its role is to support valid comparison under the trial's allocation mechanism.
What does a binary endpoint mean?
A binary endpoint reduces each participant's outcome to two categories. For DLT, the categories are whether the participant experienced a qualifying DLT during the specified first-cycle window. For ORR, the categories are whether the participant achieved a confirmed complete or partial response under the registered RECIST definition.
Why does the DLT time frame matter?
The DLT endpoint is explicitly limited to up to 28 days in Cycle 1. A toxicity occurring later does not automatically become a DLT merely because it is clinically important. The DLT definition is tied to both the severity criteria and the specified observation period.
Why is ORR different from tumor shrinkage as a continuous measurement?
RECIST converts radiologic measurements into categorical response states. ORR then counts participants who reach confirmed complete or partial response. This is useful for a clear clinical endpoint, but it discards information about the precise amount of shrinkage among responders and about participants whose tumors change without meeting the response threshold.
Why should a 15-arm trial not automatically use one pairwise p-value?
A 15-arm design creates multiple possible comparisons. The more independent statistical tests that are performed, the greater the chance of obtaining at least one apparently positive result by chance alone. A valid confirmatory analysis therefore needs a prespecified comparison structure and appropriate handling of multiplicity when multiple hypotheses are tested.
Why can the AE endpoint not be interpreted like the DLT endpoint?
Although both are binary endpoints, they have different definitions and observation windows. DLT is a narrowly defined first-cycle toxicity endpoint, whereas the AE endpoint captures adverse events over approximately 91 months through the final database cutoff. The same statistical form does not make two endpoints clinically equivalent.
Why does an empty statistical-analysis field matter?
It means that the registry extract does not provide the formal treatment-effect estimates, confidence intervals, or p-values that would be needed for a quantitative comparative results section. It would be statistically misleading to create those values from the sample size, endpoint descriptions, or general knowledge of the trial.
8. Safety Results
The ClinicalTrials.gov record includes a partial arm-level listing of serious adverse events, expressed as affected participants over participants at risk. The following entries are reproduced exactly from the registry extract.
| Study arm | Serious adverse events |
|---|---|
| Solid Tumors: Pembrolizumab 1 mg/kg Q2W | 1/4 |
| Solid Tumors: Pembrolizumab 3 mg/kg Q2W | 1/3 |
| Solid Tumors: Pembrolizumab 10 mg/kg Q2W | 3/10 |
| Solid Tumors: Pembrolizumab Titration Co | 0/4 |
| Solid Tumors: Pembrolizumab Titration Co | 2/3 |
| Solid Tumors: Pembrolizumab Titration Co | 2/6 |
The registry-reported arm-level safety field continues beyond these entries, but the ClinicalTrials.gov record terminate partway through a subsequent melanoma entry. The incomplete portion is therefore not reconstructed or interpreted as a complete arm-level safety table.
The affected/at-risk format is a direct way to express the observed frequency of serious adverse events. For example, 1/4 means one affected participant among four participants at risk in that listed arm. It does not by itself provide a comparative treatment effect, a confidence interval, or evidence that one arm differs from another.
9. Primary Endpoint 1: Dose-Limiting Toxicities
Registered endpoint
NCI-CTCAE v.4.0 · Solid Tumors, Parts A and A1
Binary endpoint; results posted in the ClinicalTrials.gov record.
The endpoint counts participants experiencing dose-limiting toxicities according to the registered NCI-CTCAE v.4.0 definition during the first 28-day cycle. The ClinicalTrials.gov record contains no formal estimate, confidence interval, or p-value for this endpoint.
The clinically meaningful quantity would ordinarily be the proportion of participants experiencing a qualifying DLT in each relevant arm. A higher proportion would indicate more frequent occurrence of the narrowly defined first-cycle toxicity endpoint. It would not mean that all adverse events were DLTs, because the DLT definition is restricted by severity, duration, attribution, and the first-cycle observation window.
A confidence interval around a DLT proportion would describe uncertainty around the estimated frequency. It would not describe the probability that an individual participant will experience toxicity. A p-value, if a formal comparison were prespecified, would address evidence against a null comparison; it would not measure the size or clinical importance of the toxicity difference.
10. Primary Endpoint 2: Adverse Events
Registered endpoint
Up to approximately 91 months
Final database cutoff: 05-Nov-2018
The registered endpoint counts participants who experienced an adverse event through approximately 91 months, using the final database cutoff of 05-Nov-2018. The registry-reported statistical-analysis array contains no formal comparison or effect estimate for this endpoint.
An AE proportion answers a broad safety question: how many participants experienced at least one event under the registry's AE definition during the specified observation period. Because this is a participant-level binary endpoint, it does not capture the number of separate events per participant, their severity distribution, or their timing unless those features are analyzed separately.
The very long observation window also matters. A participant followed for a long period has more opportunity to experience an AE than a participant observed for a shorter period. A simple binary incidence proportion therefore has to be interpreted in the context of the trial's defined follow-up and analysis population.
11. Primary Endpoint 3: Melanoma Overall Response Rate
Registered endpoint
Melanoma participants · Parts B plus D
Integrated Radiology and Oncology assessment · approximately 53 months
The melanoma ORR endpoint uses confirmed complete response or partial response under RECIST 1.1, with the study-specific modification to target-lesion tracking. The ClinicalTrials.gov record does not provide a formal statistical estimate, confidence interval, or p-value.
The natural estimand is the proportion of melanoma participants in the relevant analysis population who achieved a confirmed complete or partial response. ORR therefore focuses on the occurrence of a predefined tumor-response state rather than survival time.
A response proportion does not tell us how long a response lasted, how long participants survived, or whether every participant experienced some degree of tumor shrinkage. Those are different clinical questions requiring different endpoints.
12. Primary Endpoint 4: NSCLC Overall Response Rate
Registered endpoint
NSCLC participants · Parts C plus F
Independent Review Committee assessment · approximately 53 months
The NSCLC ORR endpoint is based on confirmed complete or partial response according to RECIST 1.1, with assessment by an Independent Review Committee. The ClinicalTrials.gov record contains no formal comparison, confidence interval, or p-value.
The response endpoint is binary, so the core result would normally be an observed response proportion for the prespecified analysis population. The confidence interval would quantify uncertainty around that proportion. If multiple arms were compared, the statistical interpretation would additionally depend on which comparisons were prespecified and whether multiplicity was controlled.
Independent review is relevant to the measurement process, but it does not itself turn ORR into a time-to-event endpoint or establish a treatment effect. The endpoint remains a categorical response measure.
13. Results Interpretation: What Can and Cannot Be Concluded From the Supplied Data
What is documented
The study was randomized, parallel, unmasked, phase 1, had 1260 participants and 15 arms, and registered four binary primary endpoints.
What is not reported
The statistical-analysis field contains no formal estimates, confidence intervals, p-values, treatment-effect models, or multiplicity procedures.
Why that distinction matters
Knowing the endpoint definition does not reveal the observed event rate, and knowing the enrollment does not permit reconstruction of an unreported treatment comparison.
Educational implication
This trial is useful for studying how a randomized multi-arm phase 1 study can combine toxicity and response endpoints with very different definitions and observation windows.
14. Limitations
- No formal statistical analyses reported: no formal statistical analyses were posted, so quantitative treatment comparisons cannot be reproduced from the ClinicalTrials.gov record.
- Multi-arm structure: the study has 15 arms, but the ClinicalTrials.gov record does not identify the full arm-level allocation, sample sizes, or prespecified comparison hierarchy.
- Different endpoint populations: the DLT endpoint concerns solid-tumor Parts A and A1, while the two ORR endpoints concern melanoma Parts B plus D and NSCLC Parts C plus F.
- Different observation windows: DLT is assessed through the first 28-day cycle, ORR through approximately 53 months, and the AE endpoint through approximately 91 months.
- Incomplete arm-level safety extract: the registry-reported serious-AE field ends during a later melanoma entry, preventing a complete reconstruction of all serious-AE results.
- Binary endpoint compression: binary endpoints simplify complex clinical information into event/no-event or response/no-response categories.
- Multiplicity: a 15-arm study can generate many potential comparisons. Without the prespecified testing strategy, it is not appropriate to assign confirmatory meaning to hypothetical pairwise p-values.
- Analysis populations: the ClinicalTrials.gov record does not specify the complete denominators or detailed analysis-population rules for the posted endpoint results.
- Unmasked design: the registry describes the study as having no masking. For endpoints involving clinical assessment, knowledge of treatment assignment can be relevant to potential assessment bias, although the ClinicalTrials.gov record does not quantify such an effect.
15. Why This Trial Matters Statistically
KEYNOTE-001 provides a useful statistical teaching case because it combines randomization, multiple parallel study arms, binary toxicity endpoints, radiologic response endpoints, different assessment populations, and markedly different observation windows.
| Concept | How it appears in KEYNOTE-001 |
|---|---|
| Randomization | The study allocation is identified as randomized. |
| Multi-arm design | The study contains 15 arms. |
| Binary endpoints | All four registered primary endpoints are classified as binary. |
| Dose-limiting toxicity | DLTs are evaluated during the first 28 days in Cycle 1 using NCI-CTCAE v.4.0. |
| Adverse-event analysis | Any AE is followed through approximately 91 months. |
| Objective response | ORR is based on confirmed CR or PR under RECIST 1.1. |
| Independent review | NSCLC ORR is assessed by an Independent Review Committee. |
| Multiple comparisons | Fifteen arms create a potentially large set of pairwise comparisons, making prespecified comparison strategy important. |
| Different estimands | DLT, AE, melanoma ORR, and NSCLC ORR answer different clinical questions. |
16. Endpoint Definitions: A Statistical Reading
DLT is an early safety estimand
The DLT endpoint is deliberately narrow. Its 28-day window focuses attention on early treatment-related toxicity. Statistically, the endpoint can be viewed as an incidence proportion over a fixed short interval, provided the analysis population and exposure rules are defined.
AE is a broad cumulative safety estimand
The AE endpoint has a much longer observation window. Its definition is also broad, covering unfavorable and unintended changes temporally associated with use of the sponsor's product whether or not considered related. A participant can therefore meet the endpoint for an event that is not judged treatment-related.
ORR is a categorical efficacy estimand
ORR asks whether a participant achieves a confirmed CR or PR. The endpoint therefore requires both a measurement system and response-classification rules. RECIST 1.1 supplies those rules, while the study's lesion-count modification defines how tumor burden was followed.
Different endpoints should not be collapsed into one statistic
A DLT rate, an AE rate, and an ORR are all proportions, but their numerators, denominators, observation windows, and clinical meanings differ. Statistical similarity in mathematical form does not imply equivalence in interpretation.
17. Randomization and Multi-Arm Inference
Randomization is the principal design feature identified by the registry that supports causal comparison among the randomized study arms. In a parallel randomized study, each participant's assigned arm is determined by the trial's allocation mechanism rather than by an investigator selecting treatment after observing the participant's outcome.
With 15 arms, however, the phrase "the treatment effect" is incomplete without specifying the comparison. There can be multiple pairwise contrasts, combinations of arms, or dose-related contrasts. A statistical analysis must therefore identify the estimand and comparison of interest before the p-value or confidence interval can be interpreted.
For a study with k arms, this formula illustrates why multiplicity becomes important as the number of arms increases. It does not imply that every possible comparison was actually performed in KEYNOTE-001.
The ClinicalTrials.gov record does not state that all possible pairwise comparisons were performed. Therefore, the formula is presented only as a statistical illustration of the multi-arm problem, not as a description of the trial's actual testing program.
18. Binary Confidence Intervals
For a binary endpoint, a confidence interval complements the observed proportion by showing how precisely the study estimates the underlying event or response probability. The exact interval method should be selected according to the prespecified analysis plan and sample size rather than assumed from the endpoint label alone.
Estimate
The observed proportion describes the frequency of participants meeting the endpoint definition in the analyzed population.
Precision
The confidence interval describes statistical uncertainty around the estimated proportion.
Comparison
A treatment comparison requires a defined contrast between study arms, not merely two separate proportions.
P-value
If used, a p-value addresses a prespecified null hypothesis. It is not a measure of effect magnitude or clinical importance.
19. Statistical Interpretation of Safety Counts
The serious-adverse-event entries posted on ClinicalTrials.gov for several solid-tumor arms illustrate why denominators must remain visible. A count such as 3/10 contains more information than the numerator alone because it establishes the population at risk in that listed arm.
| Listed arm | Affected | At risk | Statistical reading |
|---|---|---|---|
| Pembrolizumab 1 mg/kg Q2W | 1 | 4 | One serious-AE-affected participant among four at risk. |
| Pembrolizumab 3 mg/kg Q2W | 1 | 3 | One serious-AE-affected participant among three at risk. |
| Pembrolizumab 10 mg/kg Q2W | 3 | 10 | Three serious-AE-affected participants among ten at risk. |
| Pembrolizumab Titration Co | 0 | 4 | No affected participants among four at risk. |
| Pembrolizumab Titration Co | 2 | 3 | Two affected participants among three at risk. |
| Pembrolizumab Titration Co | 2 | 6 | Two affected participants among six at risk. |
20. Clinical Interpretation vs Statistical Interpretation
Clinical interpretation
The registered endpoints focus on early dose-limiting toxicity, cumulative adverse events, and objective tumor response in defined disease-specific participant groups.
Statistical interpretation
Each primary endpoint is binary, so the core analysis concerns event or response proportions. Formal comparative inference requires the relevant arm denominators, prespecified contrasts, uncertainty estimates, and testing strategy.
21. What the Available Data Do Not Establish
- The ClinicalTrials.gov record does not establish a numerical treatment-effect estimate for any of the four primary endpoints.
- The ClinicalTrials.gov record does not establish a confidence interval for any primary endpoint.
- The ClinicalTrials.gov record does not establish a p-value for any primary endpoint.
- The ClinicalTrials.gov record does not establish which of the 15 arms served as a comparator for each endpoint.
- The ClinicalTrials.gov record does not establish an allocation ratio among the 15 arms.
- The ClinicalTrials.gov record does not provide the complete arm-level sample-size structure.
- The ClinicalTrials.gov record does not provide a complete serious-adverse-event table for every arm.
- The ClinicalTrials.gov record does not provide a formal multiplicity procedure, interim-analysis plan, missing-data method, imputation strategy, Bayesian method, or non-inferiority margin.
22. Related Tutorials
Learn more about the methods and statistical ideas used to understand this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: KEYNOTE-001, NCT01295827.
- PubMed: PMID 25034862.
- PubMed: PMID 25977344.
- PubMed: PMID 27117531.
- PubMed: PMID 37699333.
- PubMed: PMID 35101941.
Continue through Clinical Biostats
Connect this trial's statistical concepts to deeper tutorials, statistical calculators, and other clinical-trial analyses.
25. Record Summary
KEYNOTE-001 is a useful example of how statistical analysis must follow the structure of the clinical question rather than simply the size of the trial. The registry describes a randomized, parallel phase 1 study with 15 arms and 1260 participants, four binary primary endpoints, and results posted. Those endpoints span early dose-limiting toxicity, long-term adverse-event occurrence, and disease-specific objective response assessed using RECIST 1.1.
The most important statistical lesson is that a binary endpoint is only the starting point. Its interpretation depends on the event definition, observation window, analysis population, study-arm structure, and prespecified comparison. DLT and AE are both safety endpoints but operate on very different time frames; melanoma and NSCLC ORR are both response endpoints but apply to different participant groups and use different assessment organizations.
Because the registry-reported statistical-analysis field contains no formal estimates, confidence intervals, or p-values, a statistically responsible analysis must stop at the information actually available rather than reconstructing numerical treatment effects from incomplete inputs. That distinction between registered endpoint information, reported numerical results, and statistical interpretation is itself an important part of rigorous clinical-trial analysis.