This page separates reported registry results from statistical interpretation. The ClinicalTrials.gov record contains posted outcome measures and safety counts, but no formal statistical analyses.
1. Trial at a Glance
KEYNOTE-012 was a phase 1, non-randomized, unmasked, parallel study evaluating pembrolizumab in participants with advanced solid tumors. The registry lists cancer and solid tumor as the conditions and reports an enrollment of 297 participants across 5 arms.
| Feature | KEYNOTE-012 |
|---|---|
| Phase | Phase 1 |
| Conditions | Cancer; Solid Tumor |
| Brief title | Study of Pembrolizumab (MK-3475) in Participants With Advanced Solid Tumors (MK-3475-012/KEYNOTE-012) |
| Allocation | NON_RANDOMIZED |
| Design model | PARALLEL |
| Masking | NONE |
| Primary purpose | TREATMENT |
| Enrollment | 297 |
| Arms | 5 |
| Intervention | Pembrolizumab (biological) |
| Primary endpoint type | Binary |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | INDUSTRY |
| Status | COMPLETED |
2. Clinical Question
The registry describes KEYNOTE-012 as a study of pembrolizumab in participants with advanced solid tumors. Because the allocation is explicitly non-randomized and no comparator intervention is listed in the ClinicalTrials.gov record, the statistical question is different from that of a randomized controlled trial.
Population
Participants with advanced solid tumors; the registered conditions are Cancer and Solid Tumor.
Intervention
Pembrolizumab, identified in the registry as a biological intervention.
Comparator
No randomized comparator is identified in the ClinicalTrials.gov record. Allocation is NON_RANDOMIZED.
Primary statistical question
How frequently do participants experience the registered adverse-event and RECIST 1.1 response outcomes under the study's defined assessment framework?
This distinction matters. In a non-randomized study, a response rate can describe the observed frequency of response among an analysis population, but it does not by itself estimate a causal treatment effect relative to a randomized control group.
3. Trial Design
Study chronology
Study start
The registry lists 2013-05-07 as the study start date.
Primary completion
The registry lists 2016-04-26 as the primary completion date.
Current registry status
the ClinicalTrials.gov record identifies the study status as COMPLETED.
4. Registered Arms and Study Structure
the ClinicalTrials.gov record reports 5 arms and an enrollment of 297. It does not provide arm-level enrollment counts or a detailed intervention schedule in the ClinicalTrials.gov record. The arm-level safety information identifies the following cohort labels.
Triple Negative Breast Cancer
- Serious AEs: 13 affected of 32 at risk
Head & Neck Cancer
- Serious AEs: 27 affected of 60 at risk
Urothelial Cancer
- Serious AEs: 21 affected of 33 at risk
Gastric Cancer
- Serious AEs: 17 affected of 39 at risk
Head & Neck Cancer Expansion
- Serious AEs: 60 affected of 132 at risk
5. Primary Endpoints
The registry lists 4 primary endpoints, all classified in the ClinicalTrials.gov record as binary endpoints. Results are posted for each of the four registered primary endpoints, but no formal statistical analyses were posted.
| Registered primary endpoint | Time frame | Type |
|---|---|---|
| Number of Participants Experiencing Adverse Events (AEs) | Serious AEs: Up to 90 days after last dose of treatment (Up to 28 months); nonserious AEs: Up to 30 days after last dose of treatment (Up to 26 months) - through final analysis (FA) cutoff date 26 Apr 2016 (Cohorts: A, B, B2, D) & 01 Sep 2015 (Cohort C) | Binary |
| Number of Participants Discontinuing From Study Treatment Due to an AE | Up to last dose of study treatment (Up to approximately 25 months) - through FA cutoff date 26 Apr 2016 (Cohorts: A, B, B2, D) & 01 Sep 2015 (Cohort C) | Binary |
| Overall Response Evaluation Criteria in Solid Tumors Version 1.1 (RECIST 1.1) Response Rate Based on Blinded Independent Central Radiology (BICR) Review (Cohorts A, B, C, and D) | Every 8 weeks until disease progression (Cohorts A, B, D: Up to ~ 35 months; Cohort C: Up to ~ 28 months) - through FA cutoff date 26 Apr 2016 (Cohorts: A, B, D) & 01 Sep 2015 (Cohort C) | Binary |
| Overall RECIST 1.1 Response Rate Based on BICR Review for Participants in Cohort B2 | Every 8 weeks until disease progression (Up to ~ 35 months) - through FA cutoff date 26 Apr 2016 | Binary |
How the response endpoint is defined
The registry defines Overall Response Rate (ORR) as the percentage of participants who experienced a Complete Response (CR) or a Partial Response (PR). CR is defined as disappearance of all target lesions, while PR is defined as at least a 30% decrease in the sum of diameters of target lesions. Response was assessed using RECIST 1.1 based on BICR evaluation.
6. Endpoint Definitions and Statistical Meaning
| Endpoint | What the binary outcome represents | Statistical interpretation |
|---|---|---|
| Participants experiencing AEs | Whether a participant experienced an adverse event under the registry's AE definition and specified follow-up windows. | A participant-level event proportion can describe observed safety burden. |
| Discontinuation due to an AE | Whether treatment discontinuation occurred because of an AE during the specified treatment period. | A binary proportion can summarize the frequency of treatment discontinuation attributed to AEs. |
| RECIST 1.1 response rate, Cohorts A-D | Whether a participant experienced CR or PR under BICR assessment. | ORR is an observed response proportion, not a time-to-event estimate. |
| RECIST 1.1 response rate, Cohort B2 | Whether a participant in Cohort B2 experienced CR or PR under BICR assessment. | The same binary response framework applies specifically to Cohort B2. |
Because the ClinicalTrials.gov record does not contain formal statistical analyses, there is no registry-reported hazard ratio, odds ratio, risk ratio, confidence interval, or p-value for these endpoints. Those quantities should not be reverse-engineered from the counts or inferred from the fact that results were posted.
7. Planned Analysis
Binary adverse-event endpoints
The number of participants experiencing AEs and the number discontinuing treatment because of an AE are binary outcomes. For a binary endpoint, a basic descriptive analysis would report the number of participants with the event and the corresponding proportion within the relevant analysis population.
Here, x is the number of participants with the specified outcome and n is the number of participants in the relevant analysis population. A proportion describes observed frequency; it does not by itself establish causality.
RECIST 1.1 response rate
ORR is also a binary endpoint at the participant level: each participant is classified according to whether a CR or PR was observed under the specified RECIST 1.1 and BICR framework. A descriptive response analysis would therefore summarize responders and the corresponding response proportion for the specified cohort or analysis population.
Confidence intervals for response proportions
For a single binary response proportion, a confidence interval can quantify sampling uncertainty around the observed proportion. Exact binomial or other binomial-based intervals are commonly used when sample sizes are limited or when a proportion is close to the boundary of 0 or 1. The ClinicalTrials.gov record does not identify which confidence-interval method was used.
Why a comparative p-value is not reported
A conventional between-arm treatment comparison requires a clearly defined comparator and an analysis population that supports the comparison. The ClinicalTrials.gov record identifies KEYNOTE-012 as non-randomized and does not provide a randomized comparator. Consequently, a formal treatment-versus-control p-value cannot be responsibly constructed from the ClinicalTrials.gov record.
8. Statistical Methods Explained
Why are the primary endpoints binary?
Each registered primary endpoint can be represented as a yes/no participant-level outcome: experiencing an AE, discontinuing because of an AE, achieving CR or PR under RECIST 1.1, or achieving CR or PR in Cohort B2. Binary endpoints are naturally summarized using counts and proportions.
What does ORR measure?
ORR measures the proportion of participants who achieved either a CR or PR according to the specified response criteria. It does not measure how long a response lasts, overall survival, or the probability that a particular participant would have responded under an alternative treatment.
Why does BICR matter statistically?
Blinded Independent Central Radiology review provides a standardized radiologic assessment framework. Blinding the central review can reduce the opportunity for knowledge of treatment or other clinical information to influence radiologic classification. The registry specifically identifies BICR review for the response endpoints.
Why is non-randomization important?
Randomization creates the basis for a direct causal comparison between assigned treatment groups. KEYNOTE-012 is identified as NON_RANDOMIZED, so an observed response proportion describes the study population under the study's treatment framework rather than automatically estimating a causal effect relative to a control group.
How should serious AE counts be interpreted?
A serious AE count records how many participants were affected among those at risk for the cohort. It is a safety-frequency measure. It should not be interpreted as proof that pembrolizumab caused every event because the registry's AE definition explicitly states that an AE did not necessarily have a causal relationship with study treatment.
Why does the AE follow-up window matter?
The registered endpoint uses different windows for serious and nonserious AEs: serious AEs are followed up to 90 days after the last dose, while nonserious AEs are followed up to 30 days after the last dose. Changing the observation window can change the number of events captured, so safety frequencies should always be interpreted with their defined follow-up period.
Why should the response rate not be converted into a treatment effect?
A response proportion is an absolute descriptive quantity. Without a randomized comparator or an appropriate external-control design, subtracting or dividing response rates does not automatically provide an unbiased estimate of the causal effect of treatment.
9. Safety Results
the ClinicalTrials.gov record provides serious adverse-event counts by cohort. These are the available arm-level safety results in the ClinicalTrials.gov record.
| Cohort | Condition | Serious AEs affected | At risk | Observed proportion |
|---|---|---|---|---|
| Cohort A | Triple Negative Breast Cancer | 13 | 32 | 13/32 |
| Cohort B | Head & Neck Cancer | 27 | 60 | 27/60 |
| Cohort C | Urothelial Cancer | 21 | 33 | 21/33 |
| Cohort D | Gastric Cancer | 17 | 39 | 17/39 |
| Cohort B2 | Head & Neck Cancer Expansion | 60 | 132 | 60/132 |
The registry defines an AE as any untoward medical occurrence in a participant administered a study treatment that did not necessarily have to have a causal relationship with the treatment. An AE could be an unfavorable or unintended sign, symptom, disease, or abnormal laboratory finding temporally associated with study treatment, whether or not considered related.
10. Clinical Biostats Interpretation of the Safety Data
The reported serious-AE counts identify the number of participants affected and the number at risk in each named cohort. For example, Cohort B2 has 60 affected participants among 132 at risk. This describes the observed frequency within that cohort.
These counts do not establish that pembrolizumab caused the events. The registry's own AE definition explicitly states that an AE did not necessarily have a causal relationship with study treatment. They also do not provide a randomized estimate of comparative safety.
The affected count cannot be interpreted without its corresponding at-risk population. A count of 60 has a different statistical meaning depending on whether the denominator is 60, 132, or another population. The ClinicalTrials.gov record therefore report both affected and at-risk counts.
The ClinicalTrials.gov record does not report confidence intervals for these serious-AE frequencies. A confidence interval could be calculated for a binomial proportion in an independent analysis, but presenting one here as a trial-reported result would incorrectly add information that is not contained in the ClinicalTrials.gov record.
11. Response Assessment
The registered response endpoints use RECIST 1.1 and Blinded Independent Central Radiology (BICR) Review. For Cohorts A, B, C, and D, response was assessed every 8 weeks until disease progression, with the registry specifying different approximate maximum time frames by cohort. Cohort B2 has its own registered response-rate endpoint with assessments every 8 weeks until disease progression.
Complete Response
The registry defines CR as disappearance of all target lesions.
Partial Response
The registry defines PR as at least a 30% decrease in the sum of diameters of target lesions.
ORR
ORR is the percentage of participants experiencing either CR or PR.
Assessment schedule
Response was assessed every 8 weeks until disease progression under the registered endpoint definitions.
The formula describes the statistical structure of ORR. The ClinicalTrials.gov record does not provide the responder counts or formal statistical estimates needed to populate a trial-specific ORR calculation.
12. Why No Formal Efficacy Comparison Is Reported Here
The registry indicates that results are posted for both RECIST response endpoints, but the ClinicalTrials.gov record contains no entries in the formal statistical-analysis section. This is important because the presence of a posted outcome does not imply that a comparative hypothesis test, confidence interval, or effect estimate is available.
For a non-randomized phase 1 study, response rates can be clinically and statistically informative as descriptive evidence of observed tumor response. However, a response rate alone does not answer the counterfactual question: what would the same participants' response probability have been under another treatment?
A formal comparative analysis would require additional information, such as a prespecified comparator, an appropriate external-control framework, or another design that supports the intended inference. None of those additional elements are contained in the ClinicalTrials.gov record, so this page does not add them.
13. Statistical Methodology
Descriptive analysis of binary outcomes
The four registered primary endpoints are binary. The most direct descriptive summary is the number and percentage of participants experiencing the specified event. For safety endpoints, the denominator should correspond to the relevant population at risk and follow the registered observation period.
Binomial confidence intervals
When a binary endpoint is summarized as a proportion, a binomial confidence interval can quantify uncertainty around the observed proportion. Several methods are available, including exact binomial intervals and score-based intervals. The ClinicalTrials.gov record does not specify which method, if any, was used for the posted results.
RECIST response analysis
RECIST 1.1 response is classified at the participant level according to radiologic findings. Because the primary response endpoints are binary, a response analysis naturally reduces each participant to a response classification for the endpoint's defined assessment framework.
Independent central review
The response endpoints are based on BICR review. From a statistical-design perspective, an independent review process can standardize outcome classification and reduce the potential influence of clinical treatment knowledge on radiologic assessment.
Safety analysis
Safety is summarized according to whether an adverse event occurred during the defined follow-up period. The distinction between serious and nonserious AE windows is part of the endpoint definition and should be retained when interpreting event frequencies.
A response or AE proportion describes what was observed in a defined population. A causal treatment effect requires a design and analysis that support a counterfactual comparison.
14. Randomization, Stratification, and Analysis Populations
The ClinicalTrials.gov record explicitly identify the allocation as NON_RANDOMIZED. They do not provide stratification factors or detailed analysis-population definitions beyond the cohort-level safety denominators.
| Feature | What the ClinicalTrials.gov record supports |
|---|---|
| Randomization | None; allocation is registered as NON_RANDOMIZED. |
| Stratification | No stratification factors are reported. |
| Analysis populations | No complete efficacy or safety population definitions are reported beyond the reported at-risk denominators for serious AE counts. |
| Comparator | No randomized comparator is reported. |
| Masking | NONE. |
| Design model | PARALLEL. |
These design features determine what statistical conclusions are supportable. In particular, the absence of randomization means that an observed response rate should not be described as a randomized treatment effect.
15. Missing Data and Follow-Up Considerations
The ClinicalTrials.gov record provides endpoint definitions and follow-up windows but do not describe a specific missing-data or imputation strategy. That omission is consequential for binary response endpoints because participants who are not evaluable for response may affect the choice of analysis population and denominator.
For safety outcomes, the registry supplies affected and at-risk counts for the named cohorts. For response, the ClinicalTrials.gov record does not provide the complete responder and denominator information needed to reconstruct the posted numerical response results.
Response missingness
A binary response analysis must define how participants without an evaluable response assessment are handled. The ClinicalTrials.gov record does not specify such a rule.
Safety observation windows
Different AE follow-up windows can capture different event experiences, so safety estimates must remain tied to their registered definitions.
16. Multiplicity and Multiple Endpoints
KEYNOTE-012 has 4 registered primary endpoints, all binary. Multiple primary endpoints raise an important statistical issue: if each endpoint is tested independently at a conventional significance level, the probability of at least one false-positive finding can exceed the nominal level for a single test.
The ClinicalTrials.gov record does not identify an alpha-allocation strategy, hierarchical testing procedure, multiplicity adjustment, or formal hypothesis-testing framework. Accordingly, no multiplicity-adjusted inference is presented on this page.
| Endpoint family | Registered role | Statistical issue |
|---|---|---|
| Adverse events | Primary endpoint | Binary safety outcome; interpretation depends on defined observation window. |
| Discontinuation due to AE | Primary endpoint | Binary treatment-tolerability outcome. |
| RECIST 1.1 response, Cohorts A-D | Primary endpoint | Binary tumor-response outcome based on BICR. |
| RECIST 1.1 response, Cohort B2 | Primary endpoint | Binary tumor-response outcome specific to the expansion cohort. |
17. Interim Analysis, Crossover, and Bayesian Methods
Interim analysis
The ClinicalTrials.gov record does not identify an interim-analysis plan or alpha-spending method.
Crossover
The ClinicalTrials.gov record does not identify a crossover design or crossover analysis.
Factorial design
The registry identifies a parallel design; no factorial structure is reported.
Bayesian methods
The ClinicalTrials.gov record does not identify a Bayesian analysis or prior distribution.
These omissions are not reasons to infer that the methods were absent from the underlying protocol or statistical analysis plan. They mean only that the ClinicalTrials.gov record does not provide enough information to describe those methods as trial-specific analyses.
18. Primary Endpoint Interpretation
The first primary endpoint measures the number of participants experiencing AEs under the registry definition. Serious AEs are followed up to 90 days after the last dose, while nonserious AEs are followed up to 30 days after the last dose. This endpoint is fundamentally a safety-frequency measure.
The second primary endpoint measures participants discontinuing study treatment due to an AE through the specified treatment period. It addresses treatment discontinuation rather than tumor response.
The third primary endpoint measures CR or PR using RECIST 1.1 based on BICR review. It is assessed every 8 weeks until disease progression under the registered time frame.
The fourth primary endpoint applies the same CR-or-PR response framework to participants in Cohort B2, with assessment every 8 weeks until disease progression and a registered time frame extending up to approximately 35 months.
19. Limitations
- Non-randomized design: the registry identifies allocation as NON_RANDOMIZED, so observed outcomes cannot automatically be interpreted as randomized causal treatment effects.
- No formal statistical analyses reported: no formal statistical analyses were posted despite results being posted.
- Incomplete numerical efficacy information: the ClinicalTrials.gov record defines the response endpoints but do not provide the numerical response estimates, confidence intervals, or p-values.
- No comparator information: the ClinicalTrials.gov record does not identify a randomized comparator arm.
- No stratification information: the ClinicalTrials.gov record does not identify stratification factors.
- No specified missing-data method: the ClinicalTrials.gov record does not state how participants without evaluable response assessments were handled.
- Multiple primary endpoints: four primary binary endpoints are registered, but the ClinicalTrials.gov record does not describe multiplicity control.
- Cohort-specific denominators: serious-AE counts are reported using different at-risk denominators across cohorts and should not be compared as though they arose from equal-sized randomized groups.
- AE causality: the registry definition explicitly states that an AE does not necessarily have a causal relationship with study treatment.
20. Why This Trial Matters Statistically
KEYNOTE-012 is useful as a statistical teaching case because it illustrates how the appropriate analysis depends on the study design and endpoint structure. It combines a non-randomized phase 1 framework with binary safety outcomes and centrally reviewed binary tumor-response outcomes.
| Concept | How it appears in KEYNOTE-012 |
|---|---|
| Non-randomized design | Allocation is registered as NON_RANDOMIZED. |
| Parallel design | The design model is registered as PARALLEL. |
| Binary endpoints | All 4 registered primary endpoints are classified as binary. |
| Safety analysis | Serious AE counts are reported by named cohort. |
| Response analysis | RECIST 1.1 response is based on BICR review. |
| Repeated assessment | Response is assessed every 8 weeks until disease progression. |
| Denominator selection | Safety results are reported using cohort-specific at-risk populations. |
| Multiplicty | Four registered primary endpoints require attention to the overall inferential framework. |
| Causal inference | Non-randomization limits direct causal interpretation of observed response proportions. |
| Registry versus analysis | Results are posted, but no formal statistical-analysis entries are reported. |
21. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: KEYNOTE-012, NCT01848834.
- PubMed: Publication indexed under PMID 27157491.
- PubMed: Publication indexed under PMID 27138582.
- PubMed: Publication indexed under PMID 27247226.
- PubMed: Publication indexed under PMID 28081914.
- PubMed: Publication indexed under PMID 27646946.
The linked PubMed records are provided as publication references associated with the ClinicalTrials.gov record. Numerical trial claims on this page are restricted to the ClinicalTrials.gov record.
Continue through the Clinical Biostats statistical library
Explore tutorials and calculators covering binary endpoints, response rates, confidence intervals, safety analysis, and clinical-trial design.
24. Record Summary
KEYNOTE-012 is a completed phase 1, non-randomized, unmasked, parallel study of pembrolizumab in participants with advanced solid tumors. The ClinicalTrials.gov record reports an enrollment of 297 participants across 5 arms and identify 4 binary primary endpoints covering adverse events, discontinuation due to adverse events, and RECIST 1.1 response in specified cohorts.
The statistical lesson is primarily methodological: because the study is non-randomized and the registry-reported statistical-analysis section is empty, the available evidence should be separated into descriptive outcome reporting and causal or comparative inference. Serious-AE counts can be described using their cohort-specific denominators, while RECIST 1.1 response can be understood as a binary endpoint assessed by BICR. Formal effect estimates, confidence intervals, and p-values should not be attributed to the trial when they are not present in the ClinicalTrials.gov record.