← Clinical Trials
Advanced Solid Tumors Phase 1 Non-Randomized NCT01848834

KEYNOTE-012: Complete Statistical Analysis of Pembrolizumab in Advanced Solid Tumors

An independent statistical review of the phase 1 KEYNOTE-012 study of pembrolizumab in participants with advanced solid tumors, focusing on its non-randomized parallel design, registered binary endpoints, response assessment, safety reporting, and the statistical methods appropriate for interpreting the posted registry data.

Trial status: COMPLETED  ·  Start: 07 May 2013  ·  Primary completion: 26 Apr 2016
Scope of this record

This page separates reported registry results from statistical interpretation. The ClinicalTrials.gov record contains posted outcome measures and safety counts, but no formal statistical analyses.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-012 was a phase 1, non-randomized, unmasked, parallel study evaluating pembrolizumab in participants with advanced solid tumors. The registry lists cancer and solid tumor as the conditions and reports an enrollment of 297 participants across 5 arms.

297
Enrollment
Participants
5
Arms
Non-randomized
4
Primary Endpoints
Registered
10
Outcomes Posted
ClinicalTrials.gov
FeatureKEYNOTE-012
PhasePhase 1
ConditionsCancer; Solid Tumor
Brief titleStudy of Pembrolizumab (MK-3475) in Participants With Advanced Solid Tumors (MK-3475-012/KEYNOTE-012)
AllocationNON_RANDOMIZED
Design modelPARALLEL
MaskingNONE
Primary purposeTREATMENT
Enrollment297
Arms5
InterventionPembrolizumab (biological)
Primary endpoint typeBinary
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeINDUSTRY
StatusCOMPLETED

2. Clinical Question

The registry describes KEYNOTE-012 as a study of pembrolizumab in participants with advanced solid tumors. Because the allocation is explicitly non-randomized and no comparator intervention is listed in the ClinicalTrials.gov record, the statistical question is different from that of a randomized controlled trial.

Population

Participants with advanced solid tumors; the registered conditions are Cancer and Solid Tumor.

Intervention

Pembrolizumab, identified in the registry as a biological intervention.

Comparator

No randomized comparator is identified in the ClinicalTrials.gov record. Allocation is NON_RANDOMIZED.

Primary statistical question

How frequently do participants experience the registered adverse-event and RECIST 1.1 response outcomes under the study's defined assessment framework?

This distinction matters. In a non-randomized study, a response rate can describe the observed frequency of response among an analysis population, but it does not by itself estimate a causal treatment effect relative to a randomized control group.

3. Trial Design

01
Enroll 297 participants
02
Study arms 5 arms
03
Treatment Pembrolizumab
04
Assess Response and safety
05
Report Binary outcomes
Allocation
NON_RANDOMIZED. Participants were not assigned through a randomized treatment allocation in the ClinicalTrials.gov record.
Design model
PARALLEL. The registry identifies the study as a parallel design with 5 arms.
Masking
NONE. The ClinicalTrials.gov record identifies no masking.
Primary purpose
TREATMENT. The registry classifies the primary purpose as treatment.

Study chronology

07 May 2013

Study start

The registry lists 2013-05-07 as the study start date.

26 Apr 2016

Primary completion

The registry lists 2016-04-26 as the primary completion date.

Completed

Current registry status

the ClinicalTrials.gov record identifies the study status as COMPLETED.

4. Registered Arms and Study Structure

the ClinicalTrials.gov record reports 5 arms and an enrollment of 297. It does not provide arm-level enrollment counts or a detailed intervention schedule in the ClinicalTrials.gov record. The arm-level safety information identifies the following cohort labels.

COHORT A

Triple Negative Breast Cancer

  • Serious AEs: 13 affected of 32 at risk
COHORT B

Head & Neck Cancer

  • Serious AEs: 27 affected of 60 at risk
COHORT C

Urothelial Cancer

  • Serious AEs: 21 affected of 33 at risk
COHORT D

Gastric Cancer

  • Serious AEs: 17 affected of 39 at risk
COHORT B2

Head & Neck Cancer Expansion

  • Serious AEs: 60 affected of 132 at risk
Important design distinction: these cohort-level safety counts should not be treated as randomized treatment-group comparisons. The registry identifies the overall allocation as NON_RANDOMIZED, and the ClinicalTrials.gov record does not provide a randomized control arm against which these cohorts can be compared.

5. Primary Endpoints

The registry lists 4 primary endpoints, all classified in the ClinicalTrials.gov record as binary endpoints. Results are posted for each of the four registered primary endpoints, but no formal statistical analyses were posted.

Registered primary endpointTime frameType
Number of Participants Experiencing Adverse Events (AEs) Serious AEs: Up to 90 days after last dose of treatment (Up to 28 months); nonserious AEs: Up to 30 days after last dose of treatment (Up to 26 months) - through final analysis (FA) cutoff date 26 Apr 2016 (Cohorts: A, B, B2, D) & 01 Sep 2015 (Cohort C) Binary
Number of Participants Discontinuing From Study Treatment Due to an AE Up to last dose of study treatment (Up to approximately 25 months) - through FA cutoff date 26 Apr 2016 (Cohorts: A, B, B2, D) & 01 Sep 2015 (Cohort C) Binary
Overall Response Evaluation Criteria in Solid Tumors Version 1.1 (RECIST 1.1) Response Rate Based on Blinded Independent Central Radiology (BICR) Review (Cohorts A, B, C, and D) Every 8 weeks until disease progression (Cohorts A, B, D: Up to ~ 35 months; Cohort C: Up to ~ 28 months) - through FA cutoff date 26 Apr 2016 (Cohorts: A, B, D) & 01 Sep 2015 (Cohort C) Binary
Overall RECIST 1.1 Response Rate Based on BICR Review for Participants in Cohort B2 Every 8 weeks until disease progression (Up to ~ 35 months) - through FA cutoff date 26 Apr 2016 Binary

How the response endpoint is defined

The registry defines Overall Response Rate (ORR) as the percentage of participants who experienced a Complete Response (CR) or a Partial Response (PR). CR is defined as disappearance of all target lesions, while PR is defined as at least a 30% decrease in the sum of diameters of target lesions. Response was assessed using RECIST 1.1 based on BICR evaluation.

6. Endpoint Definitions and Statistical Meaning

EndpointWhat the binary outcome representsStatistical interpretation
Participants experiencing AEs Whether a participant experienced an adverse event under the registry's AE definition and specified follow-up windows. A participant-level event proportion can describe observed safety burden.
Discontinuation due to an AE Whether treatment discontinuation occurred because of an AE during the specified treatment period. A binary proportion can summarize the frequency of treatment discontinuation attributed to AEs.
RECIST 1.1 response rate, Cohorts A-D Whether a participant experienced CR or PR under BICR assessment. ORR is an observed response proportion, not a time-to-event estimate.
RECIST 1.1 response rate, Cohort B2 Whether a participant in Cohort B2 experienced CR or PR under BICR assessment. The same binary response framework applies specifically to Cohort B2.

Because the ClinicalTrials.gov record does not contain formal statistical analyses, there is no registry-reported hazard ratio, odds ratio, risk ratio, confidence interval, or p-value for these endpoints. Those quantities should not be reverse-engineered from the counts or inferred from the fact that results were posted.

7. Planned Analysis

No formal statistical analyses are reported. The trial data report that results are posted for all 4 registered primary endpoints, but no formal statistical analyses were posted.

Binary adverse-event endpoints

The number of participants experiencing AEs and the number discontinuing treatment because of an AE are binary outcomes. For a binary endpoint, a basic descriptive analysis would report the number of participants with the event and the corresponding proportion within the relevant analysis population.

Binary endpoint framework
p = x / n

Here, x is the number of participants with the specified outcome and n is the number of participants in the relevant analysis population. A proportion describes observed frequency; it does not by itself establish causality.

RECIST 1.1 response rate

ORR is also a binary endpoint at the participant level: each participant is classified according to whether a CR or PR was observed under the specified RECIST 1.1 and BICR framework. A descriptive response analysis would therefore summarize responders and the corresponding response proportion for the specified cohort or analysis population.

Confidence intervals for response proportions

For a single binary response proportion, a confidence interval can quantify sampling uncertainty around the observed proportion. Exact binomial or other binomial-based intervals are commonly used when sample sizes are limited or when a proportion is close to the boundary of 0 or 1. The ClinicalTrials.gov record does not identify which confidence-interval method was used.

Why a comparative p-value is not reported

A conventional between-arm treatment comparison requires a clearly defined comparator and an analysis population that supports the comparison. The ClinicalTrials.gov record identifies KEYNOTE-012 as non-randomized and does not provide a randomized comparator. Consequently, a formal treatment-versus-control p-value cannot be responsibly constructed from the ClinicalTrials.gov record.

8. Statistical Methods Explained

Why are the primary endpoints binary?

Each registered primary endpoint can be represented as a yes/no participant-level outcome: experiencing an AE, discontinuing because of an AE, achieving CR or PR under RECIST 1.1, or achieving CR or PR in Cohort B2. Binary endpoints are naturally summarized using counts and proportions.

What does ORR measure?

ORR measures the proportion of participants who achieved either a CR or PR according to the specified response criteria. It does not measure how long a response lasts, overall survival, or the probability that a particular participant would have responded under an alternative treatment.

Why does BICR matter statistically?

Blinded Independent Central Radiology review provides a standardized radiologic assessment framework. Blinding the central review can reduce the opportunity for knowledge of treatment or other clinical information to influence radiologic classification. The registry specifically identifies BICR review for the response endpoints.

Why is non-randomization important?

Randomization creates the basis for a direct causal comparison between assigned treatment groups. KEYNOTE-012 is identified as NON_RANDOMIZED, so an observed response proportion describes the study population under the study's treatment framework rather than automatically estimating a causal effect relative to a control group.

How should serious AE counts be interpreted?

A serious AE count records how many participants were affected among those at risk for the cohort. It is a safety-frequency measure. It should not be interpreted as proof that pembrolizumab caused every event because the registry's AE definition explicitly states that an AE did not necessarily have a causal relationship with study treatment.

Why does the AE follow-up window matter?

The registered endpoint uses different windows for serious and nonserious AEs: serious AEs are followed up to 90 days after the last dose, while nonserious AEs are followed up to 30 days after the last dose. Changing the observation window can change the number of events captured, so safety frequencies should always be interpreted with their defined follow-up period.

Why should the response rate not be converted into a treatment effect?

A response proportion is an absolute descriptive quantity. Without a randomized comparator or an appropriate external-control design, subtracting or dividing response rates does not automatically provide an unbiased estimate of the causal effect of treatment.

9. Safety Results

the ClinicalTrials.gov record provides serious adverse-event counts by cohort. These are the available arm-level safety results in the ClinicalTrials.gov record.

CohortConditionSerious AEs affectedAt riskObserved proportion
Cohort ATriple Negative Breast Cancer133213/32
Cohort BHead & Neck Cancer276027/60
Cohort CUrothelial Cancer213321/33
Cohort DGastric Cancer173917/39
Cohort B2Head & Neck Cancer Expansion6013260/132

The registry defines an AE as any untoward medical occurrence in a participant administered a study treatment that did not necessarily have to have a causal relationship with the treatment. An AE could be an unfavorable or unintended sign, symptom, disease, or abnormal laboratory finding temporally associated with study treatment, whether or not considered related.

Serious AE counts by cohort
Cohort A
13/32
Cohort B
27/60
Cohort C
21/33
Cohort D
17/39
Cohort B2
60/132
Interpretation caution: the chart is a descriptive visualization of the registry-reported affected/at-risk counts. It should not be read as a randomized comparison because the trial is non-randomized and the cohort denominators differ.

10. Clinical Biostats Interpretation of the Safety Data

What the counts mean

The reported serious-AE counts identify the number of participants affected and the number at risk in each named cohort. For example, Cohort B2 has 60 affected participants among 132 at risk. This describes the observed frequency within that cohort.

What the counts do not mean

These counts do not establish that pembrolizumab caused the events. The registry's own AE definition explicitly states that an AE did not necessarily have a causal relationship with study treatment. They also do not provide a randomized estimate of comparative safety.

Why denominators matter

The affected count cannot be interpreted without its corresponding at-risk population. A count of 60 has a different statistical meaning depending on whether the denominator is 60, 132, or another population. The ClinicalTrials.gov record therefore report both affected and at-risk counts.

Why no confidence interval is shown

The ClinicalTrials.gov record does not report confidence intervals for these serious-AE frequencies. A confidence interval could be calculated for a binomial proportion in an independent analysis, but presenting one here as a trial-reported result would incorrectly add information that is not contained in the ClinicalTrials.gov record.

11. Response Assessment

The registered response endpoints use RECIST 1.1 and Blinded Independent Central Radiology (BICR) Review. For Cohorts A, B, C, and D, response was assessed every 8 weeks until disease progression, with the registry specifying different approximate maximum time frames by cohort. Cohort B2 has its own registered response-rate endpoint with assessments every 8 weeks until disease progression.

Complete Response

The registry defines CR as disappearance of all target lesions.

Partial Response

The registry defines PR as at least a 30% decrease in the sum of diameters of target lesions.

ORR

ORR is the percentage of participants experiencing either CR or PR.

Assessment schedule

Response was assessed every 8 weeks until disease progression under the registered endpoint definitions.

Binary response framework
ORR = (number with CR or PR) / (number in the response analysis population)

The formula describes the statistical structure of ORR. The ClinicalTrials.gov record does not provide the responder counts or formal statistical estimates needed to populate a trial-specific ORR calculation.

12. Why No Formal Efficacy Comparison Is Reported Here

The registry indicates that results are posted for both RECIST response endpoints, but the ClinicalTrials.gov record contains no entries in the formal statistical-analysis section. This is important because the presence of a posted outcome does not imply that a comparative hypothesis test, confidence interval, or effect estimate is available.

For a non-randomized phase 1 study, response rates can be clinically and statistically informative as descriptive evidence of observed tumor response. However, a response rate alone does not answer the counterfactual question: what would the same participants' response probability have been under another treatment?

A formal comparative analysis would require additional information, such as a prespecified comparator, an appropriate external-control framework, or another design that supports the intended inference. None of those additional elements are contained in the ClinicalTrials.gov record, so this page does not add them.

Registry-versus-analysis distinction: ClinicalTrials.gov reports that the outcomes were posted and identifies their definitions and time frames. The ClinicalTrials.gov record does not provide formal statistical analyses. This page therefore preserves the distinction between what was measured and what statistical inference can be supported from the ClinicalTrials.gov record.

13. Statistical Methodology

Descriptive analysis of binary outcomes

The four registered primary endpoints are binary. The most direct descriptive summary is the number and percentage of participants experiencing the specified event. For safety endpoints, the denominator should correspond to the relevant population at risk and follow the registered observation period.

Binomial confidence intervals

When a binary endpoint is summarized as a proportion, a binomial confidence interval can quantify uncertainty around the observed proportion. Several methods are available, including exact binomial intervals and score-based intervals. The ClinicalTrials.gov record does not specify which method, if any, was used for the posted results.

RECIST response analysis

RECIST 1.1 response is classified at the participant level according to radiologic findings. Because the primary response endpoints are binary, a response analysis naturally reduces each participant to a response classification for the endpoint's defined assessment framework.

Independent central review

The response endpoints are based on BICR review. From a statistical-design perspective, an independent review process can standardize outcome classification and reduce the potential influence of clinical treatment knowledge on radiologic assessment.

Safety analysis

Safety is summarized according to whether an adverse event occurred during the defined follow-up period. The distinction between serious and nonserious AE windows is part of the endpoint definition and should be retained when interpreting event frequencies.

Core statistical distinction
Descriptive proportion ≠ causal treatment effect

A response or AE proportion describes what was observed in a defined population. A causal treatment effect requires a design and analysis that support a counterfactual comparison.

14. Randomization, Stratification, and Analysis Populations

The ClinicalTrials.gov record explicitly identify the allocation as NON_RANDOMIZED. They do not provide stratification factors or detailed analysis-population definitions beyond the cohort-level safety denominators.

FeatureWhat the ClinicalTrials.gov record supports
RandomizationNone; allocation is registered as NON_RANDOMIZED.
StratificationNo stratification factors are reported.
Analysis populationsNo complete efficacy or safety population definitions are reported beyond the reported at-risk denominators for serious AE counts.
ComparatorNo randomized comparator is reported.
MaskingNONE.
Design modelPARALLEL.

These design features determine what statistical conclusions are supportable. In particular, the absence of randomization means that an observed response rate should not be described as a randomized treatment effect.

15. Missing Data and Follow-Up Considerations

The ClinicalTrials.gov record provides endpoint definitions and follow-up windows but do not describe a specific missing-data or imputation strategy. That omission is consequential for binary response endpoints because participants who are not evaluable for response may affect the choice of analysis population and denominator.

For safety outcomes, the registry supplies affected and at-risk counts for the named cohorts. For response, the ClinicalTrials.gov record does not provide the complete responder and denominator information needed to reconstruct the posted numerical response results.

Response missingness

A binary response analysis must define how participants without an evaluable response assessment are handled. The ClinicalTrials.gov record does not specify such a rule.

Safety observation windows

Different AE follow-up windows can capture different event experiences, so safety estimates must remain tied to their registered definitions.

16. Multiplicity and Multiple Endpoints

KEYNOTE-012 has 4 registered primary endpoints, all binary. Multiple primary endpoints raise an important statistical issue: if each endpoint is tested independently at a conventional significance level, the probability of at least one false-positive finding can exceed the nominal level for a single test.

The ClinicalTrials.gov record does not identify an alpha-allocation strategy, hierarchical testing procedure, multiplicity adjustment, or formal hypothesis-testing framework. Accordingly, no multiplicity-adjusted inference is presented on this page.

Endpoint familyRegistered roleStatistical issue
Adverse eventsPrimary endpointBinary safety outcome; interpretation depends on defined observation window.
Discontinuation due to AEPrimary endpointBinary treatment-tolerability outcome.
RECIST 1.1 response, Cohorts A-DPrimary endpointBinary tumor-response outcome based on BICR.
RECIST 1.1 response, Cohort B2Primary endpointBinary tumor-response outcome specific to the expansion cohort.

17. Interim Analysis, Crossover, and Bayesian Methods

Interim analysis

The ClinicalTrials.gov record does not identify an interim-analysis plan or alpha-spending method.

Crossover

The ClinicalTrials.gov record does not identify a crossover design or crossover analysis.

Factorial design

The registry identifies a parallel design; no factorial structure is reported.

Bayesian methods

The ClinicalTrials.gov record does not identify a Bayesian analysis or prior distribution.

These omissions are not reasons to infer that the methods were absent from the underlying protocol or statistical analysis plan. They mean only that the ClinicalTrials.gov record does not provide enough information to describe those methods as trial-specific analyses.

18. Primary Endpoint Interpretation

Adverse events

The first primary endpoint measures the number of participants experiencing AEs under the registry definition. Serious AEs are followed up to 90 days after the last dose, while nonserious AEs are followed up to 30 days after the last dose. This endpoint is fundamentally a safety-frequency measure.

Discontinuation due to an AE

The second primary endpoint measures participants discontinuing study treatment due to an AE through the specified treatment period. It addresses treatment discontinuation rather than tumor response.

RECIST 1.1 response in Cohorts A-D

The third primary endpoint measures CR or PR using RECIST 1.1 based on BICR review. It is assessed every 8 weeks until disease progression under the registered time frame.

RECIST 1.1 response in Cohort B2

The fourth primary endpoint applies the same CR-or-PR response framework to participants in Cohort B2, with assessment every 8 weeks until disease progression and a registered time frame extending up to approximately 35 months.

19. Limitations

20. Why This Trial Matters Statistically

KEYNOTE-012 is useful as a statistical teaching case because it illustrates how the appropriate analysis depends on the study design and endpoint structure. It combines a non-randomized phase 1 framework with binary safety outcomes and centrally reviewed binary tumor-response outcomes.

ConceptHow it appears in KEYNOTE-012
Non-randomized designAllocation is registered as NON_RANDOMIZED.
Parallel designThe design model is registered as PARALLEL.
Binary endpointsAll 4 registered primary endpoints are classified as binary.
Safety analysisSerious AE counts are reported by named cohort.
Response analysisRECIST 1.1 response is based on BICR review.
Repeated assessmentResponse is assessed every 8 weeks until disease progression.
Denominator selectionSafety results are reported using cohort-specific at-risk populations.
MultiplictyFour registered primary endpoints require attention to the overall inferential framework.
Causal inferenceNon-randomization limits direct causal interpretation of observed response proportions.
Registry versus analysisResults are posted, but no formal statistical-analysis entries are reported.

21. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Sources

The linked PubMed records are provided as publication references associated with the ClinicalTrials.gov record. Numerical trial claims on this page are restricted to the ClinicalTrials.gov record.

Continue through the Clinical Biostats statistical library

Explore tutorials and calculators covering binary endpoints, response rates, confidence intervals, safety analysis, and clinical-trial design.

24. Record Summary

KEYNOTE-012 is a completed phase 1, non-randomized, unmasked, parallel study of pembrolizumab in participants with advanced solid tumors. The ClinicalTrials.gov record reports an enrollment of 297 participants across 5 arms and identify 4 binary primary endpoints covering adverse events, discontinuation due to adverse events, and RECIST 1.1 response in specified cohorts.

The statistical lesson is primarily methodological: because the study is non-randomized and the registry-reported statistical-analysis section is empty, the available evidence should be separated into descriptive outcome reporting and causal or comparative inference. Serious-AE counts can be described using their cohort-specific denominators, while RECIST 1.1 response can be understood as a binary endpoint assessed by BICR. Formal effect estimates, confidence intervals, and p-values should not be attributed to the trial when they are not present in the ClinicalTrials.gov record.

Clinical Biostats methodology: A trial-results page should distinguish the endpoint that was registered, the population and observation window used to define it, the analysis actually reported, and the statistical inference that the design can support. When formal analysis results are unavailable, the appropriate approach is to explain the analysis that would ordinarily be used without presenting an unreported estimate as though it were a trial result.