← Clinical Trials
Advanced / Metastatic NSCLC Phase 3 Progression-Free Survival NCT04303780

CodeBreaK 200: Complete Statistical Analysis of Sotorasib in KRAS G12C-Mutated Advanced NSCLC

An independent statistical analysis of the randomized phase 3 CodeBreaK 200 trial comparing AMG 510 (proposed INN sotorasib) with docetaxel in participants with KRAS p.G12C-mutated advanced or metastatic non-small cell lung cancer (NSCLC).

Trial status: COMPLETED  ·  Enrollment: 345  ·  Primary completion: 02 August 2022
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial facts are restricted to the ClinicalTrials.gov record. The official registry record is the authoritative source for the trial record: ClinicalTrials.gov NCT04303780.

1. Trial at a Glance

CodeBreaK 200 was a randomized, parallel-group, phase 3 trial comparing AMG 510 with docetaxel in advanced or metastatic NSCLC with KRAS p.G12C mutation. The registry reports one primary time-to-event endpoint: progression-free survival (PFS).

345
Enrollment
2 treatment arms
3
Phase
Phase 3
0.663
PFS HR
95% CI 0.509–0.864
0.002
P-value
Stratified log-rank
FeatureCodeBreaK 200
TrialCodeBreaK 200
NCT identifierNCT04303780
PhasePhase 3
PopulationKRAS p.G12C-mutated / advanced metastatic NSCLC
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment345
InterventionsAMG 510 and docetaxel
Primary endpointProgression-free survival (PFS)
Primary endpoint typeTime-to-event
Statistical method reportedLog-rank test
Effect measureHazard ratio
Registry hypothesis typeEquivalence
Lead sponsorAmgen

2. Clinical Question

The primary statistical question is whether the time from randomization to disease progression or death differs between AMG 510 and docetaxel in participants with KRAS p.G12C-mutated advanced or metastatic NSCLC.

Population

Participants with KRAS p.G12C-mutated advanced or metastatic non-small cell lung cancer (NSCLC).

Intervention

AMG 510 (proposed INN sotorasib).

Comparator

Docetaxel.

Primary question

How does progression-free survival compare between the randomized AMG 510 and docetaxel groups?

3. Trial Design

01
Randomize345 enrolled
02
Two armsAMG 510 vs docetaxel
03
FollowTime-to-event follow-up
04
Assess PFSProgression or death
05
CompareStratified log-rank
ARM 1

AMG 510

  • AMG 510 (drug)
  • Proposed INN: sotorasib
ARM 2

Docetaxel

  • Docetaxel (drug)

The registry describes the study as randomized, with a parallel design and no masking. The primary purpose is treatment. These features matter statistically because randomization establishes the intended basis for comparing the two treatment strategies, while the parallel structure means participants are followed within their randomized treatment groups rather than moving through a sequence of randomized treatment periods.

04 June 2020

Study start

The registry lists 04 June 2020 as the study start date.

02 August 2022

Primary analysis data cut-off

The registered PFS time frame extends to the primary analysis data cut-off date of 02 August 2022.

4. Endpoints

EndpointDefinition / assessmentTime frame
Progression-free Survival (PFS) PFS was defined as the time from randomization (baseline) until disease progression or death from any cause, whichever occurred first for all participants. Progression was based on blinded independent central review (BICR) of disease response per Response Evaluation Criteria in Solid Tumors version 1.1 (RECIST v1.1). Baseline up to primary analysis data cut-off date (02 August 2022); max time on study as of primary analysis data cut off was 24.3 months

The ClinicalTrials.gov record identifies PFS as the sole registered primary endpoint and classifies it as a time-to-event endpoint. The event can therefore occur in two ways: documented disease progression or death from any cause, with whichever occurs first defining the PFS event.

Endpoint interpretation: PFS is not a measurement of tumor shrinkage alone. A participant can contribute a PFS event because of progression, because of death before documented progression, or remain event-free until censoring under the prespecified analysis framework.

5. Statistical Methodology

Analysis population

The primary PFS analysis was measured in the Full Analysis Set (FAS), which included all randomized participants. The use of all randomized participants is consistent with the intention-to-treat principle identified in the analysis text posted on ClinicalTrials.gov for the trial.

Time-to-event analysis

PFS was analyzed as a time-to-event outcome. The registry reports the log-rank test as the statistical method and specifies that the P-value was calculated using a stratified log-rank test.

Primary PFS comparison
AMG 510 vs Docetaxel  →  stratified log-rank test  →  hazard ratio

The hazard ratio summarizes the relative instantaneous event rate between the randomized groups over the analyzed follow-up. The log-rank test addresses the statistical comparison of the time-to-event experience between groups.

Hazard ratio

The reported effect measure was the hazard ratio (HR). The estimate for AMG 510 relative to docetaxel was 0.663.

Interpretation of the reported HR
HR = 0.663  →  estimated event hazard ratio for AMG 510 relative to docetaxel

Because the reported HR is below 1.0, the registry interpretation is that AMG 510 had a lower average event rate and longer PFS relative to docetaxel under the analyzed time-to-event framework.

Confidence interval

The two-sided 95% confidence interval for the hazard ratio was 0.509 to 0.864. A confidence interval describes statistical uncertainty around the estimated treatment effect; it does not describe the range of PFS outcomes experienced by individual participants.

Hypothesis type

The registry data classifies the hypothesis type as Equivalence. The ClinicalTrials.gov record does not provide an equivalence margin, so an equivalence criterion cannot be reconstructed from this record alone. In particular, the reported HR and P-value should not be used to invent or infer a numerical equivalence margin.

Important distinction: an equivalence label in the registry's hypothesis field is not itself an equivalence margin. A formal equivalence interpretation requires the prespecified margin and the corresponding confidence-interval criterion. Those details are not included in the ClinicalTrials.gov record.

6. Results: Progression-Free Survival

The registry contains a formal statistical analysis for the primary endpoint. The analysis population was the Full Analysis Set, including all randomized participants. AMG 510 was compared with docetaxel using a log-rank method, with the P-value calculated using a stratified log-rank test.

Primary endpoint: progression-free survival

HR 0.663

95% CI: 0.509–0.864   ·   P = 0.002

Analysis population: Full Analysis Set, including all randomized participants.

Primary PFS resultReported value
ComparisonAMG 510 vs Docetaxel
Analysis populationFull Analysis Set (FAS)
EndpointProgression-free Survival (PFS)
Effect measureHazard ratio
Hazard ratio0.663
95% CI0.509–0.864
CI typeTwo-sided
P-value0.002
TestStratified log-rank test
Clinical Biostats interpretation

The HR of 0.663 means that, within the reported time-to-event analysis, the estimated instantaneous rate of experiencing a PFS event was about 66.3% for AMG 510 relative to docetaxel. Expressed as a simple relative-hazard interpretation, this corresponds to approximately a 33.7% lower estimated event hazard for AMG 510 relative to docetaxel.

The HR does not mean that 33.7% of participants had their disease progression prevented, that every participant experienced exactly a 33.7% reduction in risk, or that the absolute probability of progression or death was reduced by 33.7 percentage points. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.

The two-sided 95% CI of 0.509–0.864 expresses uncertainty around the estimated hazard ratio. The interval is entirely below 1.0, so the compatible effect estimates in this statistical framework are below the no-difference hazard-ratio value of 1.0. The interval does not indicate the range of individual patient outcomes.

The P = 0.002 is evidence against the null hypothesis represented by the specified statistical comparison; it is not a measure of the magnitude of the treatment effect. The size of the effect is communicated by the HR, while its statistical uncertainty is communicated by the confidence interval.

Because PFS is a censored time-to-event endpoint, the analysis also depends on how progression and censoring are defined and handled. The registry specifies BICR assessment of progression using RECIST v1.1. A hazard-ratio interpretation also relies on the appropriateness of the underlying proportional-hazards framework; the ClinicalTrials.gov record does not provide enough information to assess that assumption directly.

Educational note: a Kaplan-Meier curve is not reconstructed here from the single reported hazard ratio, confidence interval, and P-value. Valid reconstruction would require sufficiently detailed event and censoring information. The ClinicalTrials.gov record does not contain that underlying participant-level information.

7. Understanding the PFS Analysis

Why is PFS a time-to-event endpoint?

Each participant contributes not only whether an event occurred, but also when it occurred or, if no event was observed during available follow-up, how long the participant was followed without an event. This preserves more information than reducing every participant to a simple event/no-event indicator at a fixed time point.

Why use a log-rank test?

The log-rank test is designed for comparing survival or time-to-event distributions between groups. It incorporates the timing of events and accommodates right-censored observations. In this trial, the registry specifically reports a stratified log-rank test for the P-value.

Why report a hazard ratio as well as a P-value?

The two quantities answer different questions. The P-value addresses the strength of statistical evidence under the specified testing framework, whereas the hazard ratio describes the estimated relative event rate. Reporting the HR together with its confidence interval provides substantially more information about the estimated magnitude and precision than a P-value alone.

What does HR 0.663 mean?

It means the estimated hazard of a PFS event for AMG 510 relative to docetaxel was 0.663 under the reported analysis. A value below 1 indicates a lower estimated event hazard in the AMG 510 group. It should not be translated into an absolute probability without additional survival information.

What does the 95% CI of 0.509–0.864 mean?

The interval quantifies uncertainty around the estimated HR under the statistical model and sampling framework. It indicates that the estimate is not known exactly. It does not mean that 95% of individual patients have treatment effects between 0.509 and 0.864.

Why does the Full Analysis Set matter?

The FAS included all randomized participants. An analysis based on randomized assignment preserves the principal advantage created by randomization: baseline differences between treatment groups are addressed through the random allocation process rather than by selectively analyzing only participants who completed treatment.

8. Randomization and Intention-to-Treat Analysis

Randomization is the central design feature supporting the comparison of AMG 510 with docetaxel. Participants were assigned to treatment groups before the comparative outcome was analyzed. The registry-reported statistical analysis identifies the Full Analysis Set as including all randomized participants and separately identifies intention-to-treat analysis as an analysis concept.

What randomization accomplishes

Randomization creates the intended basis for a fair comparison by assigning participants to treatment groups through a randomized allocation process rather than according to observed prognosis or treatment response.

What ITT preserves

Analyzing participants according to randomized assignment helps preserve the treatment comparison created by randomization, even when participants' subsequent experiences differ.

Importantly, randomization does not guarantee numerically identical groups on every characteristic. Its statistical value comes from the probability mechanism governing assignment and the resulting basis for causal comparison.

9. Crossover and Treatment Switching

The ClinicalTrials.gov record reports serious adverse events for a third category described as Docetaxel, Then Switched to AMG 510, with 26 affected participants among 46 at risk. This indicates that treatment switching is part of the reported trial data and should be recognized when interpreting outcomes.

Reported treatment categorySerious adverse events
Docetaxel67 / 151
AMG 51091 / 169
Docetaxel, Then Switched to AMG 51026 / 46

Treatment switching creates an important distinction between the effect of being randomized to a treatment strategy and the effect of actually receiving a particular treatment over time. For a randomized efficacy analysis, maintaining the randomized comparison is generally important because switching can otherwise change the estimand being studied.

Interpretation caution: the ClinicalTrials.gov record establishes that a switched-treatment category was reported, but it does not provide enough information to calculate a switching-adjusted PFS or overall-survival effect. No such adjusted estimate is added here.

10. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment category. These safety counts should be interpreted separately from the primary PFS efficacy analysis. Safety summaries answer an exposure-related question—what serious adverse events were reported among participants at risk—rather than the time-to-event efficacy question addressed by the primary PFS analysis.

Safety categoryAffected / at risk
Docetaxel67 / 151
AMG 51091 / 169
Docetaxel, Then Switched to AMG 51026 / 46

The denominators differ across the reported categories, so these counts should not be interpreted as though they came from a single common denominator. The switched group is also a distinct treatment-history category and should not simply be combined with either randomized treatment arm without a prespecified analysis framework.

11. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS records the time until progression or death and can include censored observations. The log-rank test is specifically designed to compare time-to-event distributions between groups while incorporating the timing of observed events and the available follow-up for censored participants. CodeBreaK 200 reports a stratified log-rank test for the primary PFS comparison.

Why is the hazard ratio useful?

A hazard ratio compresses the relative difference in event rates over follow-up into a single effect measure. Here, the HR of 0.663 indicates a lower estimated event hazard for AMG 510 relative to docetaxel. It is more informative than a P-value alone because it describes the direction and magnitude of the estimated relative effect.

Why is a confidence interval needed around the HR?

A point estimate is only one estimate from the observed data. The 95% CI of 0.509–0.864 shows the statistical uncertainty around the HR of 0.663. A narrower interval would generally indicate greater precision; a wider interval would indicate less precision.

Why does the analysis use the Full Analysis Set?

The Full Analysis Set included all randomized participants. This maintains the randomized comparison and is closely aligned with the intention-to-treat principle identified in the analysis text. Restricting efficacy analysis only to participants who completed treatment could introduce post-randomization selection.

What does P = 0.002 tell us?

The P-value describes the strength of evidence against the null hypothesis under the specified statistical test and assumptions. It does not say that there is a 0.2% probability that the null hypothesis is true, and it does not measure clinical importance. Effect size and precision are better conveyed by the HR and its confidence interval.

Why does censoring matter?

Some participants may not experience progression or death before their available follow-up ends. Rather than treating these observations as if the event never occurred, time-to-event methods retain the information available up to the censoring point. The validity of the resulting estimates depends on the prespecified endpoint and censoring framework.

12. Equivalence as a Statistical Concept

The ClinicalTrials.gov record labels the hypothesis type as Equivalence. Equivalence trials are statistically different from conventional superiority trials because the objective is generally to establish that two strategies are sufficiently close within a prespecified clinically acceptable margin rather than simply to demonstrate a difference.

What would be required for a full equivalence reconstruction?
Estimated effect + 95% CI + prespecified equivalence margin

The ClinicalTrials.gov record provides the estimated HR and its two-sided 95% CI but does not provide a numerical equivalence margin. Therefore, this page reports the registry's hypothesis classification without assigning an unreported margin or constructing an equivalence conclusion beyond the registry-reported evidence.

This distinction is important because the presence of a P-value does not replace the prespecified margin in an equivalence framework. The statistical question must be tied to the clinically meaningful boundary established before the outcome data are evaluated.

13. What the Hazard Ratio Does — and Does Not — Mean

Statistical interpretation

A PFS HR of 0.663 indicates a lower estimated instantaneous rate of progression or death for AMG 510 relative to docetaxel in the reported analysis. The corresponding simple relative interpretation is approximately a 33.7% lower estimated event hazard.

It does not mean that 33.7% of participants avoided progression, that every participant received exactly a 33.7% benefit, or that the absolute probability of progression or death decreased by 33.7 percentage points.

Why the confidence interval matters

The two-sided 95% CI of 0.509–0.864 describes uncertainty around the estimated HR. It communicates precision that the point estimate alone cannot provide. It is not a patient-level prediction interval and should not be interpreted as the range of outcomes for individual participants.

Why the P-value should not be used as an effect-size measure

The P-value of 0.002 describes evidence under the specified hypothesis-testing framework. A smaller P-value does not automatically represent a larger treatment effect. The HR communicates the estimated relative effect, while the confidence interval communicates its statistical precision.

14. PFS, Censoring, and the Kaplan-Meier Framework

The registered PFS endpoint is inherently suited to Kaplan-Meier estimation because participants can have different lengths of follow-up and some may be censored before experiencing a PFS event. Kaplan-Meier methods estimate the probability of remaining event-free over time while retaining information from censored observations up to their censoring time.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at time ti, while ni is the number at risk immediately before that event time.

The ClinicalTrials.gov record does not report a Kaplan-Meier median PFS or time-specific survival probabilities. Those quantities therefore are not added to this page. The reported primary result is the hazard ratio with its confidence interval and P-value.

15. Analysis Population and Estimand

The primary PFS analysis was conducted in the Full Analysis Set, defined in the ClinicalTrials.gov record as all randomized participants. This makes the analysis fundamentally a randomized treatment comparison rather than a comparison restricted to participants who remained on assigned treatment.

ElementReported trial specification
OutcomeProgression-free Survival
PopulationFull Analysis Set
Population definitionAll randomized participants
Treatment comparisonAMG 510 vs Docetaxel
Analysis methodLog-rank
Specific testing frameworkStratified log-rank test
Effect measureHazard ratio

Once treatment switching occurs, the interpretation of the randomized comparison becomes particularly important. An intention-to-treat-style analysis asks about the consequences of assignment to the randomized strategy, whereas a treatment-received analysis asks a different question. The ClinicalTrials.gov record does not provide a separate causal estimand or switching-adjusted analysis, so this page does not construct one.

16. Limitations and Interpretation Issues

17. Why This Trial Matters Statistically

CodeBreaK 200 is a useful teaching example because its primary analysis brings together the central elements of randomized time-to-event analysis: randomization, an event definition involving progression or death, censoring, a Full Analysis Set, a stratified log-rank test, a hazard ratio, and a confidence interval.

ConceptHow it appears in CodeBreaK 200
RandomizationThe registry describes the allocation as randomized.
Parallel designThe trial uses a parallel design model with AMG 510 and docetaxel arms.
Intention-to-treat principleThe analysis text identifies intention-to-treat analysis and the FAS includes all randomized participants.
Time-to-event endpointPFS is defined from randomization until progression or death.
BICR assessmentProgression was based on blinded independent central review using RECIST v1.1.
Log-rank testThe reported primary analysis uses a log-rank method, with a stratified log-rank P-value.
Hazard ratioThe primary effect estimate is HR 0.663.
Confidence intervalThe two-sided 95% CI is 0.509–0.864.
P-valueThe reported P-value is 0.002.
Equivalence frameworkThe registry identifies the hypothesis type as Equivalence, but the ClinicalTrials.gov record does not report a numerical margin.
Treatment switchingThe registry reports serious adverse events for a Docetaxel, Then Switched to AMG 510 category.
Safety analysisSerious adverse events are reported by treatment or treatment-history category.

18. A Statistical Reading of the Primary Result

The most compact way to summarize the primary result is:

AMG 510 vs docetaxel

HR 0.663

95% CI 0.509–0.864   ·   P = 0.002

Three separate pieces of information should be read together. First, the direction of the estimate is below 1.0, indicating a lower estimated PFS event hazard for AMG 510 relative to docetaxel. Second, the confidence interval quantifies uncertainty around the point estimate. Third, the P-value describes statistical evidence under the specified stratified log-rank testing framework.

None of these quantities independently describes an absolute treatment benefit. To communicate absolute PFS benefit, one would ordinarily also want Kaplan-Meier estimates at clinically meaningful time points and/or median PFS, but those numerical results are not included in the ClinicalTrials.gov record and therefore are not reproduced here.

Interpretation principle: the strongest statistical reading of a clinical-trial result combines the design, analysis population, endpoint definition, effect estimate, uncertainty interval, and testing method. Reporting only the P-value would omit much of the information needed to understand the result.

19. Related Statistical Concepts

The CodeBreaK 200 primary analysis connects naturally to several foundational clinical-trial concepts:

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore the time-to-event methods, effect measures, confidence intervals, and trial-design concepts illustrated by CodeBreaK 200.

23. Record Summary

CodeBreaK 200 provides a compact example of randomized clinical-trial time-to-event analysis. The study enrolled 345 participants and compared AMG 510 with docetaxel in a randomized, parallel, unmasked phase 3 design. Its registered primary endpoint was progression-free survival, defined from randomization until disease progression or death from any cause, with progression assessed by blinded independent central review using RECIST v1.1.

The formal primary analysis used the Full Analysis Set, including all randomized participants, and a stratified log-rank test. The reported hazard ratio was 0.663, with a two-sided 95% confidence interval of 0.509–0.864 and a P-value of 0.002. The HR below 1.0 indicates a lower estimated PFS event hazard for AMG 510 relative to docetaxel; the confidence interval describes the statistical uncertainty around that estimate, while the P-value describes evidence under the specified testing framework.

The registry also identifies the hypothesis type as Equivalence, but the ClinicalTrials.gov record does not contain a numerical equivalence margin. Treatment switching is represented in the reported serious-adverse-event categories, including a Docetaxel, Then Switched to AMG 510 group. These features illustrate why a complete statistical interpretation must distinguish the randomized efficacy comparison from subsequent treatment history and must avoid inferring unreported equivalence criteria.

Clinical Biostats methodology: The purpose of this page is not merely to reproduce a trial result. It is to show how the endpoint definition, randomized analysis population, time-to-event method, hazard ratio, confidence interval, P-value, treatment switching, and statistical framework fit together—while keeping reported evidence distinct from educational interpretation.