This page separates reported trial results from statistical interpretation. Numerical results and trial facts are restricted to the ClinicalTrials.gov record. The official registry record is the authoritative source for the trial record: ClinicalTrials.gov NCT04303780.
1. Trial at a Glance
CodeBreaK 200 was a randomized, parallel-group, phase 3 trial comparing AMG 510 with docetaxel in advanced or metastatic NSCLC with KRAS p.G12C mutation. The registry reports one primary time-to-event endpoint: progression-free survival (PFS).
| Feature | CodeBreaK 200 |
|---|---|
| Trial | CodeBreaK 200 |
| NCT identifier | NCT04303780 |
| Phase | Phase 3 |
| Population | KRAS p.G12C-mutated / advanced metastatic NSCLC |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 345 |
| Interventions | AMG 510 and docetaxel |
| Primary endpoint | Progression-free survival (PFS) |
| Primary endpoint type | Time-to-event |
| Statistical method reported | Log-rank test |
| Effect measure | Hazard ratio |
| Registry hypothesis type | Equivalence |
| Lead sponsor | Amgen |
2. Clinical Question
The primary statistical question is whether the time from randomization to disease progression or death differs between AMG 510 and docetaxel in participants with KRAS p.G12C-mutated advanced or metastatic NSCLC.
Population
Participants with KRAS p.G12C-mutated advanced or metastatic non-small cell lung cancer (NSCLC).
Intervention
AMG 510 (proposed INN sotorasib).
Comparator
Docetaxel.
Primary question
How does progression-free survival compare between the randomized AMG 510 and docetaxel groups?
3. Trial Design
AMG 510
- AMG 510 (drug)
- Proposed INN: sotorasib
Docetaxel
- Docetaxel (drug)
The registry describes the study as randomized, with a parallel design and no masking. The primary purpose is treatment. These features matter statistically because randomization establishes the intended basis for comparing the two treatment strategies, while the parallel structure means participants are followed within their randomized treatment groups rather than moving through a sequence of randomized treatment periods.
Study start
The registry lists 04 June 2020 as the study start date.
Primary analysis data cut-off
The registered PFS time frame extends to the primary analysis data cut-off date of 02 August 2022.
4. Endpoints
| Endpoint | Definition / assessment | Time frame |
|---|---|---|
| Progression-free Survival (PFS) | PFS was defined as the time from randomization (baseline) until disease progression or death from any cause, whichever occurred first for all participants. Progression was based on blinded independent central review (BICR) of disease response per Response Evaluation Criteria in Solid Tumors version 1.1 (RECIST v1.1). | Baseline up to primary analysis data cut-off date (02 August 2022); max time on study as of primary analysis data cut off was 24.3 months |
The ClinicalTrials.gov record identifies PFS as the sole registered primary endpoint and classifies it as a time-to-event endpoint. The event can therefore occur in two ways: documented disease progression or death from any cause, with whichever occurs first defining the PFS event.
5. Statistical Methodology
Analysis population
The primary PFS analysis was measured in the Full Analysis Set (FAS), which included all randomized participants. The use of all randomized participants is consistent with the intention-to-treat principle identified in the analysis text posted on ClinicalTrials.gov for the trial.
Time-to-event analysis
PFS was analyzed as a time-to-event outcome. The registry reports the log-rank test as the statistical method and specifies that the P-value was calculated using a stratified log-rank test.
The hazard ratio summarizes the relative instantaneous event rate between the randomized groups over the analyzed follow-up. The log-rank test addresses the statistical comparison of the time-to-event experience between groups.
Hazard ratio
The reported effect measure was the hazard ratio (HR). The estimate for AMG 510 relative to docetaxel was 0.663.
Because the reported HR is below 1.0, the registry interpretation is that AMG 510 had a lower average event rate and longer PFS relative to docetaxel under the analyzed time-to-event framework.
Confidence interval
The two-sided 95% confidence interval for the hazard ratio was 0.509 to 0.864. A confidence interval describes statistical uncertainty around the estimated treatment effect; it does not describe the range of PFS outcomes experienced by individual participants.
Hypothesis type
The registry data classifies the hypothesis type as Equivalence. The ClinicalTrials.gov record does not provide an equivalence margin, so an equivalence criterion cannot be reconstructed from this record alone. In particular, the reported HR and P-value should not be used to invent or infer a numerical equivalence margin.
6. Results: Progression-Free Survival
The registry contains a formal statistical analysis for the primary endpoint. The analysis population was the Full Analysis Set, including all randomized participants. AMG 510 was compared with docetaxel using a log-rank method, with the P-value calculated using a stratified log-rank test.
Primary endpoint: progression-free survival
95% CI: 0.509–0.864 · P = 0.002
Analysis population: Full Analysis Set, including all randomized participants.
| Primary PFS result | Reported value |
|---|---|
| Comparison | AMG 510 vs Docetaxel |
| Analysis population | Full Analysis Set (FAS) |
| Endpoint | Progression-free Survival (PFS) |
| Effect measure | Hazard ratio |
| Hazard ratio | 0.663 |
| 95% CI | 0.509–0.864 |
| CI type | Two-sided |
| P-value | 0.002 |
| Test | Stratified log-rank test |
The HR of 0.663 means that, within the reported time-to-event analysis, the estimated instantaneous rate of experiencing a PFS event was about 66.3% for AMG 510 relative to docetaxel. Expressed as a simple relative-hazard interpretation, this corresponds to approximately a 33.7% lower estimated event hazard for AMG 510 relative to docetaxel.
The HR does not mean that 33.7% of participants had their disease progression prevented, that every participant experienced exactly a 33.7% reduction in risk, or that the absolute probability of progression or death was reduced by 33.7 percentage points. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.
The two-sided 95% CI of 0.509–0.864 expresses uncertainty around the estimated hazard ratio. The interval is entirely below 1.0, so the compatible effect estimates in this statistical framework are below the no-difference hazard-ratio value of 1.0. The interval does not indicate the range of individual patient outcomes.
The P = 0.002 is evidence against the null hypothesis represented by the specified statistical comparison; it is not a measure of the magnitude of the treatment effect. The size of the effect is communicated by the HR, while its statistical uncertainty is communicated by the confidence interval.
Because PFS is a censored time-to-event endpoint, the analysis also depends on how progression and censoring are defined and handled. The registry specifies BICR assessment of progression using RECIST v1.1. A hazard-ratio interpretation also relies on the appropriateness of the underlying proportional-hazards framework; the ClinicalTrials.gov record does not provide enough information to assess that assumption directly.
7. Understanding the PFS Analysis
Why is PFS a time-to-event endpoint?
Each participant contributes not only whether an event occurred, but also when it occurred or, if no event was observed during available follow-up, how long the participant was followed without an event. This preserves more information than reducing every participant to a simple event/no-event indicator at a fixed time point.
Why use a log-rank test?
The log-rank test is designed for comparing survival or time-to-event distributions between groups. It incorporates the timing of events and accommodates right-censored observations. In this trial, the registry specifically reports a stratified log-rank test for the P-value.
Why report a hazard ratio as well as a P-value?
The two quantities answer different questions. The P-value addresses the strength of statistical evidence under the specified testing framework, whereas the hazard ratio describes the estimated relative event rate. Reporting the HR together with its confidence interval provides substantially more information about the estimated magnitude and precision than a P-value alone.
What does HR 0.663 mean?
It means the estimated hazard of a PFS event for AMG 510 relative to docetaxel was 0.663 under the reported analysis. A value below 1 indicates a lower estimated event hazard in the AMG 510 group. It should not be translated into an absolute probability without additional survival information.
What does the 95% CI of 0.509–0.864 mean?
The interval quantifies uncertainty around the estimated HR under the statistical model and sampling framework. It indicates that the estimate is not known exactly. It does not mean that 95% of individual patients have treatment effects between 0.509 and 0.864.
Why does the Full Analysis Set matter?
The FAS included all randomized participants. An analysis based on randomized assignment preserves the principal advantage created by randomization: baseline differences between treatment groups are addressed through the random allocation process rather than by selectively analyzing only participants who completed treatment.
8. Randomization and Intention-to-Treat Analysis
Randomization is the central design feature supporting the comparison of AMG 510 with docetaxel. Participants were assigned to treatment groups before the comparative outcome was analyzed. The registry-reported statistical analysis identifies the Full Analysis Set as including all randomized participants and separately identifies intention-to-treat analysis as an analysis concept.
What randomization accomplishes
Randomization creates the intended basis for a fair comparison by assigning participants to treatment groups through a randomized allocation process rather than according to observed prognosis or treatment response.
What ITT preserves
Analyzing participants according to randomized assignment helps preserve the treatment comparison created by randomization, even when participants' subsequent experiences differ.
Importantly, randomization does not guarantee numerically identical groups on every characteristic. Its statistical value comes from the probability mechanism governing assignment and the resulting basis for causal comparison.
9. Crossover and Treatment Switching
The ClinicalTrials.gov record reports serious adverse events for a third category described as Docetaxel, Then Switched to AMG 510, with 26 affected participants among 46 at risk. This indicates that treatment switching is part of the reported trial data and should be recognized when interpreting outcomes.
| Reported treatment category | Serious adverse events |
|---|---|
| Docetaxel | 67 / 151 |
| AMG 510 | 91 / 169 |
| Docetaxel, Then Switched to AMG 510 | 26 / 46 |
Treatment switching creates an important distinction between the effect of being randomized to a treatment strategy and the effect of actually receiving a particular treatment over time. For a randomized efficacy analysis, maintaining the randomized comparison is generally important because switching can otherwise change the estimand being studied.
10. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment category. These safety counts should be interpreted separately from the primary PFS efficacy analysis. Safety summaries answer an exposure-related question—what serious adverse events were reported among participants at risk—rather than the time-to-event efficacy question addressed by the primary PFS analysis.
| Safety category | Affected / at risk |
|---|---|
| Docetaxel | 67 / 151 |
| AMG 510 | 91 / 169 |
| Docetaxel, Then Switched to AMG 510 | 26 / 46 |
The denominators differ across the reported categories, so these counts should not be interpreted as though they came from a single common denominator. The switched group is also a distinct treatment-history category and should not simply be combined with either randomized treatment arm without a prespecified analysis framework.
11. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS records the time until progression or death and can include censored observations. The log-rank test is specifically designed to compare time-to-event distributions between groups while incorporating the timing of observed events and the available follow-up for censored participants. CodeBreaK 200 reports a stratified log-rank test for the primary PFS comparison.
Why is the hazard ratio useful?
A hazard ratio compresses the relative difference in event rates over follow-up into a single effect measure. Here, the HR of 0.663 indicates a lower estimated event hazard for AMG 510 relative to docetaxel. It is more informative than a P-value alone because it describes the direction and magnitude of the estimated relative effect.
Why is a confidence interval needed around the HR?
A point estimate is only one estimate from the observed data. The 95% CI of 0.509–0.864 shows the statistical uncertainty around the HR of 0.663. A narrower interval would generally indicate greater precision; a wider interval would indicate less precision.
Why does the analysis use the Full Analysis Set?
The Full Analysis Set included all randomized participants. This maintains the randomized comparison and is closely aligned with the intention-to-treat principle identified in the analysis text. Restricting efficacy analysis only to participants who completed treatment could introduce post-randomization selection.
What does P = 0.002 tell us?
The P-value describes the strength of evidence against the null hypothesis under the specified statistical test and assumptions. It does not say that there is a 0.2% probability that the null hypothesis is true, and it does not measure clinical importance. Effect size and precision are better conveyed by the HR and its confidence interval.
Why does censoring matter?
Some participants may not experience progression or death before their available follow-up ends. Rather than treating these observations as if the event never occurred, time-to-event methods retain the information available up to the censoring point. The validity of the resulting estimates depends on the prespecified endpoint and censoring framework.
12. Equivalence as a Statistical Concept
The ClinicalTrials.gov record labels the hypothesis type as Equivalence. Equivalence trials are statistically different from conventional superiority trials because the objective is generally to establish that two strategies are sufficiently close within a prespecified clinically acceptable margin rather than simply to demonstrate a difference.
The ClinicalTrials.gov record provides the estimated HR and its two-sided 95% CI but does not provide a numerical equivalence margin. Therefore, this page reports the registry's hypothesis classification without assigning an unreported margin or constructing an equivalence conclusion beyond the registry-reported evidence.
This distinction is important because the presence of a P-value does not replace the prespecified margin in an equivalence framework. The statistical question must be tied to the clinically meaningful boundary established before the outcome data are evaluated.
13. What the Hazard Ratio Does — and Does Not — Mean
A PFS HR of 0.663 indicates a lower estimated instantaneous rate of progression or death for AMG 510 relative to docetaxel in the reported analysis. The corresponding simple relative interpretation is approximately a 33.7% lower estimated event hazard.
It does not mean that 33.7% of participants avoided progression, that every participant received exactly a 33.7% benefit, or that the absolute probability of progression or death decreased by 33.7 percentage points.
The two-sided 95% CI of 0.509–0.864 describes uncertainty around the estimated HR. It communicates precision that the point estimate alone cannot provide. It is not a patient-level prediction interval and should not be interpreted as the range of outcomes for individual participants.
The P-value of 0.002 describes evidence under the specified hypothesis-testing framework. A smaller P-value does not automatically represent a larger treatment effect. The HR communicates the estimated relative effect, while the confidence interval communicates its statistical precision.
14. PFS, Censoring, and the Kaplan-Meier Framework
The registered PFS endpoint is inherently suited to Kaplan-Meier estimation because participants can have different lengths of follow-up and some may be censored before experiencing a PFS event. Kaplan-Meier methods estimate the probability of remaining event-free over time while retaining information from censored observations up to their censoring time.
Here, di represents events at time ti, while ni is the number at risk immediately before that event time.
The ClinicalTrials.gov record does not report a Kaplan-Meier median PFS or time-specific survival probabilities. Those quantities therefore are not added to this page. The reported primary result is the hazard ratio with its confidence interval and P-value.
15. Analysis Population and Estimand
The primary PFS analysis was conducted in the Full Analysis Set, defined in the ClinicalTrials.gov record as all randomized participants. This makes the analysis fundamentally a randomized treatment comparison rather than a comparison restricted to participants who remained on assigned treatment.
| Element | Reported trial specification |
|---|---|
| Outcome | Progression-free Survival |
| Population | Full Analysis Set |
| Population definition | All randomized participants |
| Treatment comparison | AMG 510 vs Docetaxel |
| Analysis method | Log-rank |
| Specific testing framework | Stratified log-rank test |
| Effect measure | Hazard ratio |
Once treatment switching occurs, the interpretation of the randomized comparison becomes particularly important. An intention-to-treat-style analysis asks about the consequences of assignment to the randomized strategy, whereas a treatment-received analysis asks a different question. The ClinicalTrials.gov record does not provide a separate causal estimand or switching-adjusted analysis, so this page does not construct one.
16. Limitations and Interpretation Issues
- Single reported primary endpoint: the ClinicalTrials.gov record identifies PFS as the registered primary endpoint. No additional primary efficacy endpoint is added.
- No median PFS reported in the ClinicalTrials.gov record: the ClinicalTrials.gov record contains the HR, confidence interval, and P-value but not a median PFS estimate, so no median is reported.
- No time-specific PFS rates: the ClinicalTrials.gov record does not provide Kaplan-Meier PFS probabilities at particular time points.
- Hazard-ratio assumptions: a hazard ratio is a model-based relative time-to-event measure. The ClinicalTrials.gov record does not provide enough information to assess whether proportional hazards held throughout follow-up.
- Censoring: PFS analysis depends on appropriate handling of participants who do not experience progression or death during observed follow-up. The endpoint definition alone does not provide all details needed to independently audit the censoring mechanism.
- Treatment switching: the trial data reports a Docetaxel, Then Switched to AMG 510 category. Switching can make the distinction between randomized-treatment and treatment-received effects important.
- Equivalence classification: the registry identifies the hypothesis type as Equivalence, but the ClinicalTrials.gov record does not provide the numerical equivalence margin. A formal margin-based equivalence reconstruction is therefore not possible from this record.
- Safety denominators: serious adverse events are reported using different affected/at-risk denominators for the treatment-history categories. These counts should not be treated as directly comparable percentages without the appropriate analysis population and exposure definition.
- Limited reported results: the ClinicalTrials.gov record contains one formal statistical analysis. Subgroup estimates, additional efficacy endpoints, and other analyses are not added when they are absent from the ClinicalTrials.gov record.
17. Why This Trial Matters Statistically
CodeBreaK 200 is a useful teaching example because its primary analysis brings together the central elements of randomized time-to-event analysis: randomization, an event definition involving progression or death, censoring, a Full Analysis Set, a stratified log-rank test, a hazard ratio, and a confidence interval.
| Concept | How it appears in CodeBreaK 200 |
|---|---|
| Randomization | The registry describes the allocation as randomized. |
| Parallel design | The trial uses a parallel design model with AMG 510 and docetaxel arms. |
| Intention-to-treat principle | The analysis text identifies intention-to-treat analysis and the FAS includes all randomized participants. |
| Time-to-event endpoint | PFS is defined from randomization until progression or death. |
| BICR assessment | Progression was based on blinded independent central review using RECIST v1.1. |
| Log-rank test | The reported primary analysis uses a log-rank method, with a stratified log-rank P-value. |
| Hazard ratio | The primary effect estimate is HR 0.663. |
| Confidence interval | The two-sided 95% CI is 0.509–0.864. |
| P-value | The reported P-value is 0.002. |
| Equivalence framework | The registry identifies the hypothesis type as Equivalence, but the ClinicalTrials.gov record does not report a numerical margin. |
| Treatment switching | The registry reports serious adverse events for a Docetaxel, Then Switched to AMG 510 category. |
| Safety analysis | Serious adverse events are reported by treatment or treatment-history category. |
18. A Statistical Reading of the Primary Result
The most compact way to summarize the primary result is:
AMG 510 vs docetaxel
95% CI 0.509–0.864 · P = 0.002
Three separate pieces of information should be read together. First, the direction of the estimate is below 1.0, indicating a lower estimated PFS event hazard for AMG 510 relative to docetaxel. Second, the confidence interval quantifies uncertainty around the point estimate. Third, the P-value describes statistical evidence under the specified stratified log-rank testing framework.
None of these quantities independently describes an absolute treatment benefit. To communicate absolute PFS benefit, one would ordinarily also want Kaplan-Meier estimates at clinically meaningful time points and/or median PFS, but those numerical results are not included in the ClinicalTrials.gov record and therefore are not reproduced here.
19. Related Statistical Concepts
The CodeBreaK 200 primary analysis connects naturally to several foundational clinical-trial concepts:
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Calculators
22. Sources
- ClinicalTrials.gov: CodeBreaK 200, NCT04303780.
- Linked publication: PubMed PMID 36764316.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to explore the time-to-event methods, effect measures, confidence intervals, and trial-design concepts illustrated by CodeBreaK 200.
23. Record Summary
CodeBreaK 200 provides a compact example of randomized clinical-trial time-to-event analysis. The study enrolled 345 participants and compared AMG 510 with docetaxel in a randomized, parallel, unmasked phase 3 design. Its registered primary endpoint was progression-free survival, defined from randomization until disease progression or death from any cause, with progression assessed by blinded independent central review using RECIST v1.1.
The formal primary analysis used the Full Analysis Set, including all randomized participants, and a stratified log-rank test. The reported hazard ratio was 0.663, with a two-sided 95% confidence interval of 0.509–0.864 and a P-value of 0.002. The HR below 1.0 indicates a lower estimated PFS event hazard for AMG 510 relative to docetaxel; the confidence interval describes the statistical uncertainty around that estimate, while the P-value describes evidence under the specified testing framework.
The registry also identifies the hypothesis type as Equivalence, but the ClinicalTrials.gov record does not contain a numerical equivalence margin. Treatment switching is represented in the reported serious-adverse-event categories, including a Docetaxel, Then Switched to AMG 510 group. These features illustrate why a complete statistical interpretation must distinguish the randomized efficacy comparison from subsequent treatment history and must avoid inferring unreported equivalence criteria.