Introduction
The 3+3 design is one of the best-known dose-escalation designs used in Phase I clinical trials.
Its purpose is fundamentally different from the purpose of a conventional Phase II efficacy design.
In an early Phase I study, investigators are generally trying to determine how the risk of treatment-related toxicity changes as the dose increases. The study may also characterize pharmacokinetics, pharmacodynamics, and preliminary safety, but the dose-escalation component is primarily concerned with finding a dose that can be taken forward safely.
The traditional 3+3 design approaches this problem using small sequential cohorts.
Three patients are initially treated at a dose. Depending on the number of dose-limiting toxicities, the trial either escalates to a higher dose, expands the current dose to six patients, or stops escalation.
What Is the Goal of a Phase I Dose-Escalation Study?
The central question in dose escalation is:
How high can the dose be increased while maintaining an acceptable level of treatment-related toxicity?
A traditional Phase I oncology study often defines a target level of toxicity and attempts to identify a dose near that target.
One commonly used concept is the maximum tolerated dose, or MTD.
A simplified definition is:
However, it is important to recognize that the precise definition of MTD is protocol-specific.
The 3+3 rules provide an operational method for selecting the dose to take forward. They should not be interpreted as a universally valid statistical definition of the biologically optimal dose.
The Dose-Toxicity Relationship
Let:
where \(d\) denotes the administered dose.
In many drug-development settings, investigators expect the probability of toxicity to increase as dose increases.
Conceptually:
The true dose-toxicity curve is unknown when the Phase I study begins.
The trial therefore observes toxicity at one dose and uses that information to determine whether it is reasonable to move to the next dose.
What Is a Dose-Limiting Toxicity?
The central outcome in a traditional 3+3 design is the dose-limiting toxicity, commonly abbreviated DLT.
A DLT is a prespecified treatment-related toxicity that is considered serious enough to prevent further escalation at that dose.
The exact definition depends on the protocol and therapeutic area.
For example, a protocol might define DLTs using:
- Severity grade of an adverse event
- Duration of the toxicity
- Relationship to study treatment
- Requirement for medical intervention
- Delay in treatment administration
- Failure to recover within a specified period
- Specific laboratory abnormalities
In oncology studies, DLT assessment is often performed during a prespecified DLT evaluation window, such as the first treatment cycle.
Why Are DLTs Defined Before the Trial?
The dose-escalation rules depend directly on whether an observed event counts as a DLT.
Therefore, the DLT definition should be established before dose escalation begins.
A protocol should generally specify:
- Which toxicities qualify
- The severity thresholds
- The attribution criteria
- The observation window
- How recurrent toxicities are handled
- How treatment interruptions affect DLT assessment
- How missing or unevaluable patients are handled
- Whether certain expected toxicities are excluded
Without a clear DLT definition, the apparent simplicity of the 3+3 algorithm can become misleading.
The Basic 3+3 Structure
The traditional design starts with a cohort of three patients at the starting dose.
The basic decision rules are:
| DLTs Among 3 Patients | Action |
|---|---|
| 0 DLTs | Escalate to the next dose |
| 1 DLT | Expand the cohort to 6 patients |
| 2 or 3 DLTs | Stop escalation |
If one DLT is observed among the initial three patients, three additional patients are treated at the same dose.
The six-patient cohort then follows a second set of rules.
| Total DLTs Among 6 | Action |
|---|---|
| 1 or fewer DLTs | Escalate to the next dose |
| 2 or more DLTs | Stop escalation |
The 3+3 Algorithm as a Flowchart
A Simple Example of Escalation
Suppose the study has five planned dose levels:
| Dose Level | Dose |
|---|---|
| 1 | 10 mg |
| 2 | 20 mg |
| 3 | 40 mg |
| 4 | 80 mg |
| 5 | 120 mg |
The trial begins at 10 mg.
Three patients receive 10 mg.
Suppose none experiences a DLT.
The next cohort therefore receives 20 mg.
If the 20 mg cohort also has zero DLTs, the trial proceeds to 40 mg.
The process continues until a dose produces sufficient evidence of excessive toxicity.
Example 1: Zero DLTs
Consider the following sequence:
| Dose | Patients | DLTs | Decision |
|---|---|---|---|
| 10 mg | 3 | 0 | Escalate |
| 20 mg | 3 | 0 | Escalate |
| 40 mg | 3 | 0 | Escalate |
| 80 mg | 3 | 0 | Escalate |
The design has observed no DLTs at any of the first four dose levels.
The next cohort can therefore receive 120 mg, assuming 120 mg is the next prespecified dose level.
Example 2: One DLT Among Three
Suppose the 40 mg cohort produces:
The trial does not immediately escalate.
Instead, three additional patients are treated at 40 mg.
Suppose none of those three additional patients experiences a DLT.
The complete cohort is then:
Because there is only one DLT among six patients, the design escalates to the next dose.
| Dose | Initial Cohort | Expanded Cohort | Total DLTs | Decision |
|---|---|---|---|---|
| 40 mg | 1/3 | 0/3 | 1/6 | Escalate |
Example 3: One DLT Among Six
Now suppose the additional three patients include one DLT.
The total at the dose becomes:
The trial does not escalate.
The dose is considered too toxic for continued escalation under the traditional 3+3 rule.
Example 4: Two DLTs Among the First Three
Suppose a dose produces:
There is no need to treat three additional patients at that dose under the traditional rule.
The escalation stops immediately.
This reflects strong evidence, within the rule-based framework, that the current dose may be too toxic to justify further escalation.
| Result | Action |
|---|---|
| 0/3 | Escalate |
| 1/3 | Expand to 6 |
| 2/3 | Stop escalation |
| 3/3 | Stop escalation |
The Six-Patient Rule
The six-patient rule is one of the most important parts of the design.
If the first three patients produce exactly one DLT, three additional patients are treated at the same dose.
The combined result determines whether escalation can continue.
| Total DLTs Among 6 | Interpretation | Action |
|---|---|---|
| 0/6 | No observed DLTs | Escalate |
| 1/6 | One observed DLT | Escalate |
| 2/6 | Two observed DLTs | Stop escalation |
| 3/6 | Three observed DLTs | Stop escalation |
| 4/6 | Four observed DLTs | Stop escalation |
| 5/6 | Five observed DLTs | Stop escalation |
| 6/6 | Six observed DLTs | Stop escalation |
What Happens When Escalation Stops?
Suppose the study reaches a dose at which the 3+3 rules indicate excessive toxicity.
The next step depends on the specific protocol.
A common interpretation is that the previous lower dose is selected as the recommended dose or MTD candidate, because it was the highest dose that met the prespecified 3+3 tolerability criterion.
However, the precise terminology matters.
A Complete Worked Example
Consider a hypothetical oncology dose-escalation study with six planned dose levels.
| Dose Level | Dose |
|---|---|
| 1 | 10 mg |
| 2 | 20 mg |
| 3 | 40 mg |
| 4 | 80 mg |
| 5 | 120 mg |
| 6 | 160 mg |
Suppose the observed DLT outcomes are:
| Dose | Patients | DLTs | 3+3 Decision |
|---|---|---|---|
| 10 mg | 3 | 0 | Escalate |
| 20 mg | 3 | 0 | Escalate |
| 40 mg | 3 | 1 | Expand |
| 40 mg | 3 additional | 0 | Escalate |
| 80 mg | 3 | 0 | Escalate |
| 120 mg | 3 | 1 | Expand |
| 120 mg | 3 additional | 1 | Stop escalation |
Let's walk through the study.
Step 1: 10 mg
The first three patients receive 10 mg.
Because there are zero DLTs, the trial escalates to 20 mg.
Step 2: 20 mg
Three patients receive 20 mg.
Again, the trial escalates.
Step 3: 40 mg
Three patients receive 40 mg.
One patient experiences a DLT.
The study therefore expands the 40 mg cohort.
Step 4: Complete the 40 mg Cohort
Three additional patients receive 40 mg.
None experiences a DLT.
The complete experience at 40 mg is:
The study therefore escalates to 80 mg.
Step 5: 80 mg
Three patients receive 80 mg.
No DLTs are observed:
The study therefore escalates to 120 mg.
Step 6: 120 mg
Three patients receive 120 mg.
One patient experiences a DLT:
The cohort expands to six patients.
Step 7: Complete the 120 mg Cohort
Three additional patients receive 120 mg.
One additional DLT occurs.
The total is therefore:
Under the traditional 3+3 rule, escalation stops.
The previous dose, 80 mg, is therefore the dose that met the traditional tolerability criterion immediately below the dose at which escalation stopped.
The Entire Trial in One Diagram
Why Is the Design Called "3+3"?
The name comes directly from its cohort structure.
The study begins with 3 patients at a dose.
If exactly one DLT is observed, 3 additional patients are treated at that same dose.
Thus:
The notation is therefore a shorthand for the cohort-expansion mechanism.
The Design Is Not Always Exactly Three Patients
The phrase "3+3" can create the mistaken impression that every dose must have exactly three or six patients.
That is not necessarily the case in actual clinical operations.
The protocol may contain provisions for:
- Patients who are not evaluable for DLT
- Replacement patients
- Enrollment holds
- Sentinel dosing
- Staggered administration
- Safety review between cohorts
- Additional patients at a selected dose
These operational provisions must be defined separately from the core 3+3 decision rules.
What Is a DLT Evaluation Window?
The 3+3 algorithm requires investigators to determine whether each patient has experienced a DLT before making the escalation decision.
Therefore, the protocol generally specifies an observation period.
For example, a study might define the DLT window as the first treatment cycle.
If the cycle is 28 days, the investigator may need to wait until the relevant patients have completed the DLT evaluation period before escalating.
This creates an important operational distinction:
A mathematically simple dose-escalation design can therefore still require substantial calendar time.
Sentinel Patients
Some early-phase protocols use sentinel dosing, particularly when there are important safety uncertainties.
For example, the first patient at a new dose may be treated and observed for a predefined period before the remaining patients in the cohort are treated.
This is an operational safety measure rather than a fundamental component of the traditional 3+3 statistical rule.
What Does "MTD" Actually Mean?
The term maximum tolerated dose is frequently used in Phase I development, but it deserves careful interpretation.
Conceptually, suppose there is a target DLT probability:
A modern dose-escalation framework might explicitly model the probability of DLT as a function of dose and attempt to identify a dose whose toxicity probability is close to \(\phi\).
The traditional 3+3 design does not do this explicitly.
Instead, its rules classify observed cohorts into escalation or stopping regions.
The Target Toxicity Rate
To understand modern dose-escalation methods, it is useful to introduce a target toxicity probability.
Let:
for example.
This means the investigator may consider a dose with approximately 25% DLT probability to be near the desired toxicity level.
The traditional 3+3 design does not explicitly state:
and then estimate which dose satisfies that condition.
Instead, the 3+3 rules implicitly impose a particular decision behavior.
The Statistical Meaning of 0/3
One of the most important insights into the 3+3 design comes from recognizing how little information 0 DLTs among three patients actually provide.
Suppose the true DLT probability at a dose is:
The probability of observing zero DLTs among three patients is:
Therefore:
So even when the true DLT probability is 30%, there is approximately a 34.3% probability of observing zero DLTs in three patients.
This illustrates why: 0/3 does not mean the dose has a low toxicity probability.
The Statistical Meaning of 0/3 at a Higher Toxicity Rate
Suppose the true DLT probability is 50%.
The probability of seeing zero DLTs among three patients is:
Thus, even when half of patients would experience a DLT in the long run, there is a 12.5% chance of observing 0/3 in a particular cohort.
The 3+3 design will escalate after that 0/3 result.
Confidence Bounds After 0/3
The same issue can be expressed using an exact confidence bound.
If zero DLTs are observed among \(n\) patients, the approximate one-sided upper 95% confidence bound for the DLT probability is obtained from:
so:
For \(n=3\):
Thus, observing 0 DLTs in only three patients is compatible, at this confidence level, with a DLT probability as high as roughly 63%.
This does not mean that the true probability is 63%. It illustrates the enormous uncertainty associated with such a small sample.
Why the 3+3 Design Can Escalate Quickly
Consider a sequence in which every cohort has zero DLTs.
The design may progress:
The algorithm is deliberately simple.
But simplicity comes at the cost of not explicitly estimating the probability of toxicity at untested or partially tested doses.
Why the 3+3 Design Is Easy to Implement
One reason the design became so widely used is that its rules are easy for a clinical team to communicate.
At every dose, the question is essentially:
- How many patients were evaluated?
- How many DLTs occurred?
- Does that count satisfy the escalation rule?
No complex statistical model is required during the trial.
The clinical team can therefore implement the design using a simple decision table.
Decision Table for the Traditional 3+3 Design
| Patients at Dose | DLTs | Action |
|---|---|---|
| 3 | 0 | Escalate |
| 3 | 1 | Add 3 patients at same dose |
| 3 | 2 | Stop escalation |
| 3 | 3 | Stop escalation |
| 6 | 0 | Escalate |
| 6 | 1 | Escalate |
| 6 | 2+ | Stop escalation |
What Happens at the First Dose?
The starting dose is usually selected before the 3+3 algorithm begins.
It is not generally the role of the 3+3 design itself to determine the starting dose.
Starting-dose selection may incorporate:
- Nonclinical toxicology
- Pharmacology
- Exposure margins
- Human equivalent dose considerations
- Prior clinical experience
- Mechanism of action
- Pharmacokinetic modeling
- Previous studies
Once the starting dose and dose levels have been selected, the 3+3 algorithm determines how observed DLTs influence escalation.
Dose Levels Must Be Prespecified
Suppose the planned dose levels are:
The 3+3 design does not normally mean:
"If 10 mg is safe, choose any dose above 10 mg."
Rather, the protocol specifies the candidate dose sequence before enrollment.
The next dose is then selected according to the prespecified escalation plan.
Dose Increments
The size of the dose increments is a separate design consideration.
A sequence might use:
The relative increments become smaller at higher doses in this example.
Alternatively, the study may use fixed percentage increments, fixed absolute increments, or another scientifically justified sequence.
The 3+3 algorithm itself does not determine the dose spacing.
Intrapatient Dose Escalation
Traditional 3+3 designs generally focus on escalation between cohorts rather than routinely escalating individual patients within a cohort.
For example, a patient who starts at 20 mg would not necessarily be increased to 40 mg simply because the 20 mg cohort has had no DLTs.
The protocol should explicitly state whether intrapatient dose escalation is allowed.
Interpatient Escalation
The classic 3+3 approach is primarily an interpatient dose-escalation design.
That means:
is generally assigned to a new cohort after the current dose has satisfied the escalation criteria.
What Is the Recommended Phase II Dose?
The MTD and recommended Phase II dose are not necessarily identical.
The recommended Phase II dose may incorporate:
- DLTs
- Non-DLT toxicity
- Chronic or cumulative toxicity
- Pharmacokinetics
- Pharmacodynamics
- Exposure-response relationships
- Preliminary efficacy
- Target engagement
- Feasibility of chronic administration
Therefore, a dose that satisfies a traditional MTD rule may not automatically be the final development dose.
Why 3+3 Is Not an Efficacy Design
A common misunderstanding is that the dose with the best response rate should be selected by the 3+3 algorithm.
That is not what the traditional design does.
The primary escalation signal is DLT occurrence.
A dose may have:
- Excellent preliminary activity with no DLTs
- No observed activity with no DLTs
- Excellent activity with unacceptable toxicity
- Modest activity with acceptable toxicity
The 3+3 toxicity rules do not by themselves resolve these broader benefit-risk questions.
A Second Worked Example
Suppose the dose sequence is:
The observed outcomes are:
| Dose | DLTs | Action |
|---|---|---|
| 5 mg | 0/3 | Escalate |
| 10 mg | 0/3 | Escalate |
| 20 mg | 1/3 | Expand |
| 20 mg | 1/6 total | Escalate |
| 40 mg | 2/3 | Stop |
The study therefore stops escalation at 40 mg.
The preceding dose, 20 mg, is the dose that satisfied the traditional tolerability criterion immediately below the dose that caused the stopping signal.
Why 2/3 Is Such a Strong Signal
Suppose the target toxicity probability is around 20%.
Under a simple binomial model, the probability of observing at least two DLTs among three patients is:
For \(p=0.20\):
Thus, even at a 20% underlying DLT probability, there is about a 10.4% chance of observing 2 or more DLTs among three patients.
This illustrates both the conservatism and the variability of small-cohort decision rules.
The Binomial Distribution Behind 3+3
At a particular dose, let:
where:
- \(X\) = number of DLTs
- \(n\) = number of evaluable patients
- \(p\) = true probability of DLT at that dose
The probability of observing exactly \(x\) DLTs is:
The 3+3 rules effectively divide the possible observed DLT counts into decision regions.
Decision Regions After Three Patients
After six patients following an initial 1/3 result:
Probability of Escalating After 3 Patients
If the true DLT probability is \(p\), the probability of immediately escalating after three patients is:
For example, if \(p=0.10\):
Thus, when the true DLT probability is 10%, approximately 72.9% of three-patient cohorts will produce zero DLTs and therefore trigger immediate escalation.
Probability of Expanding the Cohort
The probability of seeing exactly one DLT among three patients is:
or:
At \(p=0.10\):
So there is approximately a 24.3% chance of cohort expansion at a dose with a 10% DLT probability.
Probability of Stopping After Three Patients
The probability of observing at least two DLTs among three patients is:
At \(p=0.10\):
Thus, at a true DLT probability of 10%, approximately 2.8% of three-patient cohorts will stop immediately because of at least two DLTs.
A Useful Probability Table
The following table illustrates how the first-stage 3+3 decision behaves as the underlying DLT probability changes.
| True DLT Probability | P(0/3) | P(1/3) | P(≥2/3) |
|---|---|---|---|
| 5% | 85.7% | 13.5% | 0.7% |
| 10% | 72.9% | 24.3% | 2.8% |
| 20% | 51.2% | 38.4% | 10.4% |
| 30% | 34.3% | 44.1% | 21.6% |
| 40% | 21.6% | 43.2% | 35.2% |
| 50% | 12.5% | 37.5% | 50.0% |
This table reveals an important feature of the design.
As the underlying DLT probability rises, the probability of observing a stopping signal increases.
However, the relationship is not deterministic.
Even a dose with a relatively high true toxicity probability can produce a small number of DLTs by chance.
The 3+3 Design Is a Decision Rule, Not a Dose-Toxicity Model
This distinction is central to understanding the design.
A dose-toxicity model attempts to estimate:
and potentially predict \(p(d)\) at doses that have not yet been tested.
The traditional 3+3 design instead asks:
Given the number of DLTs observed in this small cohort, should the trial escalate, expand, or stop?
That is a simpler problem.
Why Modern Dose-Escalation Designs Exist
The limitations of small rule-based cohorts motivated the development and adoption of alternative dose-escalation approaches.
Examples include:
- Continual reassessment method (CRM)
- Modified continual reassessment methods
- Bayesian optimal interval designs
- Keyboard designs
- mTPI-type designs
- Model-assisted interval designs
- Other Bayesian model-based approaches
These approaches differ in how they use accumulating toxicity information.
3+3 Versus CRM
| Feature | Traditional 3+3 | CRM |
|---|---|---|
| Primary structure | Rule-based | Model-based |
| Uses explicit dose-toxicity model | No | Yes |
| Target DLT probability | Implicit | Explicit |
| Uses information across doses | Limited | Yes |
| Cohort size | Typically 3, then 6 | Can be flexible |
| Statistical complexity | Low | Higher |
The choice between these approaches is a design decision involving statistical, clinical, operational, and regulatory considerations.
3+3 Versus BOIN
The Bayesian optimal interval design, commonly abbreviated BOIN, is a model-assisted approach.
Rather than relying solely on the traditional 3+3 thresholds, BOIN defines decision boundaries around a target toxicity probability.
Conceptually, if the observed toxicity rate is below a lower boundary, the design escalates.
If it is near the target, the design stays.
If it is above an upper boundary, the design de-escalates.
This allows the design to incorporate an explicit target toxicity probability.
3+3 Versus a Dose-Expansion Cohort
The dose-escalation portion of a Phase I study should also be distinguished from a later dose-expansion cohort.
A dose-expansion cohort may enroll additional patients at a selected dose to obtain more information about:
- Safety
- Pharmacokinetics
- Pharmacodynamics
- Preliminary efficacy
- Biomarker relationships
The 3+3 rule itself does not imply that the entire development program ends once the MTD candidate has been identified.
What Is the Dose-Limiting Toxicity Rate?
At dose \(d\), define:
The observed DLT rate is:
where:
- \(x_d\) = observed DLTs at dose \(d\)
- \(n_d\) = number of evaluable patients at dose \(d\)
For example:
for one DLT among six patients.
But an observed DLT rate is not the same thing as the true DLT probability.
What If a Patient Is Not DLT-Evaluable?
This is a common practical issue.
Suppose three patients are enrolled at a dose, but one patient withdraws before completing the DLT evaluation window for a reason unrelated to toxicity.
The protocol must specify whether that patient is:
- Considered evaluable
- Replaced
- Followed until the evaluation period is complete
- Handled according to another prespecified rule
The correct handling depends on the protocol.
What If a DLT Occurs After Escalation?
Suppose the study escalates from 40 mg to 80 mg after observing 0/3 at 40 mg.
A patient at 40 mg subsequently experiences a DLT that becomes apparent after the escalation decision.
The protocol should specify how this late information is handled.
In a real clinical trial, safety monitoring does not simply stop because the 3+3 algorithm has moved to the next dose.
A dose-escalation committee may need to reassess the accumulated safety information before proceeding further.
Safety Review Committees
Many Phase I programs use a formal safety review process.
The reviewing group may include:
- Investigators
- Medical monitors
- Clinical pharmacologists
- Biostatisticians
- Safety physicians
- Other subject-matter experts
The committee reviews DLT information and other safety information before authorizing further escalation.
This is important because the 3+3 algorithm is based on a limited endpoint and does not automatically capture every clinically meaningful safety signal.
DLT Versus Overall Safety
A dose can satisfy the formal DLT rule while still producing clinically important toxicity.
For example, a study could observe no formal DLTs but see:
- Frequent grade 2 toxicities
- Repeated treatment interruptions
- Persistent laboratory abnormalities
- Cumulative toxicity
- Unexpected adverse-event patterns
The clinical team may decide that escalation should not continue even though the narrow 3+3 DLT rule would technically permit it.
Why the 3+3 Design Is Sometimes Called Conservative
The traditional 3+3 design is often characterized as conservative because it may stop escalation after relatively little information.
For example, 2 DLTs among three patients immediately stop escalation.
Likewise, one DLT among six patients is enough to prevent escalation only if another DLT appears; the design distinguishes sharply between 1/6 and 2/6.
This can reduce the chance of exposing many patients to clearly toxic doses, but it can also result in stopping at a dose below the dose that a more statistically efficient design might ultimately select.
Why Small Cohorts Are Attractive
Small cohorts have several practical advantages.
- Fewer patients are exposed at each new dose.
- Escalation decisions can be made relatively quickly.
- The design is operationally simple.
- Clinical teams can easily communicate the rules.
- No complex statistical model is required for real-time decisions.
These advantages help explain the historical popularity of the design.
Why Small Cohorts Are Statistically Difficult
The corresponding disadvantage is that three or six patients provide limited information.
For example, suppose the true DLT probabilities at two adjacent doses are:
A small number of patients at each dose may not reliably distinguish those two toxicity probabilities.
The observed data can easily appear safer or more toxic than the underlying truth because of random sampling variation.
3+3 Does Not Estimate the MTD With a Confidence Interval
Another important limitation is that the traditional algorithm does not normally produce an estimated MTD together with a confidence interval.
For example, the output might simply be:
That does not imply that the true DLT probability at 80 mg is known precisely.
A statistical model-based design can instead estimate the dose-toxicity relationship and quantify uncertainty.
Common Misconception: "The MTD Has 1/6 DLTs"
A common oversimplification is:
"The MTD is the dose with exactly one DLT among six patients."
That is not generally the correct interpretation.
A dose with 0/3 may be escalated.
A dose with 1/3 is expanded.
A dose with 1/6 can permit escalation.
A dose with 2/6 stops escalation.
The selected dose depends on the sequence of decisions and the protocol's definition of the recommended dose.
Common Misconception: "0/3 Means Safe"
This is also incorrect.
Zero observed DLTs means:
It does not mean:
The true DLT probability may be nonzero, potentially substantially so.
Common Misconception: "3+3 Finds the Maximum Safe Dose"
The design identifies a dose according to its prespecified decision rules.
That is not necessarily equivalent to finding the absolute highest dose that would be tolerated by the population.
The true dose-toxicity curve remains uncertain.
Common Misconception: "The Next Dose Is Always Higher"
Under the basic escalation algorithm, a cohort with acceptable DLT experience leads to escalation.
However, clinical safety review can interrupt or modify escalation when new information becomes available.
The protocol should define how such circumstances are handled.
Common Misconception: "3+3 Is a Statistical Test"
The 3+3 design is better understood as a rule-based dose-escalation algorithm.
It is not equivalent to conducting a conventional hypothesis test at each dose.
There is generally no:
- Null hypothesis test at each dose
- p-value used to trigger escalation
- Confidence interval used as the primary decision rule
- Likelihood-based dose estimate after each cohort
The design instead uses the observed DLT count and prespecified rules.
What Does the 3+3 Design Actually Optimize?
A crucial question is:
What statistical criterion is the traditional 3+3 design optimizing?
The answer is: not a formally specified statistical optimization criterion in the way modern optimal designs are defined.
The design was developed as a practical rule-based approach to dose escalation.
It does not explicitly minimize:
- Expected sample size
- Probability of overdosing
- Mean squared error of the selected dose
- Probability of selecting the target dose
- Expected number of patients treated above the target toxicity level
Modern designs can explicitly optimize or evaluate such criteria.
Key Operating Characteristics for Modern Dose-Escalation Designs
When evaluating a dose-escalation design, investigators may examine:
- Probability of selecting the target dose
- Probability of selecting an overly toxic dose
- Probability of selecting an underdosed level
- Number of patients treated at overly toxic doses
- Total sample size
- Probability of early termination
- Average dose received
- Probability of stopping without identifying a recommended dose
These quantities can provide a more complete picture of design performance than the simple 3+3 decision table.
Simulation of the 3+3 Design
Because the 3+3 design is sequential, simulation is a useful way to understand its behavior.
Suppose the true DLT probabilities across six dose levels are:
A simulation can repeatedly generate patient-level DLT outcomes and apply the 3+3 rules.
After thousands of simulated trials, investigators can estimate:
- How often each dose is selected
- How often escalation stops early
- How many patients receive each dose
- How often an overly toxic dose is selected
- How often the target dose is selected
R Implementation of the Basic 3+3 Rule
A simple R function can encode the core decision rule.
three_plus_three <- function(dlt3) {
if (dlt3 == 0) {
return("Escalate")
}
if (dlt3 == 1) {
return("Expand to 6")
}
if (dlt3 >= 2) {
return("Stop escalation")
}
}
For the six-patient expansion:
three_plus_three_6 <- function(dlt6) {
if (dlt6 <= 1) {
return("Escalate")
}
if (dlt6 >= 2) {
return("Stop escalation")
}
}
These functions capture the core algorithm but do not replace the full clinical protocol.
Simulating a DLT Cohort in R
Suppose the true DLT probability at a particular dose is 20%.
set.seed(123) p_dlt <- 0.20 dlt_outcomes <- rbinom( n = 3, size = 1, prob = p_dlt ) dlt_outcomes sum(dlt_outcomes)
The simulated total number of DLTs can then be passed through the 3+3 rule.
dlt_count <- sum(dlt_outcomes) three_plus_three(dlt_count)
Simulating Many 3+3 Trials
A larger simulation can estimate how often each outcome occurs.
simulate_3plus3 <- function(p_dlt, nsim = 10000) {
outcomes <- character(nsim)
for (i in seq_len(nsim)) {
dlt3 <- rbinom(
n = 3,
size = 1,
prob = p_dlt
)
dlt_count <- sum(dlt3)
outcomes[i] <- three_plus_three(
dlt_count
)
}
prop.table(table(outcomes))
}
simulate_3plus3(0.10)
simulate_3plus3(0.30)
simulate_3plus3(0.50)
This demonstrates an important feature of simulation: the same underlying DLT probability can produce different decisions from one trial to another because of random variation.
A More Complete Cohort Simulation
A simplified simulation of sequential dose escalation can be written as:
simulate_dose_escalation <- function(
p_dlt,
max_dose = length(p_dlt)
) {
dose <- 1
while (dose <= max_dose) {
dlt_first <- rbinom(
3,
size = 1,
prob = p_dlt[dose]
)
n_dlt <- sum(dlt_first)
if (n_dlt == 0) {
dose <- dose + 1
next
}
if (n_dlt >= 2) {
return(max(1, dose - 1))
}
dlt_second <- rbinom(
3,
size = 1,
prob = p_dlt[dose]
)
total_dlt <- n_dlt + sum(dlt_second)
if (total_dlt <= 1) {
dose <- dose + 1
} else {
return(max(1, dose - 1))
}
}
max_dose
}
This is a simplified educational implementation. A production trial simulation would need substantially more detail concerning patient evaluability, DLT timing, cohort enrollment, dose delays, stopping rules, and the actual protocol definition.
Example Simulation Scenario
Suppose the true dose-toxicity probabilities are:
| Dose | True DLT Probability |
|---|---|
| 10 mg | 5% |
| 20 mg | 8% |
| 40 mg | 12% |
| 80 mg | 20% |
| 120 mg | 30% |
| 160 mg | 45% |
A simulation of thousands of trials can show how frequently the 3+3 algorithm selects each dose.
The results may demonstrate that the selected dose is not always the dose whose true toxicity probability is closest to a prespecified target.
That distinction is central when comparing traditional rule-based designs with model-based or model-assisted methods.
Why Dose Selection Is Random Even When the True Curve Is Fixed
Imagine that the true toxicity probabilities are known by the investigator:
The investigator still does not know which DLT outcomes will occur in a particular trial.
One trial may observe:
while another may observe:
The same underlying dose-toxicity curve can therefore produce different selected doses.
Stopping Rules Beyond DLTs
A complete Phase I protocol may contain stopping criteria unrelated to the formal 3+3 algorithm.
Examples include:
- Unexpected serious adverse events
- Drug-related deaths
- Severe organ toxicity
- Pharmacokinetic exposure exceeding predefined limits
- Unexpected accumulation
- Specific laboratory abnormalities
- Other clinical safety signals
These rules should be specified separately and integrated into the overall safety-monitoring framework.
Pharmacokinetics Can Change the Interpretation
The dose-toxicity relationship is not necessarily determined by nominal dose alone.
Two patients receiving the same dose can have very different exposures.
Let:
where \(E_i\) represents exposure for patient \(i\).
If pharmacokinetics are nonlinear, increasing the dose by 50% may increase exposure by substantially more or less than 50%.
This is one reason why Phase I dose selection may incorporate PK data in addition to DLT counts.
Pharmacodynamics Can Also Matter
For some therapies, the biologically active dose may be identified through a pharmacodynamic marker.
For example, investigators may observe that target engagement reaches a plateau before the DLT rate becomes unacceptable.
In such a situation, simply selecting the highest tolerated dose may not be the scientifically optimal strategy.
3+3 and Combination Therapies
Dose escalation becomes substantially more complicated when two or more agents are being combined.
For a single agent, the dose can be represented as:
For a two-drug combination, the dose becomes a pair:
The number of possible combinations can grow quickly.
Traditional 3+3 logic can become inefficient when many dose combinations must be evaluated.
3+3 in First-in-Human Studies
The traditional design has historically been particularly associated with first-in-human oncology studies.
The underlying rationale is straightforward:
- Start conservatively.
- Treat a small number of patients.
- Observe toxicity.
- Escalate when the observed toxicity is acceptable.
- Expand when uncertainty is greater.
- Stop when toxicity becomes excessive.
The simplicity is attractive in settings where patient safety is paramount and the amount of clinical information is initially very limited.
Why 3+3 Remains Commonly Recognized
The design has several features that make it easy to communicate across clinical teams.
The decision rules can be summarized in one sentence:
That simplicity is a genuine operational advantage.
However, operational simplicity should not be confused with statistical optimality.
Advantages of the Traditional 3+3 Design
- Simple: the rules are easy to understand.
- Transparent: the next action follows directly from the observed DLT count.
- Small cohorts: relatively few patients are treated at each new dose.
- Operationally familiar: clinical teams have extensive experience implementing the framework.
- No complex real-time model: dose decisions do not require fitting a statistical model after each cohort.
- Easy to document: the decision table can be incorporated directly into a protocol.
Limitations of the Traditional 3+3 Design
- Small sample sizes: three or six patients provide limited information about the true toxicity probability.
- No explicit dose-toxicity model: the design does not estimate the full dose-toxicity relationship.
- Limited use of information: information from previous doses is not incorporated in a formal statistical model.
- Potential inefficiency: the design can treat patients at doses that are substantially below or above the desired toxicity level.
- No explicit target probability: the traditional rules do not directly target a chosen DLT probability such as 20% or 25%.
- Variable selection: the selected dose can be sensitive to random DLT outcomes in very small cohorts.
When the 3+3 Design May Be Appropriate
The traditional design may be considered when:
- The clinical team values operational simplicity.
- The number of candidate dose levels is relatively small.
- The protocol has a clearly defined DLT endpoint.
- Cohort-based enrollment is operationally feasible.
- The development program is comfortable with a rule-based approach.
- There is extensive clinical familiarity with the framework.
The design should nevertheless be selected deliberately rather than simply because it is familiar.
When a Model-Based or Model-Assisted Design May Be Attractive
Alternative approaches may be particularly useful when:
- The study has many dose levels.
- A target DLT probability is explicitly defined.
- Efficient use of all accumulating toxicity information is important.
- The study has a complicated dose-toxicity relationship.
- Multiple dose combinations are being considered.
- There is a strong interest in formally quantifying dose-selection probabilities.
3+3 Is Not "Wrong"
It is important to avoid an overly simplistic conclusion that the traditional 3+3 design is simply a bad statistical method.
The design solves a particular operational problem with a simple set of rules.
Its limitations arise because the rules deliberately sacrifice statistical complexity and information efficiency for simplicity and transparency.
Whether that tradeoff is appropriate depends on the trial.
A Practical Dose-Escalation Workflow
What Should Be Specified in the Protocol?
A well-written dose-escalation section should make the algorithm reproducible.
At minimum, specify:
- Starting dose
- Candidate dose levels
- Dose increments
- Cohort size
- DLT definition
- DLT observation window
- Rules for evaluability
- Rules for replacing unevaluable patients
- Escalation criteria
- Expansion criteria
- Stopping criteria
- De-escalation criteria, if applicable
- Definition of the MTD or selected dose
- Rules for dose interruptions and delays
- Safety review procedures
- Criteria for dose expansion
What Should Be Specified in the Statistical Analysis Plan?
The Statistical Analysis Plan should describe how DLT data will be summarized and how the dose-escalation results will be analyzed.
Relevant components include:
- DLT analysis population
- Definition of evaluability
- Handling of missing DLT assessments
- DLT incidence by dose
- Number and percentage of patients with DLTs
- Nature and severity of DLTs
- Timing of DLTs
- Relationship to treatment
- Exposure information
- Final dose-escalation decision
Example of a Dose-Escalation Summary Table
| Dose | N | DLTs | DLT Rate | Decision |
|---|---|---|---|---|
| 10 mg | 3 | 0 | 0% | Escalate |
| 20 mg | 3 | 0 | 0% | Escalate |
| 40 mg | 6 | 1 | 16.7% | Escalate |
| 80 mg | 3 | 0 | 0% | Escalate |
| 120 mg | 6 | 2 | 33.3% | Stop escalation |
Interpreting the Worked Example Statistically
The observed DLT rate at 120 mg is:
or approximately 33.3%.
The observed rate at 80 mg is:
These observed rates might appear to suggest a substantial difference.
However, the sample sizes are very small.
The correct interpretation is therefore: the observed data triggered the prespecified 3+3 stopping rule at 120 mg.
It would be inappropriate to claim that the true DLT probability at 80 mg is zero based on 0/3.
Why the Selected Dose May Be Uncertain
Suppose the true DLT probabilities at 80 mg and 120 mg were actually:
A small sample could still generate 0/3 at 80 mg and 2/6 at 120 mg.
Another random realization might generate 1/3 at 80 mg and 0/3 at 120 mg.
The selected dose can therefore be influenced substantially by sampling variation.
The Importance of Simulation
For any proposed dose-escalation strategy, simulation can answer questions that the simple decision table cannot.
For example:
- How often does the design select the target dose?
- How often does it select an overly toxic dose?
- How many patients are exposed above the target?
- How often does the trial stop too early?
- How many patients are enrolled on average?
These are particularly useful when comparing 3+3 with alternative designs.
A Conceptual Comparison of Design Philosophies
| Design Philosophy | Core Question |
|---|---|
| Traditional 3+3 | Do observed DLTs permit escalation? |
| Model-based | What dose has a DLT probability near the target? |
| Model-assisted | Which dose decision is supported by the observed toxicity interval? |
These are different statistical philosophies rather than merely different implementations of the same algorithm.
Key Takeaways for Biostatisticians
When reviewing a 3+3 protocol, the most important questions are not limited to whether the decision table is correct.
A biostatistician should also ask:
- What is the DLT definition?
- What is the target clinical context?
- What is the dose sequence?
- What is the DLT observation window?
- How are unevaluable patients handled?
- What happens if late safety information emerges?
- How is the MTD defined?
- How is the recommended Phase II dose selected?
- Are PK and PD data incorporated?
- What simulation was performed?
- Are alternative dose-escalation designs appropriate?
Common Mistakes
- Confusing DLTs with all adverse events. The 3+3 algorithm operates on protocol-defined DLTs, not every adverse event.
- Forgetting the 1/3 expansion rule. One DLT among three does not automatically mean stop and does not automatically mean escalate.
- Ignoring the total six-patient result. After a 1/3 finding, the final decision is based on the combined six-patient experience.
- Calling 0/3 "proof of safety." Zero observed DLTs provides limited information about the true toxicity probability.
- Assuming the selected dose is known with precision. Small cohorts create substantial uncertainty.
- Confusing MTD with recommended Phase II dose. The final development dose can incorporate information beyond DLT counts.
- Ignoring evaluability rules. Missing or incomplete DLT observations can change the actual operation of the design.
- Changing the dose sequence after observing data without a prespecified framework. Such adaptations can alter the statistical properties of the design.
- Ignoring non-DLT toxicity. A clinically important toxicity does not necessarily have to satisfy the formal DLT definition to affect a dose-escalation decision.
- Assuming 3+3 is statistically optimal. The design is simple and transparent, but it does not explicitly optimize dose selection under a formal statistical criterion.
3+3 Design: The Rules You Should Memorize
One-Page Mental Model
| Observation | What It Means Operationally |
|---|---|
| 0/3 DLTs | Current dose passes → move up |
| 1/3 DLTs | Uncertain → add three patients |
| 2/3 DLTs | Too much toxicity signal → stop |
| 3/3 DLTs | Strong toxicity signal → stop |
| 1/6 DLTs | Still acceptable under traditional rule → move up |
| 2/6 DLTs | Too much toxicity signal → stop |
Final Worked Example Summary
| Dose | Initial Cohort | Expansion | Total | Decision |
|---|---|---|---|---|
| 10 mg | 0/3 | — | 0/3 | Escalate |
| 20 mg | 0/3 | — | 0/3 | Escalate |
| 40 mg | 1/3 | 0/3 | 1/6 | Escalate |
| 80 mg | 0/3 | — | 0/3 | Escalate |
| 120 mg | 1/3 | 1/3 | 2/6 | Stop |
The traditional interpretation is that escalation stops at 120 mg and the preceding dose, 80 mg, becomes the MTD candidate according to the protocol's 3+3 definition.
The Most Important Concept
The most important thing to understand about the 3+3 design is that it is a simple sequential decision algorithm, not a precise statistical estimator of the dose-toxicity curve.
The algorithm uses small cohorts and prespecified DLT rules:
- Zero DLTs → escalate.
- One DLT → expand to six.
- Two or more DLTs → stop escalation.
- After expansion, one or fewer DLTs → escalate.
- After expansion, two or more DLTs → stop escalation.
Its strength is simplicity.
Its limitation is that very small cohorts provide relatively little statistical information about the true probability of toxicity.
For this reason, modern Phase I development may consider model-based or model-assisted approaches when a more explicit statistical framework for dose selection is desired.
References
Storer, B.E. (1989).
Design and analysis of phase I clinical trials.
Biometrics, 45, 925–937.
O'Quigley, J., Pepe, M. & Fisher, L. (1990).
Continual reassessment method: a practical design for phase 1 clinical
trials in cancer.
Biometrics, 46, 33–48.
Le Tourneau, C., Lee, J.J. & Siu, L.L. (2009).
Dose escalation methods in phase I cancer clinical trials.
Journal of the National Cancer Institute, 101, 708–720.
Iasonos, A., Wilton, A.S., Riedel, E.R., Seshan, V.E. & Spriggs, D.R.
(2008).
A comprehensive comparison of the continual reassessment method to
the standard 3+3 dose escalation rule.
Clinical Trials, 5, 465–477.
Neuenschwander, B., Branson, M. & Gsponer, T. (2008).
Critical aspects of the Bayesian approach to phase I cancer trials.
Statistics in Medicine, 27, 2420–2439.
Yuan, Y., Lee, J.J. & Tang, L.L. (2016).
Bayesian phase I/II adaptive designs for targeted agents.
Journal of the American Statistical Association.
Liu, S. & Yuan, Y. (2015).
Bayesian optimal interval designs for phase I clinical trials.
Journal of the Royal Statistical Society: Series C, 64, 507–523.
Iasonos, A. & O'Quigley, J. (2017).
Adaptive dose-finding designs.
Journal of Biopharmaceutical Statistics.