Introduction
Sample size determination is traditionally introduced through concepts such as type I error, power, effect size, and statistical significance.
A Bayesian clinical trial can use a fundamentally different framework. Instead of designing the study primarily around the probability of rejecting a null hypothesis, the investigator may define a clinically meaningful posterior decision criterion such as:
where: \(\theta\) is the parameter of interest, \(\theta_0\) is a clinically relevant threshold, and \(c\) is the posterior probability required to make a decision.
The sample size is then chosen so that the planned study has a sufficiently high probability of satisfying this Bayesian criterion under one or more plausible assumptions about the true treatment effect.
Bayesian Sample Size Is Not Simply "Power With a Prior"
A common misconception is that Bayesian sample size determination is obtained by taking an ordinary frequentist power calculation and adding a prior distribution.
That is not generally the case.
A Bayesian design may instead ask:
- What is the probability that the posterior probability exceeds a decision threshold?
- What is the probability that the posterior credible interval is sufficiently narrow?
- What is the probability that the posterior expected benefit exceeds a clinically relevant value?
- What is the expected utility of each possible sample size?
- What sample size gives an adequate probability of making the desired decision?
These are different design objectives and can produce different sample sizes.
The Bayesian Sample Size Problem
Suppose the parameter of interest is:
and the investigator wants to determine the smallest sample size \(n\) such that a prespecified Bayesian criterion is satisfied.
A generic decision criterion can be written:
where:
- \(\theta_0\) = clinically relevant threshold
- \(Y\) = observed data
- \(c\) = required posterior probability
The sample size problem becomes:
where:
- \(\theta_A\) = assumed true value used for planning
- \(A\) = desired assurance
Power Versus Assurance
The distinction between power and assurance is central.
In a conventional design, power is typically:
In a Bayesian design using a posterior probability criterion, assurance can be defined as:
Thus, power is replaced by a probability that the future Bayesian decision criterion will be satisfied.
| Concept | Frequentist Design | Bayesian Design |
|---|---|---|
| Primary planning quantity | Power | Assurance |
| Decision criterion | Reject or fail to reject \(H_0\) | Posterior probability threshold |
| Prior distribution | Not part of the inferential model | Explicitly incorporated |
| Uncertainty after study | Confidence interval / test statistic | Posterior distribution / credible interval |
| Sample size objective | Achieve target power | Achieve target assurance or precision |
A Simple Binary Endpoint
To make the calculation concrete, consider a single-arm Phase II study with a binary response endpoint.
Let:
where: \(p\) is the true response probability and \(X\) is the number of responders among \(n\) patients.
Suppose a response rate of 20% is considered the minimum clinically relevant benchmark.
The Bayesian decision criterion might be:
The study is therefore considered successful when there is at least a 95% posterior probability that the true response rate exceeds 20%.
Choosing a Prior Distribution
For a binomial endpoint, a convenient conjugate prior is the beta distribution:
with density proportional to:
The prior mean is:
and the prior variance is:
The quantity:
is often interpreted as the prior effective sample size for this conjugate binomial model.
Prior Effective Sample Size
Consider two beta priors:
and:
Both have mean 0.50:
and:
Actually, these two priors do not have the same mean. This illustrates why both the prior location and prior strength must be considered.
For:
the prior contributes only two units of effective sample size.
For:
the prior contributes 50 units.
A strong informative prior can therefore materially change the sample size required to achieve a specified posterior criterion.
The Posterior Distribution
With:
and:
the posterior distribution is:
This conjugacy makes the binary example particularly useful for illustrating Bayesian sample size determination.
Posterior Probability of Exceeding a Threshold
Suppose the clinical threshold is:
After observing \(x\) responses, the posterior probability of clinical interest is:
where \(F\) is the beta cumulative distribution function.
The Bayesian success rule is therefore:
For a common choice:
the trial succeeds if the posterior probability that the response rate exceeds 20% is at least 95%.
Worked Example: Bayesian Sample Size for Response Rate
Suppose a Phase II study evaluates a new treatment using objective response rate.
The investigators define:
| Parameter | Planning Value |
|---|---|
| Clinically relevant threshold \(p_0\) | 20% |
| Planning response rate \(p_A\) | 40% |
| Prior distribution | \(\operatorname{Beta}(1,1)\) |
| Posterior probability criterion | \(P(p>0.20\mid data)\ge0.95\) |
| Target assurance | 80% |
The design problem is:
This is the Bayesian analogue of a conventional sample size problem.
Step 1: Specify the Prior
We use a uniform prior:
Its mean is:
and its effective sample size is:
This is a weakly informative prior for illustration because it contributes relatively little information compared with a typical clinical trial.
Step 2: Specify the Bayesian Decision Rule
The treatment will be considered successful if:
This is fundamentally different from requiring:
The observed response rate alone is not the Bayesian decision criterion.
The entire posterior distribution is used.
Step 3: Determine the Minimum Number of Responses Required
For a proposed sample size \(n\), calculate the posterior probability for every possible response count:
For each value of \(x\):
Then calculate:
The smallest response count satisfying the 95% posterior probability criterion defines the Bayesian success boundary for that sample size.
Step 4: Consider a Sample Size of 24
Suppose:
If eight responses are observed:
the posterior distribution is:
The posterior probability that the response rate exceeds 20% is:
which is approximately:
Therefore, eight or more responses among 24 patients satisfy the Bayesian decision criterion.
| Responses \(X\) | Posterior Distribution | \(P(p>0.20\mid X)\) | Decision |
|---|---|---|---|
| 6 | \(\operatorname{Beta}(7,19)\) | Below 0.95 | Do not declare success |
| 7 | \(\operatorname{Beta}(8,18)\) | Below 0.95 | Do not declare success |
| 8 | \(\operatorname{Beta}(9,17)\) | 0.9532 | Declare success |
| 9+ | Correspondingly higher posterior probability | >0.95 | Declare success |
Step 5: Calculate Assurance
The sample size is not determined merely by identifying the response threshold. We must also determine how likely it is that the trial will reach that threshold.
Assume that the true response probability for planning purposes is:
Under this assumption:
Because the Bayesian decision requires at least eight responses, assurance is:
Therefore:
The resulting assurance is approximately:
or:
Thus, if the true response rate is 40%, approximately 80.8% of hypothetical future trials would satisfy the Bayesian success criterion.
Step 6: Verify Smaller Sample Sizes
Sample size determination requires searching smaller candidate values as well.
For example, with \(n=23\), eight responses are still required to meet the posterior probability criterion.
The assurance becomes approximately:
Therefore:
| Sample Size | Minimum Responses | Assurance at \(p=0.40\) |
|---|---|---|
| 20 | 7 | Approximately 75.0% |
| 21 | 8 | Approximately 65.0% |
| 22 | 8 | Approximately 71.0% |
| 23 | 8 | Approximately 76.3% |
| 24 | 8 | Approximately 80.8% |
Thus, 24 is the smallest sample size in this search that achieves at least 80% assurance.
The Complete Bayesian Design
| Design Component | Value |
|---|---|
| Endpoint | Binary response |
| Parameter | Response probability \(p\) |
| Clinically relevant threshold | \(p_0=0.20\) |
| Planning true response rate | \(p_A=0.40\) |
| Prior | \(\operatorname{Beta}(1,1)\) |
| Posterior criterion | \(P(p>0.20\mid data)\ge0.95\) |
| Target assurance | 80% |
| Required sample size | 24 |
| Success boundary | \(X\ge8\) |
| Actual assurance | Approximately 80.8% |
Why This Is a Bayesian Sample Size Calculation
Notice what happened in the example.
We did not begin by specifying a conventional null hypothesis and alternative hypothesis and then calculating a type I error and power.
Instead, we specified:
- A prior distribution for the response probability.
- A clinically relevant threshold.
- A posterior probability required for a successful conclusion.
- A planning value for the true response probability.
- A desired assurance.
The sample size was then selected so that the future posterior decision would have a sufficiently high probability of being successful.
Assurance Is a Prior-Predictive Concept
The term assurance is sometimes confusing because it is not simply a posterior probability.
A posterior probability is calculated after observing the data:
Assurance is calculated before the data are observed:
The outer probability therefore describes the distribution of possible future datasets.
Posterior probability: "Given the data we observed, how probable is it that the treatment exceeds the clinical threshold?"
Assurance: "Before observing the data, how likely is the planned study to produce a posterior probability that meets our decision criterion?"
Prior Predictive Versus Conditional Assurance
There are several ways to define Bayesian assurance.
The worked example used a fixed planning value:
This is sometimes called conditional assurance.
An alternative is to average over a distribution of plausible true values.
For example, if the planning distribution for the true response probability is \(\pi(p)\), then assurance can be written:
This integrates the probability of success over the assumed distribution of possible true response rates.
Conditional Assurance
Conditional assurance asks:
It is particularly useful when investigators have a specific clinically important response rate that they want the trial to detect reliably.
For example:
means that the design is being evaluated under a true response rate of 40%.
Prior-Predictive Assurance
A fully Bayesian planning framework can instead assign a probability distribution to the unknown true response rate itself.
Suppose:
Then the prior predictive distribution of the future response count is:
For the beta-binomial model, this has a closed-form expression:
where \(B(\cdot,\cdot)\) is the beta function.
Prior-predictive assurance is then:
where \(\mathcal S\) is the set of response counts satisfying the posterior decision criterion.
Why the Choice of Planning Distribution Matters
Suppose an investigator believes the true response rate is most likely around 40%.
A design based on:
will have different assurance than a design based on:
or:
Therefore, there is no universally correct Bayesian sample size without specifying the assumptions under which the sample size is being evaluated.
Posterior Precision as an Alternative Criterion
Not every Bayesian trial should be designed around a probability of success.
Another common objective is posterior precision.
For example, the investigator may require:
where the interval is a 95% posterior credible interval and \(w\) is the maximum acceptable width.
For example, the requirement might be:
This is a fundamentally different sample size objective from achieving a posterior probability of efficacy.
Decision-Based Versus Precision-Based Bayesian Sample Size
| Objective | Example Criterion |
|---|---|
| Posterior efficacy | \(P(\theta>\theta_0\mid data)\ge0.95\) |
| Posterior superiority | \(P(\theta_A>\theta_B\mid data)\ge0.95\) |
| Posterior non-inferiority | \(P(\theta_A-\theta_B>-\Delta\mid data)\ge0.975\) |
| Posterior precision | 95% credible interval width \(\le w\) |
| Decision utility | Expected utility exceeds a prespecified level |
Bayesian Sample Size for Two Treatment Groups
The same framework extends naturally to randomized trials.
Suppose:
A Bayesian decision criterion might be:
where \(\Delta\) is the clinically meaningful treatment difference.
The sample size is then selected to provide sufficient assurance that this posterior probability criterion will be satisfied under a clinically relevant true treatment effect.
Bayesian Non-Inferiority Sample Size
A similar framework can be used for non-inferiority.
Suppose the treatment effect is:
and the non-inferiority margin is:
A Bayesian criterion could be:
The corresponding assurance criterion might be:
This provides a direct Bayesian interpretation of the sample size objective.
Bayesian Sample Size for Continuous Outcomes
For a continuous endpoint, suppose:
A Bayesian analysis may place a prior on \(\mu\):
The sample size might be selected so that the posterior probability:
is achieved with a prespecified assurance.
Alternatively, the design might require a sufficiently narrow posterior credible interval for \(\mu\).
Bayesian Sample Size for Time-to-Event Outcomes
For survival endpoints, the design may be based on a posterior probability criterion for a hazard ratio:
or:
Sample size may then be determined by simulation because the distribution of future events and posterior quantities is substantially more complicated than in the beta-binomial example.
Simulation-Based Bayesian Sample Size Determination
Many realistic Bayesian trial designs cannot be solved analytically.
Simulation is then used.
The general algorithm is:
Monte Carlo Estimate of Assurance
Suppose \(M\) simulated trials are generated for a candidate sample size. Let:
Then the estimated assurance is:
For example, if 8,120 of 10,000 simulated trials satisfy the criterion:
or approximately 81.2% assurance.
Monte Carlo Error Matters
Because simulation produces an estimate rather than an exact probability, Monte Carlo error should be considered.
If the estimated assurance is approximately \(\widehat A\), its Monte Carlo standard error is approximately:
For \(M=10,000\) and \(\widehat A=0.80\):
Thus the simulation estimate has a Monte Carlo standard error of approximately 0.4 percentage points.
Frequentist Operating Characteristics Can Still Be Evaluated
A Bayesian design does not prevent investigators from examining frequentist operating characteristics.
For example, suppose the Bayesian decision rule is:
The investigator can still calculate the probability of declaring success when the true response rate is exactly 20%:
This quantity is sometimes described as a frequentist type I error associated with the Bayesian decision rule.
It is important to distinguish this from the Bayesian posterior probability.
Posterior Probability Is Not Type I Error
Suppose the final analysis produces:
This does not mean that there is a 3% type I error.
It means that, conditional on the specified model, prior, and observed data, the posterior probability that \(p\) exceeds 20% is 97%.
A frequentist operating characteristic such as type I error is instead a long-run probability evaluated across repeated hypothetical datasets under a specified true parameter value.
Sensitivity to the Prior
Bayesian sample size calculations can be sensitive to the prior, particularly when the prior is informative relative to the amount of new clinical data.
Suppose the prior is:
and the study observes \(x\) responses among \(n\) patients.
The posterior is:
A stronger prior therefore contributes more information to the posterior.
This can reduce the amount of new information required to satisfy a posterior criterion.
Prior Sensitivity Analysis
A robust Bayesian sample size assessment should examine multiple plausible priors.
For example:
| Prior | Interpretation |
|---|---|
| \(\operatorname{Beta}(1,1)\) | Weakly informative / approximately uniform |
| \(\operatorname{Beta}(2,3)\) | Mild prior concentration around 40% |
| \(\operatorname{Beta}(4,6)\) | More concentrated prior around 40% |
| Historical-data prior | Prior informed by relevant external evidence |
The sample size should be evaluated under each scientifically plausible prior rather than relying exclusively on the prior that produces the smallest sample size.
Prior-Data Conflict
An informative prior can also conflict with the observed data.
For example, suppose historical studies strongly support a response rate near 40%, but the new trial produces a response rate near 10%.
The posterior will reflect both sources of information.
The degree to which the historical evidence influences the posterior depends on its effective strength and relevance.
Discounting Historical Information
In some Bayesian clinical trial designs, historical information may be discounted.
Conceptually, an effective prior might contribute less information than the raw historical dataset.
For example, a historical dataset containing 100 patients might be represented by a prior with an effective sample size substantially smaller than 100 if there is uncertainty about exchangeability.
This is one reason that prior elicitation and prior-data compatibility assessment are important parts of Bayesian design.
Sample Size and Prior Information
The relationship can be summarized conceptually as:
Therefore, stronger relevant prior information can sometimes reduce the amount of new information needed.
However, this should not be interpreted as a general justification for using strong priors simply to obtain smaller trials.
What Happens If the Prior Is Very Informative?
Suppose the prior effective sample size is very large.
Then the posterior may remain close to the prior even after substantial new data are collected.
This creates an important design consideration: the trial may appear to have high assurance because the prior is doing much of the inferential work.
Investigators should therefore examine how much of the posterior conclusion comes from the new trial versus the prior.
Effective Sample Size of the Prior
For a beta prior:
a simple measure of prior effective sample size is:
After observing \(n\) new patients:
This is a useful conceptual diagnostic, although effective sample size in more complex Bayesian models requires more sophisticated definitions.
Bayesian Sample Size Is Often a Design Optimization Problem
The investigator may not simply want the smallest possible sample size.
There may be competing objectives:
- High assurance of a successful conclusion
- High posterior precision
- Low patient exposure
- High probability of identifying an ineffective treatment
- Robustness to prior uncertainty
- High probability of correctly selecting the best treatment
- Acceptable operational complexity
The final design can therefore be viewed as an optimization problem.
A General Bayesian Sample Size Workflow
Exact Enumeration Versus Simulation
For simple conjugate models, exact enumeration is often possible.
For the beta-binomial example, all possible values of \(X\) can be evaluated:
For each value, calculate the posterior probability and determine whether the Bayesian criterion is satisfied.
The assurance is then obtained by summing the probabilities of the successful response counts.
For complex models, simulation is usually more practical.
| Method | Best Suited For |
|---|---|
| Exact enumeration | Simple conjugate models and discrete outcomes |
| Numerical integration | Moderately complex analytical models |
| Monte Carlo simulation | Complex posterior models |
| MCMC-based simulation | Hierarchical or otherwise computationally complex Bayesian models |
R Implementation: Define the Design
p0 <- 0.20 pA <- 0.40 a0 <- 1 b0 <- 1 posterior_cutoff <- 0.95 target_assurance <- 0.80
Here:
p0is the clinically relevant threshold.pAis the planning true response probability.a0andb0define the beta prior.posterior_cutoffis the required posterior probability.target_assuranceis the desired probability of satisfying the Bayesian criterion.
R Implementation: Posterior Probability
posterior_prob <- function(x, n,
a = a0,
b = b0,
threshold = p0) {
1 - pbeta(
threshold,
shape1 = a + x,
shape2 = b + n - x
)
}
For a given number of responses \(x\), this function calculates:
R Implementation: Determine the Bayesian Success Boundary
success_boundary <- function(n) {
posterior_values <- sapply(
0:n,
function(x)
posterior_prob(x, n)
)
x_success <- which(
posterior_values >= posterior_cutoff
) - 1
if (length(x_success) == 0) {
return(NA)
}
min(x_success)
}
For the worked example:
success_boundary(24) # 8
Thus eight responses among 24 patients are sufficient to satisfy the posterior probability criterion.
R Implementation: Calculate Assurance
assurance <- function(n, p_true = pA) {
x_boundary <- success_boundary(n)
if (is.na(x_boundary)) {
return(0)
}
pbinom(
x_boundary - 1,
size = n,
prob = p_true,
lower.tail = FALSE
)
}
For the worked example:
assurance(24) # approximately 0.8081
Search for the Required Sample Size
results <- data.frame( n = 1:100 ) results$boundary <- sapply( results$n, success_boundary ) results$assurance <- sapply( results$n, assurance ) results[ results$assurance >= target_assurance, ][1, ]
The first qualifying design is approximately:
n # 24 boundary # 8 assurance # approximately 0.8081
A Compact Exact Search Function
find_bayesian_n <- function(
p0,
pA,
a0,
b0,
posterior_cutoff = 0.95,
target_assurance = 0.80,
max_n = 200) {
for (n in 1:max_n) {
posterior_values <- sapply(
0:n,
function(x) {
1 - pbeta(
p0,
a0 + x,
b0 + n - x
)
}
)
successful_x <- which(
posterior_values >= posterior_cutoff
) - 1
if (length(successful_x) == 0)
next
boundary <- min(successful_x)
assurance <- pbinom(
boundary - 1,
size = n,
prob = pA,
lower.tail = FALSE
)
if (assurance >= target_assurance) {
return(
data.frame(
n = n,
success_boundary = boundary,
assurance = assurance
)
)
}
}
NULL
}
The design can then be obtained with:
find_bayesian_n( p0 = 0.20, pA = 0.40, a0 = 1, b0 = 1, posterior_cutoff = 0.95, target_assurance = 0.80 )
which returns approximately:
n success_boundary assurance 24 8 0.8081
Evaluating the Induced Frequentist Type I Error
Although the design is Bayesian, the decision rule can be evaluated under a frequentist null response rate.
For the worked example, the Bayesian success rule is:
when \(n=24\).
If the true response rate is exactly 20%, the probability of falsely declaring success is:
This is approximately:
This quantity is not required to be exactly 5%, because 5% was not used as the primary Bayesian design criterion.
R Code for the Induced Type I Error
n <- 24 boundary <- 8 type1_error <- pbinom( boundary - 1, size = n, prob = p0, lower.tail = FALSE ) type1_error
This evaluates:
Posterior Probability After Different Results
It is useful to examine how the posterior conclusion changes with the observed number of responses.
| Responses | Posterior | \(P(p>0.20\mid X)\) | Decision |
|---|---|---|---|
| 5 | \(\operatorname{Beta}(6,20)\) | Below 0.95 | No success |
| 6 | \(\operatorname{Beta}(7,19)\) | Below 0.95 | No success |
| 7 | \(\operatorname{Beta}(8,18)\) | Below 0.95 | No success |
| 8 | \(\operatorname{Beta}(9,17)\) | 0.9532 | Success |
| 9 | \(\operatorname{Beta}(10,16)\) | Higher | Success |
| 10+ | Increasingly favorable | Higher | Success |
Observed Response Rate Versus Posterior Probability
An important Bayesian design principle is that the observed response rate and the posterior probability answer different questions.
For eight responses among 24:
The observed response rate is therefore approximately 33.3%.
But the Bayesian decision is based on:
The second quantity incorporates uncertainty about the true response rate and the prior distribution.
Why the Posterior Threshold Is Not the Same as the Clinical Threshold
Another common mistake is to assume that:
That is not necessarily true.
The posterior distribution incorporates uncertainty, so the posterior probability depends on the entire observed dataset and the prior.
Likewise, an observed response rate above 20% does not automatically imply that the posterior probability exceeds 95%.
Bayesian Sample Size Under Different Assumed True Response Rates
The required sample size can be evaluated under several possible true response rates.
For the worked decision rule:
assurance will increase as the true response rate becomes larger.
| True Response Rate | Interpretation |
|---|---|
| 10% | Treatment performs substantially below the clinical threshold |
| 20% | Treatment is exactly at the clinical threshold |
| 30% | Moderately promising treatment |
| 40% | Planning alternative in the worked example |
| 50% | Highly active treatment |
The sample size calculation should ideally be evaluated across this entire range rather than only at one assumed value.
Assurance Curves
A useful operating characteristic is the assurance curve:
This function shows how the probability of a successful Bayesian conclusion changes as the true treatment effect changes.
For the worked example:
while the assurance will generally be lower when the true response rate is near or below 20% and higher when the true response rate is substantially above 40%.
Assurance Is Not Guaranteed Probability of Clinical Benefit
Suppose the design has 80% assurance at \(p=0.40\).
This does not mean:
nor does it mean that the treatment has an 80% probability of being clinically beneficial.
It means that under repeated hypothetical trials generated with the true response probability set to 40%, approximately 80% would satisfy the prespecified posterior decision criterion.
Bayesian Sample Size With Multiple Scenarios
A strong design analysis often considers several scenarios simultaneously.
For example:
| Scenario | True Response Rate | Desired Assessment |
|---|---|---|
| Unfavorable | 10% | Probability of incorrectly declaring success |
| Threshold | 20% | Induced false-positive operating characteristic |
| Intermediate | 30% | Assurance in gray zone |
| Primary planning | 40% | Target assurance |
| Very favorable | 50% | Probability of success under strong activity |
This gives decision-makers a much more complete understanding of how the design behaves.
Bayesian Sample Size and Clinical Interpretation
The most important design quantity should ultimately have a clinical meaning.
For example:
should represent a response rate below which the treatment would not be considered sufficiently promising.
Similarly:
should represent a response rate that would meaningfully change the development decision.
Common Mistakes
- Calling posterior probability "power." A posterior probability and frequentist power are different quantities.
- Ignoring the prior. The prior is part of the Bayesian model and can affect the required sample size.
- Choosing an informative prior only because it reduces sample size. Prior information should be scientifically justified.
- Reporting assurance without specifying the planning scenario. Assurance depends on the assumed true parameter value or its planning distribution.
- Confusing assurance with the probability that the treatment works. Assurance is a probability over future datasets under specified planning assumptions.
- Using the observed effect size as the posterior decision criterion. The posterior distribution, not merely the point estimate, determines the Bayesian decision.
- Ignoring prior sensitivity. Informative priors can materially affect sample size and posterior conclusions.
- Failing to evaluate frequentist operating characteristics. Even Bayesian designs can benefit from examining repeated-sampling behavior.
- Using simulation without enough Monte Carlo replicates. Simulation error can be important when candidate sample sizes are close.
- Failing to prespecify the Bayesian decision rule. Changing the posterior probability threshold after seeing the data can alter the design's operating characteristics.
Protocol Considerations
A Bayesian sample size calculation should be fully documented before enrollment.
At minimum, the protocol should specify:
- Primary estimand
- Primary endpoint
- Statistical model
- Prior distribution
- Rationale for the prior
- Clinically meaningful threshold
- Posterior decision criterion
- Planning treatment effect
- Target assurance
- Maximum sample size
- Method used to calculate assurance
- Prior sensitivity analyses
- Relevant frequentist operating characteristics
- Handling of missing data
- Rules for unevaluable patients
- Simulation details, when applicable
Interim Analyses in Bayesian Designs
Bayesian trials can naturally accommodate interim analyses because posterior probabilities can be updated as new data accumulate.
For example, a study might evaluate:
after 20, 30, 40, and 50 patients.
The trial could stop for success if:
or stop for futility if:
provided that these rules were prespecified and their operating characteristics were evaluated during design.
Bayesian Monitoring and Sample Size Are Connected
If interim stopping rules are introduced, the design is no longer simply a fixed-sample calculation.
The sample size analysis should then evaluate:
- Probability of early success
- Probability of early futility
- Probability of reaching the maximum sample size
- Expected sample size
- Probability of a final successful conclusion
- Frequentist false-positive behavior
Simulation is often the most practical way to evaluate these quantities.
Bayesian Sample Size With Adaptive Designs
The same principles extend to more complicated adaptive trials.
Examples include:
- Response-adaptive randomization
- Sample size re-estimation
- Adaptive dose selection
- Multiple treatment arms
- Platform trials
- Bayesian hierarchical models
- Borrowing of historical information
In these settings, the probability distribution of future data depends on the adaptive rules themselves.
Consequently, simulation becomes increasingly important.
Decision-Theoretic Bayesian Sample Size
A more advanced Bayesian approach uses expected utility.
Suppose:
is the utility associated with making decision \(d\) when the true parameter is \(\theta\).
The expected utility before observing the data can be written:
where \(C(n)\) represents the cost or burden associated with enrolling additional patients.
The optimal sample size can then be defined as:
This approach explicitly recognizes that larger sample sizes have both benefits and costs.
When Decision-Theoretic Methods Become Useful
Expected-utility approaches can be attractive when:
- Patient exposure has substantial ethical or economic consequences.
- Multiple possible clinical decisions exist.
- The cost of delaying development is important.
- The consequences of false-positive and false-negative decisions are asymmetric.
- There are several competing treatment-development options.
These methods are considerably more complex than the simple assurance approach, but they provide a natural framework for Bayesian decision-making.
Bayesian Versus Frequentist Sample Size: Conceptual Comparison
| Feature | Frequentist | Bayesian |
|---|---|---|
| Parameter treatment | Fixed but unknown | Random through a prior distribution |
| Prior information | Not part of the inferential model | Explicitly modeled |
| Primary planning quantity | Power | Assurance, precision, or expected utility |
| Decision criterion | p-value / confidence interval / test | Posterior probability / credible interval / utility |
| Future-data planning | Repeated-sampling distribution | Predictive or conditional distribution |
| Sample size result | Usually tied to \(\alpha\), power, and effect size | Tied to prior, posterior criterion, and assurance |
What Makes Bayesian Sample Size Determination Difficult?
The mathematics of the simple beta-binomial example are relatively easy.
The difficult part is usually the design specification.
Investigators must decide:
- What prior is scientifically appropriate?
- How much historical information should be borrowed?
- What posterior probability constitutes sufficient evidence?
- What effect should be used for planning?
- How much assurance is sufficient?
- What is the cost of an additional patient?
- What operating characteristics are clinically important?
These are scientific and decision-making questions, not merely mathematical questions.
A Practical Bayesian Sample Size Workflow
Worked Example Summary
The complete binary-endpoint example can be summarized as follows.
| Component | Value |
|---|---|
| Endpoint | Binary response |
| Parameter | Response probability \(p\) |
| Clinical threshold | \(p_0=0.20\) |
| Planning response rate | \(p_A=0.40\) |
| Prior | \(\operatorname{Beta}(1,1)\) |
| Posterior criterion | \(P(p>0.20\mid data)\ge0.95\) |
| Target assurance | 80% |
| Required sample size | 24 |
| Minimum responses for success | 8 |
| Posterior probability at 8 responses | Approximately 95.32% |
| Assurance at \(p=0.40\) | Approximately 80.81% |
| Induced frequentist false-positive probability at \(p=0.20\) | Approximately 3.27% |
The Most Important Concept
The most important conceptual point is that Bayesian sample size determination is fundamentally a future-posterior design problem.
The investigator does not know what the future data will look like.
Instead, the investigator considers all plausible future datasets and asks:
For a binary endpoint, this can sometimes be solved exactly by enumerating every possible response count.
For more complex clinical trial designs, the same principle is usually implemented through simulation.
The design therefore connects three distributions:
Bottom Line
References
Berry, S.M., Carlin, B.P., Lee, J.J. & Muller, P. (2010).
Bayesian Adaptive Methods for Clinical Trials.
CRC Press.
Spiegelhalter, D.J., Abrams, K.R. & Myles, J.P. (2004).
Bayesian Approaches to Clinical Trials and Health-Care Evaluation.
Wiley.
O'Hagan, A., Stevens, J.W. & Campbell, M.J. (2005).
Assurance in clinical trial design.
Pharmaceutical Statistics, 4, 187–201.
Berry, S.M. (2006).
Bayesian clinical trials.
Nature Reviews Drug Discovery, 5, 27–36.
Berry, D.A. (2006).
Bayesian clinical trials.
Nature Reviews Drug Discovery, 5, 27–36.
Neuenschwander, B., Branson, M. & Gsponer, T. (2008).
Critical aspects of the Bayesian approach to phase I cancer trials.
Statistics in Medicine, 27, 2420–2439.
Thall, P.F. & Wathen, J.K. (2007).
Practical Bayesian adaptive randomisation in clinical trials.
European Journal of Cancer, 43, 859–866.
Hobbs, B.P., Carlin, B.P. & Mandrekar, S.J. (2013).
Bayesian group sequential designs for clinical trials.
Statistical Methods in Medical Research.