Tutorials › AP Statistics › Balancing Errors When Designing a Study

Inference errors and practical significance · Tutorial 796 of 1000

Balancing Errors When Designing a Study

See how researchers can choose a tolerable Type I error rate and plan sample size to manage the chance of missing a meaningful difference.

Intermediate 10 min read

What You'll Learn

  • Explain how changing alpha affects Type I error risk and power for a fixed sample size.
  • Describe how increasing sample size can reduce Type II error risk without changing the chosen alpha.
  • Understand why beta and power must be tied to a specified true mean.
  • Use practical consequences to help justify a study’s alpha and sample-size plan.
  • Evaluate approximate power to decide whether a proposed design meets its goals.

Designing for Both Kinds of Error

A mean test can lead to two kinds of incorrect decisions: rejecting a true null hypothesis or failing to reject a false one. In “Type I Error in a Test About Means” and “Type II Error in a Test About Means,” you learned what these errors mean. In “How Significance Level Affects Power and Errors,” you saw that changing \(\alpha\) affects both error risks. This tutorial puts those ideas together as a study-design problem.

Before collecting data, researchers can choose a significance level and plan a sample size. A smaller \(\alpha\) makes it less likely that the test rejects a true null hypothesis, but, for a fixed sample size and specified true difference, it generally also makes rejection harder when the null is false. That can increase \(\beta\), the Type II error probability, and reduce power. A larger sample can often offset some of that loss of power without changing the chosen \(\alpha\).

Definition: For a specified true population mean that makes \(H_0\) false, \(\beta\) is the probability that the test fails to reject \(H_0\), and power is \(1-\beta\). These probabilities depend on the true mean, the test direction, \(\alpha\), sample size, and variability. There is not one \(\beta\) that applies to every possible alternative.

This last point is central to planning. A statement such as “the study has 80% power” is incomplete unless it identifies the true mean, or meaningful difference from the null, for which that power is intended. As in “How Effect Size and Variability Affect Power,” planning also requires an estimate of variability. A study may have high power for a large difference but low power for a smaller one.

Design principle: Choose \(\alpha\) to reflect how much Type I error risk is acceptable, then plan \(n\) so the test has adequate power for a specified, practically meaningful alternative. Increasing \(n\) can reduce \(\beta\) while keeping the selected \(\alpha\) unchanged.

How Alpha and Sample Size Work Together

For a fixed sample size, lowering \(\alpha\) moves the rejection boundary farther into the tail under \(H_0\). This reduces the chance of a Type I error, but a sample statistic must be more extreme to reject. If the true mean is fixed at a particular alternative value, that stricter boundary usually makes rejection less likely, increasing \(\beta\).

Increasing \(n\), by contrast, generally reduces the standard error of the sample mean. For a one-sample mean test, the planning standard error is approximately \(\sigma_{\text{planning}}/\sqrt{n}\). With a smaller standard error, the sampling distribution under the specified alternative is narrower, making it more likely that the sample mean will cross the rejection boundary. The test can therefore gain power without raising the nominal Type I error rate, provided the same \(\alpha\) and a suitable test are used.

$$ SE_{\text{planning}}=\frac{\sigma_{\text{planning}}}{\sqrt{n}}, \qquad \text{Power}=1-\beta $$

The design choices are related, but they are not interchangeable. Lowering \(\alpha\) directly reduces the planned Type I error rate; increasing \(n\) does not. Increasing \(n\) generally helps reduce Type II error for a specified alternative; lowering \(\alpha\) does not. When a smaller \(\alpha\) is needed, a larger sample may help preserve the desired power.

There is no universal best \(\alpha\) or sample size. A researcher should consider the consequences of both errors, how important it is to detect a meaningful difference, and the practical limits on collecting data. A larger sample can take more time, money, or participants. Conversely, too small a sample may leave a study poorly equipped to detect an important effect.

A Planning Sequence

1
Define the question and parameter.
Identify the population mean, the null value, and the direction of the alternative. Use the design to determine which population the study can address.
2
Identify a meaningful alternative.
Specify a true mean or difference that would matter in context. Power and \(\beta\) are evaluated for this alternative, not for every possible false null.
3
Choose an acceptable alpha.
Consider the consequences of falsely rejecting \(H_0\). State the chosen significance level before examining the study’s results.
4
Plan the sample size and assess power.
Use a reasonable planning estimate of variability to assess whether the proposed \(n\) gives adequate power at the chosen \(\alpha\). If not, consider whether a larger sample is feasible.

For a t test, an approximate planning calculation can locate the rejection boundary with a t critical value and then use a normal model for the sample mean under the specified true mean. This is a planning approximation, not an exact guarantee. The examples below show the method; the key design lesson is the comparison among error risks, \(\alpha\), and \(n\).

Worked Example: A Stricter Alpha at the Same Sample Size

Worked Example: Testing a New Cooling Method

A fictional food-science team plans to test whether a new cooling method increases the mean time, in minutes, that a product stays within a target temperature range. Let \(\mu\) be the population mean time for products treated with the new method. The team will test \(H_0:\mu=40\) versus \(H_a:\mu>40\), using \(n=36\) products. A planning standard deviation is 6 minutes, and the team wants to understand the effect of choosing \(\alpha=0.10\) or \(\alpha=0.01\) if the true mean is 42 minutes.

State. Compare the approximate power and Type II error probability for the two significance levels at the same sample size and specified true mean.

Plan. The observations should come from an appropriate random sample or randomized experiment. They should be independent; if sampled without replacement, the 10% condition should be met. The data should show no severe skewness or outliers; the sample size of 36 supports use of a t procedure when there are no serious departures. For planning, use \(\sigma_{\text{planning}}=6\) minutes and a normal approximation for the sample mean after locating each t rejection boundary.

Do. The planning standard error is \(6/\sqrt{36}=1\) minute, and the degrees of freedom are \(35\). For a one-sided test, the t critical values are approximately 1.3062 at \(\alpha=0.10\) and 2.4377 at \(\alpha=0.01\).

$$ \begin{aligned} \alpha=0.10:\quad \bar{x}_{\text{boundary}}&=40+1.3062(1)=41.3062,\\ z&=\frac{41.3062-42}{1}=-0.6938,\\ \text{Power}&\approx P(Z>-0.6938)=0.7561,\qquad \beta\approx 1-0.7561=0.2439.\\[4pt] \alpha=0.01:\quad \bar{x}_{\text{boundary}}&=40+2.4377(1)=42.4377,\\ z&=\frac{42.4377-42}{1}=0.4377,\\ \text{Power}&\approx P(Z>0.4377)=0.3308,\qquad \beta\approx 1-0.3308=0.6692. \end{aligned} $$

Conclude. With \(n=36\), tightening \(\alpha\) from 0.10 to 0.01 lowers the planned Type I error rate but also lowers approximate power for a true mean of 42 minutes, from 0.7561 to 0.3308. For this specified alternative, the estimated Type II error probability rises from 0.2439 to 0.6692. The estimates are approximate and apply to the stated planning assumptions.

Worked Example: Using a Larger Sample to Offset a Stricter Alpha

Worked Example: Measuring Improvement in a Training Program

A fictional rehabilitation team plans to test whether a training program increases a quantitative mobility score. Let \(\mu\) be the population mean improvement score, with \(H_0:\mu=0\) and \(H_a:\mu>0\). A planning standard deviation is 8 score points, and the team considers a true mean improvement of 4 points meaningful. It will use \(\alpha=0.01\). Compare \(n=25\) with \(n=100\).

State. The goal is to see how increasing sample size affects approximate power and \(\beta\) for the same true mean improvement and the same \(\alpha\).

Plan. Improvement scores should come from an appropriate random sample or randomized experiment. Observations from different participants should be independent; check the 10% condition if participants are sampled without replacement. The score distribution should not have severe skewness or outliers, particularly for the smaller sample. Use the one-sample t test’s one-sided critical value and the planning standard error, then approximate the sample mean’s distribution by a normal model centered at 4 points.

Do. With \(n=25\), \(df=24\) and \(SE_{\text{planning}}=8/\sqrt{25}=1.6\). The one-sided critical value at \(\alpha=0.01\) is approximately 2.4922. The rejection boundary is \(0+2.4922(1.6)=3.9875\) points. Under a true mean of 4 points, \(z=(3.9875-4)/1.6=-0.0078\). Thus, approximate power is \(P(Z>-0.0078)=0.5031\), and \(\beta\approx1-0.5031=0.4969\).

With \(n=100\), \(df=99\) and \(SE_{\text{planning}}=8/\sqrt{100}=0.8\). The one-sided critical value at \(\alpha=0.01\) is approximately 2.3646. The rejection boundary is \(0+2.3646(0.8)=1.8917\) points. Under the same true mean of 4 points, \(z=(1.8917-4)/0.8=-2.6354\). Therefore, approximate power is \(P(Z>-2.6354)=0.9958\), and \(\beta\approx1-0.9958=0.0042\).

Conclude. For this specified true improvement, increasing the sample size from 25 to 100 raises approximate power from 0.5031 to 0.9958 while keeping \(\alpha=0.01\). The larger sample reduces the standard error, improving the chance of detecting the meaningful improvement without raising the planned Type I error rate.

Worked Example: Is a Proposed Design Powerful Enough?

Worked Example: Monitoring Nutrient Levels in a Stream

A fictional environmental team plans to test whether the mean nutrient level in a stream exceeds a reference value of 20 units. Let \(\mu\) be the population mean nutrient level, and use \(H_0:\mu=20\) versus \(H_a:\mu>20\). The team plans to sample \(n=36\) sites, use \(\alpha=0.05\), and expects a standard deviation of 9 units. It considers a true mean of 23 units important to detect and wants at least 80% power for that alternative.

State. Assess whether the proposed design meets the 80% power goal for a true mean of 23 units, and identify an appropriate design response if it does not.

Plan. Sites should be selected through an appropriate random sampling process. Measurements at distinct sites should be independent; if sampling without replacement, the 10% condition should be met. The distribution should not have severe skewness or outliers; \(n=36\) can support a t procedure when departures are not serious. Use the one-sided t critical value for \(df=35\), then estimate power with a normal model for the sample mean and the planning standard deviation.

Do. The planning standard error is \(9/\sqrt{36}=1.5\) units. At \(\alpha=0.05\) with \(df=35\), the one-sided t critical value is approximately 1.6896. The rejection boundary is \(20+1.6896(1.5)=22.5344\) units. Under the specified true mean of 23, \(z=(22.5344-23)/1.5=-0.3104\). Approximate power is \(P(Z>-0.3104)=0.6219\), so \(\beta\approx1-0.6219=0.3781\).

Conclude. The proposed design has approximate power of 0.6219 for detecting a true mean of 23 units, below the team’s 0.80 goal. If the \(\alpha=0.05\) limit reflects an acceptable Type I error risk, the team should consider increasing the sample size or improving measurement precision rather than raising \(\alpha\) automatically. It should recalculate power for any revised design.

Common Mistakes and AP Exam Tips

  • Claiming that a larger sample lowers alpha: It does not. The significance level is chosen for the test. A larger \(n\) generally reduces \(\beta\) and raises power for a specified alternative; it does not change the selected \(\alpha\).
  • Discussing beta without naming an alternative: Type II error probability depends on which false null is true. A full response states the specified population mean or meaningful difference for which \(\beta\) or power is being discussed.
  • Assuming lower alpha is always better: A lower \(\alpha\) reduces Type I error risk, but may make it harder to detect a real difference at a fixed sample size. Explain the tradeoff and consider whether sample size can be increased.
  • Treating a power target as a guarantee: Power is a long-run probability under specified conditions, not a promise that a particular sample will detect an effect. Approximate planning values also depend on the planning estimate of variability and the approximation used.
  • Choosing a sample size without considering design conditions: A large \(n\) does not fix biased sampling, dependence, or serious problems with the data distribution. State the appropriate sampling or assignment process, independence, 10% condition when relevant, and distribution check for the planned t procedure.
  • Changing alpha just to reach a power target: Raising \(\alpha\) can increase power, but it also raises the long-run Type I error rate. If false positives are costly, explain why a larger sample or better measurement may be preferable.

A strong AP response connects each design choice to its consequence: “At the same sample size and for this specified true mean, lowering \(\alpha\) reduces Type I error risk but lowers power,” or “Increasing \(n\) can improve power for this alternative while retaining the chosen \(\alpha\).” For a numerical planning result, identify the assumed true mean, show the boundary and probability calculation, and label the result approximate when using a planning standard deviation and normal model.

Key takeaway: Manage both error risks deliberately. Choose \(\alpha\) in light of the consequences of a Type I error, then plan \(n\) to provide adequate power for a specified, meaningful alternative while meeting the conditions for the intended mean test.

Check Your Understanding

Unless a question says otherwise, assume the test direction and specified true alternative stay fixed.

  1. For a fixed sample size and true alternative, what generally happens to power and Type I error risk when \(\alpha\) is lowered?
  2. Why does increasing sample size not lower the significance level chosen for a test?
  3. A power statement gives no true mean or effect size. What important information is missing?
  4. A proposed design misses its power target, but its \(\alpha\) was chosen to limit costly false positives. Name one design change to consider and explain why raising \(\alpha\) is not an automatic fix.
  5. What conditions should be checked before relying on an approximate power calculation for a one-sample t test?