Tutorials › AP Statistics › Type II Error in a Test About Means

Inference errors and practical significance · Tutorial 787 of 1000

Type II Error in a Test About Means

See how a t test can miss a real difference in means, and learn what a Type II error means in context.

Intermediate 9 min read

What You'll Learn

  • Define a Type II error for tests about population means and mean differences
  • Distinguish a Type II error from simply failing to reject a null hypothesis
  • Identify when a test’s failure would be a Type II error if the true mean were known
  • Connect a matching confidence interval to a possible Type II error
  • Explain how sample size, variability, effect size, and alpha affect the chance of a Type II error

When a Test Misses a Real Difference

In “Type I Error in a Test About Means,” you learned that rejecting a true null hypothesis is a Type I error. This tutorial considers the other possible testing error: a test may fail to reject the null hypothesis even though the null hypothesis is false. The sample may not provide strong enough evidence to detect a real difference in a population mean.

A test decision is based on sample data, but the truth of a hypothesis concerns a population parameter. Because different random samples can produce different sample means and standard deviations, a valid test will not always make the same decision. As discussed in “Why Small Samples Miss Real Differences,” a real population difference can be difficult to detect when a sample estimate is imprecise.

Definition: A Type II error occurs when a hypothesis test fails to reject a false null hypothesis. In a test about means, this means failing to reject a null claim about a population mean or a difference between population means when that claim is actually false.

For instance, suppose a test evaluates \(H_0:\mu=50\) against \(H_a:\mu\ne50\). If the true population mean is 54, then \(H_0\) is false. If the test fails to reject \(H_0\), it has made a Type II error. If the true mean is exactly 50, the same decision is not a Type II error; it is simply a failure to reject a true null hypothesis.

A test’s result alone cannot tell us whether a Type II error occurred, because the true population mean is ordinarily unknown. Use conditional wording: “If the true mean is 54, failing to reject the null hypothesis would be a Type II error.” Do not say that every failure to reject is an error, or that the test has accepted or proved the null hypothesis. The distinction between “fail to reject” and “accept” is explained in “Fail to Reject H0 Wording That Earns Full Credit.”

Beta, Power, and the Particular Alternative

The probability of a Type II error is often denoted by \(\beta\). Unlike the significance level \(\alpha\), which describes the test’s long-run Type I error rate when the null hypothesis is true, \(\beta\) concerns a specified situation in which the null hypothesis is false. Its value depends on which alternative mean is actually true, along with the test procedure and its conditions. For a composite alternative such as \(H_a:\mu\ne50\), there is not one universal \(\beta\) for every possible mean different from 50.

Definition: For a specified true alternative, \(\beta\) is the probability that the test fails to reject \(H_0\). The power of the test for that alternative is the probability that the test rejects \(H_0\) when that alternative is true; power equals \(1-\beta\).

A single sample’s p-value is not \(\beta\). The p-value measures how unusual the observed test statistic would be if \(H_0\) were true. Beta describes how often the test would fail to reject across repeated samples when a particular alternative value is true. Determining an exact \(\beta\) generally requires specifying that alternative and the sampling model. The examples here focus on identifying a possible Type II error from a test decision, rather than calculating \(\beta\).

Several features affect the chance of a Type II error. When the real difference from the null value is larger, it is generally easier for a test to detect it. Greater sample variability makes sample means less precise, while a larger sample generally reduces the standard error and helps distinguish a real difference from random variation. For a fixed design and alternative, increasing \(\alpha\) generally makes rejection easier and can reduce \(\beta\), but it also raises the long-run Type I error rate. These are tendencies, not guarantees for an individual sample.

Key distinction: A large p-value does not show that \(H_0\) is true, and it does not tell you the probability that a Type II error occurred. It means the sample did not provide convincing evidence against \(H_0\) at the chosen significance level. Whether that failure was a Type II error depends on the unknown truth about the population parameter.

Worked Examples: Identifying a Possible Type II Error

Worked Example: Battery Capacity Around a Target

A fictional quality-control team tests whether the mean capacity of a type of battery differs from 50 amp-hours. A random sample of 16 batteries has \(\bar{x}=52.4\) amp-hours and \(s=8\) amp-hours. Use a two-sided test at \(\alpha=0.05\). Then describe what a Type II error would mean if the actual population mean were 54 amp-hours.

State: Let \(\mu\) be the true mean capacity, in amp-hours, for the population represented by the sample. The hypotheses are \(H_0:\mu=50\) and \(H_a:\mu\ne50\), with \(\alpha=0.05\).

Plan: Use a one-sample t test because one quantitative sample is compared with a fixed benchmark. Assume the batteries were selected using an appropriate random sample, capacities from distinct batteries are independent, and the population contains at least 160 batteries so the 10% condition is met. Assume the sample’s capacity distribution has no extreme outliers or severe skew. The degrees of freedom are \(df=n-1=16-1=15\).

Do: First calculate the standard error and test statistic:

$$ SE=\frac{s}{\sqrt{n}} =\frac{8}{\sqrt{16}} =2\text{ amp-hours} $$
$$ t=\frac{\bar{x}-\mu_0}{SE} =\frac{52.4-50}{2} =1.20 $$

For a two-sided test with \(df=15\) at \(\alpha=0.05\), the critical values are approximately \(-2.131\) and \(2.131\). Since \(1.20\) is between these values, fail to reject \(H_0\).

Conclude: The data do not provide convincing evidence that the population mean battery capacity differs from 50 amp-hours. If the true population mean were actually 54 amp-hours, then \(H_0\) would be false, and this failure to reject would be a Type II error. We cannot determine from this sample alone whether that error occurred.

The corresponding 95% one-sample t interval also illustrates the decision. The critical value is \(t^*=2.131\), so the margin of error is \(2.131(2)=4.262\) amp-hours. The interval is \(52.4\pm4.262\), or \((48.138,56.662)\) amp-hours. Rounded to one decimal place, it is \((48.1,56.7)\) amp-hours. It includes the null value of 50, matching the test’s failure to reject. It also includes the hypothetical true mean of 54. If 54 were the actual population mean, this sample’s interval would include the null value and the test would fail to reject the false null.

Worked Example: Change in Study Time After a Program

In a fictional study, 12 students record their weekly study time before and after a planning program. Define each difference as after-program time minus before-program time. The sample mean difference is \(\bar{x}_d=2.1\) hours per week, and the standard deviation of differences is \(s_d=3.5\) hours. A two-sided paired t test of no mean change fails to reject at \(\alpha=0.05\). Explain the decision and identify the possible Type II error if the true mean increase were 4 hours per week.

State: Let \(\mu_d\) be the true mean change in weekly study time, in hours, for the population represented by the students. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\), with \(\alpha=0.05\).

Plan: Use a paired t test because each student is measured twice; the analysis is on the within-student differences, defined as after minus before. Assume the students were selected using an appropriate random process, the differences for distinct students are independent, and the population contains at least 120 students so the 10% condition is met. Assume the distribution of differences has no extreme outliers or severe skew. Here \(df=12-1=11\).

Do: The standard error and test statistic are:

$$ SE=\frac{s_d}{\sqrt{n}} =\frac{3.5}{\sqrt{12}} \approx1.010\text{ hours} $$
$$ t=\frac{\bar{x}_d-0}{SE} =\frac{2.1}{3.5/\sqrt{12}} \approx2.08 $$

For a two-sided test with \(df=11\) at \(\alpha=0.05\), the critical values are approximately \(-2.201\) and \(2.201\). Since \(2.08\) is between them, fail to reject \(H_0\).

Conclude: The data do not provide convincing evidence that the population mean after-before difference in weekly study time is nonzero. If the true mean increase were 4 hours per week, the null hypothesis of no mean change would be false. Failing to reject it would then be a Type II error. The sample’s estimated increase of 2.1 hours does not establish that the true increase is 4 hours; 4 hours is a hypothetical value used to describe the error.

In a paired analysis, the relevant variability is the standard deviation of the differences, \(s_d\), not the separate standard deviations of the before and after measurements. The error concerns a claim about \(\mu_d\), the population mean of those differences.

Worked Example: Comparing Two Mean Response Times

A fictional technology team compares response times for two independent systems. For a random sample of 20 requests on System 1, the mean response time is \(\bar{x}_1=18.6\) milliseconds and the standard deviation is \(s_1=4\) milliseconds. For 20 independent requests on System 2, \(\bar{x}_2=17.0\) milliseconds and \(s_2=4\) milliseconds. Test for a difference in population means at \(\alpha=0.05\), and describe a Type II error if the true mean difference, System 1 minus System 2, is 3.5 milliseconds.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean response times, in milliseconds, for the populations represented by the two samples. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\), with \(\alpha=0.05\).

Plan: Use an unpooled two-sample t test because the response is quantitative and the groups are independent. Assume the requests were selected using an appropriate random process, observations are independent within and between groups, and each population contains at least 200 requests so the 10% condition is met for both samples. Assume neither group’s response-time distribution has extreme outliers or severe skew. With equal sample sizes and standard deviations, the unpooled degrees of freedom are \(df=38\).

Do: The observed difference in sample means is \(18.6-17.0=1.6\) milliseconds. Calculate its standard error and the test statistic:

$$ SE=\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}} =\sqrt{\frac{4^2}{20}+\frac{4^2}{20}} =\sqrt{1.6} \approx1.265\text{ milliseconds} $$
$$ t=\frac{(\bar{x}_1-\bar{x}_2)-0}{SE} =\frac{18.6-17.0}{\sqrt{1.6}} \approx1.265 $$

For a two-sided test with \(df=38\) at \(\alpha=0.05\), the critical values are approximately \(-2.024\) and \(2.024\). Since \(1.265\) is between these values, fail to reject \(H_0\).

Conclude: The data do not provide convincing evidence that the population mean response times differ between the two systems. If the true difference \(\mu_1-\mu_2\) were 3.5 milliseconds, then the null hypothesis would be false. Failing to reject it would be a Type II error. The sample result by itself does not reveal whether the true difference is zero or 3.5 milliseconds.

Common Mistakes and AP Exam Tips

  • Calling every failure to reject a Type II error: It is a Type II error only if \(H_0\) is actually false. A full-credit description says what the error would mean under a stated true alternative.
  • Writing “accept \(H_0\)” or “prove there is no difference”: A large p-value does not establish the null claim. Say that the data do not provide convincing evidence for the alternative, in context.
  • Treating the p-value as \(\beta\): The p-value is calculated assuming \(H_0\) is true. Beta is a probability of failing to reject for a specified false null, across repeated samples under that alternative.
  • Claiming a known Type II error from one sample: The truth about the population mean or mean difference is usually unknown. Use conditional phrasing such as, “If the true mean difference were 3.5 milliseconds, this failure to reject would be a Type II error.”
  • Leaving the parameter and context vague: Identify the population mean or mean difference, name the response, and include its units. For two groups, state the order of subtraction so the direction is clear.
  • Ignoring the procedure’s conditions: A Type II error discussion follows a valid test setup. Check the random process, independence, 10% condition when appropriate, and the data-shape condition for the specific one-sample, paired, or two-sample t procedure.

A concise AP response can state the test decision first, then explain the possible error conditionally: “We fail to reject the null hypothesis. If the true population mean were actually different from the null value, this decision would be a Type II error.” That statement distinguishes the observed decision from the unknown truth and keeps the conclusion tied to the parameter and context.

Key takeaway: In a t test about means, a Type II error is failing to reject a false null claim about a population mean or mean difference. A sample can fail to provide convincing evidence even when a real difference exists; the test result alone cannot tell us whether that happened.

Check Your Understanding

For each situation, distinguish the test decision from what would make it a Type II error.

  1. A one-sample t test of \(H_0:\mu=40\) fails to reject. What must be true about the actual population mean for this decision to be a Type II error?
  2. A paired t test defines differences as after minus before and fails to reject \(H_0:\mu_d=0\). Describe a Type II error if the true mean difference is an increase of 2 minutes.
  3. Why is a p-value of 0.30 not the probability that a Type II error occurred?
  4. For a specified true difference and procedure, what does \(\beta\) represent? How is power related to \(\beta\)?
  5. Name two design or data features that can affect the chance of a Type II error, and describe the general direction of their effects.