Alpha Is a Long-Run Error Rate
In “Costs of Wrong Decisions in Mean Tests,” you considered the practical consequences of Type I and Type II errors. This tutorial focuses on a statistical quantity that sets the rate of one of those errors: the significance level, \(\alpha\). For a t test about a population mean, alpha specifies how often the test procedure will reject a true null hypothesis over many repetitions, when the assumptions of the procedure hold.
The word “rate” matters. Alpha describes the behavior of a testing procedure across repeated samples under a stated null model. It is not the probability that the null hypothesis is true, and it does not tell us the probability that a particular rejection is wrong after seeing the data. As in “Why a t Test Never Proves the Null Mean,” a test decision does not reveal the population mean with certainty.
A t test uses a t distribution to set a rejection region. The location and size of that region depend on the alternative hypothesis, the chosen \(\alpha\), and the degrees of freedom. If the test is two-sided, extreme results in either direction can lead to rejection. If it is one-sided, only results in the direction specified by the alternative can lead to rejection.
For a standard t test of a single null value, the procedure has a Type I error probability of \(\alpha\) when that null value is true and the t model conditions are met. For a one-sided test whose null hypothesis includes a range of values, the probability of rejection is controlled at no more than \(\alpha\) for values in that null range; it reaches \(\alpha\) at the boundary value. In introductory problems, the stated null is commonly the boundary value being tested.
How a t Test Sets the Error Rate
Consider a two-sided one-sample t test of \(H_0:\mu=\mu_0\) against \(H_a:\mu\ne\mu_0\). The test statistic measures how far the sample mean is from the null mean in standard-error units:
Under \(H_0\), and when the one-sample t conditions are met, this statistic follows a t distribution with \(n-1\) degrees of freedom. The test rejects for sufficiently large positive or negative values of \(t\). The two tails together have probability \(\alpha\) under the null model. For a one-sided test, the rejection region is in just the direction specified by \(H_a\), and that tail has probability \(\alpha\).
A smaller alpha makes the rejection region more extreme: stronger sample evidence is required to reject \(H_0\). That reduces the chance of a Type I error when \(H_0\) is true, but it can also make it harder for a test to detect a false null. The next tutorial, “What Power of a Test Means,” develops that second consequence.
Worked Examples: Alpha in t Tests
Worked Example: A Two-Sided Test of Mean Battery Life
A fictional manufacturer wants to test whether the mean battery life for a certain model differs from 20 hours. Let \(\mu\) be the true mean battery life, in hours, for the population represented by the sample. A random sample of 16 batteries has \(\bar{x}=22.4\) hours and \(s=4.0\) hours. Use a two-sided one-sample t test at \(\alpha=0.05\).
State. The hypotheses are \(H_0:\mu=20\) and \(H_a:\mu\ne20\), where \(\mu\) is the population mean battery life in hours. A Type I error would be concluding that the mean battery life differs from 20 hours when it actually equals 20 hours.
Plan and conditions. A one-sample t test is appropriate because one quantitative sample is compared with a fixed value. Assume the batteries were selected using an appropriate random process. The measurements should be independent; if sampling without replacement, 16 batteries should be less than 10% of the population represented. Because the sample size is small, check that the battery-life distribution has no strong skewness or outliers. Those shape details are not supplied here and would need to be checked before using the test.
Do. The standard error, test statistic, and degrees of freedom are:
For a two-sided test with 15 degrees of freedom and \(\alpha=0.05\), the critical values are approximately \(-2.131\) and \(2.131\). Because \(2.40>2.131\), reject \(H_0\). The data provide convincing evidence that the population mean battery life differs from 20 hours.
Conclude about alpha. If the true population mean were 20 hours, this test would reject \(H_0\) in 5% of repeated samples, assuming the conditions and t model hold. That 5% is the test’s long-run Type I error rate. It does not mean that there is a 5% chance that the true mean is 20 hours, nor does it mean there is a 5% chance this particular rejection is wrong. If the mean really is 20 hours, this rejection is a Type I error; if it is not, the rejection is not a Type I error.
Worked Example: A One-Sided Test of Filling Time
A fictional packaging team investigates whether a machine’s mean filling time is below 12 seconds. Let \(\mu\) be the true mean filling time, in seconds, for the machine’s output under the conditions of interest. A random sample of 10 fills has \(\bar{x}=11.1\) seconds and \(s=1.5\) seconds. Test \(H_0:\mu=12\) against \(H_a:\mu<12\) at \(\alpha=0.05\).
State. The parameter is the population mean filling time. A Type I error would be concluding that the mean is below 12 seconds when it is actually 12 seconds.
Plan and conditions. Use a one-sample t test because one quantitative sample is compared with a fixed benchmark. Assume the fills were randomly sampled or generated under a suitable random process. The observations should be independent; if sampling without replacement, 10 should be less than 10% of the population represented. Since \(n=10\) is small, check that the sample data show no strong skewness or outliers. Summary statistics alone cannot confirm this shape condition.
Do. The standard error and test statistic are:
For a lower-tail test with 9 degrees of freedom and \(\alpha=0.05\), the critical value is approximately \(-1.833\). Since \(-1.897<-1.833\), reject \(H_0\). The data provide convincing evidence that the population mean filling time is below 12 seconds.
Conclude about alpha. Under the boundary null value \(\mu=12\), 5% of repeated samples would produce a t statistic in the lower-tail rejection region, if the t conditions hold. That is the Type I error rate for this test at the null value. Because the alternative is one-sided, the entire 5% rejection probability is in the lower tail; the test does not use 2.5% in each tail.
The Connection Between Alpha and Confidence Intervals
A two-sided t test at level \(\alpha\) is linked to a matching \(100(1-\alpha)\%\) confidence interval for the population mean. The test rejects \(H_0:\mu=\mu_0\) exactly when the matching interval does not include \(\mu_0\), apart from rounding at the boundary. This is the test-and-interval connection introduced in “Connecting Confidence Intervals to Test Decisions.”
The interval connection gives another way to understand the error rate. Across many samples, a matching \(100(1-\alpha)\%\) t interval captures the true mean in about \(100(1-\alpha)\%\) of repetitions. Equivalently, it misses the true mean in about \(\alpha\) of repetitions. When the null mean is the true mean, missing it corresponds to the matching two-sided test rejecting that true null.
Worked Example: A 90% Interval and Its Matching Test
A fictional environmental team estimates the mean reading from a monitoring process. Let \(\mu\) be the true mean reading, in units, for the population of measurements represented by the sample. A random sample of 16 readings has \(\bar{x}=51.6\) units and \(s=3.2\) units. Construct the matching 90% t interval and test \(H_0:\mu=50\) against \(H_a:\mu\ne50\) at \(\alpha=0.10\).
State and plan. The parameter is the population mean reading. A one-sample t interval estimates this mean, and its matching two-sided t test evaluates the null value of 50 units. Assume the readings came from an appropriate random sample, are independent, and satisfy the 10% condition if sampled without replacement. Since \(n=16\) is small, the distribution of readings should have no strong skewness or outliers; check the data for these features.
Do. The standard error is \(3.2/\sqrt{16}=0.8\). With \(df=15\), the 90% critical value is approximately \(t^*=1.753\). Thus the margin of error is \(1.753(0.8)=1.4024\) units, and the interval is:
The null value, 50 units, is outside the interval, so the matching two-sided test rejects \(H_0\) at \(\alpha=0.10\). The test statistic is \(t=(51.6-50)/0.8=2.00\), with 15 degrees of freedom; this is beyond the two-sided critical values \(\pm1.753\). The data provide convincing evidence that the population mean reading differs from 50 units.
Conclude about alpha. If the true mean were 50 units, the matching 90% interval procedure would miss that true mean in 10% of repeated samples. Equivalently, the matching test would reject the true null value in 10% of repeated samples. For the particular interval calculated here, the true mean is either inside or outside it; the 10% describes the long-run performance of the procedure, not a probability assigned to this one interval.
Common Mistakes and AP Exam Tips
- Calling alpha the probability that \(H_0\) is true: Alpha is conditional on \(H_0\) being true. It does not assign a probability to the truth of the hypothesis.
- Calling alpha the probability this decision is wrong: The Type I error rate is a long-run property of the test. Whether a particular rejection is a Type I error depends on the unknown truth about the population mean.
- Putting alpha in the wrong tail: A one-sided alternative determines the direction of the rejection region. For a two-sided test, the total tail area is \(\alpha\), with \(\alpha/2\) in each tail for the usual symmetric t procedure.
- Confusing alpha with the p-value: Alpha is the chosen cutoff for the test. The p-value is calculated from the sample result under the null model. Compare the p-value with alpha to make the test decision, as discussed in “One-Sided and Two-Sided P-Values Compared.”
- Assuming a 95% interval gives a 5% probability for this interval to miss: The 5% describes the long-run miss rate of the interval method when its conditions hold. Do not treat the fixed population mean as randomly moving in or out of a calculated interval.
- Forgetting that the interval and test must match: The direct connection requires the corresponding two-sided test and confidence interval. A one-sided test does not match a two-sided interval at the same confidence level.
For a full-credit AP explanation, name the population mean and the null claim, identify alpha as the probability of rejecting that true null in repeated use of the test, and keep that conditional statement distinct from what is known about a particular sample. When using an interval, state its confidence level and explain the matching test connection without claiming a probability for the fixed parameter after the interval is observed.
Check Your Understanding
Answer each question by distinguishing a long-run rate from a statement about one particular sample or interval.
- A one-sample t test uses \(\alpha=0.01\) to test \(H_0:\mu=30\) against \(H_a:\mu\ne30\). If the true mean is 30, what does the 0.01 describe?
- In a lower-tail t test at \(\alpha=0.05\), where is the rejection region, and what would a Type I error mean in context?
- A matching 95% two-sided t interval excludes the null mean. What test decision follows at \(\alpha=0.05\), and how is that related to Type I error in repeated samples?
- Why is it incorrect to say, “There is a 5% chance that the null hypothesis is true,” when a test uses \(\alpha=0.05\)?
- A matching 90% t interval is calculated from one sample. Does the 10% long-run miss rate tell you the probability that this particular interval misses the true mean? Explain.