When a Correctly Run Test Makes the Wrong Decision
In “Using Confidence Intervals to Judge Practical Importance,” you compared a confidence interval with a practical threshold. This tutorial turns to a different question: what can go wrong when a t test evaluates a claim about a population mean? Even when a test follows its procedure, a sample can lead it to reject a null hypothesis that is actually true.
A hypothesis test uses sample data to decide whether they provide convincing evidence against a null hypothesis. But a sample varies from one random sample to another. When the null hypothesis is true, that variation can still produce a sample statistic far enough from the null value to trigger rejection. That mistaken rejection is called a Type I error.
For example, suppose a test evaluates \(H_0:\mu=100\) hours against a two-sided alternative. If the true population mean battery life is 100 hours, but the test rejects \(H_0\), the test has made a Type I error. If the true mean is not 100 hours, rejecting that null hypothesis is not a Type I error. In practice, the true population mean is usually unknown, so we cannot tell from one test’s result whether its rejection was actually an error.
What Alpha Controls
The significance level, \(\alpha\), is chosen before using the test result to make a decision. It sets the rejection threshold. When the null hypothesis is true and the assumptions for the t procedure are satisfied, \(\alpha\) is the probability that the procedure rejects \(H_0\). In other words, it controls the long-run Type I error rate for that test.
The long-run language matters. Imagine repeatedly taking random samples from a population whose true mean equals the null value and applying the same valid test at \(\alpha=0.05\). Over many repetitions, about 5% of those tests would reject the true null hypothesis. The rate is a property of the test procedure under the null, not a verdict about the truth of the null in an individual study.
For a two-sided t test at \(\alpha=0.05\), the rejection region is split between the two tails of the t distribution under \(H_0\): 2.5% of the area is in each tail. If the calculated t statistic falls far enough into either tail, the test rejects. As discussed in “How Alpha Changes the Conclusion,” changing \(\alpha\) changes the decision threshold, not the p-value calculated from the data.
The stated rate relies on using an appropriate test and meeting its conditions. The random process, independence, 10% condition when sampling without replacement, and data-shape conditions should be checked for the chosen t procedure, as in “Matching Procedures to Conditions.” If those assumptions are not reasonable, the advertised Type I error rate may not be achieved.
Worked Examples: Finding the Error in Context
Worked Example: Testing a Battery-Life Claim
A fictional manufacturer claims that its rechargeable batteries have a mean lifetime of 100 hours. A random sample of 25 batteries has a mean lifetime of \(\bar{x}=104.8\) hours and a standard deviation of \(s=10\) hours. Test the claim against the possibility that the population mean is different from 100 hours, using \(\alpha=0.05\). Explain what a Type I error would mean in this setting.
State: Let \(\mu\) be the true mean battery lifetime, in hours, for the population represented by this sample. The hypotheses are \(H_0:\mu=100\) and \(H_a:\mu\ne100\). The significance level is \(\alpha=0.05\).
Plan: Use a one-sample t test because one quantitative sample is being compared with a fixed benchmark. Assume the batteries were selected using an appropriate random sample, lifetimes from distinct batteries are independent, and the population contains at least 250 batteries so the 10% condition is met. Assume the sample’s distribution of lifetimes has no extreme outliers or severe skew; with \(n=25\), this supports using a t procedure. The test has \(df=n-1=24\).
Do: Calculate the standard error and test statistic:
For a two-sided test with \(t=2.40\) and \(df=24\), the p-value is approximately \(0.02451\). Since \(0.02451<0.05\), reject \(H_0\).
Conclude: The sample provides convincing evidence that the population mean battery lifetime differs from 100 hours. The sample mean is above 100 hours, but this two-sided test’s conclusion is that the mean differs from 100, not specifically that it is greater. A Type I error would occur if the true population mean were exactly 100 hours and this test nevertheless rejected \(H_0\). The observed rejection alone does not tell us whether that error actually occurred.
The Type I error is defined by the relationship between the decision and the true population value. It is not defined by whether the sample result looks surprising, or by whether the p-value is smaller than some particular number. Here, if the true mean is 100 hours, rejecting is an error; if it is not 100 hours, rejecting is not a Type I error.
Worked Example: A Matching Interval Misses a True Mean
A fictional packaging line is expected to fill containers in a mean of 30 seconds. A random sample of 16 containers has a mean fill time of \(\bar{x}=33.5\) seconds and a standard deviation of \(s=4\) seconds. Construct a 95% confidence interval for the population mean fill time. Then explain how the interval relates to a two-sided test of \(H_0:\mu=30\) at \(\alpha=0.05\).
State: Let \(\mu\) be the true mean fill time, in seconds, for the population represented by the sample. The interval estimates \(\mu\); the null value for the comparison is 30 seconds.
Plan: Use a one-sample t interval. Assume the containers are an appropriate random sample, measurements from different containers are independent, and the population contains at least 160 containers so the 10% condition is met. Assume the sample’s distribution has no extreme outliers or severe skew; this check is especially important for a sample of 16. Here \(df=16-1=15\), and the 95% critical value is \(t^*\approx2.131\).
Do: Calculate the standard error and margin of error:
The interval is \(33.5\pm2.131\), giving endpoints \(31.369\) and \(35.631\) seconds. Rounded to one decimal place, the 95% confidence interval is \((31.4,35.6)\) seconds.
Conclude: We are 95% confident that the population mean fill time is between about 31.4 and 35.6 seconds. The interval excludes 30 seconds, so it agrees with rejecting the matching two-sided t test of \(H_0:\mu=30\) at \(\alpha=0.05\). If the true mean were actually 30 seconds, this interval would have missed the true mean, and the matching test’s rejection would be a Type I error.
As covered in “Connecting Confidence Intervals to Test Decisions,” a two-sided test at level \(\alpha\) matches a \(100(1-\alpha)\%\) t confidence interval: the test rejects the null value when that value is outside the interval. Thus, for a 5% two-sided test, the matching 95% interval fails to include the true mean in the same cases that the test rejects that true mean. Across repeated samples under the null, this happens at the test’s Type I error rate.
Worked Example: Comparing Mean Yields in Two Fields
In a fictional study, randomly selected plots in two independent groups receive different soil treatments. The first group has \(n_1=25\), \(\bar{x}_1=56\) kilograms per plot, and \(s_1=10\) kilograms. The second group has \(n_2=25\), \(\bar{x}_2=50\) kilograms per plot, and \(s_2=10\) kilograms. Test whether the population mean yields differ, using \(\alpha=0.05\), and describe a Type I error.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean yields, in kilograms per plot, for the populations represented by the two groups. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\), at \(\alpha=0.05\).
Plan: Use an unpooled two-sample t test because the groups are independent and the response is quantitative. Assume plots were selected or assigned by an appropriate random process, observations within and between groups are independent, and each population is at least 250 plots so each 25-plot sample meets the 10% condition. Assume each group’s yield distribution has no extreme outliers or severe skew. For these equal sample sizes and standard deviations, the unpooled degrees of freedom are 48.
Do: The estimated difference is \(\bar{x}_1-\bar{x}_2=56-50=6\) kilograms per plot. Calculate its standard error and the t statistic:
For a two-sided test with \(df=48\) and \(\alpha=0.05\), the critical values are approximately \(-2.011\) and \(2.011\). Since \(2.121>2.011\), reject \(H_0\).
Conclude: The data provide convincing evidence that the population mean yields differ between the two treatment groups. A Type I error would mean the true population mean yields were equal, but the test rejected that equality. The test compares the mean yields in the two groups; the decision itself still does not reveal whether the null hypothesis is true.
Common Mistakes and AP Exam Tips
- Saying alpha is the chance that \(H_0\) is true: Alpha describes the test’s long-run rejection rate when \(H_0\) is true and the procedure’s conditions hold. It is not a probability assigned to the null hypothesis.
- Claiming a particular rejection is definitely a Type I error: The true parameter is generally unknown. State the conditional description: “If the null hypothesis is true, rejecting it would be a Type I error.”
- Calling every rejection an error: A rejection is a Type I error only when the null hypothesis is true. If it is false, rejection is the intended decision when the data provide sufficient evidence against it.
- Confusing the p-value with alpha: The p-value is calculated from the sample under the assumption that \(H_0\) is true; alpha is the chosen decision threshold. As discussed in “Strength of Evidence on a Continuum,” the p-value measures evidence, while the alpha comparison determines the decision.
- Using a confidence interval without matching the test: The direct connection described here is for a two-sided t test and its matching confidence interval. Name the confidence level and test level when explaining why exclusion of the null value corresponds to rejection.
- Forgetting context and units: A complete explanation identifies the population mean or mean difference and the measurement units. For a two-sample test, identify which group’s mean is first so the direction of the estimated difference is clear.
A strong AP response separates the decision from the possible error. First state whether the test rejects or fails to reject based on the p-value and \(\alpha\). Then explain what a Type I error would mean in context, using conditional wording. Do not say that a rejection proves the alternative, or that a particular rejection is known to be false.
Check Your Understanding
Answer each question using the meaning of Type I error in mean inference.
- A one-sample t test of \(H_0:\mu=72\) rejects at \(\alpha=0.05\). What would have to be true for this decision to be a Type I error?
- In a valid test at \(\alpha=0.01\), what does the 0.01 describe when the null mean is true?
- A matching 95% confidence interval for a population mean excludes the null value. What decision does the corresponding two-sided test make at \(\alpha=0.05\)?
- Why can’t a researcher determine from a single rejected test alone whether a Type I error occurred?
- In a two-sample t test, the null hypothesis states that the population mean difference is zero. Describe a Type I error in terms of the two populations.