Tutorials › AP Statistics › Common Errors in One-Sample t Tests

One-sample t hypothesis tests · Tutorial 678 of 1000

Common Errors in One-Sample t Tests

Practice spotting and correcting common one-sample t test errors, from selecting the reference distribution and tail to writing a cautious conclusion in context.

Intermediate 9 min read

What You'll Learn

  • Explain why a one-sample t test uses a t distribution when the population standard deviation is unknown.
  • Calculate the correct degrees of freedom and recognize the effect of using the wrong value.
  • Match the p-value tail to the alternative hypothesis, even when the sample result points in the opposite direction.
  • Distinguish failing to reject a null hypothesis from proving that it is true.
  • Write test conclusions that state the decision and describe the evidence in context.

Small Choices Can Change a Test’s Conclusion

A one-sample t test has a clear sequence: state hypotheses, check conditions, calculate a test statistic, find the p-value for the stated alternative, and draw a conclusion in context. A mistake at any point can change the reported evidence—or even the test decision. This tutorial focuses on four errors: using a z distribution instead of t, using incorrect degrees of freedom, choosing the wrong tail, and claiming that a failure to reject proves the null hypothesis.

As covered in “The One-Sample t Test Statistic” and “Finding a P-Value Using tcdf,” the statistic compares the sample mean with the null mean in standard-error units. Because the population standard deviation \(\sigma\) is usually unknown and is estimated by the sample standard deviation \(s\), the reference distribution is a t distribution.

Formula: For a one-sample t test of \(H_0:\mu=\mu_0\), calculate \(t=(\bar{x}-\mu_0)/(s/\sqrt{n})\) and use \(df=n-1\). The alternative hypothesis determines which tail area—or areas—give the p-value.

Error 1: Using z Instead of t

A common shortcut is to treat a test statistic as a standard Normal z statistic, especially when the sample seems “large enough.” But if \(\sigma\) is unknown and the test uses \(s\), use the t distribution with \(df=n-1\). A t distribution has heavier tails than the standard Normal distribution, particularly when the degrees of freedom are small. As the degrees of freedom increase, the t distribution gets closer to the standard Normal distribution; that similarity does not change which procedure is appropriate.

Using z can make a p-value too small and may lead to rejecting \(H_0\) when the correct t test would fail to reject it. The example below shows this decision-changing error. “Effect of Sample Size on a One-Sample t Test” explains that the reference t distribution also changes with the degrees of freedom.

Error 2: Using the Wrong Degrees of Freedom

For a one-sample t test, the degrees of freedom are \(n-1\), not \(n\). They are also not the number of measurements minus two; that count belongs to a different setting. The degrees of freedom identify which t distribution to use for the p-value. “Why the t Test Uses n Minus 1 Degrees of Freedom” explains why one degree of freedom is lost when the sample mean is estimated.

An incorrect \(df\) can shift the p-value even if the t statistic and alternative are correct. That shift can matter especially when the p-value is near \(\alpha\). Always calculate \(df=n-1\) from the sample size for the test at hand; do not copy a value from another test or calculator screen without checking it.

Error 3: Choosing the Wrong Tail

The alternative hypothesis determines the direction of the p-value. For \(H_a:\mu<\mu_0\), use the area to the left of the observed t statistic. For \(H_a:\mu>\mu_0\), use the area to the right. For \(H_a:\mu\ne\mu_0\), count outcomes at least as far from zero in either direction. The observed sample result does not change the alternative after the data are collected.

If a study asks whether a mean is lower but the sample mean happens to be higher, the p-value is still the lower-tail probability specified by \(H_a\). It may be large, which is evidence against the alternative in the sense that the observed result is not in the direction claimed. Do not switch to the upper tail just because it gives a smaller number.

Error 4: Saying the Null Hypothesis Is True

A test decision is either reject \(H_0\) or fail to reject \(H_0\). If the p-value is greater than \(\alpha\), the data do not provide convincing evidence for the alternative at that significance level. They do not prove that the null value is correct. A sample may be consistent with \(H_0\) while also being compatible with other population means.

This distinction matters in the wording of a conclusion. Write “fail to reject \(H_0\)” rather than “accept \(H_0\),” “prove \(H_0\),” or “the population mean equals \(\mu_0\).” A careful conclusion describes what the data provide evidence for, or do not provide convincing evidence for, in context.

Worked Examples

Worked Example: A z Shortcut Changes the Decision

A fictional food-packaging facility randomly samples 25 containers from a production run of 500. The sample mean fill is 52 grams, and the sample standard deviation is 5 grams. Assume the fill amounts in the run are approximately Normally distributed. Test at \(\alpha=0.05\) whether the true mean fill differs from 50 grams.

State. Let \(\mu\) be the true mean fill, in grams, of containers in this production run. Test \(H_0:\mu=50\) grams against \(H_a:\mu\ne50\) grams at \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test for a population mean. The containers were randomly sampled, supporting the Random condition. For sampling without replacement, the 10% condition is met because \(25\leq0.10(500)=50\), supporting independence. The sample size is below 30, so the large-sample route is not met; however, the stated approximately Normal distribution supports the Normal/Large Sample condition. Thus, a one-sample t test is appropriate.

Do. The standard error is \(s/\sqrt{n}\), and the degrees of freedom are \(n-1\):

$$ SE=\frac{5}{\sqrt{25}}=1\text{ gram}, \qquad t=\frac{52-50}{5/\sqrt{25}}=\frac{2}{1}=2.0000, \qquad df=25-1=24. $$

For a two-sided alternative, the p-value is twice the upper-tail area beyond \(2.0000\) for a t distribution with 24 degrees of freedom. The p-value is \(0.05694\), rounded. Since \(0.05694>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean fill in this production run differs from 50 grams. It would be incorrect to replace the t distribution with the standard Normal distribution just because the calculation of the statistic looks similar. Doing so gives a two-sided p-value of about \(0.0455\), which is below \(0.05\) and would incorrectly change the decision to reject \(H_0\).

This example also shows why degrees of freedom must be checked. If \(df=9\) is used by mistake with the same statistic \(t=2.0000\), the two-sided p-value is about \(0.07655\), not \(0.05694\). The wrong degrees of freedom produce the wrong p-value even though the t statistic has not changed. Here, both t-based p-values lead to failing to reject at 0.05, but that will not always be true when a p-value is close to the cutoff.

Worked Example: Do Not Change the Tail to Match the Sample

A fictional transit team is interested in whether the mean time for a particular maintenance task is less than 12 minutes. A random sample of 9 tasks from a set of 200 has a mean time of 10 minutes and a standard deviation of 3 minutes. Assume the population distribution of task times is approximately Normal. Test the stated claim at \(\alpha=0.05\).

State. Let \(\mu\) be the true mean maintenance time, in minutes, for tasks in this setting. Test \(H_0:\mu=12\) minutes against \(H_a:\mu<12\) minutes at \(\alpha=0.05\). The lower-tail alternative is specified by the question and must be used even though the sample mean is below 12.

Plan and check conditions. Use a one-sample t test. Random sampling supports the Random condition. The 10% condition is met because \(9\leq0.10(200)=20\), supporting independence for sampling without replacement. Since \(n=9<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition.

Do. Calculate the standard error, t statistic, and degrees of freedom:

$$ SE=\frac{3}{\sqrt{9}}=1\text{ minute}, \qquad t=\frac{10-12}{3/\sqrt{9}}=-2.0000, \qquad df=9-1=8. $$

Because \(H_a:\mu<12\), use the area to the left of \(-2.0000\) for \(df=8\). The p-value is about \(0.0403\), rounded. Since \(0.0403<0.05\), reject \(H_0\).

Conclude. The sample provides convincing evidence that the true mean maintenance time is less than 12 minutes. A wrong-tail calculation would use the area to the right of \(-2.0000\), about \(0.9597\). That is not the p-value for the stated alternative. For comparison, the two-sided p-value would be about \(0.0805\), but a two-sided test answers a different question from the lower-tailed test specified here.

Worked Example: A Large P-Value Does Not Prove Equality

A fictional community garden randomly selects 16 tomato plants from a group of 800. Their mean height is 18 centimeters, and their sample standard deviation is 4 centimeters. Assume the plant heights in this group are approximately Normally distributed. Test whether the true mean height differs from 18 centimeters at \(\alpha=0.05\).

State. Let \(\mu\) be the true mean height, in centimeters, of tomato plants in this group. Test \(H_0:\mu=18\) centimeters against \(H_a:\mu\ne18\) centimeters at \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test. Random selection supports the Random condition. The 10% condition is met because \(16\leq0.10(800)=80\), supporting independence. Since \(n=16<30\), the large-sample route is not met; the stated approximately Normal distribution supports the Normal/Large Sample condition.

Do. The standard error and test statistic are

$$ SE=\frac{4}{\sqrt{16}}=1\text{ centimeter}, \qquad t=\frac{18-18}{4/\sqrt{16}}=0.0000, \qquad df=16-1=15. $$

For a two-sided test, the p-value is the probability of a t statistic at least as far from zero as 0, in either direction. Every possible t statistic is at least 0 units from zero, so the p-value is \(1.0000\). Because \(1.0000>0.05\), fail to reject \(H_0\).

Conclude. This sample does not provide convincing evidence that the true mean height differs from 18 centimeters. It does not prove that the true mean height is exactly 18 centimeters. The appropriate conclusion is about the strength of evidence in this sample, not certainty about the population mean.

Common Mistakes and AP Exam Tips

When reviewing a written solution, check the reasoning in order. A correctly calculated statistic cannot rescue a p-value found from the wrong distribution or tail, and a correct p-value does not justify an overconfident conclusion.

  • Using z because the sample is not small. If the population standard deviation is unknown and \(s\) is used, use a t distribution with \(df=n-1\). A large sample makes t more Normal-like; it does not change the procedure.
  • Using \(df=n\) or carrying over another test’s df. Write \(df=n-1\) beside the statistic before finding the p-value. For the 25-container example, the correct value is 24.
  • Picking a tail after looking at the sample mean. Choose the tail from \(H_a\), not from whichever direction makes the result look more significant. A lower alternative always uses the left-tail area.
  • Doubling every p-value—or never doubling one. Double a one-tail area for a two-sided test when using a symmetric t distribution. Do not double for a one-sided alternative.
  • Writing “accept” or “prove” \(H_0\). When \(p>\alpha\), say “fail to reject \(H_0\)” and state that the data do not provide convincing evidence for \(H_a\). Do not claim that the null value has been established.
  • Leaving the conclusion out of context. A full-credit conclusion names the decision and describes evidence about the population mean using the variable, population, and direction relevant to the question.
  • Skipping the conditions because the calculator gave an answer. Check randomness, independence (including the 10% condition when applicable), and the Normal/Large Sample condition. A calculator output does not verify the study design or the data shape.
Key takeaway: For a one-sample t test using \(s\), use the t distribution with \(df=n-1\), and use the tail specified by \(H_a\). Reject or fail to reject based on the p-value and \(\alpha\); failing to reject does not prove \(H_0\) true.

Check Your Understanding

For each item, identify the error or explain why the reasoning is appropriate.

  1. A one-sample test uses \(n=18\), with \(\sigma\) unknown and \(s\) reported. What degrees of freedom should be used, and why is using \(df=18\) an error?
  2. A test has \(H_a:\mu>40\), but the observed t statistic is negative. Which tail gives the p-value? Should the alternative be changed after seeing the statistic?
  3. For a two-sided test, a student uses the standard Normal distribution because \(n=80\). Explain why this is not the standard one-sample t procedure when \(\sigma\) is unknown.
  4. A test has \(p=0.18\) and \(\alpha=0.05\). Write an appropriate decision and a cautious conclusion in context, without claiming that \(H_0\) is true.
  5. For \(H_a:\mu\ne\mu_0\), a student reports just the upper-tail area beyond \(|t|\). What adjustment is needed, and why?