Tutorials › AP Statistics › Worked Example: Two-Sided Test for a Mean

One-sample t hypothesis tests · Tutorial 674 of 1000

Worked Example: Two-Sided Test for a Mean

Learn how a two-sided alternative determines the test setup and how to calculate and interpret the doubled tail area.

Intermediate 10 min read

What You'll Learn

  • State hypotheses for testing whether a population mean differs from a target.
  • Explain why a two-sided p-value counts extreme results in both directions.
  • Calculate a one-sample t statistic and use \(df=n-1\).
  • Double the appropriate tail area to find a two-sided p-value.
  • Check conditions and write a conclusion in context.
  • Avoid choosing the alternative based on the observed sample mean.

Testing Whether a Mean Differs From a Target

In “Worked Example: Testing a Claim of a Higher Mean,” the question was specifically whether a population mean exceeded a stated value, so the p-value came from one tail. Here, the question is broader: could the population mean be either higher or lower than the target? That requires a two-sided alternative and a p-value that accounts for extreme results in both directions.

Let \(\mu_0\) be the target value. For a two-sided test, the hypotheses are \(H_0:\mu=\mu_0\) and \(H_a:\mu\ne\mu_0\). As in “The One-Sample t Test Statistic,” calculate \(t=(\bar{x}-\mu_0)/(s/\sqrt{n})\), and use \(df=n-1\). The sign of \(t\) tells whether the sample mean is above or below the target; the two-sided p-value reflects how far the sample mean is from the target in either direction.

Definition: For a two-sided one-sample t test, the p-value is the probability, assuming \(H_0\) is true, of obtaining a t statistic at least as far from 0 as the observed statistic, in either direction. With a symmetric t distribution, this equals twice the tail area beyond \(|t|\).

If the observed t statistic is positive, the relevant tail area is to the right of \(t\), and the equally extreme area in the other direction is to the left of \(-t\). If the observed statistic is negative, use the area to the left of \(t\) and double it. In either case, the calculation can be written using the positive value \(|t|\):

$$ \text{two-sided p-value} =2P(T\geq |t|) =2P(T\leq -|t|), \qquad df=n-1. $$

The doubling works because the t distribution under the null hypothesis is symmetric around 0. It does not mean that every two-sided p-value is twice the p-value from any other test: it is twice the single tail area beyond the observed statistic’s absolute value. A two-sided alternative must also match the question being asked, not be chosen after seeing which direction the sample mean happens to go.

The Four Steps for a Two-Sided t Test

The four-step structure used in earlier mean-inference tutorials still applies. In the Plan step, name the one-sample t test and check its conditions. In the Do step, find the statistic and double the correct tail area. In the Conclude step, compare the p-value with \(\alpha\) and describe the evidence for a difference in context.

1
State.
Define \(\mu\) in context, write \(H_0:\mu=\mu_0\) and \(H_a:\mu\ne\mu_0\), and state the significance level \(\alpha\).
2
Plan and check conditions.
Name the one-sample t test. Check the Random condition, the 10% condition when sampling without replacement, and the Normal/Large Sample condition.
3
Do.
Calculate \(t\), use \(df=n-1\), and find twice the tail area beyond \(|t|\).
4
Conclude.
Compare the two-sided p-value with \(\alpha\), state whether you reject or fail to reject \(H_0\), and explain what the evidence says about the population mean in context.

As emphasized in “Why Conditions Matter in Mean Inference,” a calculator can produce a p-value even when the study design or data shape does not support the inference. Check the evidence for each condition before treating the result as a sound basis for a population claim.

Worked Examples

Worked Example: Is a Diagnostic Restart Time Different From Its Target?

A fictional technology team randomly selects 16 routers from a production batch of 240. The routers’ mean diagnostic restart time is \(\bar{x}=52\) seconds, with sample standard deviation \(s=3.2\) seconds. Assume restart times in this batch follow an approximately Normal distribution. At \(\alpha=0.05\), test whether the true mean restart time differs from the 50-second target.

State. Let \(\mu\) be the true mean diagnostic restart time, in seconds, for all routers in this production batch. The target is 50 seconds, so the hypotheses are \(H_0:\mu=50\) seconds and \(H_a:\mu\ne50\) seconds. Use \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test for a population mean. The routers were randomly selected, supporting the Random condition. For the 10% condition, \(0.10(240)=24\), and \(16\leq24\); the sample is no more than 10% of the batch, so independence is reasonable. Since \(n=16<30\), the large-sample route is not met. The stated approximately Normal population supports the Normal/Large Sample condition. The degrees of freedom are \(df=16-1=15\).

Do. Calculate the standard error and test statistic:

$$ \frac{s}{\sqrt{n}} =\frac{3.2}{\sqrt{16}} =\frac{3.2}{4} =0.8\text{ seconds}, \qquad t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{52-50}{3.2/\sqrt{16}} =\frac{2}{0.8} =2.50. $$

The observed statistic is positive, so the one-tail area beyond \(|t|=2.50\) is to the right. With 15 degrees of freedom, that area is \(\operatorname{tcdf}(2.50,1\mathrm{E}99,15)\approx0.01225\). The two-sided p-value is twice that tail area:

$$ p=2(0.0122529)\approx0.0245058\approx0.02451. $$

The p-value is about 0.02451, rounded to five decimal places. Since \(0.02451<0.05\), reject \(H_0\).

Conclude. The sample provides convincing evidence that the true mean diagnostic restart time for routers in this batch differs from 50 seconds.

The sample mean is above 50 seconds, but the two-sided conclusion is about a difference in either direction. The p-value is calculated assuming the population mean is 50 seconds; it is not the probability that \(H_0\) is true.

Worked Example: A Lower Sample Mean and a Two-Sided Test

A fictional farm technician randomly selects 25 soil plots from a set of 400 plots to examine soil pH. The sample mean is \(\bar{x}=6.1\) and the sample standard deviation is \(s=1.0\). Assume the pH values follow an approximately Normal distribution. Test whether the true mean pH differs from the target value of 6.5, using \(\alpha=0.05\).

State. Let \(\mu\) be the true mean soil pH for all 400 plots in this set. The hypotheses are \(H_0:\mu=6.5\) and \(H_a:\mu\ne6.5\), with \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test. The plots were randomly selected, supporting the Random condition. For the 10% condition, \(0.10(400)=40\), and \(25\leq40\), so the sample is no more than 10% of the plots. Since \(n=25<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. Use \(df=25-1=24\).

Do. The estimated standard error and test statistic are

$$ \frac{s}{\sqrt{n}} =\frac{1.0}{\sqrt{25}} =\frac{1.0}{5} =0.2, \qquad t=\frac{6.1-6.5}{1.0/\sqrt{25}} =\frac{-0.4}{0.2} =-2.00. $$

Because this is a two-sided test, use the tail area beyond \(|t|=2.00\), then double it. For \(df=24\), \(\operatorname{tcdf}(2.00,1\mathrm{E}99,24)\approx0.0285\), so

$$ p=2(0.0285)\approx0.0570. $$

The two-sided p-value is approximately 0.0570, rounded to four decimal places. Since \(0.0570>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean soil pH for these plots differs from 6.5.

The sample mean is below 6.5, but the alternative was not “lower than 6.5.” The two-sided test counts evidence for a difference in either direction, so the same procedure would apply if the sample mean had been equally far above 6.5.

Worked Example: A Small Sample With a Difference in the Lower Direction

A fictional school randomly selects 9 laptops from an inventory of 90 and records the time each takes to start. The sample mean is \(\bar{x}=19.2\) seconds, and the sample standard deviation is \(s=1.2\) seconds. Assume laptop startup times in this inventory are approximately Normally distributed. At \(\alpha=0.05\), test whether the true mean startup time differs from 20 seconds.

State. Let \(\mu\) be the true mean startup time, in seconds, for all laptops in this inventory. The hypotheses are \(H_0:\mu=20\) seconds and \(H_a:\mu\ne20\) seconds. Set \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test. The 9 laptops were randomly selected, supporting the Random condition. For the 10% condition, \(0.10(90)=9\), and \(9\leq9\), so the sample is no more than 10% of the inventory. Since \(n=9<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. The degrees of freedom are \(df=9-1=8\).

Do. The standard error and observed t statistic are

$$ \frac{s}{\sqrt{n}} =\frac{1.2}{\sqrt{9}} =\frac{1.2}{3} =0.4\text{ seconds}, \qquad t=\frac{19.2-20}{1.2/\sqrt{9}} =\frac{-0.8}{0.4} =-2.00. $$

For a two-sided test, double the area to the left of \(-2.00\), or equivalently double the area to the right of \(2.00\). With \(df=8\), the one-tail area is approximately 0.0403:

$$ p=2\operatorname{tcdf}(2.00,1\mathrm{E}99,8) \approx2(0.0403) \approx0.0805. $$

The two-sided p-value is approximately 0.0805, rounded to four decimal places. Since \(0.0805>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean startup time for laptops in this inventory differs from 20 seconds.

A t statistic of \(-2.00\) is two standard errors below the target. The two-sided p-value also includes outcomes at least two standard errors above the target, which is why the relevant area is doubled.

Common Mistakes and AP Exam Tips

A full-credit response keeps the hypotheses, p-value calculation, decision, and conclusion aligned. For a two-sided test, make clear that the evidence is being assessed for departures above or below the null value.

  • Using only one tail. A two-sided alternative \(H_a:\mu\ne\mu_0\) requires counting results at least as extreme in either direction. Find the tail beyond \(|t|\) and double it.
  • Doubling the wrong area. Do not double the area between 0 and \(t\), or the larger tail area. Double the small tail beyond the observed statistic’s absolute value.
  • Ignoring the sign without explaining why. The sign indicates the direction of the sample’s difference, but the two-sided p-value is based on \(|t|\) because equally extreme outcomes on either side count.
  • Choosing the alternative after seeing the data. The question determines whether the alternative is two-sided or one-sided. A lower sample mean does not by itself make the test lower-tailed.
  • Using \(s\) instead of \(s/\sqrt{n}\). The test statistic uses the estimated standard error of the sample mean, not the sample standard deviation by itself.
  • Forgetting the degrees of freedom. For a one-sample t test, use \(df=n-1\). The tail area depends on the degrees of freedom as well as the test statistic.
  • Writing “accept the null.” If \(p>\alpha\), say “fail to reject \(H_0\).” This does not prove that the population mean equals the target.
  • Giving a conclusion without context. State whether the sample provides convincing evidence that the named population mean differs from the target. Do not turn a conclusion about a mean into a claim about every individual.
  • Misinterpreting the p-value. The p-value is calculated assuming \(H_0\) is true. It is not the probability that the null hypothesis is true or the probability that the observed sample mean was caused by chance.
Key takeaway: For \(H_0:\mu=\mu_0\) versus \(H_a:\mu\ne\mu_0\), calculate \(t=(\bar{x}-\mu_0)/(s/\sqrt{n})\), use \(df=n-1\), and double the tail area beyond \(|t|\). Compare that two-sided p-value with \(\alpha\), then state the evidence for a difference in context.

Check Your Understanding

Use \(\alpha=0.05\) for each question. Assume any unstated t test conditions are supported unless the question asks you to evaluate them.

  1. For \(H_0:\mu=12\) versus \(H_a:\mu\ne12\), a sample produces \(t=1.80\) with \(df=10\). Which tail area should be doubled?
  2. A random sample of 18 observations is drawn without replacement from a population of 250. Does the 10% condition hold? Show the comparison.
  3. A sample has \(n=16\), \(\bar{x}=48\), and \(s=4\). For a test of \(H_0:\mu=50\), calculate the standard error and t statistic.
  4. A two-sided test gives \(p=0.032\). State the decision and explain what the result means in context, assuming the test concerns whether a population mean differs from its target.
  5. Why does a two-sided test with observed \(t=-2.00\) also count outcomes at least \(2.00\) units above zero on the t scale?