Tutorials › AP Statistics › The Logic of a Significance Test for a Mean

One-sample t hypothesis tests · Tutorial 664 of 1000

The Logic of a Significance Test for a Mean

See how a significance test uses the null hypothesis as a starting assumption and the p-value to assess how surprising the sample mean is.

Intermediate 9 min read

What You'll Learn

  • Explain what a significance test assumes about the population mean while calculating a p-value.
  • Describe a p-value as a probability about results from repeated sampling under the null hypothesis.
  • Match the tail area used for a p-value to a lower-sided, upper-sided, or two-sided alternative.
  • Distinguish a small p-value from the probability that the null hypothesis is true.
  • Use a p-value and a stated significance level to reach a cautious conclusion in context.
  • Identify why statistical significance does not by itself measure practical importance.

A Significance Test Starts by Assuming the Null Hypothesis

In “Defining the Parameter \(\mu\) in Context,” you identified the population mean a test is about. In “Stating Hypotheses for a Population Mean” and “One-Sided Versus Two-Sided Alternatives for a Mean,” you wrote a null hypothesis and chose an alternative. Now consider the central question of a significance test: If the null hypothesis were true, how surprising would a sample result like this one be?

A test begins by treating \(H_0\) as true. That is a working assumption for measuring how unusual the sample evidence would be under the null model; it is not a declaration that the null is known to be true. If results that are at least as extreme as the observed result would be rare under that assumption, the data provide evidence against \(H_0\). If such results would not be rare, the data do not give strong evidence against it.

Definition: A significance test assesses how well sample data agree with a null hypothesis. It calculates a p-value: assuming \(H_0\) is true, the probability of obtaining a result at least as extreme as the observed result in the direction or directions specified by \(H_a\).

A sample mean \(\bar{x}\) varies from sample to sample, even when the population mean really equals the null value \(\mu_0\). The random variable \(X\) here represents the sample mean that would result from taking a sample according to the study’s method. Under \(H_0\), the center of its sampling distribution is \(\mu_0\). The observed \(\bar{x}\) is one outcome from that process, not the population mean itself.

The p-value describes a probability about possible sample results under the assumption \(H_0\) is true. It is not the probability that \(H_0\) is true, and it is not the probability that chance alone caused the observed result. A small p-value says that the observed result, or a more extreme one in the relevant direction, would be unusual under the null model. It does not identify the cause of the result.

“At Least as Extreme” Depends on the Alternative

The alternative hypothesis tells us what kind of departure from \(\mu_0\) counts as evidence. The p-value includes outcomes at least as extreme as the observed result in the direction specified by \(H_a\). For a two-sided alternative, departures in either direction count. This is why the alternative must be chosen from the question of interest, not selected after seeing which direction the sample mean went.

Alternative hypothesisResults counted for the p-valuePlain-language question
\(H_a:\mu<\mu_0\)Sample means at least as low as the observed \(\bar{x}\)If \(H_0\) were true, how often would a result this low or lower occur?
\(H_a:\mu>\mu_0\)Sample means at least as high as the observed \(\bar{x}\)If \(H_0\) were true, how often would a result this high or higher occur?
\(H_a:\mu\ne\mu_0\)Results at least as far from \(\mu_0\) as the observed result, in either directionIf \(H_0\) were true, how often would a result this far from the null value occur?

In a one-sample t test, the calculator expresses the observed result on a t scale and finds the appropriate tail area from a t distribution. The next tutorial develops the one-sample t test statistic. For now, the important idea is that the calculator’s reported p-value is the probability of a result at least as extreme as the observed one, assuming \(H_0\) is true and using the alternative to choose the tail area.

A p-value depends on both the observed sample result and how much sample-to-sample variation is expected. A sample mean that looks far from \(\mu_0\) in the original units may not be especially unusual if sample means commonly vary by that amount. Conversely, a smaller difference may be unusual when sample means typically vary very little. The p-value combines these ideas into a probability under the null model; it is not simply the numerical distance between \(\bar{x}\) and \(\mu_0\).

Key distinction: The null hypothesis supplies the assumption used to calculate the p-value. The alternative hypothesis determines which departures count as “at least as extreme.” The p-value then measures how often such results would occur if the null assumption were true.

How the p-Value Guides a Decision

A significance level, written \(\alpha\), is a cutoff chosen for deciding when a p-value is small enough to reject \(H_0\). For example, if \(\alpha=0.05\), a p-value no larger than 0.05 is considered statistically significant at the 5% level. The significance level should be set before examining the sample results; it is not the probability that the conclusion will be wrong in this particular test.

When \(p\leq\alpha\), reject \(H_0\). The data provide convincing evidence for the alternative claim, in context. When \(p>\alpha\), fail to reject \(H_0\). The data do not provide convincing evidence for the alternative claim. “Fail to reject” does not mean accept \(H_0\), prove it, or establish that the population mean equals \(\mu_0\). It means the sample has not supplied sufficiently strong evidence against the null at the chosen cutoff.

Decision rule: Compare the p-value with the stated significance level \(\alpha\). If \(p\leq\alpha\), reject \(H_0\); if \(p>\alpha\), fail to reject \(H_0\). In either case, describe the evidence in context and do not claim that a test proves a hypothesis.

Statistical significance and practical importance are also different. A very small p-value does not tell us that the difference from \(\mu_0\) is large enough to matter in the setting. A test answers a question about how compatible the data are with \(H_0\); judging practical importance also requires understanding the measurement, the size of the difference, and the consequences of that difference.

Worked Examples

Worked Example: Testing a Mean Fill Amount

A fictional beverage company claims its bottles contain a mean of 50 milliliters of concentrate. A technician takes a random sample of 25 bottles. The sample mean is 52 milliliters and the sample standard deviation is 10 milliliters. The question is whether the mean fill amount is greater than 50 milliliters. Use \(\alpha=0.05\). A calculator’s one-sample T-Test reports \(t=1.00\), \(df=24\), and \(p=0.1636\).

State. Let \(\mu\) be the true mean amount of concentrate, in milliliters, in bottles from this production process. The hypotheses are \(H_0:\mu=50\) milliliters and \(H_a:\mu>50\) milliliters. The upper-sided alternative matches the question asking whether the mean is greater than 50.

Plan and check conditions. The one-sample t test is appropriate if its conditions are supported. The bottles were randomly sampled, supporting the Random condition. The process produced 6,000 bottles that day, and \(25\leq0.10(6000)=600\), so the 10% condition supports independence when sampling without replacement. Since \(n=25<30\), the large-sample route is not met; the company reports that bottle fill amounts are approximately Normally distributed, which supports the Normal/Large Sample condition.

Do. The calculator’s output gives an observed standardized result of \(t=1.00\) with 24 degrees of freedom. Because the alternative is upper-sided, the p-value is the area to the right of 1.00:

$$ p=P(T_{24}\geq1.00)=\operatorname{tcdf}(1.00,1\text{E}99,24)=0.1636\text{ (rounded)}. $$

If the true mean were 50 milliliters, results at least this far above the null value in the direction of the alternative would occur about 16.36% of the time under the t model. Since \(0.1636>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean concentrate amount in bottles from this production process is greater than 50 milliliters. This does not establish that the mean is exactly 50 milliliters.

Worked Example: Testing Whether Packages Are Lighter

A fictional shipping supplier labels packages as having a mean mass of 78 ounces. A random sample of 16 packages has a mean mass of 74 ounces and a sample standard deviation of 8 ounces. The supplier wants to know whether the true mean package mass is below 78 ounces. The one-sample T-Test output is \(t=-2.00\), \(df=15\), and \(p=0.03197\). Use \(\alpha=0.05\).

State. Let \(\mu\) be the true mean mass, in ounces, of packages produced by this supplier during the period being examined. The hypotheses are \(H_0:\mu=78\) ounces and \(H_a:\mu<78\) ounces.

Plan and check conditions. The supplier selected the 16 packages at random, supporting the Random condition. The supplier produced 800 packages in the period, and \(16\leq0.10(800)=80\), so the 10% condition supports independence for sampling without replacement. Because \(n=16<30\), the large-sample route is not met. A quality report states that package masses are approximately Normally distributed with no pronounced outliers, supporting the Normal/Large Sample condition.

Do. The lower-sided alternative means the p-value is the area to the left of the observed \(t=-2.00\):

$$ p=P(T_{15}\leq-2.00)=\operatorname{tcdf}(-1\text{E}99,-2.00,15)=0.03197\text{ (rounded)}. $$

Assuming the true mean is 78 ounces, a result this low or lower would occur about 0.03197 of the time under the t model. Since \(0.03197<0.05\), reject \(H_0\).

Conclude. The data provide convincing evidence that the true mean mass of packages produced by this supplier during the period is below 78 ounces. This is evidence for the lower-mean alternative, not proof that every package is underweight.

Worked Example: A Two-Sided Question About Battery Life

A fictional tablet maker advertises a mean battery life of 20 hours under a specified test setting. An independent lab randomly selects 25 tablets from a shipment and measures battery life under that same setting. The sample mean is 21 hours and the sample standard deviation is 5 hours. The lab asks whether the true mean differs from 20 hours in either direction. The calculator’s one-sample T-Test reports \(t=1.00\), \(df=24\), and two-sided \(p=0.3273\). Use \(\alpha=0.05\).

State. Let \(\mu\) be the true mean battery life, in hours under the specified test setting, for tablets in this shipment. The hypotheses are \(H_0:\mu=20\) hours and \(H_a:\mu\ne20\) hours. Since either a higher or lower mean would matter to the question, the alternative is two-sided.

Plan and check conditions. The lab randomly selected the 25 tablets, supporting the Random condition. The shipment contained 1,200 tablets, and \(25\leq0.10(1200)=120\), supporting the 10% condition and independence for sampling without replacement. Here \(n=25<30\), so the large-sample route is not met. The lab’s production documentation describes battery-life measurements as approximately Normally distributed, supporting the Normal/Large Sample condition.

Do. For a two-sided alternative, outcomes at least as far from the null value as the observed result count in both tails. The calculator’s p-value is twice the area in the upper tail beyond \(t=1.00\):

$$ p=2P(T_{24}\geq1.00)=2\operatorname{tcdf}(1.00,1\text{E}99,24)=0.3273\text{ (rounded)}. $$

If the true mean were 20 hours, a result at least this far from 20 in either direction would occur about 32.73% of the time under the t model. Because \(0.3273>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean battery life for tablets in this shipment, under the specified test setting, differs from 20 hours. The result does not prove that the true mean is exactly 20 hours.

Common Mistakes and AP Exam Tips

A complete answer explains the logic behind the probability, not just the calculator result. These are common errors to avoid:

  • Calling the p-value the chance that \(H_0\) is true. A p-value is calculated assuming \(H_0\) is true. It is a probability of sample results under that assumption, not a probability assigned to the hypothesis.
  • Using the wrong tail. The alternative determines what “at least as extreme” means. For a lower-sided test, count low results; for an upper-sided test, count high results; for a two-sided test, count departures in either direction.
  • Choosing the alternative after seeing \(\bar{x}\). The research question determines the alternative. Choosing a one-sided alternative just because the sample mean went that way changes the test after seeing the data.
  • Thinking a large p-value proves the null. A large p-value means the data do not provide convincing evidence against \(H_0\) at the chosen significance level. It does not demonstrate that \(H_0\) is true.
  • Using “accept \(H_0\)” as the conclusion. AP-standard wording is “fail to reject \(H_0\).” Then explain whether the data provide convincing evidence for the alternative in context.
  • Confusing statistical significance with importance. A small p-value indicates incompatibility with the null model, not necessarily a large or consequential difference in the original units.
  • Leaving out the assumption in the interpretation. State that the probability is calculated assuming \(H_0\) is true. A p-value without that condition is not fully explained.

For full-credit communication, identify the direction or directions counted, include the assumption that \(H_0\) is true, compare the p-value with \(\alpha\), and state the conclusion about the population mean in context. For example: “Assuming the true mean package mass is 78 ounces, the probability of a sample result this low or lower is 0.03197. Since this is less than \(\alpha=0.05\), we reject \(H_0\). The data provide convincing evidence that the true mean package mass is below 78 ounces.”

Key takeaway: A significance test asks how surprising the sample result would be if \(H_0\) were true. The p-value counts results at least as extreme as the observed one in the direction or directions specified by \(H_a\). A small p-value can provide convincing evidence against \(H_0\), but it does not prove the alternative or measure practical importance.

Check Your Understanding

For each question, focus on the assumption behind the p-value and on what results count as at least as extreme.

  1. A test has \(H_a:\mu<\mu_0\). If the observed sample mean is below \(\mu_0\), which tail contains results counted for the p-value?
  2. In your own words, explain what a p-value of 0.03 means when calculated for a test of a population mean.
  3. A two-sided test has \(p=0.08\) and \(\alpha=0.05\). State the decision using AP-standard wording and explain what it does not prove.
  4. Why is it incorrect to say that a p-value of 0.02 means there is a 2% chance that \(H_0\) is true?
  5. A sample mean is only slightly above the null value, but its p-value is small. What does that suggest about the expected variation of sample means under \(H_0\)?