Tutorials › AP Statistics › The Logic of a Significance Test

One-proportion hypothesis tests · Tutorial 461 of 1000

The Logic of a Significance Test

Follow the reasoning of a significance test, from treating the null claim as a working model to deciding what surprising data do—and do not—show.

Intermediate 10 min read

What You'll Learn

  • Explain why a significance test begins by assuming the null claim is true.
  • Describe how a random process produces data to compare with that claim.
  • Define a p-value as the chance, under the null model, of results at least as extreme as those observed.
  • Distinguish statistically significant evidence from proof that a claim is false.
  • Explain why failing to reject a null claim does not establish that it is true.

The Logic: Assume, Observe, Compare

In Exam Practice on Conditions for One-Proportion Inference, you checked whether a study’s data and sampling process support a one-proportion procedure. Now we turn to the reasoning behind a significance test. The central question is not simply whether a sample proportion differs from a claimed value. It is whether a difference this large, or larger, would be surprising if the claim were true.

A significance test starts with a claim about a population proportion. For the test’s reasoning, treat that claim as true for the moment. Use it to describe what results would be expected from the chance process that produced the data. Then compare the actual result with those expected results. If the actual result would be unusual under the claim, it provides evidence against that claim.

Definition: A significance test assesses whether sample data provide convincing evidence against a stated null claim. It does this by assuming the null claim is true, considering the results that chance could produce under that assumption, and judging whether the observed result is unusually far from what the null predicts.

The claim treated as true during this reasoning is called the null hypothesis, often written \(H_0\). In this tutorial, we focus on the logic of the test; a later tutorial develops how to write null and alternative hypotheses for a proportion. The null hypothesis is a working assumption for calculating what outcomes would be expected. It is not a conclusion that the claim has already been proved.

For a one-proportion population claim, data should come from an appropriate random sample. In an experiment, random assignment supports inference about treatment effects, but by itself does not justify generalizing to a population. As discussed in Conditions for a Survey With Nonresponse and Observational Data and Limits on Generalization, how the data are collected matters: a large but self-selected group does not automatically represent a population. A surprising result from a biased process is not a reliable basis for a population conclusion.

How a P-Value Describes Surprise

The p-value measures how surprising the observed data are under the null hypothesis. More precisely, it is the probability, assuming the null hypothesis is true and the test’s model and conditions apply, of obtaining a result at least as extreme as the one observed in the direction or directions considered by the test.

Definition: A p-value is a probability calculated under the assumption that the null hypothesis is true. A small p-value means that the observed result, or a more extreme result, would be uncommon under that assumption.

“More extreme” is defined in relation to the question being asked. If the concern is that a proportion is higher than the null claim, results with an even higher sample proportion count as more extreme. If the concern is that the proportion differs in either direction, results at least as far from the null value in either direction count. The precise comparison depends on the test.

A small p-value is evidence against the null claim because the data would be unusual if that claim were true. A large p-value means the data are not especially unusual under the null model. It does not mean the null claim is probably true, and a small p-value does not give the probability that the null claim is true. A p-value is conditional on the null assumption; it is not a probability about which hypothesis is correct.

Before looking at the data, a researcher often chooses a significance level, written \(\alpha\). It is the cutoff used to decide whether the p-value is small enough to call the result statistically significant. For example, at \(\alpha=0.05\), a p-value at or below 0.05 leads to rejecting the null hypothesis; a p-value above 0.05 leads to failing to reject it. The significance level is a decision rule, not a line that separates certainty from uncertainty.

1
Assume the null claim.
Use the claimed population proportion to describe the chance process that would produce the data if the claim were true.
2
Collect the data.
Observe a statistic, such as the sample proportion \(\hat{p}\), from a suitable random process.
3
Measure how unusual the result is.
Find the chance, assuming the null claim, of getting the observed result or a result at least as extreme. This chance is the p-value.
4
Make a careful decision.
Compare the p-value with the chosen \(\alpha\). Reject the null if the p-value is small enough; otherwise, fail to reject it.
5
Explain the evidence in context.
Describe what the result suggests about the claim. A test does not prove either that the null is true or that it is false.

Worked Examples

Worked Example: Is a Coin Fair?

A student tosses a coin 10 times and gets 9 heads. Could results like this be convincing evidence against the claim that the coin is fair? Use a two-sided comparison: results at least as far from 5 heads as the observed result count as equally or more extreme.

State: Let \(p\) be the long-run proportion of heads for this coin. The null claim is \(H_0:p=0.50\). We will assess whether 9 heads in 10 tosses is unusual if the coin has a 0.50 chance of heads on each toss.

Plan: Assume the coin is fair and the tosses are independent. Let \(X\) be the random variable counting heads in 10 tosses. Under the null claim, each toss has probability 0.50 of heads, so we can count the outcomes at least as far from the expected count of 5 as 9 is. This is an exact chance calculation for the coin-toss model, not a one-proportion \(z\)-test; with only 10 tosses, the \(z\)-test’s expected-count condition would not be met.

Do: Nine heads is 4 away from 5. Outcomes at least that far away are \(X\leq1\) or \(X\geq9\). There are \(2^{10}=1{,}024\) equally likely sequences of 10 fair tosses. The number of sequences with 0 or 1 head is \(1+10\); by symmetry, there are also \(10+1\) sequences with 9 or 10 heads. Thus:

$$ P(X\leq1\text{ or }X\geq9) =\frac{1+10+10+1}{1{,}024} =\frac{22}{1{,}024} \approx 0.0215 $$

Conclude: If the coin were fair and the tosses independent, the chance of getting a result at least this far from 5 heads, in either direction, would be about 0.0215. At a significance level of 0.05, that is statistically significant evidence against the fair-coin claim. It does not prove the coin is unfair: unusual results can occur by chance, and a test does not tell us what caused this result.

Worked Example: A Defect Claim and a Result That Is Not Very Surprising

A quality engineer is assessing a claim that 20% of packages from a production process are defective. Imagine checking 10 independently produced packages and finding 4 defective. For this illustration, assume the null claim is true and each package has an independent 0.20 chance of being defective. Is 4 or more defects surprising evidence that the defect proportion is higher than claimed?

State: Let \(p\) be the process’s long-run proportion of defective packages. The null claim is \(H_0:p=0.20\). The observed result is 4 defective packages out of 10.

Plan: The question asks whether the result is unusually high, so count outcomes of 4 or more defects. Let \(X\) be the random variable counting defective packages. Under the stated independent-trial model and null claim, calculate \(P(X\geq4)\). This exact-model illustration is not a one-proportion \(z\)-test.

Do: The chance of 4 or more defects is one minus the chance of 3 or fewer. A calculator’s binomial cumulative probability gives:

$$ P(X\geq4) =1-P(X\leq3) =1-\operatorname{binomcdf}(10,0.20,3) \approx 1-0.8791 =0.1209 $$

Conclude: If the defect proportion were 0.20 and the independent-trial model applied, a result of 4 or more defects would occur about 12.09% of the time. This is not especially unusual at the 0.05 significance level, so the result does not provide convincing evidence that the defect proportion is higher than 0.20. It also does not establish that the claim is true; this small sample may simply provide limited evidence.

Worked Example: A Large P-Value Does Not Prove a Claim

Suppose a company claims that half of its app users prefer a new screen design. In a hypothetical random sample of 10 users, 6 prefer the new design. To illustrate the logic, assume independent responses and compare outcomes at least as far from 5 preferences as the observed count.

State: Let \(p\) be the proportion of the company’s app users who prefer the new design. The null claim is \(H_0:p=0.50\). The observed count is 6 out of 10.

Plan: Six is 1 away from the null model’s expected count of 5. For a two-sided comparison, results at least as far away are \(X\leq4\) or \(X\geq6\). Under the assumed independent-response model with \(p=0.50\), every sequence is equally likely.

Do: There are \(2^{10}=1{,}024\) possible response sequences. Exactly 252 have 5 preferences, since there are \(\binom{10}{5}=252\) ways to choose which 5 users prefer the design. Every other result is at least 1 away from 5:

$$ P(X\leq4\text{ or }X\geq6) =1-\frac{252}{1{,}024} =\frac{772}{1{,}024} \approx0.7539 $$

Conclude: A result at least this far from the null expectation would be common if the proportion were 0.50. Six preferences out of 10 therefore do not provide convincing evidence against the claim in this illustration. The large p-value does not show that exactly half of all users prefer the design; it says only that this sample result is not unusual under the null model. Also, because the null expected counts are \(10(0.50)=5\) and \(10(0.50)=5\), this small sample does not satisfy the Large Counts condition for a one-proportion \(z\)-test.

Common Mistakes and What Full Credit Requires

  • Claiming the p-value is the probability that the null is true. It is the probability of results at least as extreme as the observed data, assuming the null model is true—not the probability that the null claim is correct.
  • Describing a p-value without the null assumption. A complete interpretation says “assuming the null claim is true” and identifies the result being counted as equally or more extreme.
  • Thinking a small p-value proves the null is false. A small p-value provides evidence against the null claim. It does not rule out chance variation or identify the reason for the result.
  • Treating “fail to reject” as “accept” or “prove.” If the p-value is not small, say there is not convincing evidence against the null. Do not say the null has been proved or that there is no difference.
  • Calling every difference statistically significant. A sample proportion can differ from the null value just by chance. The test asks whether the result is unusual under the null, not merely whether the two numbers are unequal.
  • Forgetting the data-collection process. The test’s reasoning relies on a suitable random process and an appropriate model. As in the earlier condition-check tutorials, a small p-value cannot repair biased selection or an unmet condition.
  • Confusing statistical significance with practical importance. A result can be statistically significant but small in real-world terms. The p-value measures how unusual the data are under the null model; it does not measure the size or importance of an effect.
AP Exam Tip: Write the p-value interpretation conditionally and in context: “Assuming the null claim is true, the probability of getting [the observed result] or a result at least as extreme is [p-value].” Then state whether that probability gives convincing evidence against the claim. Do not describe the p-value as the probability that the null is true.

Key Takeaway

A significance test asks whether chance variation under a null claim could reasonably produce the data observed. The null claim supplies the model; the data supply the evidence; and the p-value describes how unusual that evidence would be if the null model were true.

Key takeaway: A small p-value can provide convincing evidence against a null claim, but it does not prove the claim false. A large p-value means the data are not unusually inconsistent with the null; it does not prove the claim true. Always connect the test’s conclusion to the study’s population and data-collection process.

Check Your Understanding

Answer each question using the logic of assuming a null claim and comparing the data with results expected under it.

  1. In your own words, why does a significance test begin by assuming the null claim is true?
  2. A p-value is 0.03. Explain what this probability means if the null claim is true. What does it not mean?
  3. At significance level 0.05, a test has p-value 0.12. Give the appropriate decision and explain why it does not prove the null claim.
  4. Why must “at least as extreme” be defined in relation to the question being asked?
  5. A large voluntary-response poll produces a very small p-value. Explain why that result alone does not ensure a sound conclusion about the whole population.