The Logic: Assume, Observe, Compare
In Exam Practice on Conditions for One-Proportion Inference, you checked whether a study’s data and sampling process support a one-proportion procedure. Now we turn to the reasoning behind a significance test. The central question is not simply whether a sample proportion differs from a claimed value. It is whether a difference this large, or larger, would be surprising if the claim were true.
A significance test starts with a claim about a population proportion. For the test’s reasoning, treat that claim as true for the moment. Use it to describe what results would be expected from the chance process that produced the data. Then compare the actual result with those expected results. If the actual result would be unusual under the claim, it provides evidence against that claim.
The claim treated as true during this reasoning is called the null hypothesis, often written \(H_0\). In this tutorial, we focus on the logic of the test; a later tutorial develops how to write null and alternative hypotheses for a proportion. The null hypothesis is a working assumption for calculating what outcomes would be expected. It is not a conclusion that the claim has already been proved.
For a one-proportion population claim, data should come from an appropriate random sample. In an experiment, random assignment supports inference about treatment effects, but by itself does not justify generalizing to a population. As discussed in Conditions for a Survey With Nonresponse and Observational Data and Limits on Generalization, how the data are collected matters: a large but self-selected group does not automatically represent a population. A surprising result from a biased process is not a reliable basis for a population conclusion.
How a P-Value Describes Surprise
The p-value measures how surprising the observed data are under the null hypothesis. More precisely, it is the probability, assuming the null hypothesis is true and the test’s model and conditions apply, of obtaining a result at least as extreme as the one observed in the direction or directions considered by the test.
“More extreme” is defined in relation to the question being asked. If the concern is that a proportion is higher than the null claim, results with an even higher sample proportion count as more extreme. If the concern is that the proportion differs in either direction, results at least as far from the null value in either direction count. The precise comparison depends on the test.
A small p-value is evidence against the null claim because the data would be unusual if that claim were true. A large p-value means the data are not especially unusual under the null model. It does not mean the null claim is probably true, and a small p-value does not give the probability that the null claim is true. A p-value is conditional on the null assumption; it is not a probability about which hypothesis is correct.
Before looking at the data, a researcher often chooses a significance level, written \(\alpha\). It is the cutoff used to decide whether the p-value is small enough to call the result statistically significant. For example, at \(\alpha=0.05\), a p-value at or below 0.05 leads to rejecting the null hypothesis; a p-value above 0.05 leads to failing to reject it. The significance level is a decision rule, not a line that separates certainty from uncertainty.
Use the claimed population proportion to describe the chance process that would produce the data if the claim were true.
Observe a statistic, such as the sample proportion \(\hat{p}\), from a suitable random process.
Find the chance, assuming the null claim, of getting the observed result or a result at least as extreme. This chance is the p-value.
Compare the p-value with the chosen \(\alpha\). Reject the null if the p-value is small enough; otherwise, fail to reject it.
Describe what the result suggests about the claim. A test does not prove either that the null is true or that it is false.
Worked Examples
Worked Example: Is a Coin Fair?
A student tosses a coin 10 times and gets 9 heads. Could results like this be convincing evidence against the claim that the coin is fair? Use a two-sided comparison: results at least as far from 5 heads as the observed result count as equally or more extreme.
State: Let \(p\) be the long-run proportion of heads for this coin. The null claim is \(H_0:p=0.50\). We will assess whether 9 heads in 10 tosses is unusual if the coin has a 0.50 chance of heads on each toss.
Plan: Assume the coin is fair and the tosses are independent. Let \(X\) be the random variable counting heads in 10 tosses. Under the null claim, each toss has probability 0.50 of heads, so we can count the outcomes at least as far from the expected count of 5 as 9 is. This is an exact chance calculation for the coin-toss model, not a one-proportion \(z\)-test; with only 10 tosses, the \(z\)-test’s expected-count condition would not be met.
Do: Nine heads is 4 away from 5. Outcomes at least that far away are \(X\leq1\) or \(X\geq9\). There are \(2^{10}=1{,}024\) equally likely sequences of 10 fair tosses. The number of sequences with 0 or 1 head is \(1+10\); by symmetry, there are also \(10+1\) sequences with 9 or 10 heads. Thus:
Conclude: If the coin were fair and the tosses independent, the chance of getting a result at least this far from 5 heads, in either direction, would be about 0.0215. At a significance level of 0.05, that is statistically significant evidence against the fair-coin claim. It does not prove the coin is unfair: unusual results can occur by chance, and a test does not tell us what caused this result.
Worked Example: A Defect Claim and a Result That Is Not Very Surprising
A quality engineer is assessing a claim that 20% of packages from a production process are defective. Imagine checking 10 independently produced packages and finding 4 defective. For this illustration, assume the null claim is true and each package has an independent 0.20 chance of being defective. Is 4 or more defects surprising evidence that the defect proportion is higher than claimed?
State: Let \(p\) be the process’s long-run proportion of defective packages. The null claim is \(H_0:p=0.20\). The observed result is 4 defective packages out of 10.
Plan: The question asks whether the result is unusually high, so count outcomes of 4 or more defects. Let \(X\) be the random variable counting defective packages. Under the stated independent-trial model and null claim, calculate \(P(X\geq4)\). This exact-model illustration is not a one-proportion \(z\)-test.
Do: The chance of 4 or more defects is one minus the chance of 3 or fewer. A calculator’s binomial cumulative probability gives:
Conclude: If the defect proportion were 0.20 and the independent-trial model applied, a result of 4 or more defects would occur about 12.09% of the time. This is not especially unusual at the 0.05 significance level, so the result does not provide convincing evidence that the defect proportion is higher than 0.20. It also does not establish that the claim is true; this small sample may simply provide limited evidence.
Worked Example: A Large P-Value Does Not Prove a Claim
Suppose a company claims that half of its app users prefer a new screen design. In a hypothetical random sample of 10 users, 6 prefer the new design. To illustrate the logic, assume independent responses and compare outcomes at least as far from 5 preferences as the observed count.
State: Let \(p\) be the proportion of the company’s app users who prefer the new design. The null claim is \(H_0:p=0.50\). The observed count is 6 out of 10.
Plan: Six is 1 away from the null model’s expected count of 5. For a two-sided comparison, results at least as far away are \(X\leq4\) or \(X\geq6\). Under the assumed independent-response model with \(p=0.50\), every sequence is equally likely.
Do: There are \(2^{10}=1{,}024\) possible response sequences. Exactly 252 have 5 preferences, since there are \(\binom{10}{5}=252\) ways to choose which 5 users prefer the design. Every other result is at least 1 away from 5:
Conclude: A result at least this far from the null expectation would be common if the proportion were 0.50. Six preferences out of 10 therefore do not provide convincing evidence against the claim in this illustration. The large p-value does not show that exactly half of all users prefer the design; it says only that this sample result is not unusual under the null model. Also, because the null expected counts are \(10(0.50)=5\) and \(10(0.50)=5\), this small sample does not satisfy the Large Counts condition for a one-proportion \(z\)-test.
Common Mistakes and What Full Credit Requires
- Claiming the p-value is the probability that the null is true. It is the probability of results at least as extreme as the observed data, assuming the null model is true—not the probability that the null claim is correct.
- Describing a p-value without the null assumption. A complete interpretation says “assuming the null claim is true” and identifies the result being counted as equally or more extreme.
- Thinking a small p-value proves the null is false. A small p-value provides evidence against the null claim. It does not rule out chance variation or identify the reason for the result.
- Treating “fail to reject” as “accept” or “prove.” If the p-value is not small, say there is not convincing evidence against the null. Do not say the null has been proved or that there is no difference.
- Calling every difference statistically significant. A sample proportion can differ from the null value just by chance. The test asks whether the result is unusual under the null, not merely whether the two numbers are unequal.
- Forgetting the data-collection process. The test’s reasoning relies on a suitable random process and an appropriate model. As in the earlier condition-check tutorials, a small p-value cannot repair biased selection or an unmet condition.
- Confusing statistical significance with practical importance. A result can be statistically significant but small in real-world terms. The p-value measures how unusual the data are under the null model; it does not measure the size or importance of an effect.
Key Takeaway
A significance test asks whether chance variation under a null claim could reasonably produce the data observed. The null claim supplies the model; the data supply the evidence; and the p-value describes how unusual that evidence would be if the null model were true.
Check Your Understanding
Answer each question using the logic of assuming a null claim and comparing the data with results expected under it.
- In your own words, why does a significance test begin by assuming the null claim is true?
- A p-value is 0.03. Explain what this probability means if the null claim is true. What does it not mean?
- At significance level 0.05, a test has p-value 0.12. Give the appropriate decision and explain why it does not prove the null claim.
- Why must “at least as extreme” be defined in relation to the question being asked?
- A large voluntary-response poll produces a very small p-value. Explain why that result alone does not ensure a sound conclusion about the whole population.