A Large P-Value Does Not Prove the Null
In Failing to Reject the Null Hypothesis Correctly, we learned that when a p-value is greater than the preselected significance level, we fail to reject \(H_0\). That decision describes what the sample evidence supports; it does not certify that the null hypothesis is true.
This distinction matters especially with a small sample. A small sample often has substantial sample-to-sample variability, so even results that would be unusual if the null were true may not be unusual enough to meet the test’s rejection rule. A real difference from the null value can therefore go undetected.
A significance test does not compare the probabilities that \(H_0\) and \(H_a\) are true. As discussed in What a P-Value Really Measures, a p-value is calculated assuming \(H_0\) is true. It measures how often the null model would produce results at least as extreme as the observed result, according to the alternative hypothesis. It is not the probability that \(H_0\) is true.
A small sample gives us a useful way to see the issue. For a one-proportion \(z\)-test, the Large Counts condition may not hold, so we should not use the Normal approximation for the test. Instead, the examples below use binomial probabilities to illustrate how much—or how little—a small sample can tell us. These binomial-test examples are an optional extension, not one of the standard inference procedures expected on the AP Statistics exam. A binomial model requires a fixed number of trials, two possible outcomes on each trial, a common probability of success, and independent trials. When sampling without replacement, the 10% condition supports treating trials as approximately independent.
A Small Sample Can Leave Many Outcomes Plausible
Suppose we are testing whether a population proportion is greater than a benchmark. With a small sample, the observed number of successes can vary quite a bit just by chance. The binomial model lets us calculate how often a result as high as the observed one would occur if the null proportion were correct.
For a right-tailed question, “at least as extreme” means the observed number of successes or more. If that probability is greater than \(\alpha\), the result does not meet the chosen standard for rejecting \(H_0\). But that only tells us that the result is not sufficiently unusual under the null model; it does not tell us that the null model is true.
Worked Examples: What a Small Sample Can and Cannot Show
Worked Example: A New Checkout Option
A fictional shop is considering whether more than half of its customers would choose a new self-checkout feature. A random sample of 10 customers from a list of 5,000 includes 6 who say they would choose it. The shop set \(\alpha=0.05\) before collecting the sample. Use a binomial-model tail probability to assess the result and explain what it does—and does not—show.
State: Let \(p\) be the true proportion of customers on this shop’s list who would choose the feature. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\).
Plan: The customers were randomly sampled, meeting the Random condition. The sample was drawn without replacement from 5,000 customers, and \(10\leq0.10(5{,}000)=500\), so the 10% condition is met. The number of trials, 10, is fixed; each trial has two outcomes, and the success probability is treated as common. The 10% condition supports treating the trials as approximately independent. Under \(H_0\), the expected success and failure counts are \(10(0.50)=5\) and \(10(0.50)=5\), respectively. Both are below 10, so the Large Counts condition for a one-proportion \(z\)-test is not met. We should not use a one-proportion \(z\)-test here. Instead, we can use the binomial model to calculate the relevant tail probability directly.
Do: If \(H_0\) is true, the random variable \(X\), the number of sampled customers who would choose the feature, is approximately binomial with \(n=10\) and \(p=0.50\). Since \(H_a\) is right-tailed and 6 customers chose the feature, the p-value is \(P(X\geq6)\):
The p-value is greater than \(0.05\), so this result does not provide convincing evidence against \(H_0\). The decision is to fail to reject \(H_0\).
Conclude: At the 0.05 significance level, the sample does not provide convincing evidence that more than half of the customers on this shop’s list would choose the new self-checkout feature.
That conclusion does not mean the proportion is 50%. To see why, suppose for illustration that the true proportion were actually 70%. With \(n=10\), the probability of getting 9 or 10 successes under that proportion is about \(0.1493\). Under the null proportion of 50%, the probability of 9 or 10 successes is about \(0.0107\). A test using the binomial tail probability at the 0.05 level would reject for 9 or 10 successes, but not for 8: under \(H_0\), \(P(X\geq8)=56/1024\approx0.0547\). So even if the true proportion were 70%, the sample would fail to reject about \(1-0.1493=0.8507\), or 85.07%, of the time. A small sample can often fail to detect a real difference.
Worked Example: Text Alerts in a Small Town
A fictional town’s emergency office wants to know whether more than 20% of residents have signed up for text alerts. A random sample of 8 residents from a list of 4,000 includes 3 who have signed up. Use the binomial model to assess the evidence at \(\alpha=0.05\).
State: Let \(p\) be the true proportion of residents on this town’s list who have signed up for text alerts. The hypotheses are \(H_0:p=0.20\) and \(H_a:p>0.20\).
Plan: The residents were randomly sampled, so the Random condition is met. Because \(8\leq0.10(4{,}000)=400\), the 10% condition is met. The sample has a fixed size, each resident either has or has not signed up, and the binomial model treats the success probability as common and the trials as approximately independent. Under \(H_0\), the expected counts are \(8(0.20)=1.6\) sign-ups and \(8(0.80)=6.4\) non-sign-ups. These are not both at least 10, so a one-proportion \(z\)-test is not appropriate. We use a binomial-model tail calculation instead.
Do: If \(H_0\) is true, the random variable \(X\), the number of sampled residents who signed up, is approximately binomial with \(n=8\) and \(p=0.20\). For the right-tailed alternative, the p-value is the probability of 3 or more sign-ups:
Since \(0.2031>0.05\), fail to reject \(H_0\). For perspective, if the true proportion were 40%, the same binomial model gives \(P(X\geq3)\approx0.6846\). Thus, 3 or more sign-ups is not an outcome that rules out either the null value or that illustrative higher value.
Conclude: At the 0.05 significance level, the sample does not provide convincing evidence that more than 20% of the residents on this town’s list have signed up for text alerts. It would be incorrect to conclude that exactly 20% have signed up.
Worked Example: A Small Sample of Seedlings
A fictional greenhouse wants to know whether more than half of a certain type of seedling survives its first week after transplanting. In a random sample of 6 seedlings, 4 survive. Consider the test at \(\alpha=0.05\), using a binomial model.
State: Let \(p\) be the true proportion of this type of seedling in the greenhouse’s growing conditions that survives its first week after transplanting. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\).
Plan: The seedlings were randomly selected from the greenhouse’s stock, so the Random condition is met for representing that stock. The sample is a small fraction of the stock, so the 10% condition is reasonable. The number of trials is fixed, survival is a success and non-survival is a failure, and the model treats outcomes as independent with a common survival proportion. Under \(H_0\), the expected success and failure counts are \(6(0.50)=3\) and \(6(0.50)=3\). The Large Counts condition fails, so we use binomial probabilities rather than a one-proportion \(z\)-test.
Do: Under \(H_0\), the random variable \(X\), the number of seedlings that survive, is approximately binomial with \(n=6\) and \(p=0.50\). The right-tailed p-value for 4 or more survivors is:
Because \(0.3438>0.05\), fail to reject \(H_0\). If the true survival proportion were 75%, the probability of 4 or more survivors would be about \(0.8306\). With this sample size, even that higher proportion would frequently produce 4 or more survivors, a result that still does not cross the 0.05 rejection threshold. In fact, under \(H_0\), only 6 survivors would produce a p-value at or below 0.05, since \(P(X\geq5)=7/64\approx0.1094\), while \(P(X=6)=1/64\approx0.0156\).
Conclude: At the 0.05 significance level, the sample does not provide convincing evidence that more than half of this type of seedling in the greenhouse’s growing conditions survives its first week after transplanting. The result does not prove that the survival proportion is 50%.
Common Mistakes and AP Exam Tips
- Saying “accept \(H_0\)”: A large p-value does not establish that the null hypothesis is true. Say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for \(H_a\).
- Turning a failure to find evidence into evidence of no difference: “The proportion is exactly 50%” claims more than the test supports. A result may be inconclusive because the sample is small, even if the true proportion differs from the null value.
- Using the \(z\)-test when Large Counts fails: A small sample can make the Normal approximation unsuitable. Check the Large Counts condition using the null proportion, as in Checking the Success-Failure Condition for Tests. Do not ignore a failed condition just to get a test statistic.
- Interpreting the p-value as a probability about \(H_0\): A p-value of \(0.3770\) does not mean there is a 37.70% chance that \(H_0\) is true. It describes the probability of results at least as extreme as the observed one under the null model.
- Claiming that a non-significant result proves the alternative is false: Failing to reject means the evidence did not meet the chosen threshold. It does not show that the alternative is impossible.
- Forgetting what the sample represents: Keep the conclusion about the population or process represented by the sample. A random sample from one shop’s list does not automatically support a claim about all customers everywhere.
Key Takeaway
A small sample may provide too little information to distinguish the null value from a real departure. When the p-value exceeds the significance level, fail to reject \(H_0\)—but do not treat that decision as proof. The conclusion is about the strength of the sample evidence, not certainty about the population proportion.
Check Your Understanding
Focus on what the test decision says about the evidence, and what it cannot establish about the null hypothesis.
- A test has p-value \(0.22\) and \(\alpha=0.05\). State the decision and explain why it does not prove \(H_0\).
- Why might a test based on 8 observations fail to detect a real difference from the null proportion?
- A student says, “The p-value is 0.30, so there is a 30% chance that the null hypothesis is true.” Identify the error.
- Under \(H_0\), the expected success count in a proposed one-proportion \(z\)-test is 4. Should the student proceed with that \(z\)-test? Explain.
- Write a contextual conclusion for a right-tailed test that fails to reject \(H_0:p=0.40\), where \(p\) is the proportion of residents on a specified town list who use a public library each month.