When the Data Do Not Give Convincing Evidence for the Alternative
In Rejecting the Null Hypothesis Correctly, we saw that a sufficiently small p-value can lead us to reject \(H_0\) and describe the evidence for \(H_a\). This tutorial addresses the other outcome: when the p-value is larger than the significance level chosen before the test, we fail to reject \(H_0\).
Failing to reject is not a claim that the null hypothesis has been shown to be true. It means that the sample did not provide strong enough evidence, by the test’s preselected rule, to support the alternative hypothesis. In a contextual conclusion, the key phrase is “there is not convincing evidence” for the claim in \(H_a\).
The direction of the conclusion still comes from the alternative hypothesis. If \(H_a:p<p_0\), say there is not convincing evidence that the population proportion is below \(p_0\). If \(H_a:p>p_0\), say there is not convincing evidence that it is above \(p_0\). If \(H_a:p\ne p_0\), say there is not convincing evidence that it differs from \(p_0\).
As discussed in What a P-Value Really Measures and Interpreting a P-Value in Context, the p-value is calculated under the assumption that \(H_0\) is true. A large p-value means that the observed sample result, or a result at least as extreme in the direction or directions specified by \(H_a\), would not be especially unusual under that assumption. It is not the probability that \(H_0\) is true, and it does not prove there is no difference from the null value.
A Reliable Structure for a Fail-to-Reject Conclusion
Use the same logical structure as for a rejection, but change the evidence statement. First identify the decision from the comparison between the p-value and \(\alpha\). Then state that the data do not provide convincing evidence for the alternative, translating \(H_a\) into words about the population proportion \(p\).
If the p-value is greater than the significance level selected in advance, say “Fail to reject \(H_0\).”
Say that the data do not provide convincing evidence for \(H_a\). Do not say the data prove \(H_0\).
Use the direction in \(H_a\): lower, higher, or different from the null value. Name the population and characteristic represented by \(p\).
Refer only to the population represented by the sample or the process studied.
The phrase “not convincing evidence” is deliberately measured. It does not mean the sample exactly matched the null value, or that no evidence at all points toward the alternative. It says that the evidence did not meet the test’s standard for rejecting \(H_0\). In particular, a sample proportion can be above the null value even when a right-tailed test fails to reject; the observed difference simply may not be sufficiently unusual under the null model.
The significance level matters to the decision. A p-value above \(\alpha=0.05\), for example, leads to failing to reject at that 0.05 level. The p-value itself does not change when the significance level changes, but the decision rule can. State the level when it helps make the conclusion clear.
Worked Examples: Failing to Reject in Context
Worked Example: A Lower Composting Rate
A fictional community-garden coordinator wonders whether fewer than 30% of the garden’s plot renters compost food scraps. A random sample of 100 renters from a list of 1,800 includes 27 who compost. The coordinator chose \(\alpha=0.05\) before collecting the sample. Conduct and conclude a one-proportion \(z\)-test.
State: Let \(p\) be the true proportion of renters on this community garden’s list who compost food scraps. The hypotheses are \(H_0:p=0.30\) and \(H_a:p<0.30\). The alternative asks whether the population proportion is lower than 30%.
Plan: Use a one-proportion \(z\)-test. The renters were randomly sampled, so the Random condition is met. The sample was selected without replacement from 1,800 renters, and \(100\leq0.10(1{,}800)=180\), so the 10% condition is met. Under \(H_0\), the expected number who compost is \(np_0=100(0.30)=30\), and the expected number who do not is \(n(1-p_0)=100(0.70)=70\). Both are at least 10, so the Large Counts condition is met.
Do: The observed sample proportion is \(\hat{p}=27/100=0.27\). Using the null proportion, calculate the standard error and test statistic:
Because \(H_a\) is left-tailed, the p-value is the area to the left of the test statistic. Using the unrounded statistic, the p-value is approximately \(0.2563\), rounded to four decimal places. Since \(0.2563>0.05\), fail to reject \(H_0\).
Conclude: At the 0.05 significance level, the data do not provide convincing evidence that fewer than 30% of the renters on this community garden’s list compost food scraps.
The sample proportion, 27%, is below the 30% benchmark, but that direction alone does not establish convincing evidence of a lower population proportion. The test decision reflects both the observed difference and the amount of sample-to-sample variability expected under \(H_0\).
Worked Example: A Higher Rate of Choosing a Digital Receipt
A fictional retailer investigates whether more than 65% of its customers choose a digital receipt at checkout. A random sample of 120 customers from a list of 6,000 includes 82 who choose a digital receipt. The retailer set \(\alpha=0.05\) in advance.
State: Let \(p\) be the true proportion of customers on this retailer’s list who choose a digital receipt at checkout. The hypotheses are \(H_0:p=0.65\) and \(H_a:p>0.65\).
Plan: Use a one-proportion \(z\)-test. The customers were randomly sampled, meeting the Random condition. The sample was drawn without replacement from 6,000 customers, and \(120\leq0.10(6{,}000)=600\), so the 10% condition is met. Under \(H_0\), the expected digital-receipt count is \(120(0.65)=78\), and the expected count who do not choose one is \(120(0.35)=42\). Both counts are at least 10, so the Large Counts condition is met.
Do: Here \(\hat{p}=82/120\approx0.6833\). Calculate the null standard error and test statistic:
For this right-tailed test, the p-value is the area to the right of the unrounded \(z\) statistic, approximately \(0.2220\), rounded to four decimal places. Since \(0.2220>0.05\), fail to reject \(H_0\).
Conclude: At the 0.05 significance level, the data do not provide convincing evidence that more than 65% of the customers on this retailer’s list choose a digital receipt at checkout.
Although the sample proportion is higher than 65%, the p-value indicates that a result this high or higher is not sufficiently unusual under the null model to meet the selected rejection rule. Do not turn “the sample percentage was higher” into a conclusion that the population percentage is higher.
Worked Example: A Proportion That Might Differ from a Benchmark
A fictional parks department asks whether the proportion of visitors to a regional park who bring a reusable water bottle differs from 40%. A random sample of 200 visitors from a visitor-list population of 5,000 includes 84 who bring one. The department selected \(\alpha=0.05\) before examining the results.
State: Let \(p\) be the true proportion of visitors represented by this regional park’s list who bring a reusable water bottle. The hypotheses are \(H_0:p=0.40\) and \(H_a:p\ne0.40\). Because this is a two-sided alternative, evidence in either direction could count against \(H_0\).
Plan: Use a one-proportion \(z\)-test. The visitors were randomly sampled, so the Random condition is met. The sample was drawn without replacement from 5,000 visitors, and \(200\leq0.10(5{,}000)=500\), so the 10% condition is met. Under \(H_0\), the expected number who bring a bottle is \(200(0.40)=80\), and the expected number who do not is \(200(0.60)=120\). Both are at least 10, so the Large Counts condition is met.
Do: The observed sample proportion is \(\hat{p}=84/200=0.42\). The null standard error and test statistic are:
For a two-sided test, the p-value is twice the area beyond \(\lvert z\rvert\) in one tail. Using the unrounded statistic, the p-value is approximately \(0.5637\), rounded to four decimal places. Since \(0.5637>0.05\), fail to reject \(H_0\).
Conclude: At the 0.05 significance level, the data do not provide convincing evidence that the proportion of visitors represented by this regional park’s list who bring a reusable water bottle differs from 40%.
The sample proportion is 42%, but the conclusion must reflect the two-sided alternative. It would be inaccurate to conclude that the proportion is higher than 40%; this test did not provide convincing evidence of a difference in either direction.
Common Mistakes and AP Exam Tips
- Saying “accept \(H_0\)”: A large p-value does not show that the null hypothesis is true. Say “fail to reject \(H_0\)” for the decision and “the data do not provide convincing evidence for \(H_a\)” for the contextual conclusion.
- Claiming there is no difference: A test that fails to reject has not established that the population proportion equals the null value. State only that the data do not provide convincing evidence for the alternative claim.
- Describing the sample instead of the population: “The sample had 27 composters” reports what happened in the sample. An inference conclusion should describe the population proportion \(p\) and the population represented by the sample.
- Ignoring the alternative’s direction: A left-tailed alternative concerns a lower proportion, a right-tailed alternative concerns a higher proportion, and a two-sided alternative concerns a difference. Use the alternative that was specified before examining the sample.
- Interpreting the p-value as the probability that \(H_0\) is true: A p-value is calculated assuming \(H_0\) is true. It is a probability about possible sample results under that assumption, not a probability assigned to either hypothesis.
- Calling a large p-value proof of “no effect”: The decision says the evidence was not convincing at the chosen significance level. It does not prove that the true proportion equals the benchmark or that a real departure is impossible.
Key Takeaway
When the p-value is greater than the preselected significance level, fail to reject \(H_0\). The contextual conclusion should say that the data do not provide convincing evidence for \(H_a\). This is a statement about the strength of the evidence from the test, not proof that the null hypothesis is true.
Check Your Understanding
For each question, focus on the decision and on what can—and cannot—be concluded about the population proportion.
- A test of \(H_0:p=0.25\) against \(H_a:p>0.25\) has p-value \(0.18\) and \(\alpha=0.05\). State the decision and the form of the contextual conclusion.
- A test of \(H_0:p=0.60\) against \(H_a:p\ne0.60\) has p-value \(0.41\). Why would “the proportion is exactly 60%” go beyond the test’s conclusion?
- In a left-tailed test, the sample proportion is below the null value, but the p-value is greater than \(\alpha\). What should the conclusion say, and why is the sample direction not enough to reject \(H_0\)?
- Explain why a p-value of \(0.30\) is not a 30% probability that the null hypothesis is true.
- A random sample represents visitors on one park’s visitor list. What should you consider before extending a fail-to-reject conclusion to visitors at other parks?