From Rejecting the Null to a Contextual Conclusion
In Choosing a Significance Level Before Testing, we compared a p-value with a significance level chosen in advance. When the p-value is less than or equal to \(\alpha\), the testing rule says to reject \(H_0\). But a correct response does not stop at that decision. It explains what the data provide evidence for, using the alternative hypothesis and the specific population in the question.
For a one-proportion test, \(H_0\) states a reference value for the population proportion \(p\), while \(H_a\) describes the departure being investigated. Rejecting \(H_0\) means the sample result would be sufficiently unusual under the null model, according to the chosen cutoff. It gives support to the stated alternative; it does not prove that alternative is true.
A strong conclusion connects three things: the decision, the direction or claim in \(H_a\), and the population parameter defined for the test. For example, if \(H_a:p>0.30\), the conclusion should say that the data provide convincing evidence that the population proportion is greater than 30%. It should not merely say “the result is significant,” and it should not change the claim to a different direction.
The p-value helps justify the decision, but the conclusion is not a second p-value interpretation. As explained in Interpreting a P-Value in Context, the p-value describes how often results at least as extreme as the observed result would occur under \(H_0\). It is not the probability that \(H_0\) is true, nor the probability that \(H_a\) is true.
A Reliable Structure for a Rejection Conclusion
Use the wording of \(H_a\) as your guide. For a one-sided alternative, state that there is convincing evidence for the specified increase or decrease. For a two-sided alternative, state that there is convincing evidence of a difference. Include the population and the characteristic represented by \(p\), rather than leaving the conclusion as a statement about the sample.
Say “Reject \(H_0\)” after comparing the p-value with the significance level chosen before looking at the results.
Use the standard phrase “the data provide convincing evidence.” Do not say the test proves the claim.
Translate \(H_a\) into words about the true population proportion and its characteristic.
Refer to the population represented by the sample or the process studied. Do not generalize to people or settings the data do not represent.
The decision is about evidence, not certainty. A rejection does not mean every observed difference is important in practice, either. In Statistical Significance Versus Practical Significance, we distinguished a test result meeting a statistical cutoff from a difference being large enough to matter. The conclusion should accurately describe what \(H_a\) says; any judgment about practical importance requires considering the size and consequences of the difference.
Worked Examples: Writing Rejection Conclusions
Worked Example: A Higher Share of Library Patrons
A fictional public library wants to know whether more than half of its active cardholders prefer reading e-books for leisure. A random sample of 200 cardholders from a source population of 5,000 includes 116 who prefer e-books. The library chose \(\alpha=0.05\) before collecting the data. Conduct and conclude a one-proportion \(z\)-test.
State: Let \(p\) be the true proportion of the library’s active cardholders who prefer reading e-books for leisure. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\). The research question is whether the population proportion is greater than one-half.
Plan: Use a one-proportion \(z\)-test. The cardholders were randomly sampled, so the Random condition is met. The sample was drawn without replacement from 5,000 cardholders, and \(200\leq0.10(5{,}000)=500\), so the 10% condition is met. Under \(H_0\), the expected number who prefer e-books is \(np_0=200(0.50)=100\), and the expected number who do not is \(n(1-p_0)=200(0.50)=100\). Both are at least 10, so the Large Counts condition is met.
Do: The observed sample proportion is \(116/200=0.58\). The null standard error and test statistic are:
Because the alternative is right-tailed, the p-value is the area to the right of \(z=2.263\), approximately \(0.0118\), rounded to four decimal places. Since \(0.0118<0.05\), reject \(H_0\).
Conclude: At the 0.05 significance level, the data provide convincing evidence that more than 50% of this library’s active cardholders prefer reading e-books for leisure. This conclusion supports the direction in \(H_a\); it does not prove that the true proportion exceeds 50%.
Notice that the conclusion names the cardholder population and the preference being measured. Saying “116 of the 200 sampled people preferred e-books” would accurately describe the sample, but it would not give the inferential conclusion about \(p\).
Worked Example: A Proportion That Differs from a Benchmark
A fictional transit-planning group asks whether the proportion of commuters who choose public transit differs from 20% in the population represented by a citywide commuter list. A random sample of 250 commuters, selected without replacement from 8,000, includes 66 who choose public transit. The group set \(\alpha=0.05\) in advance.
State: Let \(p\) be the true proportion of commuters represented by this citywide list who choose public transit. The hypotheses are \(H_0:p=0.20\) and \(H_a:p\ne0.20\). The alternative allows a proportion either above or below 20%.
Plan: Use a one-proportion \(z\)-test. The sample was randomly selected, meeting the Random condition. The sample is 250 of 8,000, and \(250\leq0.10(8{,}000)=800\), so the 10% condition is met. Under \(H_0\), the expected transit count is \(250(0.20)=50\), and the expected count who do not choose transit is \(250(0.80)=200\). Both counts are at least 10, so the Large Counts condition is met.
Do: Here \(\hat{p}=66/250=0.264\). Calculate the null standard error and test statistic:
This is a two-sided test, so the p-value is twice the upper-tail area beyond \(2.530\), approximately \(0.0114\), rounded to four decimal places. Since \(0.0114<0.05\), reject \(H_0\).
Conclude: At the 0.05 significance level, the data provide convincing evidence that the proportion of commuters represented by this citywide list who choose public transit differs from 20%. The sample proportion is above 20%, but the stated alternative is two-sided, so the conclusion is evidence of a difference, not a test of a specifically higher proportion.
That last distinction matters: the sample’s direction does not change the alternative that was selected for the test. A conclusion should reflect the research question and \(H_a\), not replace them after seeing the sample result.
Worked Example: Evidence That a Service Rate Is Lower
A fictional heating-equipment service team investigates whether the proportion of newly installed heat pumps that require a service visit during their first year is below a historical benchmark of 40%. A random sample of 300 installations from a list of 10,000 includes 99 that required a service visit. The team selected \(\alpha=0.01\) before examining the sample.
State: Let \(p\) be the true proportion of newly installed heat pumps in the population represented by the list that require a service visit during their first year. The hypotheses are \(H_0:p=0.40\) and \(H_a:p<0.40\).
Plan: Use a one-proportion \(z\)-test. The installations were randomly sampled, meeting the Random condition. The sample is 300 of 10,000, and \(300\leq0.10(10{,}000)=1{,}000\), so the 10% condition is met. Under \(H_0\), the expected service-visit count is \(300(0.40)=120\), and the expected count without a service visit is \(300(0.60)=180\). Both are at least 10, so the Large Counts condition is met.
Do: The sample proportion is \(\hat{p}=99/300=0.33\). The null standard error and test statistic are:
For the left-tailed alternative, the p-value is the area to the left of \(z=-2.475\), approximately \(0.0067\), rounded to four decimal places. Since \(0.0067<0.01\), reject \(H_0\).
Conclude: At the 0.01 significance level, the data provide convincing evidence that the true proportion of newly installed heat pumps in the population represented by this list that require a first-year service visit is below 40%. The result supports a lower rate in this population; it does not establish that the rate is lower for every heat-pump model or in every setting.
Common Mistakes and AP Exam Tips
- Stopping at “reject \(H_0\)”: The decision alone does not answer the research question in context. Add what the data provide convincing evidence for, using the population and characteristic in the definition of \(p\).
- Writing “the null hypothesis is false” or “the alternative is proven”: A test does not give certainty. Full-credit wording says the data provide convincing evidence for \(H_a\), assuming the conditions for the procedure are satisfied.
- Giving the wrong direction: If \(H_a:p>p_0\), conclude evidence that the proportion is greater; if \(H_a:p<p_0\), conclude evidence that it is lower. For \(H_a:p\ne p_0\), conclude evidence that it differs.
- Making the conclusion only about the sample: The sample proportion is an observed statistic. In an inference conclusion, name the population proportion \(p\) and the population that the sampling process represents.
- Claiming the p-value is the chance a hypothesis is true: The p-value is calculated assuming \(H_0\) is true. It is not the probability that \(H_0\) or \(H_a\) is true.
- Overgeneralizing: A random sample supports inference to the population represented by its sampling frame, not automatically to all people in a region or to other populations. Keep the conclusion within the study’s scope.
- Confusing statistical evidence with practical importance: Rejecting \(H_0\) does not by itself show that the difference is large or consequential. Avoid calling a result “important” unless the context provides a basis for that judgment.
Key Takeaway
A rejection is a decision based on a p-value and a preselected significance level. The written conclusion explains what that decision means: the data provide convincing evidence for the alternative hypothesis in context. Match the alternative’s direction, identify the population represented by \(p\), and avoid claiming proof or certainty.
Check Your Understanding
For each question, focus on matching the conclusion to the alternative hypothesis and stating the population claim carefully.
- A test of \(H_0:p=0.25\) against \(H_a:p>0.25\) has p-value 0.018 and \(\alpha=0.05\). What is the decision, and what form should the contextual conclusion take?
- A test of \(H_0:p=0.60\) against \(H_a:p\ne0.60\) leads to rejection. Why should the conclusion say “differs from 60%” rather than only “is greater than 60%”?
- Explain why “There is a 2% probability that \(H_0\) is true” is not a correct interpretation of a p-value of 0.02.
- A random sample represents customers on one company’s mailing list. What should you consider before generalizing a rejection conclusion to all customers of that company?
- Why does rejecting \(H_0\) not automatically show that the observed difference is practically important?