From a Test Result to a Conclusion
A two-proportion test produces a p-value, but a complete answer does more than report that number. You must compare it with the significance level \(\alpha\), make the correct decision about \(H_0\), and explain what the result indicates about the true proportions in the situation.
As in the earlier tutorial “Finding the P-Value for a Two-Proportion Test,” the alternative hypothesis determines which results count as at least as extreme as the observed difference. Once the p-value has been found, the conclusion must match that alternative. Keep the group order from the hypotheses throughout: \(p_1-p_2\) means the true proportion in Group 1 minus the true proportion in Group 2.
The Decision Rule and What It Means
The significance level \(\alpha\) is the cutoff chosen for deciding whether the sample evidence is sufficiently unusual under the null model to reject \(H_0\). Compare the p-value directly with that cutoff:
Rejecting \(H_0\) means the data provide convincing evidence in favor of the alternative hypothesis, at the chosen significance level. For a one-sided alternative, that evidence is in the direction specified by \(H_a\). For a two-sided alternative, it is evidence that the true proportions differ; the direction of the observed sample difference can be reported, but the two-sided test did not specify that direction in advance.
Failing to reject \(H_0\) means the data do not provide convincing evidence for the alternative at the chosen significance level. It does not prove that \(H_0\) is true, establish that the proportions are equal, or show that any difference is unimportant. A study may fail to find convincing evidence because the true proportions are similar, because the samples are not large enough to detect a difference, or for other reasons.
A Reliable Conclusion Structure
A strong response links the decision to the context instead of stopping at “reject” or “fail to reject.” It names the population proportions and characteristic, uses the direction or two-sided wording from \(H_a\), and avoids claiming more than the test supports.
Compare the p-value with the stated \(\alpha\), using the decision rule.
Say “reject \(H_0\)” or “fail to reject \(H_0\).” Do not say that you accept \(H_0\).
State whether the data provide convincing evidence for the one-sided or two-sided claim in \(H_a\).
Name the groups, the shared characteristic, and—when appropriate—the direction of the difference between the true proportions.
For a two-sided test, a useful conclusion is “There is convincing evidence that the true proportions differ,” followed by the group and outcome definitions. If the sample proportion in Group 1 is larger, you may add that the observed difference was in that direction. Do not turn this into a claim that the test established \(p_1>p_2\); that is a one-sided alternative.
For a one-sided test with \(H_a:p_1>p_2\), a rejection supports the claim that the true proportion in Group 1 is higher than the true proportion in Group 2. If you fail to reject, say that there is not convincing evidence for that claim. Do not switch to the opposite direction just because \(\hat{p}_1<\hat{p}_2\).
Worked Examples
Worked Example: Evidence That One Neighborhood Has a Higher Rate
Question: In a fictional study, independent random samples of 100 households are selected from each of two neighborhoods. In Neighborhood 1, 59 households report using a community compost service; in Neighborhood 2, 41 do. Researchers test at \(\alpha=0.01\) whether the true use proportion is higher in Neighborhood 1.
State: Let \(p_1\) be the true proportion of households in Neighborhood 1 that use the community compost service, and let \(p_2\) be the corresponding proportion in Neighborhood 2. The hypotheses are \(H_0:p_1=p_2\) and \(H_a:p_1>p_2\).
Plan: Each neighborhood contributes an independent random sample, and every household is classified using the same yes-or-no outcome. Assume each sample is less than 10% of its neighborhood’s households; this supports the 10% condition for sampling without replacement. The pooled proportion is \((59+41)/(100+100)=0.50\). Under the null, the expected success and failure counts are \(100(0.50)=50\) and \(100(0.50)=50\) in each group, so the Large Counts condition is met for all four expected counts. A two-proportion \(z\)-test is appropriate.
Do: The sample proportions are \(\hat{p}_1=59/100=0.59\) and \(\hat{p}_2=41/100=0.41\), so their difference is \(0.59-0.41=0.18\). The pooled standard error is:
The test statistic is \(z=0.18/0.07071\approx2.546\). The upper-tail p-value is approximately \(0.0055\) when rounded to four decimal places. Since \(0.0055<0.01\), reject \(H_0\).
Conclude: The samples provide convincing evidence that the true proportion of households using the community compost service is higher in Neighborhood 1 than in Neighborhood 2. This conclusion supports the one-sided alternative; it does not say that every household in Neighborhood 1 uses the service or establish why the neighborhoods differ.
Worked Example: A Two-Sided Test Does Not Find Convincing Evidence
Question: In a fictional survey, independent random samples of 100 customers are selected from each of two online stores. Forty-two customers from Store A and 38 from Store B say that they would recommend the store. At \(\alpha=0.05\), test whether the true recommendation proportions differ.
State: Let \(p_1\) be the true proportion of Store A customers who would recommend the store, and \(p_2\) the corresponding proportion for Store B. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\).
Plan: The two groups are independent random samples, with the same yes-or-no recommendation outcome. Assume each sample is less than 10% of its store’s customer population, so the 10% condition is met. The pooled proportion is \((42+38)/200=0.40\). The expected success counts are \(100(0.40)=40\) in each group, and the expected failure counts are \(100(0.60)=60\) in each group. All four expected counts are at least 10, so the Large Counts condition is satisfied.
Do: The sample proportions are \(\hat{p}_1=0.42\) and \(\hat{p}_2=0.38\), giving an observed difference of \(0.04\). The pooled standard error is:
Thus, \(z=0.04/0.06928\approx0.577\). For a two-sided alternative, the p-value is approximately \(2P(Z\ge0.57735)=0.5637\). Since \(0.5637>0.05\), fail to reject \(H_0\).
Conclude: The data do not provide convincing evidence that the true proportions of customers who would recommend Store A and Store B differ. Store A’s sample proportion was slightly higher, but the test does not establish that Store A’s true proportion is higher. Nor does failing to reject prove that the true proportions are equal.
Worked Example: A Test of a Treatment Difference
Question: In a fictional randomized experiment, 200 plants are randomly assigned, with 100 assigned to receive a new nutrient treatment and 100 to receive standard care. After a fixed growing period, 62 treatment plants and 48 standard-care plants meet a defined growth target. At \(\alpha=0.05\), test whether the treatment increases the proportion meeting the target.
State: Let \(p_T\) be the true proportion of plants that would meet the growth target under the new treatment, and let \(p_C\) be the true proportion under standard care. Test \(H_0:p_T=p_C\) against \(H_a:p_T>p_C\).
Plan: The plants were randomly assigned to the two conditions, supporting the randomized design for comparing treatment outcomes. Each plant contributes to only one group, so the groups are independent. The pooled proportion is \((62+48)/(100+100)=0.55\). Under \(H_0\), each group has expected successes \(100(0.55)=55\) and expected failures \(100(0.45)=45\); all four expected counts meet the Large Counts condition. A two-proportion \(z\)-test is appropriate.
Do: The sample proportions are \(\hat{p}_T=0.62\) and \(\hat{p}_C=0.48\), with observed difference \(0.14\). The pooled standard error and test statistic are:
The upper-tail p-value is approximately \(P(Z\ge1.990)=0.0233\). Since \(0.0233<0.05\), reject \(H_0\).
Conclude: The experiment provides convincing evidence that the new treatment increases the proportion of plants meeting the growth target compared with standard care. Because plants were randomly assigned, the design supports a cause-and-effect interpretation for the experimental units. The experiment does not, by itself, justify generalizing to a broader population of plants unless the way those plants were selected also supports that generalization.
Rounding, Equality, and the Meaning of the P-Value
Use enough precision when comparing the p-value with \(\alpha\), especially when the two values are close. A calculator may display more digits than you need in the final response. Do not round early in a way that changes the decision. For example, a p-value just below \(\alpha\) should not be rounded up and treated as greater than \(\alpha\); compare using the unrounded values when possible.
The p-value is calculated assuming the null model is true. It is the probability of getting a result at least as extreme as the observed result, in the direction or directions specified by \(H_a\). It is not the probability that \(H_0\) is true, and it is not the probability that chance alone caused the observed data.
The examples above also illustrate why a small p-value and a large difference are not interchangeable ideas. The p-value measures how unusual the sample result is under the null model; it does not measure the practical importance of the difference. To describe the estimated size of a difference, report \(\hat{p}_1-\hat{p}_2\) in context. A confidence interval, studied earlier in this unit, can provide plausible values for the population difference.
Common Mistakes and AP Exam Tips
- Saying “accept the null”: A large p-value does not establish that the null hypothesis is true. Write “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for the alternative.
- Leaving out the comparison with \(\alpha\): A p-value alone is not a decision. State whether it is less than, equal to, or greater than the significance level.
- Writing only the decision: “Reject \(H_0\)” does not answer the context question. Follow it with a statement about the true proportions and the shared outcome.
- Reversing the groups: Keep the parameter definitions and group order fixed. A conclusion about \(p_1-p_2\) must name Group 1 first and Group 2 second.
- Overstating a two-sided result: Rejecting \(H_0\) for \(H_a:p_1\ne p_2\) supports a difference, not a preplanned claim that \(p_1>p_2\). You may report which sample proportion was larger.
- Claiming equality after a nonsignificant result: “There is no difference” is not justified by failing to reject. State that the test did not find convincing evidence of a difference.
- Ignoring the design: A significant result from random samples can support generalization to the populations sampled. Random assignment can support a causal conclusion. Do not claim both unless the design supports both.
- Confusing statistical evidence with practical importance: A small p-value does not say whether a difference matters in practice. Describe the estimated difference separately if the question calls for its size.
Key Takeaway
A test conclusion is a linked argument: compare the p-value with \(\alpha\), make the correct decision about \(H_0\), and describe what the result indicates about the true proportions. A rejection supports the stated alternative; failing to reject is not proof of equality. In either case, preserve the group order and keep the conclusion within the scope allowed by the study design.
Check Your Understanding
For each question, use the stated alternative and significance level.
- A one-sided test has p-value \(0.032\) and \(\alpha=0.05\). What is the decision, and what should the conclusion say about the direction specified in \(H_a\)?
- A two-sided test has p-value \(0.08\) and \(\alpha=0.05\). Write a contextual conclusion that does not claim the true proportions are equal.
- A test has p-value \(0.01\) and \(\alpha=0.01\). What is the correct decision, and why?
- For \(H_a:p_1\ne p_2\), the sample proportion in Group 1 is larger and the test rejects \(H_0\). What can you conclude, and what directional claim should you avoid making?
- A significant two-proportion test comes from a randomized experiment. What does random assignment support, and what does it not automatically support?