Tutorials › AP Statistics › Writing a Conclusion for a Two-Proportion Test

Two-proportion hypothesis tests · Tutorial 535 of 1000

Writing a Conclusion for a Two-Proportion Test

Connect the p-value to alpha, state the correct decision, and explain what the evidence says about the two true proportions.

Intermediate 10 min read

What You'll Learn

  • Compare a two-proportion test’s p-value with the stated significance level.
  • State “reject” or “fail to reject” the null hypothesis accurately.
  • Write a contextual conclusion about the difference between the true proportions.
  • Distinguish evidence of a difference from evidence about which proportion is higher.
  • Explain what a nonsignificant result does—and does not—show.
  • Match the scope of a conclusion to the study’s sampling and assignment design.

From a Test Result to a Conclusion

A two-proportion test produces a p-value, but a complete answer does more than report that number. You must compare it with the significance level \(\alpha\), make the correct decision about \(H_0\), and explain what the result indicates about the true proportions in the situation.

As in the earlier tutorial “Finding the P-Value for a Two-Proportion Test,” the alternative hypothesis determines which results count as at least as extreme as the observed difference. Once the p-value has been found, the conclusion must match that alternative. Keep the group order from the hypotheses throughout: \(p_1-p_2\) means the true proportion in Group 1 minus the true proportion in Group 2.

Definition: A test conclusion combines the comparison of the p-value with \(\alpha\), the resulting decision about \(H_0\), and a statement in context about the evidence concerning the population proportions specified by \(H_a\).

The Decision Rule and What It Means

The significance level \(\alpha\) is the cutoff chosen for deciding whether the sample evidence is sufficiently unusual under the null model to reject \(H_0\). Compare the p-value directly with that cutoff:

Decision rule: If the p-value is less than or equal to \(\alpha\), reject \(H_0\). If the p-value is greater than \(\alpha\), fail to reject \(H_0\). When the p-value equals \(\alpha\), reject \(H_0\).

Rejecting \(H_0\) means the data provide convincing evidence in favor of the alternative hypothesis, at the chosen significance level. For a one-sided alternative, that evidence is in the direction specified by \(H_a\). For a two-sided alternative, it is evidence that the true proportions differ; the direction of the observed sample difference can be reported, but the two-sided test did not specify that direction in advance.

Failing to reject \(H_0\) means the data do not provide convincing evidence for the alternative at the chosen significance level. It does not prove that \(H_0\) is true, establish that the proportions are equal, or show that any difference is unimportant. A study may fail to find convincing evidence because the true proportions are similar, because the samples are not large enough to detect a difference, or for other reasons.

Formula: For hypotheses written as \(H_0:p_1=p_2\), the conclusion addresses whether there is convincing evidence for the stated alternative about \(p_1-p_2\). Use the same groups, outcome, and subtraction order in the conclusion that you used in the hypotheses.

A Reliable Conclusion Structure

A strong response links the decision to the context instead of stopping at “reject” or “fail to reject.” It names the population proportions and characteristic, uses the direction or two-sided wording from \(H_a\), and avoids claiming more than the test supports.

1
Compare.
Compare the p-value with the stated \(\alpha\), using the decision rule.
2
Decide.
Say “reject \(H_0\)” or “fail to reject \(H_0\).” Do not say that you accept \(H_0\).
3
Connect to the alternative.
State whether the data provide convincing evidence for the one-sided or two-sided claim in \(H_a\).
4
Put it in context.
Name the groups, the shared characteristic, and—when appropriate—the direction of the difference between the true proportions.

For a two-sided test, a useful conclusion is “There is convincing evidence that the true proportions differ,” followed by the group and outcome definitions. If the sample proportion in Group 1 is larger, you may add that the observed difference was in that direction. Do not turn this into a claim that the test established \(p_1>p_2\); that is a one-sided alternative.

For a one-sided test with \(H_a:p_1>p_2\), a rejection supports the claim that the true proportion in Group 1 is higher than the true proportion in Group 2. If you fail to reject, say that there is not convincing evidence for that claim. Do not switch to the opposite direction just because \(\hat{p}_1<\hat{p}_2\).

Worked Examples

Worked Example: Evidence That One Neighborhood Has a Higher Rate

Question: In a fictional study, independent random samples of 100 households are selected from each of two neighborhoods. In Neighborhood 1, 59 households report using a community compost service; in Neighborhood 2, 41 do. Researchers test at \(\alpha=0.01\) whether the true use proportion is higher in Neighborhood 1.

State: Let \(p_1\) be the true proportion of households in Neighborhood 1 that use the community compost service, and let \(p_2\) be the corresponding proportion in Neighborhood 2. The hypotheses are \(H_0:p_1=p_2\) and \(H_a:p_1>p_2\).

Plan: Each neighborhood contributes an independent random sample, and every household is classified using the same yes-or-no outcome. Assume each sample is less than 10% of its neighborhood’s households; this supports the 10% condition for sampling without replacement. The pooled proportion is \((59+41)/(100+100)=0.50\). Under the null, the expected success and failure counts are \(100(0.50)=50\) and \(100(0.50)=50\) in each group, so the Large Counts condition is met for all four expected counts. A two-proportion \(z\)-test is appropriate.

Do: The sample proportions are \(\hat{p}_1=59/100=0.59\) and \(\hat{p}_2=41/100=0.41\), so their difference is \(0.59-0.41=0.18\). The pooled standard error is:

$$ SE_{\text{pooled}}=\sqrt{0.50(1-0.50)\left(\frac{1}{100}+\frac{1}{100}\right)} =\sqrt{0.005}=0.07071 $$

The test statistic is \(z=0.18/0.07071\approx2.546\). The upper-tail p-value is approximately \(0.0055\) when rounded to four decimal places. Since \(0.0055<0.01\), reject \(H_0\).

Conclude: The samples provide convincing evidence that the true proportion of households using the community compost service is higher in Neighborhood 1 than in Neighborhood 2. This conclusion supports the one-sided alternative; it does not say that every household in Neighborhood 1 uses the service or establish why the neighborhoods differ.

Worked Example: A Two-Sided Test Does Not Find Convincing Evidence

Question: In a fictional survey, independent random samples of 100 customers are selected from each of two online stores. Forty-two customers from Store A and 38 from Store B say that they would recommend the store. At \(\alpha=0.05\), test whether the true recommendation proportions differ.

State: Let \(p_1\) be the true proportion of Store A customers who would recommend the store, and \(p_2\) the corresponding proportion for Store B. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\).

Plan: The two groups are independent random samples, with the same yes-or-no recommendation outcome. Assume each sample is less than 10% of its store’s customer population, so the 10% condition is met. The pooled proportion is \((42+38)/200=0.40\). The expected success counts are \(100(0.40)=40\) in each group, and the expected failure counts are \(100(0.60)=60\) in each group. All four expected counts are at least 10, so the Large Counts condition is satisfied.

Do: The sample proportions are \(\hat{p}_1=0.42\) and \(\hat{p}_2=0.38\), giving an observed difference of \(0.04\). The pooled standard error is:

$$ SE_{\text{pooled}}=\sqrt{0.40(0.60)\left(\frac{1}{100}+\frac{1}{100}\right)} =\sqrt{0.0048}\approx0.06928 $$

Thus, \(z=0.04/0.06928\approx0.577\). For a two-sided alternative, the p-value is approximately \(2P(Z\ge0.57735)=0.5637\). Since \(0.5637>0.05\), fail to reject \(H_0\).

Conclude: The data do not provide convincing evidence that the true proportions of customers who would recommend Store A and Store B differ. Store A’s sample proportion was slightly higher, but the test does not establish that Store A’s true proportion is higher. Nor does failing to reject prove that the true proportions are equal.

Worked Example: A Test of a Treatment Difference

Question: In a fictional randomized experiment, 200 plants are randomly assigned, with 100 assigned to receive a new nutrient treatment and 100 to receive standard care. After a fixed growing period, 62 treatment plants and 48 standard-care plants meet a defined growth target. At \(\alpha=0.05\), test whether the treatment increases the proportion meeting the target.

State: Let \(p_T\) be the true proportion of plants that would meet the growth target under the new treatment, and let \(p_C\) be the true proportion under standard care. Test \(H_0:p_T=p_C\) against \(H_a:p_T>p_C\).

Plan: The plants were randomly assigned to the two conditions, supporting the randomized design for comparing treatment outcomes. Each plant contributes to only one group, so the groups are independent. The pooled proportion is \((62+48)/(100+100)=0.55\). Under \(H_0\), each group has expected successes \(100(0.55)=55\) and expected failures \(100(0.45)=45\); all four expected counts meet the Large Counts condition. A two-proportion \(z\)-test is appropriate.

Do: The sample proportions are \(\hat{p}_T=0.62\) and \(\hat{p}_C=0.48\), with observed difference \(0.14\). The pooled standard error and test statistic are:

$$ SE_{\text{pooled}}=\sqrt{0.55(0.45)\left(\frac{1}{100}+\frac{1}{100}\right)} =\sqrt{0.00495}\approx0.07036 $$
$$ z=\frac{0.62-0.48}{0.07036}\approx1.990 $$

The upper-tail p-value is approximately \(P(Z\ge1.990)=0.0233\). Since \(0.0233<0.05\), reject \(H_0\).

Conclude: The experiment provides convincing evidence that the new treatment increases the proportion of plants meeting the growth target compared with standard care. Because plants were randomly assigned, the design supports a cause-and-effect interpretation for the experimental units. The experiment does not, by itself, justify generalizing to a broader population of plants unless the way those plants were selected also supports that generalization.

Rounding, Equality, and the Meaning of the P-Value

Use enough precision when comparing the p-value with \(\alpha\), especially when the two values are close. A calculator may display more digits than you need in the final response. Do not round early in a way that changes the decision. For example, a p-value just below \(\alpha\) should not be rounded up and treated as greater than \(\alpha\); compare using the unrounded values when possible.

The p-value is calculated assuming the null model is true. It is the probability of getting a result at least as extreme as the observed result, in the direction or directions specified by \(H_a\). It is not the probability that \(H_0\) is true, and it is not the probability that chance alone caused the observed data.

Calculation check: For a standard normal statistic reported as \(z=2.546\), the upper-tail probability is \(P(Z\ge2.546)\approx0.005448\). For a two-sided test with \(z=1.005\), the p-value is \(2P(Z\ge1.005)\approx0.3149\). These are tail probabilities for the stated statistics and alternatives; round only after the calculation.

The examples above also illustrate why a small p-value and a large difference are not interchangeable ideas. The p-value measures how unusual the sample result is under the null model; it does not measure the practical importance of the difference. To describe the estimated size of a difference, report \(\hat{p}_1-\hat{p}_2\) in context. A confidence interval, studied earlier in this unit, can provide plausible values for the population difference.

Common Mistakes and AP Exam Tips

  • Saying “accept the null”: A large p-value does not establish that the null hypothesis is true. Write “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for the alternative.
  • Leaving out the comparison with \(\alpha\): A p-value alone is not a decision. State whether it is less than, equal to, or greater than the significance level.
  • Writing only the decision: “Reject \(H_0\)” does not answer the context question. Follow it with a statement about the true proportions and the shared outcome.
  • Reversing the groups: Keep the parameter definitions and group order fixed. A conclusion about \(p_1-p_2\) must name Group 1 first and Group 2 second.
  • Overstating a two-sided result: Rejecting \(H_0\) for \(H_a:p_1\ne p_2\) supports a difference, not a preplanned claim that \(p_1>p_2\). You may report which sample proportion was larger.
  • Claiming equality after a nonsignificant result: “There is no difference” is not justified by failing to reject. State that the test did not find convincing evidence of a difference.
  • Ignoring the design: A significant result from random samples can support generalization to the populations sampled. Random assignment can support a causal conclusion. Do not claim both unless the design supports both.
  • Confusing statistical evidence with practical importance: A small p-value does not say whether a difference matters in practice. Describe the estimated difference separately if the question calls for its size.
AP Exam Tip: A full-credit conclusion typically names the decision, says whether there is convincing evidence for the alternative, and identifies the true proportions and outcome in context. Match the wording to the direction of \(H_a\), and let the study design limit any causal or population claim.

Key Takeaway

A test conclusion is a linked argument: compare the p-value with \(\alpha\), make the correct decision about \(H_0\), and describe what the result indicates about the true proportions. A rejection supports the stated alternative; failing to reject is not proof of equality. In either case, preserve the group order and keep the conclusion within the scope allowed by the study design.

Key takeaway: Reject \(H_0\) when the p-value is at most \(\alpha\), and fail to reject \(H_0\) when it is greater. Then state the evidence for the alternative in context—without claiming that the test proves the null, proves causation without random assignment, or shows that a difference is practically important.

Check Your Understanding

For each question, use the stated alternative and significance level.

  1. A one-sided test has p-value \(0.032\) and \(\alpha=0.05\). What is the decision, and what should the conclusion say about the direction specified in \(H_a\)?
  2. A two-sided test has p-value \(0.08\) and \(\alpha=0.05\). Write a contextual conclusion that does not claim the true proportions are equal.
  3. A test has p-value \(0.01\) and \(\alpha=0.01\). What is the correct decision, and why?
  4. For \(H_a:p_1\ne p_2\), the sample proportion in Group 1 is larger and the test rejects \(H_0\). What can you conclude, and what directional claim should you avoid making?
  5. A significant two-proportion test comes from a randomized experiment. What does random assignment support, and what does it not automatically support?