Tutorials › AP Statistics › Frequent Errors in Writing Test Conclusions

P-values and conclusions for proportions · Tutorial 499 of 1000

Frequent Errors in Writing Test Conclusions

Build clear, contextual test conclusions that describe evidence about the population proportion without confusing it with the sample result or claiming certainty.

Intermediate 11 min read

What You'll Learn

  • Identify when a conclusion refers to the sample statistic instead of the population parameter.
  • Name the population and characteristic in a contextual conclusion.
  • State the correct decision when a p-value is less than, equal to, or greater than alpha.
  • Replace claims of proof with appropriately cautious evidence language.
  • Distinguish a descriptive statement about the sample from an inferential conclusion.
  • Revise conclusions after failing to reject the null hypothesis without treating it as proven.

Three Errors Can Weaken an Otherwise Correct Test

A test can have a correctly stated hypothesis, a suitable procedure, and a correctly calculated p-value—and still end with a poor conclusion. The final sentence must say what the test indicates about the population parameter, in the setting of the question. It should not merely repeat what happened in the sample, leave out the population, or claim that a hypothesis has been proved.

In Writing a Complete Conclusion With Evidence, you learned to connect the decision to a contextual statement about the alternative hypothesis. This tutorial focuses on three frequent writing errors: omitting context, drawing a conclusion about the sample rather than the population proportion, and using absolute language such as “proves.” These are wording errors, but they reflect important distinctions in statistical reasoning.

Definition: A test conclusion states the decision about \(H_0\) and explains what the sample evidence indicates about the population parameter, in context. A conclusion does not claim certainty or turn a sample statistic into a known population value.

A useful conclusion names the population, the characteristic being measured, and the direction of the alternative hypothesis. It also matches the decision to the p-value and the significance level. If the p-value is at or below \(\alpha\), reject \(H_0\); if it is greater than \(\alpha\), fail to reject \(H_0\). Equality belongs with rejection.

Conclusion frame: Because the p-value is less than or equal to \(\alpha\), we reject \(H_0\); because it is greater than \(\alpha\), we fail to reject \(H_0\). In context, the data [do / do not] provide convincing evidence that the true proportion of [defined population] who [meet the stated criterion] [state the direction in \(H_a\)].

What a Strong Conclusion Needs

Each part of the conclusion has a job. “Reject” or “fail to reject” reports the test decision. “Convincing evidence” describes what that decision says about the alternative hypothesis. The contextual phrase identifies whose true proportion is at issue and what characteristic defines success. Together, these parts keep the conclusion tied to the question the test actually asked.

1
Make the decision.
Compare the p-value with the preselected \(\alpha\). Reject when the p-value is less than or equal to \(\alpha\); otherwise fail to reject.
2
Refer to the parameter.
State what the evidence indicates about the true population proportion \(p\), not merely the observed sample proportion \(\hat{p}\).
3
Restore the context.
Name the population and the precise characteristic. Match the direction—lower, higher, or different—to the alternative hypothesis.
4
Use cautious evidence language.
Say “convincing evidence” when rejecting \(H_0\), or “not convincing evidence” for the alternative when failing to reject. Do not claim proof.

This is not a requirement to copy a long template word for word. It is a check that the decision, population claim, and strength of language all agree with the test. You may include the p-value in the conclusion, but the number alone is not a contextual interpretation.

Sample Result or Population Conclusion?

The sample proportion \(\hat{p}=x/n\) describes the sample. For example, “63% of the sampled households reported using a compost bin” is a statement about the observed sample, and it can be exactly correct. A significance test, however, uses that sample result to evaluate a claim about the true proportion \(p\) in a defined population. The inferential conclusion should therefore name that population and discuss \(p\).

Do not replace the sample statement with an assertion that the population proportion equals the sample proportion. A sample proportion is an estimate; it is not generally the known value of the population parameter. Instead, a test conclusion describes whether the evidence is convincing for the alternative claim. As in Defining the Parameter in Hypothesis Statements, a clear definition of \(p\) makes it possible to say exactly whose proportion the conclusion concerns.

Worked Example: A Sample Result Is Not the Population Claim

A fictional city survey randomly selects 200 households from a registry of 4,000 households. In the sample, 126 households report using a compost bin. The question is whether more than 55% of all households represented by the registry use one. A one-proportion \(z\)-test gives a p-value of approximately \(0.0115\), and the test uses \(\alpha=0.05\).

State: Let \(p\) be the true proportion of households represented by this city registry that use a compost bin. The hypotheses are \(H_0:p=0.55\) and \(H_a:p>0.55\). The sample proportion is \(\hat{p}=126/200=0.63\).

Plan and check conditions: The survey is described as a random sample, so the Random condition is met. The sample size is \(200\), which is less than 10% of the 4,000-household registry, so the 10% condition is met. Using the null proportion, the Large Counts condition is met because \(np_0=200(0.55)=110\) and \(n(1-p_0)=200(0.45)=90\); both are at least 10.

Do: The null standard error and test statistic are

$$ SE_0=\sqrt{\frac{0.55(0.45)}{200}}=\sqrt{0.0012375}\approx0.03518,\qquad z=\frac{0.63-0.55}{0.03518}\approx2.274 $$

For the right-tailed alternative, the p-value is \(P(Z\ge2.274)\approx0.0115\), rounded to four decimal places. Since \(0.0115<0.05\), reject \(H_0\).

Conclude: The sample provides convincing evidence that more than 55% of households represented by the city registry use a compost bin. The conclusion concerns the registry’s population proportion \(p\), not just the 63% observed in the sample. It does not establish that the population proportion is exactly 63%.

Repair the error: “Exactly 63% of city households use compost bins” incorrectly treats \(\hat{p}\) as the known population proportion and leaves the sampling frame unclear. The revised conclusion identifies the registry population and states the direction supported by the test.

What Failing to Reject Does—and Does Not—Say

When the p-value is greater than \(\alpha\), fail to reject \(H_0\). That decision does not show that the null hypothesis is true, and it does not prove that the population proportion equals \(p_0\). It says that the data do not provide convincing evidence for the alternative at the selected significance level.

Be precise about which claim lacks convincing evidence. For a right-tailed test, for instance, say that there is not convincing evidence that the population proportion exceeds the null value. Do not say “there is no difference” unless that is exactly the claim being tested—and even then, failing to reject a null value does not establish equality.

Worked Example: Do Not Turn a Large P-Value into Proof

A fictional parks department randomly samples 80 residents from a registry of 2,000 residents. Of those sampled, 38 say they visited a neighborhood park in the past month. The department asks whether more than half of residents visited. A one-proportion \(z\)-test gives a p-value of approximately \(0.6724\), with \(\alpha=0.05\).

State and plan: Let \(p\) be the true proportion of residents represented by the registry who visited a neighborhood park in the past month. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\). The random sample meets the Random condition as described. Since \(80<0.10(2000)=200\), the 10% condition is met. Under \(H_0\), the expected numbers of successes and failures are \(80(0.50)=40\) and \(80(0.50)=40\), so the Large Counts condition is met.

Do: The sample proportion is \(\hat{p}=38/80=0.475\). The null standard error and test statistic are

$$ SE_0=\sqrt{\frac{0.50(0.50)}{80}}\approx0.05590,\qquad z=\frac{0.475-0.50}{0.05590}\approx-0.447 $$

For the right-tailed alternative, the p-value is \(P(Z\ge-0.447)\approx0.6724\), rounded to four decimal places. Since \(0.6724>0.05\), fail to reject \(H_0\).

Conclude: The data do not provide convincing evidence that more than half of residents represented by the registry visited a neighborhood park in the past month. The sample result was 47.5%, but that does not prove that the population proportion is 50%, nor does it prove that the proportion is less than 50%.

Repair the error: “The test proves that half of residents visited a park” is not supported. A large p-value is not evidence that the null is true; it means this sample result is not unusual under the null model in the direction tested.

Why “Proves” Is Too Strong

A test conclusion is based on sample evidence subject to chance variation and the assumptions of the procedure. Rejecting \(H_0\) is a decision supported by the data, not a demonstration that the alternative is certainly true. “Proves,” “confirms beyond doubt,” and “guarantees” overstate what a significance test can establish.

This remains true even when a p-value is very small. As explained in What a P-Value Really Measures, the p-value is calculated assuming \(H_0\) is true and describes the probability of data at least as extreme as those observed, under the test model. It is not the probability that \(H_0\) is true, and a small value does not turn evidence into proof.

Worked Example: Replace “Proves” With an Evidence Conclusion

A fictional school district randomly samples 400 students from a district enrollment list. In the sample, 224 students say they would use a proposed late bus service. The district tests whether the proportion of students represented by the list who would use the service differs from 50%. The two-sided one-proportion \(z\)-test gives \(p=0.0164\), using \(\alpha=0.05\).

State and plan: Let \(p\) be the true proportion of students represented by the enrollment list who would use the proposed late bus service. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\). The random sample meets the Random condition as described. The 400 sampled students are less than 10% of the district enrollment list, so the 10% condition is met. Under the null, expected counts are \(400(0.50)=200\) students in each response category, satisfying the Large Counts condition.

Do: The sample proportion is \(\hat{p}=224/400=0.56\). The null standard error is \(\sqrt{0.50(0.50)/400}=0.025\), so

$$ z=\frac{0.56-0.50}{0.025}=2.40 $$

For the two-sided alternative, \(P\text{-value}=2P(Z\ge2.40)\approx0.0164\), rounded to four decimal places. Since \(0.0164<0.05\), reject \(H_0\).

Conclude: The sample provides convincing evidence that the proportion of students represented by the district enrollment list who would use the proposed late bus service differs from 50%. The observed sample proportion is 56%, but the test does not prove the population proportion is 56% or prove that the alternative is true with certainty.

Repair the error: “The survey proves that 56% of district students will use the bus” has two problems. It treats the sample percentage as the exact population proportion, and it uses “proves.” It also changes the measured response—saying one would use a proposed service—into a claim about future actual use. The conclusion should preserve the test’s precise characteristic and population.

Common Mistakes and AP Exam Tip

  • Writing only “reject” or “fail to reject”: A decision without a contextual interpretation is incomplete. Add what the data indicate about the population proportion named by \(p\).
  • Writing a sample-only conclusion: “The sample had 63%” is descriptive, not the full inferential conclusion. State what evidence says about the true proportion in the defined population.
  • Claiming the parameter equals the sample proportion: Do not convert \(\hat{p}\) into a known value of \(p\). The test evaluates a population claim; it does not make the sample percentage exact for the population.
  • Leaving out the population or characteristic: “There is convincing evidence of an increase” is unclear if the reader cannot tell which population’s proportion increased or what outcome is counted.
  • Using “accept \(H_0\)” after a large p-value: Say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for \(H_a\). Do not claim that the null has been established.
  • Ignoring equality at the cutoff: The decision rule includes equality: reject when the p-value is less than or equal to \(\alpha\), and fail to reject only when it is greater than \(\alpha\).
  • Claiming proof or certainty: A small p-value may provide convincing evidence against \(H_0\), but it does not prove the alternative. Use cautious evidence language.
  • Changing the alternative in the conclusion: A test with \(H_a:p\ne p_0\) supports a claim of a difference in either direction; it does not test only for an increase or only for a decrease.
AP Exam Tip: For a full-credit conclusion, include the decision, “convincing evidence” language, and a statement about the true population proportion in context. Match the direction to \(H_a\). Reject when the p-value is less than or equal to \(\alpha\); otherwise fail to reject. Never say the test proves a claim.

Key Takeaway

A clear test conclusion is about the population parameter, not just the sample statistic. Name the population and characteristic, match the claim to the alternative hypothesis, and use “convincing evidence” rather than certainty. When the p-value exceeds \(\alpha\), report a failure to find convincing evidence for the alternative; do not treat the null hypothesis as proven.

Key takeaway: State the decision correctly, including rejection when the p-value equals \(\alpha\). Then describe what the data indicate about the true population proportion in context—without presenting the sample proportion as the population value or claiming proof.

Check Your Understanding

For each statement, identify the conclusion-writing problem and describe a more appropriate approach.

  1. A right-tailed test gives a p-value of 0.03 at \(\alpha=0.05\). The student writes, “There is a 97% chance that the alternative hypothesis is true.” What is wrong with this interpretation?
  2. A random sample finds \(\hat{p}=0.42\). The test rejects \(H_0\) in favor of \(H_a:p<0.50\). Why is “42% of the population has the characteristic” not a suitable test conclusion?
  3. A test of \(H_0:p=0.30\) against \(H_a:p\ne0.30\) has a p-value greater than \(\alpha\). Write the key idea that a contextual conclusion should convey.
  4. A test has a p-value exactly equal to its preselected \(\alpha\). Should the decision be reject or fail to reject \(H_0\)?
  5. A report says, “The result proves that more than half of residents support the plan.” Name two ways to make this wording more statistically appropriate.