Turn a Test Result Into an Exam-Ready Answer
A free-response question may give you the hypotheses, a p-value, and a significance level, then ask you to interpret the p-value and state a conclusion. These are two related but different tasks. An interpretation explains what the p-value says about sample results under the null hypothesis. A conclusion compares the p-value with \(\alpha\), reports the test decision, and describes what the evidence indicates about the population proportion in context.
As in Interpreting a P-Value in Context and Writing a Complete Conclusion With Evidence, the alternative hypothesis determines which results count as evidence against \(H_0\). The key exam skill is to carry that direction through both parts of your answer: first into the p-value interpretation, then into the conclusion.
The interpretation is not a decision, and it does not state the probability that a hypothesis is true. The conclusion is not merely a repeat of the p-value. Keeping these jobs separate helps you answer both parts clearly without overclaiming.
A Short Routine for Reading the Prompt
Before writing, identify the population proportion \(p\), the null value \(p_0\), the direction of \(H_a\), the p-value, and \(\alpha\). You do not need to recalculate a statistic when the prompt gives a p-value and asks only for its interpretation and conclusion. You do need to use the hypotheses to describe what “at least as extreme” means.
For \(H_a:p<p_0\), focus on results at or below the observed result. For \(H_a:p>p_0\), focus on results at or above it. For \(H_a:p\ne p_0\), include results at least as far from \(p_0\) in either direction.
Begin with “Assuming \(H_0\) is true...” Then describe the sample results counted by the p-value and give the probability in the setting.
Reject \(H_0\) when the p-value is less than or equal to \(\alpha\). If the p-value is greater than \(\alpha\), fail to reject \(H_0\).
Say whether the data provide convincing evidence for the alternative, naming the population and characteristic in context.
This routine builds on the decision rule in Comparing P-Value to the Significance Level and the conclusion language in the previous tutorial, Frequent Errors in Writing Test Conclusions. The worked examples show how to use the routine when the prompt provides a p-value and when you calculate one as part of a complete test.
Worked Examples
Worked Example: A Left-Tailed Test With a P-Value Near the Cutoff
Exam-style prompt: A fictional county library system randomly samples 160 cardholders from a registry of 3,200. Of those sampled, 54 say they use digital audiobooks at least once a month. The system wants to know whether fewer than 40% of cardholders use them monthly. The one-proportion \(z\)-test gives a p-value of \(0.0533\). At \(\alpha=0.05\), interpret the p-value and state a conclusion.
State: Let \(p\) be the true proportion of cardholders represented by the county library registry who use digital audiobooks at least once a month. The hypotheses are \(H_0:p=0.40\) and \(H_a:p<0.40\).
Plan and check conditions: The registry sample is described as random, meeting the Random condition. The sample size is less than 10% of the registry because \(160<0.10(3200)=320\), so the 10% condition is met. Under the null hypothesis, the expected numbers of successes and failures are \(160(0.40)=64\) and \(160(0.60)=96\). Both are at least 10, so the Large Counts condition is met.
Do: The observed sample proportion is
The test statistic and lower-tail p-value are
Interpret the p-value: Assuming the true proportion of cardholders who use digital audiobooks monthly is 40%, the probability of getting a random sample of 160 with a sample proportion of about 0.3375 or lower is approximately 0.0533. The “or lower” direction comes from \(H_a:p<0.40\).
Conclude: Since \(0.0533>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that fewer than 40% of cardholders represented by the county library registry use digital audiobooks at least once a month.
The p-value is only slightly above the cutoff, but the decision rule still says to fail to reject at \(\alpha=0.05\). It would be incorrect to say the result proves that the true proportion is 40%, or that the data provide convincing evidence for a decrease.
Worked Example: One P-Value, Two Significance Levels
Exam-style prompt: A fictional software company randomly samples 100 customers from a list of 3,000 customers eligible for an optional security feature. Fifty-nine say they have enabled it. The company tests whether more than half of eligible customers have enabled the feature. The test produces \(p\text{-value}=0.0359\). Interpret the p-value, then state the conclusion at \(\alpha=0.05\) and at \(\alpha=0.01\).
State: Let \(p\) be the true proportion of customers on the eligibility list who have enabled the security feature. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\).
Plan and check conditions: The customers were selected at random, meeting the Random condition. Because \(100<0.10(3000)=300\), the 10% condition is met. Under \(H_0\), the expected counts are \(100(0.50)=50\) enabled and \(100(0.50)=50\) not enabled, satisfying the Large Counts condition.
Do: The observed proportion and test statistic are
For the right-tailed alternative, \(P(Z\ge1.80)\approx0.0359\), rounded to four decimal places. Assuming exactly half of eligible customers have enabled the feature, a sample of 100 with a sample proportion of 0.59 or higher would occur about 3.59% of the time.
Conclude at \(\alpha=0.05\): Since \(0.0359\le0.05\), reject \(H_0\). The data provide convincing evidence that more than half of customers on the eligibility list have enabled the security feature.
Conclude at \(\alpha=0.01\): Since \(0.0359>0.01\), fail to reject \(H_0\). At this stricter significance level, the data do not provide convincing evidence that more than half of customers on the eligibility list have enabled the feature.
The p-value did not change; the significance level did. This is why an exam response must use the \(\alpha\) specified for that part rather than relying on a decision made at another cutoff.
Worked Example: A Two-Sided Alternative Requires Both Directions
Exam-style prompt: A fictional utility randomly samples 200 residential customers from a list of 5,000. Ninety-six say they have chosen paperless billing. The utility tests whether the proportion of customers on the list who choose paperless billing differs from 40%. The test gives a p-value of \(0.0209\). Interpret the p-value and state a conclusion at \(\alpha=0.05\).
State: Let \(p\) be the true proportion of customers represented by the utility’s list who have chosen paperless billing. The hypotheses are \(H_0:p=0.40\) and \(H_a:p\ne0.40\).
Plan and check conditions: The sample is random, meeting the Random condition. Since \(200<0.10(5000)=500\), the 10% condition is met. Under the null hypothesis, the expected numbers of customers with and without paperless billing are \(200(0.40)=80\) and \(200(0.60)=120\). Both are at least 10, so the Large Counts condition is met.
Do: The sample proportion and null standard error are
Thus,
Interpret the p-value: Assuming the true proportion of customers on the utility’s list who chose paperless billing is 40%, the probability of a random sample of 200 producing a sample proportion at least as far from 0.40 as 0.48 is approximately 0.0209. Because the alternative is two-sided, the results counted include outcomes at least as far below 0.40 as the observed result is above it.
Conclude: Since \(0.0209\le0.05\), reject \(H_0\). The data provide convincing evidence that the proportion of customers represented by the utility’s list who chose paperless billing differs from 40%.
The sample proportion is above 40%, but the alternative asks whether the population proportion differs from 40% in either direction. The conclusion should not change the two-sided claim into “more than 40%.”
Common Mistakes and AP Exam Tip
- Interpreting the p-value as a chance hypothesis: “There is a 5.33% chance that \(H_0\) is true” is not an interpretation. State the assumption that \(H_0\) is true, then describe the probability of sample results at least as extreme as the observed result.
- Leaving out the direction of extremeness: For a one-sided test, specify the tail matching \(H_a\). For a two-sided test, describe results far enough from the null value in either direction.
- Giving only the decision: “Reject \(H_0\)” does not finish the conclusion. Add whether the data provide convincing evidence for the alternative and state that claim in context.
- Using the wrong population statement: A conclusion about the sample’s percentage does not answer an inference question about \(p\). Name the population represented by the sample and the characteristic being counted.
- Ignoring the stated \(\alpha\): Compare the p-value with the cutoff in the specific prompt. A p-value can lead to different decisions at different significance levels.
- Changing a two-sided claim into a one-sided claim: If \(H_a:p\ne p_0\), the conclusion is about a difference, not specifically an increase or decrease—even if the sample proportion is on one side of \(p_0\).
- Forgetting equality: Reject \(H_0\) when the p-value equals \(\alpha\), as well as when it is smaller.
Key Takeaway
A strong exam response keeps the p-value interpretation and the test conclusion distinct but connected. The alternative determines which sample results the p-value counts; the comparison with \(\alpha\) determines whether to reject or fail to reject; and the conclusion states what the evidence indicates about the population proportion in context.
Check Your Understanding
For each prompt, identify the correct interpretation or conclusion language.
- A left-tailed test has \(H_a:p<0.30\) and a p-value of 0.041. In a contextual p-value interpretation, which direction of sample results should be described?
- A test has a p-value of 0.08 and uses \(\alpha=0.05\). State the decision and the general evidence conclusion.
- A two-sided test has \(H_a:p\ne0.60\), and the sample proportion is 0.66. Why should the conclusion say “differs from 60%” rather than simply “is greater than 60%”?
- A p-value is exactly 0.02, and the stated significance level is \(\alpha=0.02\). What is the test decision?
- Rewrite this p-value interpretation: “There is a 3% chance that the null hypothesis is correct.” Include the null assumption and what the probability describes.