Turning a P-Value into a Contextual Sentence
In What a P-Value Really Measures, we learned that a p-value is calculated assuming the null hypothesis is true. The next skill is putting that probability into words that describe the actual question and sample. A p-value of 0.018 is a probability of about 1.8%—not 0.018%—for sample results at least as extreme as the observed result under the null model.
A useful interpretation names the population proportion in context, the value assumed by the null hypothesis, the sample result, and what counts as “at least as extreme.” The alternative hypothesis determines the last part: an upper-tail test considers higher results, a lower-tail test considers lower results, and a two-sided test considers results at least as far from the null value in either direction.
This frame is a guide, not a script to repeat without context. Replace every bracketed phrase with details from the problem. For example, say “the proportion of library visitors who borrow an e-book,” not merely “the proportion.” Say “a sample proportion of 0.446 or higher” when the test is right-tailed and the observed sample proportion is 0.446. Those details help a reader see exactly which hypothetical outcomes the probability includes.
What “At Least This Extreme” Means in a Sentence
A p-value includes the observed result and results that are more extreme according to the alternative hypothesis. In a one-proportion \(z\)-test, the result can be described using the sample proportion \(\hat{p}\) or its test statistic \(z\). The contextual sentence should use the direction or directions stated by \(H_a\), not simply the direction in which the sample happened to differ.
- For \(H_a:p>p_0\), describe sample proportions at least as high as the observed \(\hat{p}\).
- For \(H_a:p<p_0\), describe sample proportions at least as low as the observed \(\hat{p}\).
- For \(H_a:p\ne p_0\), describe sample proportions at least as far from \(p_0\) as the observed \(\hat{p}\), in either direction.
In a two-sided test, if \(\hat{p}\) is above \(p_0\), a result at least as extreme may be above or below \(p_0\). In a one-sided test, the opposite tail does not count just because it is far away. As in Sketching the P-Value Region for a Test, the alternative hypothesis sets the rule for extremeness.
Worked Examples: Writing the Interpretation
Each example uses a one-proportion \(z\)-test. The conditions are checked before interpreting the calculated p-value, as in Four-Step Process for a One-Proportion z-Test. The focus here is the final sentence: it should match the population, benchmark, sample, and alternative.
Worked Example: A Higher Rate of E-Book Borrowing
A fictional library system reports that 40% of its visitors borrow an e-book during a particular month. A random sample of 500 visitors includes 223 who borrowed one. A test of whether the true proportion is higher than 0.40 has a p-value of 0.018. Write an interpretation in context.
State: Let \(p\) be the true proportion of visitors to this library system during the month who borrow an e-book. The hypotheses are \(H_0:p=0.40\) and \(H_a:p>0.40\).
Plan: Use a one-proportion \(z\)-test. The visitors were randomly sampled, so the Random condition is met for visitors represented by the sampling process. The sample was selected without replacement from 8,000 visitors, and \(500\leq0.10(8{,}000)=800\), so the 10% condition is met. Under \(H_0\), the expected number who borrow an e-book is \(np_0=500(0.40)=200\), and the expected number who do not is \(n(1-p_0)=500(0.60)=300\). Both are at least 10, so the Large Counts condition is met.
Do: The observed sample proportion is:
Using the null proportion, the standard error and test statistic are:
This is a right-tailed test, so the p-value is the area at or above \(z=2.100\). Using the unrounded standard error, \(\text{normalcdf}(2.0996,1\text{E}99,0,1)\approx0.0179\), which rounds to 0.018 to three decimal places.
Conclude: If the true proportion of visitors who borrowed an e-book that month were 0.40, the probability of getting a sample proportion of 0.446 or higher in a random sample of 500 visitors would be about 0.018. This describes sample results under the null assumption; it is not the probability that the true proportion is 0.40.
Worked Example: A Lower Rate of Missed Appointments
A fictional clinic’s scheduling team uses 60% as a benchmark for the proportion of scheduled appointments that are missed. In a random sample of 500 appointments from a particular period, 277 were missed. A test asks whether the true proportion is lower than 0.60 and reports a p-value of 0.018. What does that p-value mean?
State: Let \(p\) be the true proportion of appointments scheduled during this period at the clinic that were missed. The hypotheses are \(H_0:p=0.60\) and \(H_a:p<0.60\).
Plan: Use a one-proportion \(z\)-test. The appointments were randomly sampled, meeting the Random condition for appointments represented by the sampling process. The sample was selected without replacement from 12,000 appointments, and \(500\leq0.10(12{,}000)=1{,}200\), so the 10% condition is met. Under \(H_0\), the expected missed count is \(500(0.60)=300\), and the expected count not missed is \(500(0.40)=200\). Both expected counts are at least 10, so the Large Counts condition is met.
Do: The sample proportion of missed appointments is:
The null standard error and test statistic are:
Because the alternative is left-tailed, the p-value is the area at or below this observed statistic. Thus, \(\text{normalcdf}(-1\text{E}99,-2.0996,0,1)\approx0.0179\), or 0.018 when rounded to three decimal places.
Conclude: If the true proportion of appointments missed during this period were 0.60, the probability of getting a sample proportion of 0.554 or lower in a random sample of 500 appointments would be about 0.018. The phrase “or lower” is essential: this test asks about a decrease, so a high sample proportion is not counted as more extreme in the direction of the alternative.
Worked Example: A Different Rate of Reusable-Cup Use
A fictional event organizer expects half of attendees to use a reusable cup. A random sample of 2,500 attendees includes 1,309 who used one. A test asks whether the true proportion differs from 0.50 in either direction. The calculator reports a p-value of 0.018 when rounded to three decimal places. Interpret it.
State: Let \(p\) be the true proportion of attendees at this event who use a reusable cup. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\).
Plan: Use a one-proportion \(z\)-test. The attendees were randomly sampled, meeting the Random condition for attendees represented by the sampling process. The sample was selected without replacement from 40,000 attendees, and \(2{,}500\leq0.10(40{,}000)=4{,}000\), so the 10% condition is met. Under \(H_0\), the expected count using a reusable cup is \(2{,}500(0.50)=1{,}250\), and the expected count not using one is also 1,250. Both are at least 10, so the Large Counts condition is met.
Do: The sample proportion is:
The null standard error and test statistic are:
For a two-sided alternative, results at least as far from 0.50 as 0.5236 count in either direction. The two-tail calculation is:
Conclude: If the true proportion of attendees who used a reusable cup were 0.50, the probability of getting a sample proportion at least 0.0236 away from 0.50 in either direction, in a random sample of 2,500 attendees, would be about 0.0183. Rounded to three decimal places, this is 0.018. Because the alternative is two-sided, the interpretation includes both a proportion at least as high as 0.5236 and one at least as low as 0.4764.
Common Mistakes and AP Exam Tips
A p-value interpretation is short, but each phrase has a job. A strong response identifies the null assumption, the population and characteristic, the sample result or its extremeness rule, and the probability. Avoid statements that sound plausible but reverse the conditional probability or leave out the alternative’s direction.
- Reversing the condition: “There is a 1.8% probability that the true proportion is 0.40” is incorrect. The p-value is calculated under the assumption that the true proportion is 0.40; it is a probability about sample results.
- Leaving out the direction: “The probability of getting a sample proportion of 0.446 is 0.018” is not an adequate interpretation of a right-tailed p-value. Include results at least as high as the observed proportion.
- Using “exactly” instead of “at least”: A p-value is not the probability of getting only the observed result. It includes the observed result and more extreme results according to \(H_a\).
- Ignoring context: “The probability is 0.018” does not tell the reader what population or outcome is being discussed. Name the characteristic and the relevant population.
- Describing only the observed direction in a two-sided test: If \(H_a:p\ne p_0\), the p-value includes departures on either side of \(p_0\). Say “at least as far from the null proportion in either direction.”
- Calling it the chance that the result happened randomly: That phrase is vague and can confuse the probability model with a claim about causes. State which sample results the null model would produce and how often.
Key Takeaway
A p-value of 0.018 means that sample results at least as extreme as the observed result would occur about 1.8% of the time under the null model, using the alternative hypothesis to define “extreme.” Its contextual interpretation describes a hypothetical sampling outcome, not the chance that the null hypothesis is true.
Check Your Understanding
For each question, practice making the interpretation conditional on the null hypothesis and specific to the alternative.
- A right-tailed test has \(p_0=0.25\), an observed sample proportion of 0.29, and a p-value of 0.018. What sample proportions count as at least as extreme?
- A left-tailed test has a p-value of 0.018. Write a sentence frame that names the null assumption and the lower-tail direction.
- For a test with \(H_a:p\ne p_0\), what should the contextual interpretation say about the two directions?
- Why is “there is a 1.8% chance that \(H_0\) is true” not a correct interpretation of a p-value of 0.018?
- Rewrite “the probability is 0.018” as a complete p-value interpretation for a population proportion. Include context, sample size, and “at least as extreme.”