Tutorials › AP Statistics › What a P-Value Really Measures

P-values and conclusions for proportions · Tutorial 481 of 1000

What a P-Value Really Measures

Understand what a p-value says about the data if the null hypothesis is true—and why it is not the probability that the null hypothesis is true.

Intermediate 8 min read

What You'll Learn

  • Define a p-value as a conditional probability calculated under the assumption that the null hypothesis is true
  • Explain what “at least as extreme as observed” means for left-tailed, right-tailed, and two-sided tests
  • Connect a one-proportion z statistic to the null model used to calculate a p-value
  • Distinguish the probability of data under a null hypothesis from the probability that the null hypothesis is true
  • Describe what small and large p-values do and do not tell us

The Question Behind a P-Value

In Exam-Style Free Response on a One-Proportion Test, the p-value was one part of a complete test. Now we will slow down and examine what that number means. A p-value is not a measure of how likely the null hypothesis is to be true. It describes how unusual the data would be if the null hypothesis were true.

That “if” matters. To calculate a p-value, we temporarily assume the null hypothesis is true and use it to model the results that could occur from sample to sample. We then measure the probability of observing a result at least as extreme as the one in our sample, using the direction specified by the alternative hypothesis.

Definition: A p-value is the probability, assuming the null hypothesis is true, of obtaining a result at least as extreme as the observed result in the direction or directions specified by the alternative hypothesis.

In a one-proportion \(z\)-test, the null hypothesis supplies the reference proportion \(p_0\). If the conditions for the test are satisfied, the null model lets us approximate the sampling distribution of \(\hat{p}\) with a Normal model centered at \(p_0\). We can express the p-value as a conditional probability:

$$ P(\text{a result at least as extreme as observed}\mid H_0\text{ is true}) $$

The vertical bar means “given that.” It signals that this probability is calculated under the assumption that \(H_0\) is true. For the one-proportion \(z\)-test, we usually find the probability using the test statistic \(z\) and the standard Normal model. The alternative hypothesis determines which part of that model counts as “at least as extreme.”

What Counts as “At Least as Extreme”?

A result is extreme when it is far from what the null model predicts in the direction relevant to the question. This is not simply “a result farther from the sample result.” The alternative hypothesis sets the rule for which results count.

  • For \(H_a:p>p_0\), results at or above the observed statistic count. Use the upper tail.
  • For \(H_a:p<p_0\), results at or below the observed statistic count. Use the lower tail.
  • For \(H_a:p\ne p_0\), results at least as far from the null value as the observed statistic count in either direction. Use both tails.

This is why the p-value is not determined by the sample result alone. A sample proportion above \(p_0\), for example, leads to different p-value calculations depending on whether the question asks about an increase or a difference in either direction. As emphasized in Sketching the P-Value Region for a Test, follow the alternative hypothesis when identifying the p-value region.

Key idea: A p-value measures how often the null model would produce results at least as extreme as the observed result—not how often the null hypothesis itself would be true.

Three Worked Examples

Worked Example: A Right-Tailed Test of a Recycling Rate

A fictional town’s recycling team wants to know whether more than 40% of households in a neighborhood regularly separate recyclable materials. A random sample of 250 households from the neighborhood’s 5,000 households includes 115 that do so. What does the p-value for a one-proportion \(z\)-test measure?

State: Let \(p\) be the true proportion of households in the neighborhood that regularly separate recyclable materials. The question asks whether \(p\) is greater than 0.40, so \(H_0:p=0.40\) and \(H_a:p>0.40\).

Plan: Use a one-proportion \(z\)-test. The sample is random, so the Random condition is met for the neighborhood households represented by the sampling process. The sample was selected without replacement from 5,000 households, and \(250\leq0.10(5{,}000)=500\), so the 10% condition is met. Under \(H_0\), the expected success count is \(np_0=250(0.40)=100\), and the expected failure count is \(n(1-p_0)=250(0.60)=150\). Both counts are at least 10, so the Large Counts condition is met.

Do: The sample proportion is \(115/250=0.46\). Under the null hypothesis, the standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.40(0.60)}{250}} =\sqrt{0.00096} \approx0.03098 $$
$$ z=\frac{0.46-0.40}{\sqrt{0.40(0.60)/250}} \approx1.9365 $$

Because the alternative is right-tailed, results with a \(z\)-statistic of 1.9365 or greater count as at least as extreme. The p-value is \(\text{normalcdf}(1.9365,1\text{E}99,0,1)\approx0.0264\), rounded.

Interpret: If the true proportion of households that regularly separate recyclable materials were 0.40, the probability of getting a sample result at least as high as the observed 46%—or a still higher result—in a random sample of 250 households would be about 0.0264 under the test model. The p-value is not the probability that \(p=0.40\).

Worked Example: A Left-Tailed Test of Delivery Delays

A fictional technology supplier claims that 30% of its components require a delivery delay. A quality team wonders whether the proportion for a particular production period is lower. A random sample of 200 components from that period includes 48 with a delivery delay. Describe the p-value.

State: Let \(p\) be the true proportion of components from this production period that require a delivery delay. The question asks whether the proportion is lower than 0.30, so \(H_0:p=0.30\) and \(H_a:p<0.30\).

Plan: Use a one-proportion \(z\)-test. The components were randomly sampled, meeting the Random condition for the production period. The sample was selected without replacement from 3,000 components, and \(200\leq0.10(3{,}000)=300\), so the 10% condition is met. Under \(H_0\), the expected delay count is \(200(0.30)=60\), and the expected count without a delay is \(200(0.70)=140\). Both are at least 10, so the Large Counts condition is met.

Do: The sample proportion with delays is \(48/200=0.24\). Using the null proportion:

$$ SE_0=\sqrt{\frac{0.30(0.70)}{200}} =\sqrt{0.00105} \approx0.03240 $$
$$ z=\frac{0.24-0.30}{\sqrt{0.30(0.70)/200}} \approx-1.8516 $$

Because the alternative is left-tailed, results with a \(z\)-statistic of \(-1.8516\) or less count as at least as extreme. The p-value is \(\text{normalcdf}(-1\text{E}99,-1.8516,0,1)\approx0.0320\), rounded.

Interpret: If 30% of components from this production period truly require a delivery delay, the probability of getting a sample proportion as low as 0.24 or lower in a random sample of 200 components would be about 0.0320 under the test model. This probability describes sample results under the null assumption; it does not give the probability that the supplier’s claim is true.

Worked Example: A Two-Sided Test of a Battery Return Rate

A fictional battery maker uses 50% as a benchmark for the proportion of returned batteries that can be refurbished. A random sample of 400 returned batteries from a production period includes 224 that can be refurbished. The question is whether the true proportion differs from 50%, in either direction.

State: Let \(p\) be the true proportion of returned batteries from this production period that can be refurbished. The question asks whether \(p\) differs from 0.50, so \(H_0:p=0.50\) and \(H_a:p\ne0.50\).

Plan: Use a one-proportion \(z\)-test. The returned batteries were randomly sampled, meeting the Random condition for those represented by the sampling process. The sample was selected without replacement from 10,000 returned batteries, and \(400\leq0.10(10{,}000)=1{,}000\), so the 10% condition is met. Under \(H_0\), the expected refurbishable count is \(400(0.50)=200\), and the expected count not refurbishable is also 200. Both are at least 10, so the Large Counts condition is met.

Do: The sample proportion is \(224/400=0.56\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.50(0.50)}{400}}=0.025 \qquad z=\frac{0.56-0.50}{0.025}=2.4 $$

For a two-sided alternative, results at least as far from zero as \(2.4\) count in either direction. Thus:

$$ \text{p-value} =2P(Z\geq2.4) =2\text{normalcdf}(2.4,1\text{E}99,0,1) \approx0.0164 $$

Interpret: If the true refurbishable proportion were 0.50, the probability of observing a sample proportion at least 0.06 away from 0.50 in either direction, in a random sample of 400, would be about 0.0164 under the test model. The two tails are included because the research question allows either a higher or lower proportion.

What the P-Value Does—and Does Not—Say

The examples show the same basic logic with different alternatives. Each calculation begins with the null model, locates the observed result, and measures the probability in the relevant tail or tails. A small p-value means that results at least as extreme as the observed result would be unusual if the null hypothesis were true. That can provide evidence against \(H_0\), but it does not prove \(H_0\) false.

A large p-value means that the observed result is not especially unusual under the null model. It does not show that \(H_0\) is true, nor does it measure the probability that random chance alone caused the result. Statistical tests account for chance variation through their model; the p-value is a probability of data outcomes under that model.

The p-value also does not measure the size or practical importance of a difference. In the examples, the p-value is tied to both the observed departure from \(p_0\) and the sample size, through the test statistic and its null standard error. A small p-value can reflect a modest difference measured precisely, while practical importance requires considering the size and consequences of the difference in context.

Keep the condition in mind: Say, “Assuming \(H_0\) is true, the probability of a result at least as extreme as the one observed is …” Do not say, “The probability that \(H_0\) is true is …”

Common Mistakes and AP Exam Tips

  • Reversing the conditional statement. The p-value is calculated assuming \(H_0\) is true. It is not the probability that the null hypothesis is true given the data.
  • Leaving out “at least as extreme.” The p-value includes the observed result and results more extreme according to the alternative. It is not just the probability of getting exactly the observed sample proportion.
  • Using the wrong tail. The alternative hypothesis determines which results count. A two-sided alternative includes both tails, even when the observed result is on just one side of \(p_0\).
  • Calling the p-value the chance result happened “by chance.” That phrase is too vague. A full-credit explanation names the null assumption, the relevant outcomes, and the probability.
  • Treating a large p-value as proof of the null. A large p-value means the data do not provide strong evidence against \(H_0\); it does not establish that \(p=p_0\).
AP Exam Tip: When asked to interpret a p-value, include the context, the assumption that \(H_0\) is true, the sample size or sampling process, and the phrase “at least as extreme as observed.” Keep the interpretation about sample results—not the probability of a hypothesis.

Key Takeaway

The p-value connects the observed sample result to a probability model built from the null hypothesis. It tells us how unusual the result, or a more extreme result in the direction set by \(H_a\), would be if that null model were true.

Key takeaway: A p-value is a conditional probability about data: it is calculated assuming \(H_0\) is true. It is not the probability that \(H_0\) is true.

Check Your Understanding

For each question, focus on the condition in the definition of a p-value.

  1. In your own words, what does the phrase “assuming the null hypothesis is true” tell us about how a p-value is calculated?
  2. A right-tailed test has an observed \(z\)-statistic of 1.7. Which results count as at least as extreme as observed?
  3. For \(H_a:p\ne p_0\), why are results in both tails included in the p-value?
  4. A p-value is 0.08. Does this mean there is an 8% probability that \(H_0\) is true? Explain.
  5. Write a contextual p-value interpretation for a test about a population proportion. Include the null assumption and the phrase “at least as extreme as observed.”