Tutorials › AP Statistics › Misinterpretations of the P-Value

P-values and conclusions for proportions · Tutorial 490 of 1000

Misinterpretations of the P-Value

Learn to spot two common p-value misinterpretations and explain how a p-value differs from the significance level and the long-run risk of a Type I error.

Intermediate 10 min read

What You'll Learn

  • Explain why a p-value is calculated assuming the null hypothesis is true, rather than giving the probability that it is true.
  • Distinguish a p-value from the probability of making a Type I error.
  • Describe how the significance level alpha relates to the long-run Type I error rate.
  • Interpret p-values and test decisions accurately in context.
  • Explain why a smaller significance cutoff cannot increase the long-run Type I error rate.

Two Probabilities That a P-Value Does Not Give

In What a P-Value Really Measures and Interpreting a P-Value in Context, you learned to describe a p-value as a probability about sample results, calculated under the assumption that the null hypothesis is true. This conditional wording matters. It does not say how likely the null hypothesis is to be true, and it does not tell us the chance that a test decision is a Type I error.

These mix-ups often happen because a p-value, a null hypothesis, and an error probability all appear in the same test. But they answer different questions. A p-value describes how unusual the data, or more extreme data, would be under a null model. A Type I error is a particular kind of incorrect decision: rejecting a null hypothesis that is in fact true. The significance level \(\alpha\) sets a rule for deciding when to reject.

Definition: A p-value is the probability, assuming \(H_0\) is true, of obtaining a sample result at least as extreme as the observed result in the direction or directions specified by \(H_a\). It is not the probability that \(H_0\) is true, nor the probability that this test’s decision is a Type I error.

The distinction is about what is assumed and what is being treated as uncertain. To calculate a p-value, we assume the null hypothesis and use its model to ask how often results like the observed one would occur. We do not use the p-value to calculate a probability for the hypothesis itself.

A P-Value Is Not the Probability That \(H_0\) Is True

Suppose a test produces a p-value of 0.04. The correct interpretation is not “there is a 4% chance that \(H_0\) is true.” Instead, if \(H_0\) were true, the probability of getting a result at least as extreme as the observed result, according to the test’s alternative, would be 4%.

The difference can be expressed as the difference between two questions:

  • The p-value question: Assuming \(H_0\) is true, how likely are data at least this extreme?
  • The mistaken question: Given these data, how likely is it that \(H_0\) is true?

A significance test answers the first question. It does not answer the second. In an AP Statistics test, use the p-value to describe evidence against \(H_0\); do not turn it into a probability that \(H_0\) is true or false. As explained in Why We Never Accept the Null Hypothesis, failing to reject \(H_0\) does not prove it true, and rejecting it does not prove it false with certainty.

For example, a small p-value indicates that the observed result would be unusual if \(H_0\) were true. That can provide evidence against \(H_0\). It does not mean that \(H_0\) has a probability equal to the p-value. A large p-value means the result is not especially unusual under the null model; it does not establish that the null model is correct.

A P-Value Is Not the Chance of a Type I Error

A Type I error occurs when a test rejects \(H_0\) even though \(H_0\) is true. Its probability is connected to the significance level \(\alpha\), not to the p-value calculated from one particular sample.

Definition: The significance level \(\alpha\) is the preselected cutoff used to decide whether to reject \(H_0\). It controls the long-run probability of a Type I error when \(H_0\) is true. A test’s p-value is calculated from the observed data; \(\alpha\) is chosen as part of the decision rule.

As covered in Choosing a Significance Level Before Testing, choosing \(\alpha=0.05\) means using a rule designed to limit the long-run Type I error rate to about 5% or less when the null hypothesis is true and the procedure is used appropriately. It does not mean there is a 5% chance that \(H_0\) is true, nor does it mean that any particular rejection has a 5% chance of being a Type I error.

The long-run interpretation refers to repeatedly applying the same testing rule in situations where \(H_0\) is true. It describes the rate of incorrect rejections across those uses. A single test either rejects \(H_0\) or fails to reject it; its p-value is not a personal risk score for whether that decision is wrong.

For a test with a discrete set of possible results, the actual long-run Type I error rate at a chosen cutoff can be less than \(\alpha\). So it is safest to say that \(\alpha\) controls or sets an upper bound for the long-run Type I error probability, rather than claiming every procedure achieves that exact rate.

Worked Examples: Correcting the Misinterpretations

Worked Example: A Small P-Value in a Community Survey

A fictional town takes a random sample of 200 households from a list of 5,000 and asks whether they support a proposed bus route. In the sample, 120 support it. The town tests whether the true proportion of households on the list who support the route differs from 50%. Explain the p-value and distinguish it from the chance of a Type I error.

State: Let \(p\) be the true proportion of households on the town’s list that support the proposed route. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\).

Plan: The random sample meets the Random condition. Because \(200\leq0.10(5{,}000)=500\), the 10% condition is met. Under \(H_0\), the expected numbers of supporters and nonsupporters are \(200(0.50)=100\) each, so both meet the Large Counts condition. The conditions for a one-proportion \(z\)-test are met.

Do: The sample proportion is \(\hat{p}=120/200=0.60\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.50(1-0.50)}{200}}=0.03536,\qquad z=\frac{0.60-0.50}{0.03536}=2.828 $$

For a two-sided test, the p-value is twice the standard Normal upper-tail area beyond \(2.828\): \(2P(Z\geq2.828)\approx0.0047\), rounded. Assuming that 50% of households on the list support the route, results at least as far from 50% as the observed result would occur about 0.47% of the time under the test’s Normal model.

Conclude: This p-value is not a 0.47% chance that the true proportion is 50%, and it is not the chance that rejecting \(H_0\) is a Type I error. It describes how unusual the sample result would be if \(H_0\) were true. At a preselected \(\alpha=0.05\), the town would reject \(H_0\); that decision provides convincing evidence that the proportion of households on the list who support the route differs from 50%.

If the null hypothesis were true and the town repeated this testing rule many times, \(\alpha\) would describe the rule’s long-run Type I error rate or upper bound. The p-value of 0.0047 comes from this sample, not from that long-run error calculation.

Worked Example: Correcting “There Is an 18% Chance the Null Is True”

A fictional wildlife group tests whether the true proportion of tagged turtles in a specified population that return to a nesting beach is greater than 40%. Conditions for a one-proportion \(z\)-test have been verified, and the test report gives a p-value of 0.18. A volunteer says, “There is an 18% chance the true proportion is 40%.” Is that interpretation correct?

State: Let \(p\) be the true proportion of tagged turtles in the specified population that return to the nesting beach. The hypotheses are \(H_0:p=0.40\) and \(H_a:p>0.40\).

Plan: The reported p-value is for a right-tailed test. Conditions are stated to have been verified, so we can interpret the p-value using the null model for this test. No new probability model for the hypothesis itself is being calculated.

Do: The p-value of 0.18 means that, if 40% of tagged turtles in the population return, the probability of obtaining a sample result at least as high as the observed one is about 18% under the test model.

Conclude: The volunteer’s interpretation is incorrect. The p-value is not the probability that the true proportion equals 40%. It is a probability about sample results, assuming that proportion is 40%. Because 0.18 is greater than \(\alpha=0.05\), the group would fail to reject \(H_0\). The data do not provide convincing evidence that the return proportion exceeds 40%; they also do not prove that it equals 40%.

Worked Example: Correcting “The P-Value Is the Type I Error Chance”

A fictional lab tests whether the true proportion of a product’s packages that meet a quality standard differs from 95%. Its test gives a p-value of 0.03. The lab chose \(\alpha=0.05\) before collecting data. A staff member says, “There is a 3% chance our rejection is a Type I error.” Explain the mistake, and compare the decision rules for \(\alpha=0.05\) and \(\alpha=0.01\).

State: The p-value is 0.03 for the test of the stated null proportion against a two-sided alternative. A Type I error would occur if the lab rejected that true null hypothesis.

Plan: The p-value is interpreted under the assumption that \(H_0\) is true. The significance level is a cutoff chosen before seeing the data. To compare the cutoffs, consider the rejection sets: any result that meets the stricter \(0.01\) cutoff also meets the \(0.05\) cutoff.

Do: Since \(0.03\leq0.05\), the lab rejects \(H_0\) using its chosen rule. Since \(0.03>0.01\), it would fail to reject \(H_0\) under a rule with \(\alpha=0.01\). The p-value describes the chance, assuming \(H_0\) is true, of a result at least as extreme as the observed one. The Type I error probability describes the long-run behavior of a rejection rule when \(H_0\) is true.

Conclude: The staff member has mistaken the p-value for the probability that this particular rejection is a Type I error. The value 0.03 does not give that probability. With \(\alpha=0.05\), the lab’s result provides convincing evidence that the true proportion differs from 95%, provided the test conditions hold.

When \(H_0\) is true, the long-run Type I error rate for the \(0.01\) rule cannot be greater than the rate for the \(0.05\) rule, because the stricter rule’s rejection set is contained in the other’s. The rate may be equal, and the stricter rule may reject fewer null hypotheses. For a discrete test, do not assume the actual rates must equal 0.01 and 0.05.

Common Mistakes and AP Exam Tips

  • Reversing the conditional probability: “Assuming \(H_0\) is true, the chance of results this extreme is the p-value” is correct. “Given the results, the chance \(H_0\) is true is the p-value” reverses the condition and is incorrect.
  • Calling the p-value the chance of a Type I error: A p-value describes the observed data under the null model. The long-run Type I error rate is controlled by the preselected decision rule and its \(\alpha\).
  • Treating \(\alpha\) as the probability \(H_0\) is true: Alpha is not a probability assigned to the null hypothesis. It is a cutoff used to decide when to reject.
  • Claiming a decision is certainly correct: Rejecting \(H_0\) does not guarantee that \(H_0\) is false. A Type I error remains possible when a true null hypothesis is rejected.
  • Reporting that a large p-value proves the null: Failing to reject \(H_0\) means the data do not provide convincing evidence for \(H_a\); it does not establish the null claim.
  • Equating actual error rates with every chosen cutoff: For discrete tests, the actual long-run Type I error rate can be below the nominal \(\alpha\). When comparing two cutoffs, a smaller cutoff cannot increase the rejection set or the long-run error rate, but the rate can be equal.
AP Exam Tip: Use a conditional sentence for a p-value: “Assuming [the null claim] is true, the probability of obtaining [the observed result or a more extreme result in the direction(s) of the alternative] is [the p-value].” Use a separate sentence to describe the decision rule and its long-run Type I error rate.

Key Takeaway

A p-value is a probability about sample results calculated under the assumption that \(H_0\) is true. It is not the probability that \(H_0\) is true, and it is not the chance that this test’s decision is a Type I error. The significance level \(\alpha\), chosen before examining the data, controls the long-run Type I error rate of the decision rule when \(H_0\) is true.

Key takeaway: Keep the probabilities straight: the p-value concerns data under \(H_0\); \(\alpha\) concerns the long-run Type I error behavior of a testing rule under \(H_0\).

Check Your Understanding

For each statement, identify what the p-value or significance level actually describes.

  1. A one-proportion test reports \(p\text{-value}=0.02\). Write a correct general interpretation of this value, including the assumption it uses.
  2. Explain why a p-value of 0.02 does not mean there is a 2% chance that \(H_0\) is true.
  3. In a test with \(\alpha=0.05\), explain what the significance level says about Type I errors over repeated uses when \(H_0\) is true.
  4. A student says, “Our p-value was 0.08, so there is an 8% chance our failure to reject \(H_0\) is wrong.” Explain why this is not a correct interpretation.
  5. When comparing rules with \(\alpha=0.01\) and \(\alpha=0.05\), can the first rule have a greater long-run Type I error rate under a true null? Could the rates be equal? Explain.