Tutorials › AP Statistics › Common Misinterpretations of the P-Value

P-values and mean-inference conclusions · Tutorial 743 of 1000

Common Misinterpretations of the P-Value

Learn what a p-value does and does not say, including how the null hypothesis, observed t statistic, degrees of freedom, and alternative hypothesis shape its interpretation.

Intermediate 9 min read

What You'll Learn

  • Explain why a p-value is conditional on the null hypothesis rather than a probability that the null hypothesis is true
  • Distinguish the probability of a result at least as extreme as observed from the probability of the exact observed result
  • Describe how the observed t statistic, degrees of freedom, and alternative hypothesis determine a t-test p-value
  • Explain why a p-value does not measure the size or practical importance of an effect
  • Write conclusions that separate evidence against the null from proof of an alternative

What a P-Value Can—and Cannot—Tell You

In “Interpreting a P-Value From a t Test in Context,” you learned to describe a p-value as a probability calculated under a null hypothesis. That “under” matters: the calculation assumes the null hypothesis is true. A p-value describes how unusual a test statistic at least as extreme as the observed one would be under that assumption. It does not tell us the probability that the null hypothesis itself is true.

Several tempting statements reverse or stretch that meaning. A p-value is not the chance the null hypothesis is true, the chance the alternative is true, or the probability that the exact sample result will occur again. Nor does it directly measure how large or important an effect is. Careful interpretation keeps the probability attached to the test statistic and the null model that produced it.

Definition: For a t test, the p-value is the probability, assuming the null hypothesis is true, of obtaining a t statistic at least as extreme as the observed t statistic in the direction or directions specified by the alternative hypothesis. The p-value is determined by the observed t statistic, the degrees of freedom, and the alternative’s direction; the null hypothesis specifies the reference value used to calculate the statistic.

As in “What a P-Value Says About Sample Means,” “at least as extreme” is about the test statistic, not only the exact sample mean observed. A one-sided alternative counts results in its specified direction. A two-sided alternative counts results in either direction that are at least as far from the null value. The degrees of freedom matter because they identify the reference t distribution used to find those probabilities.

Misinterpretation 1: “The P-Value Is the Probability the Null Is True”

This is the most important error to correct. If a test gives \(p=0.04\), it does not mean there is a 4% chance that the null hypothesis is true. The p-value answers a conditional question: If the null hypothesis were true, how likely would a test statistic at least as extreme as the observed one be? It does not answer the reversed question: “Given the test result, how likely is the null hypothesis to be true?”

The distinction is like the difference between asking how likely a particular kind of evidence is under an assumption and asking how likely the assumption is after seeing the evidence. A standard t test calculates the first probability. It does not assign probabilities to the competing hypotheses. So avoid sentences such as “There is a 4% chance the population mean equals the null value” or “The alternative hypothesis has a 96% chance of being true.”

Worked Example: Does a Mean Processing Time Differ From a Target?

A fictional archive randomly selects 16 document requests and records each request’s processing time. The sample mean is 21 minutes, and the sample standard deviation is 4 minutes. The archive wants to test whether the true mean processing time differs from 19 minutes. Assume the 16 requests are less than 10% of all requests in the target period, and a plot shows no strong skewness or outliers. Use \(\alpha=0.05\).

1
State.
Let \(\mu\) be the true mean processing time, in minutes, for all document requests in the target period. The hypotheses are \(H_0:\mu=19\) minutes and \(H_a:\mu\ne19\) minutes.
2
Plan and check conditions.
Use a one-sample t test. The requests were randomly selected, supporting inference to the target period. The 10% condition holds, supporting independence when sampling without replacement. Since \(n=16\), examine the data’s shape; the plot shows no strong skewness or outliers, so the t procedure is reasonable.
3
Do.
Calculate the standard error, t statistic, and two-sided p-value. Because the alternative is two-sided, statistics at least as far from zero as the observed statistic count in both tails.
$$ SE_{\bar{x}}=\frac{s}{\sqrt{n}} =\frac{4}{\sqrt{16}} =1\text{ minute}, \qquad t=\frac{\bar{x}-\mu_0}{SE_{\bar{x}}} =\frac{21-19}{1} =2.00, \qquad df=16-1=15. $$

Using the t distribution with 15 degrees of freedom, the two-sided p-value is \(P(T_{15}\leq-2.00)+P(T_{15}\geq2.00)\approx0.0639\), rounded. This means that if the true mean processing time is 19 minutes, the probability of obtaining a t statistic of \(-2.00\) or less, or \(2.00\) or greater, is about 0.0639.

4
Conclude in context.
Because \(0.0639>0.05\), fail to reject \(H_0\). These data do not provide convincing evidence that the true mean processing time differs from 19 minutes.

It would be incorrect to say, “There is a 6.39% probability that the true mean processing time is 19 minutes.” The test assumed that value to calculate the p-value; it did not calculate a probability that the value is true. It would also be incorrect to say that the result proves the mean is 19 minutes. Failing to reject means these data do not provide convincing evidence against the null at the chosen significance level.

Misinterpretation 2: “The P-Value Is the Chance of the Exact Result”

A p-value does not usually describe the probability of observing the exact sample mean or exact t statistic obtained. Instead, it sums the probability of the observed statistic and the more extreme statistics counted by the alternative. For a continuous t distribution, the probability of getting one exact t value is zero; probabilities are calculated for ranges of values, such as a tail beyond the observed statistic.

The phrase “at least as extreme” is essential. For a right-tailed test, a statistic larger than the observed positive t statistic is at least as extreme in the relevant direction. For a left-tailed test, smaller statistics count. For a two-sided test, both tails count. The alternative hypothesis determines which of these outcomes belong in the p-value.

AP exam tip: Do not write “the chance of getting this result by chance.” Name the null assumption and describe the test statistics counted. For example: “Assuming the null hypothesis is true, the probability of obtaining a t statistic at least as far from zero as the observed value, in either direction, is …”

Misinterpretation 3: “The Same T Statistic Always Gives the Same P-Value”

A t statistic alone is not enough to identify a p-value. The reference t distribution depends on the degrees of freedom, and the alternative determines whether one tail or two tails are counted. In other words, the observed t statistic, degrees of freedom, and alternative’s direction all matter. This is one reason an interpretation should not report a t statistic without the relevant degrees of freedom and test direction.

Worked Example: Same T Statistic, Different Degrees of Freedom

Consider two hypothetical two-sided t tests. Both have an observed statistic of \(t=2.00\), but one uses 5 degrees of freedom and the other uses 30 degrees of freedom. Compare their p-values to see why the t statistic by itself is not enough.

For the first test, the two-sided p-value is the probability of a t statistic at most \(-2.00\) or at least \(2.00\) from a t distribution with 5 degrees of freedom:

$$ P(T_5\leq-2.00)+P(T_5\geq2.00)\approx0.1019. $$

For the second test, use the t distribution with 30 degrees of freedom:

$$ P(T_{30}\leq-2.00)+P(T_{30}\geq2.00)\approx0.0546. $$

The observed statistic and alternative are the same, but the p-values differ because the reference t distributions differ. The degrees of freedom determine which t distribution is used. Also, if either test instead had a one-sided alternative in the direction of the positive statistic, only the right tail would count: the p-values would be about \(0.0510\) for 5 degrees of freedom and \(0.0273\) for 30 degrees of freedom. Thus a contextual interpretation needs the null reference, the observed statistic, the degrees of freedom, and the alternative’s direction—not just the statistic.

Misinterpretation 4: “A Small P-Value Means a Large or Important Effect”

A p-value describes how compatible the observed test statistic is with the null model. It does not directly state how large the population difference is or whether that difference matters in practice. The test statistic compares the observed difference from the null value with its estimated standard error. A modest difference can produce a relatively small p-value when the standard error is small, while a larger observed difference can produce a less impressive p-value when the standard error is large.

The sample mean and the estimated difference from the null value are useful for describing the size of the observed result. The p-value answers a different question about evidence against the null. To judge practical importance, consider the measurement scale and what size of difference would matter in the setting. Do not use the p-value as a substitute for that context.

Worked Example: Minutes Saved by a Scheduling Tool

A fictional company randomly samples 100 users of a scheduling tool and records weekly minutes saved. The sample mean is 52 minutes, the sample standard deviation is 10 minutes, and the company tests whether the true mean differs from 50 minutes. Assume the sample is less than 10% of the target population and there are no extreme outliers. For this example, the company regards a change of 5 minutes per week as an important operational difference.

Let \(\mu\) be the true mean weekly minutes saved for users in the target population. The hypotheses are \(H_0:\mu=50\) minutes and \(H_a:\mu\ne50\) minutes. Use a one-sample t test. The random sample supports inference to the target population, the 10% condition supports independence, and the sample size of 100 is large enough for the t procedure to be reasonable in the absence of extreme outliers.

$$ SE_{\bar{x}}=\frac{10}{\sqrt{100}}=1\text{ minute}, \qquad t=\frac{52-50}{1}=2.00, \qquad df=100-1=99. $$

For a two-sided test with 99 degrees of freedom, the p-value is \(P(T_{99}\leq-2.00)+P(T_{99}\geq2.00)\approx0.0482\), rounded. At \(\alpha=0.05\), reject \(H_0\): the data provide convincing evidence that the true mean weekly minutes saved differs from 50.

The observed sample mean is 2 minutes above 50, while the company’s stated operational threshold is 5 minutes. The p-value of 0.0482 does not say that the effect is large, that it is important, or that there is a 95.18% chance the alternative hypothesis is true. It describes the probability of a t statistic at least as extreme as 2.00 under the null model. Whether a difference matters is a separate judgment tied to the setting and the size of the effect.

Misinterpretation 5: “A Large P-Value Proves the Null”

A large p-value means the observed test statistic is not especially unusual under the null model, using the reference distribution and alternative specified for the test. It does not prove that the null hypothesis is true or that the population mean equals the null value exactly. There may be little evidence against the null because the null is reasonable, because the data are variable, or because the sample does not provide precise information. A test’s failure to reject should not be rewritten as proof of equality.

Likewise, a small p-value is not proof that the alternative is true. It is evidence against the null model, provided the study design and t-procedure conditions support the analysis. The conclusion uses cautious AP Statistics language: “The data provide convincing evidence that …” or “The data do not provide convincing evidence that …” Avoid claiming certainty.

Common Mistakes and Full-Credit Communication

  • Reversing the conditional probability: “The p-value is the probability the null is true” is incorrect. Say that the probability is calculated assuming the null is true.
  • Omitting what counts as extreme: State whether the alternative is one-sided or two-sided and describe the relevant tail or tails.
  • Leaving out degrees of freedom: For a t test, the degrees of freedom determine the reference t distribution. The same t statistic can have different p-values with different degrees of freedom.
  • Calling the p-value the probability of the exact sample: It is a tail probability for statistics at least as extreme, not the probability of one exact t statistic or sample mean.
  • Treating the p-value as an effect-size measure: Report the observed difference in its units when describing size. Explain practical importance using the context, not the p-value alone.
  • Claiming a large p-value proves equality: Use “fail to reject” and “do not provide convincing evidence against the null,” not “accept the null” or “prove there is no difference.”
  • Claiming a small p-value proves the alternative: A small p-value is evidence against the null under the test model, not certainty that the alternative is true.

A strong AP response makes the probability’s conditions explicit: it names the null claim, the t statistic, the degrees of freedom, the alternative’s direction, and the outcome being counted. Then it keeps the p-value interpretation separate from the reject-or-fail-to-reject decision and the final conclusion in context.

Key takeaway: A p-value is a probability about test statistics calculated under the null hypothesis. It is not a probability that the null or alternative is true, the chance of the exact observed result, or a measure of practical importance. For a t test, interpret it using the observed t statistic, degrees of freedom, and alternative’s direction.

Check Your Understanding

For each question, keep the null assumption, the relevant t distribution, and the claim being made separate.

  1. A student says, “The p-value is 0.03, so there is a 3% chance that the null hypothesis is true.” Explain the error and state what the p-value means instead.
  2. A two-sided t test has \(t=1.8\) and 12 degrees of freedom. Which t statistics count toward its p-value?
  3. Two two-sided tests have the same observed t statistic but different degrees of freedom. Explain why their p-values can differ.
  4. A test gives a small p-value, but the observed mean difference is smaller than the difference considered important in the setting. What does the p-value establish, and what does it not establish?
  5. A test has a large p-value. Write a conclusion that avoids claiming the null hypothesis has been proved.