Tutorials › AP Statistics › Strength of Evidence on a Continuum

P-values and mean-inference conclusions · Tutorial 747 of 1000

Strength of Evidence on a Continuum

Interpret a p-value as a degree of evidence against the null hypothesis while keeping decisions, effect size, and certainty distinct.

Intermediate 9 min read

What You'll Learn

  • Explain why evidence against a null hypothesis varies along a continuum.
  • Compare p-values near a significance threshold without exaggerating their difference.
  • Describe smaller p-values as stronger evidence against the null, when tests and assumptions are comparable.
  • Separate evidence strength from the reject-or-fail-to-reject decision.
  • Explain why a p-value does not measure the size or practical importance of a mean difference.
  • Avoid treating a large p-value as proof that the null hypothesis is true.

Evidence Is More Than a Yes-or-No Decision

In “How Alpha Changes the Conclusion,” you saw that alpha sets a threshold for a decision: reject \(H_0\) when the p-value is at or below alpha, and otherwise fail to reject \(H_0\). That rule is useful, but it does not mean that evidence itself suddenly changes at the threshold. The evidence against \(H_0\) can be described as stronger or weaker along a continuum.

A p-value is calculated under the assumption that \(H_0\) is true. It is the probability of obtaining a test statistic at least as extreme as the observed statistic, in the direction specified by \(H_a\). When two p-values come from comparable tests of the same null claim, the smaller p-value indicates data that would be less usual if \(H_0\) were true, and therefore generally stronger evidence against \(H_0\).

Key idea: The p-value helps describe the strength of evidence against \(H_0\) on a continuum. Alpha supplies a decision threshold, not a boundary where the evidence changes abruptly. A p-value is not the probability that \(H_0\) is true.

For example, p-values of \(0.047\) and \(0.052\) are close. At \(\alpha=0.05\), they lead to different decisions, but the evidence they represent is not suddenly different in kind. The first is a little smaller and so indicates somewhat stronger evidence against the null claim, assuming the tests are comparable. It would be misleading to describe one as a decisive discovery and the other as no evidence at all just because they fall on opposite sides of \(0.05\).

That does not make alpha irrelevant. If a study selected \(\alpha=0.05\) before collecting data, it should still make its primary decision using that threshold. The continuum idea adds nuance to how we describe the result; it does not replace the decision rule.

Describing the Continuum Carefully

In ordinary language, a very small p-value may be described as stronger evidence against \(H_0\) than a moderate p-value, and a larger p-value may provide little evidence against \(H_0\). These descriptions are relative, not a universal scale with fixed labels. AP Statistics does not require a single set of numerical cutoffs for words such as “weak,” “moderate,” or “strong.” State the p-value and explain what it means in context rather than relying on a label alone.

Some fields use conventional ranges to summarize p-values, but those ranges do not change the definition of a p-value. Nor do they turn a decision threshold into a natural divide in the evidence. In an AP response, connect the p-value to the null claim and the observed result, and use the preselected alpha when the question asks for a decision.

Useful comparison: For comparable tests of the same null claim, a smaller p-value indicates stronger evidence against \(H_0\). This comparison does not say how large or practically important the effect is, and it does not assign a probability to either hypothesis.

The word comparable matters. P-values depend on the test, its alternative hypothesis, the data, and the conditions that justify the procedure. Do not rank the evidence in unrelated studies just by looking at their p-values if they test different claims or use substantially different designs. In particular, a one-sided and a two-sided test do not answer the same question.

A p-value also does not measure the size of an effect. A small difference in sample means can produce a small p-value when the standard error is small, while a larger difference can produce a larger p-value when the data are highly variable. To discuss the estimated size of a difference, examine the sample means and, when available, a confidence interval—not the p-value alone.

Worked Example: Two Results Near 0.05

Worked Example: Two Results Near 0.05

Two fictional studies independently take random samples of rechargeable batteries of the same model and test whether the true mean operating time differs from 12 hours. Each study uses a two-sided one-sample t test, and the procedure’s conditions are considered reasonable. The first study reports \(p=0.047\); the second reports \(p=0.052\). Consider what the results say about evidence and about a decision at \(\alpha=0.05\).

1
State.
For each study, let \(\mu\) be the true mean operating time, in hours, for batteries of this model. The hypotheses are \(H_0:\mu=12\) hours and \(H_a:\mu\ne12\) hours.
2
Plan and check conditions.
Each study uses a one-sample t test for a population mean. The batteries were randomly sampled, and each sample is less than 10% of the population of batteries of this model, so the 10% condition is reasonable. The sample distributions are described as having no severe skewness or extreme outliers, making the t procedure reasonable. The two-sided alternative matches a question about a difference in either direction.
3
Do.
The reported p-values are \(0.047\) and \(0.052\), respectively. At \(\alpha=0.05\), \(0.047\le0.05\), so the first study rejects \(H_0\). Since \(0.052>0.05\), the second study fails to reject \(H_0\).
4
Conclude in context.
At the 0.05 significance level, the first study provides convincing evidence that the true mean operating time differs from 12 hours. The second does not provide convincing evidence of a difference at that level. However, the p-values are close, so the evidence against \(H_0\) is not meaningfully separated into “evidence” and “no evidence” by the threshold alone. The first result indicates somewhat stronger evidence against the null claim than the second.

The conclusions use the alpha-based rule correctly while avoiding an exaggerated contrast. Failure to reject in the second study does not show that the true mean is exactly 12 hours. Also, because the p-values alone do not show the estimated mean difference, they cannot tell us whether any difference is practically important.

Smaller P-Values: Stronger Evidence, Not More Certainty

A very small p-value can be strong evidence against \(H_0\), but it is not proof that \(H_a\) is true. A test result is conditional on the null hypothesis and the assumptions of the procedure. If the data collection or test conditions are not reasonable, a small p-value may not deserve the interpretation that a valid test would support.

It is also important to distinguish evidence from certainty. A p-value of \(0.003\) does not mean there is a \(99.7\%\) chance that the alternative hypothesis is true. That would assign a probability to a hypothesis, which the p-value does not do. Instead, the p-value describes how unusual a test statistic at least as extreme as the observed one would be if the null hypothesis were true.

Worked Example: A Small P-Value for a Mean

A fictional city takes a random sample of households to test whether the true mean daily water use per household exceeds a conservation target of 300 liters. A valid one-sided t test reports \(p=0.003\). The sample is less than 10% of the city’s households, and the data do not show severe skewness or extreme outliers.

Let \(\mu\) be the true mean daily water use, in liters per household, for households in the city. The hypotheses are \(H_0:\mu=300\) liters and \(H_a:\mu>300\) liters. The one-sided alternative is appropriate because the research question asks whether the mean exceeds the target.

A p-value of \(0.003\) means that, if the true mean daily water use were 300 liters per household, the probability of obtaining a t statistic at least as large as the observed statistic in the direction of greater use would be \(0.003\). This is a small probability, so the data provide strong evidence against \(H_0\) and in favor of the claim that the mean exceeds 300 liters. At \(\alpha=0.05\), \(0.003\le0.05\), so reject \(H_0\).

The conclusion is evidence about the population mean, not certainty about every household’s use. It does not mean that there is a 99.7% probability the mean exceeds 300 liters, nor does the p-value state how many liters above 300 the true mean might be. Those are different questions.

A Large P-Value Is Not Proof of No Difference

A larger p-value generally indicates weaker evidence against the null hypothesis, when the test and claim are comparable. It does not establish that \(H_0\) is true, and it does not prove that two population means are equal. The result may reflect a true difference that the study was not able to detect clearly, perhaps because the sample was small or the measurements varied substantially.

This is why AP Statistics uses “fail to reject \(H_0\)” rather than “accept \(H_0\).” The test has not provided enough evidence to reject the null claim at the chosen alpha. It has not confirmed that claim as true. A confidence interval can add information by showing a range of plausible values for a population mean or difference, as you learned in earlier tutorials on interpreting confidence intervals.

Worked Example: Little Evidence Against Equal Mean Times

A fictional school district compares the mean time students in two programs spend traveling to school. Independent random samples are taken from each program. A valid two-sample t test of \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\) reports \(p=0.31\). Here, \(\mu_1\) and \(\mu_2\) are the true mean travel times, in minutes, for students in programs 1 and 2. The study planned \(\alpha=0.05\).

Because \(0.31>0.05\), the district fails to reject \(H_0\). The result provides little evidence against the claim that the two population mean travel times are equal, compared with a result having a much smaller p-value. It does not prove that the means are equal or that the difference is zero. The data do not provide convincing evidence of a difference at the planned significance level.

The p-value alone also does not tell the district whether any possible difference would matter for students. A useful follow-up would consider the sample mean difference and an appropriate confidence interval, interpreted in minutes and in context. A large p-value is not a measure of practical equivalence.

Common Mistakes and AP Exam Tips

  • Treating alpha as a cliff in the evidence: A p-value just below \(0.05\) and one just above \(0.05\) lead to different decisions at that level, but they can represent very similar degrees of evidence. State the decision accurately without pretending the evidence changed abruptly.
  • Calling a result “no evidence” after failing to reject: A larger p-value means the test provides weaker evidence against \(H_0\); it does not prove there is no difference. Say that the data do not provide convincing evidence for the alternative at the chosen alpha.
  • Interpreting the p-value as a probability about a hypothesis: Do not say there is a certain percentage chance that \(H_0\) is true. A full-credit explanation makes the probability conditional on \(H_0\) and describes test statistics at least as extreme as the observed one.
  • Using p-values to describe effect size: A p-value does not say how many units the mean differs from the null value or whether that difference matters in practice. Describe the sample result and use an interval when the question asks about plausible values.
  • Comparing p-values from unlike tests as if they were on one universal scale: The strength comparison is clearest for comparable tests of the same claim. Note differences in the alternative hypothesis or study design before making broad comparisons.
  • Using a qualitative label without a contextual interpretation: Words such as “strong” can summarize a result, but they do not replace the AP-standard explanation. Identify the null claim and explain what the p-value says under that claim.

A strong AP response can combine the continuum description with the formal decision. For example: “The p-value of \(0.047\) indicates somewhat stronger evidence against the null claim than a p-value of \(0.052\), but the results are close. Since \(0.047\le0.05\), reject \(H_0\); the data provide convincing evidence that the true mean operating time differs from 12 hours.” If the p-value exceeds alpha, use “fail to reject” and avoid claiming that the null value has been proved.

Key takeaway: Evidence against \(H_0\) varies along a continuum: smaller p-values generally mean stronger evidence for comparable tests. Alpha turns that continuum into a decision rule, but it does not create a sharp change in evidence. A p-value is neither the probability that a hypothesis is true nor a measure of effect size.

Check Your Understanding

Answer each question using the distinction between evidence strength and an alpha-based decision.

  1. Two comparable tests of the same null claim report p-values of \(0.041\) and \(0.045\). Which indicates somewhat stronger evidence against \(H_0\), and why should you avoid treating the difference as a sharp divide?
  2. A valid test reports \(p=0.052\) with a planned \(\alpha=0.05\). State the decision and explain what the p-value says under \(H_0\).
  3. A one-sided test of whether a population mean exceeds a target reports \(p=0.003\). What does this indicate about evidence, and what does it not say about the probability that the alternative is true?
  4. A two-sample t test reports \(p=0.31\) for a difference in mean travel times. Why is “the means are equal” an unjustified conclusion?
  5. Can a small p-value by itself show that a difference in population means is large or practically important? Explain what additional information would help.