Tutorials › AP Statistics › Strength of Evidence from Different P-Values

P-values and conclusions for proportions · Tutorial 489 of 1000

Strength of Evidence from Different P-Values

Learn to compare the strength of evidence different p-values provide, while keeping evidence, significance decisions, and practical importance distinct.

Intermediate 10 min read

What You'll Learn

  • Compare p-values as measures of how unusual the data are under the null hypothesis.
  • Describe the relative evidence associated with p-values of 0.001, 0.04, and 0.20.
  • Use verbal descriptions of evidence without treating them as universal cutoffs.
  • Explain why a p-value is not the probability that the null hypothesis is true.
  • Distinguish evidence strength from statistical significance and practical importance.
  • Write conclusions that describe evidence against the null in context.

How P-Values Describe Evidence

In Writing a Complete Conclusion With Evidence, we connected a p-value to a decision and a contextual conclusion. Now we will compare what different p-values suggest about the evidence against a null hypothesis. A p-value of 0.001, 0.04, or 0.20 does not change the meaning of a p-value; it changes how unusual the observed result is under the null model.

As explained in What a P-Value Really Measures, a p-value is calculated assuming the null hypothesis is true. It is the probability of getting the observed result, or a result at least as extreme, following the alternative hypothesis’s rule for extremeness. Smaller p-values mean the observed data would be less common under that null model. That is why a smaller p-value generally provides stronger evidence against the null hypothesis.

Definition: The strength of evidence against a null hypothesis describes how surprising the observed data, or more extreme data, would be if that null hypothesis were true. A smaller p-value indicates greater surprise under the null model and generally stronger evidence against it.

Imagine repeating a test many times under the assumption that its null hypothesis is true. A p-value of 0.001 means results at least as extreme as the observed result would occur about 1 time in 1,000 repetitions. A p-value of 0.04 corresponds to about 40 times in 1,000, and a p-value of 0.20 to about 200 times in 1,000. These are long-run descriptions under the null model, not probabilities that the null hypothesis is true.

P-valueHow unusual the result is under the nullCareful evidence description
0.001Very unusual; about 1 in 1,000 repetitionsStrong evidence against the null hypothesis
0.04Somewhat unusual; about 40 in 1,000 repetitionsSome evidence against the null hypothesis
0.20Not especially unusual; about 200 in 1,000 repetitionsLittle evidence against the null hypothesis

The descriptions in the table are useful guides, not fixed scientific laws. There is no universal boundary at which evidence suddenly changes from weak to strong. In many settings, a p-value of 0.04 is described as some or moderate evidence, while 0.001 is strong evidence. A p-value of 0.20 usually provides little evidence against the null. The important point is the ordering: 0.001 is stronger evidence against the null than 0.04, which is stronger than 0.20.

Compare Like With Like

A comparison is easiest to interpret when the tests address the same null hypothesis and use the same alternative and method. For example, comparing p-values from tests of the same population proportion against the same benchmark gives a clearer comparison than comparing results from unrelated questions. In either case, make sure the tests and data collection are suitable; a small p-value does not repair a biased sample or a violated condition.

Even for comparable tests, do not turn the numerical comparison into a claim that one result is a particular number of times “more important.” A p-value of 0.001 is 40 times as small as 0.04, but that does not mean it represents exactly 40 times the evidence. P-values are tail probabilities, not a direct scale of evidence or a measure of how large an effect is.

Comparison guide: First confirm what null hypothesis and alternative each test uses. Then compare how unusual the results would be under those null models. Use “stronger” or “weaker evidence against the null,” not “the null is less likely by this percentage.”

Also keep evidence strength separate from the test decision. As covered in Comparing P-Value to the Significance Level, the decision to reject or fail to reject \(H_0\) depends on comparing the p-value with the significance level \(\alpha\), which should be chosen before examining the data. The p-value describes the evidence; \(\alpha\) supplies a decision rule. A p-value of 0.04 would lead to rejection if \(\alpha=0.05\), but not if \(\alpha=0.01\).

Finally, evidence against a null hypothesis is not automatically evidence that an effect is large or important in practice. A very small departure from a benchmark can produce a small p-value when a study has substantial information. Conversely, a meaningful departure might not produce a small p-value in a study with limited information. The sample result and its practical meaning require attention alongside the p-value.

Worked Examples: Describing Evidence in Context

Worked Example: Comparing Three Possible Survey Results

A fictional school district takes a random sample of 400 families from a list of 8,000. It tests whether the true proportion of families on the list who support a proposed calendar change differs from 50%. Suppose the test conditions have been checked and are met. Consider three possible results from this same testing question: p-values of 0.001, 0.04, and 0.20. Describe and compare their evidence.

State: Let \(p\) be the true proportion of families on this district’s list who support the calendar change. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\).

Plan: The families were randomly sampled, so the Random condition is met. The sample is no more than 10% of the list because \(400\leq0.10(8{,}000)=800\). Under the null hypothesis, the expected counts are \(400(0.50)=200\) supporters and \(400(0.50)=200\) nonsupporters, both at least 10. Thus, the conditions for a one-proportion \(z\)-test are met.

Do: The three p-values are stipulated possible rounded outputs from this test, not three separate calculations. Under \(H_0\), a result at least as extreme as the one giving p-value 0.001 would occur about 0.1% of the time. The corresponding rates for 0.04 and 0.20 are about 4% and 20%.

Conclude: A p-value of 0.001 would provide strong evidence that the proportion of families on this district’s list who support the calendar change differs from 50%. A p-value of 0.04 would provide some evidence of a difference, but less than 0.001 does. A p-value of 0.20 would provide little evidence that the proportion differs from 50%; results at least that extreme are not especially unusual under the null model.

The last statement does not establish that the proportion is 50%. It says only that the result represented by p-value 0.20 is relatively unsurprising if the null hypothesis is true. If the district selected \(\alpha=0.05\) in advance, the first two p-values would lead to rejecting \(H_0\), while 0.20 would lead to failing to reject it. Those decisions do not erase the difference in evidence strength between 0.001 and 0.04.

Worked Example: A P-Value of 0.001 for a Recycling Program

A fictional city surveys a random sample of households about whether they would use a new curbside recycling service. The question is whether the true proportion of households on the city’s residential list that would use the service differs from 30%. The test report gives a two-sided p-value of 0.001, and the report confirms the conditions for a one-proportion \(z\)-test are met. Describe the evidence.

State: Let \(p\) be the true proportion of households on the city’s residential list that would use the recycling service. The hypotheses are \(H_0:p=0.30\) and \(H_a:p\ne0.30\).

Plan: The survey used a random sample, meeting the Random condition. The sample was no more than 10% of the residential list, meeting the 10% condition. The null-based expected counts of households who would and would not use the service were each at least 10, meeting the Large Counts condition.

Do: The p-value is 0.001. Assuming that 30% of households on the list would use the service, the probability of obtaining a sample result at least as far from 30% as the observed result, in either direction, is about 0.1%.

Conclude: Because a result this extreme would be very unusual if the true proportion were 30%, the data provide strong evidence that the proportion of households on the city’s residential list that would use the recycling service differs from 30%.

The conclusion follows the two-sided alternative: “differs,” not necessarily “is greater.” The p-value alone does not say how far the true proportion is from 30%, and it does not tell the city whether the program is practical or worthwhile. Those questions require additional information.

Worked Example: A P-Value of 0.20 for a Transit App

A fictional transit agency takes a random sample of riders and tests whether the true proportion of riders who use a trip-planning app is greater than 40%. The test report gives a p-value of 0.20. The random selection, 10% condition, and null-based Large Counts condition have all been verified. Explain what the result indicates.

State: Let \(p\) be the true proportion of riders in the population represented by the agency’s sampling list who use the trip-planning app. The hypotheses are \(H_0:p=0.40\) and \(H_a:p>0.40\).

Plan: The random sample supports inference to the population represented by the sampling list. The sample is no more than 10% of that population, and the expected counts of app users and nonusers under \(H_0\) are both at least 10. The conditions for a one-proportion \(z\)-test are met.

Do: For this right-tailed test, a p-value of 0.20 means that, if 40% of riders in the represented population use the app, there is about a 20% chance of obtaining a sample result at least as high as the observed result.

Conclude: The result is not especially unusual under \(H_0\), so it provides little evidence that more than 40% of riders in the represented population use the trip-planning app. If the agency chose \(\alpha=0.05\) in advance, it would fail to reject \(H_0\). It should not conclude that the true proportion is 40% or that the app has no users.

Notice that “little evidence against \(H_0\)” is not the same as evidence that \(H_0\) is true. The test did not establish equality; it found that the data were not unusually high under the null model.

Common Mistakes and AP Exam Tips

  • Treating the labels as cutoffs: Do not claim that a p-value just above a particular number suddenly becomes a different kind of evidence. “Strong,” “some,” and “little” are descriptive language; explain how unusual the result is under the null.
  • Calling 0.04 proof of the alternative: A p-value of 0.04 can provide evidence against \(H_0\), but it does not prove \(H_a\). State the evidence cautiously and in context.
  • Calling 0.20 proof of the null: A large p-value means the data do not provide much evidence against \(H_0\). It does not establish that the null value is correct.
  • Interpreting the p-value as the chance \(H_0\) is true: The probability is calculated assuming \(H_0\), not the other way around. State what data would be unusual under the null model.
  • Claiming a smaller p-value means a larger effect: The p-value reflects how unusual the data are under a model, not the size or practical importance of a difference. Describe the observed difference or practical stakes separately when those are part of the question.
  • Using evidence language as the decision rule: First compare the p-value with the preselected \(\alpha\) to decide whether to reject or fail to reject. Then describe the evidence and conclusion in context.
  • Comparing unrelated p-values as if they were identical tests: Check that the hypotheses, alternative directions, and methods are comparable. Always explain which null model each p-value assumes.
AP Exam Tip: For a p-value of 0.001, 0.04, or 0.20, connect the number to the null model: “Assuming [null claim] is true, results at least as extreme as [observed result] would occur about [p-value as a percentage] of the time.” Then describe the evidence against the null without claiming that the null is true or false with certainty.

Key Takeaway

Smaller p-values indicate data that are more unusual under the null hypothesis and generally provide stronger evidence against it. Thus, 0.001 suggests stronger evidence than 0.04, and 0.04 stronger evidence than 0.20. These comparisons are not rigid labels, proofs, or measures of practical importance. Keep evidence strength separate from the decision made by comparing the p-value with \(\alpha\).

Key takeaway: Describe each p-value by how surprising the data would be under the null model. Use verbal evidence labels as guides, not cutoffs, and never interpret a p-value as the probability that the null hypothesis is true.

Check Your Understanding

Answer in context where a setting is provided, and distinguish evidence from a test decision.

  1. Order p-values 0.001, 0.04, and 0.20 from strongest to weakest evidence against the null hypothesis. Briefly explain your ordering.
  2. A test of whether a community’s true support proportion differs from 25% gives p-value 0.04. Describe the evidence without saying that the alternative is proved.
  3. A test gives p-value 0.20. Explain why this does not show that the null hypothesis is true.
  4. For a p-value of 0.001, what does the value describe, assuming the null hypothesis is true? What does it not describe?
  5. A test has p-value 0.04. State the decision if \(\alpha=0.05\), and the decision if \(\alpha=0.01\). Why are these decisions separate from describing the evidence strength?