Tutorials › AP Statistics › Matching Decisions Across Significance Levels

P-values and conclusions for proportions · Tutorial 494 of 1000

Matching Decisions Across Significance Levels

Compare a fixed p-value with several significance levels, identify exactly where a test decision changes, and explain that change carefully in context.

Intermediate 9 min read

What You'll Learn

  • Apply the p-value decision rule at alpha levels of 0.10, 0.05, and 0.01
  • Identify the range of significance levels that produces rejection for a fixed p-value
  • Recognize that equality between the p-value and alpha leads to rejection
  • Explain how a conclusion can change as the significance level changes
  • Distinguish a comparison across stated levels from choosing alpha after seeing results
  • Write cautious, contextual conclusions for both rejection and failure to reject

One P-Value, Several Possible Decisions

In Small Samples and Large P-Values, you saw that a p-value describes how unusual the observed result, or a more extreme result, would be if the null hypothesis were true. The p-value is calculated for a particular test. Once it has been found, the decision also depends on the significance level, \(\alpha\), chosen for that test.

This tutorial holds the p-value fixed and compares the resulting decisions at \(\alpha=0.10\), \(\alpha=0.05\), and \(\alpha=0.01\). The evidence summarized by the p-value does not change between these comparisons. What changes is the decision cutoff: a smaller \(\alpha\) requires a smaller p-value to reject the null hypothesis.

Decision rule: Reject \(H_0\) when the p-value is less than or equal to \(\alpha\). If the p-value is greater than \(\alpha\), fail to reject \(H_0\). Equality counts as rejection.

For example, a p-value of 0.037 is less than 0.10 and 0.05, but greater than 0.01. The test would reject \(H_0\) at the first two levels and fail to reject it at the third. It would be wrong to say that the p-value itself changed: the same value is being compared with three different cutoffs.

This comparison is useful when a question explicitly asks how decisions differ across significance levels. In an actual study, however, \(\alpha\) should ordinarily be selected before the results are examined. As discussed in Choosing a Significance Level Before Testing, choosing a cutoff after seeing the p-value to obtain a preferred decision undermines the purpose of setting a decision rule in advance.

A Decision Profile for a Fixed P-Value

A convenient way to organize comparisons is to find the significance levels at which the decision changes. For a p-value \(p\), rejection occurs at every \(\alpha\) that is at least as large as \(p\). Failure to reject occurs at every \(\alpha\) smaller than \(p\).

$$ \text{Reject }H_0\text{ when }\alpha\geq p \qquad\text{and}\qquad \text{fail to reject }H_0\text{ when }\alpha<p $$

Thus, if the p-value is 0.037, the dividing point is 0.037: significance levels of 0.037 or larger produce rejection, while levels below 0.037 produce failure to reject. Among the three conventional levels in this tutorial, the decision changes between 0.05 and 0.01.

This dividing point can be called the test’s decision threshold across significance levels. It is not a new p-value and it does not change the interpretation of the original p-value. It is simply a way to summarize which values of \(\alpha\) lead to which decision. Two edge cases are worth noticing: if the p-value is exactly 0.05, the test rejects at \(\alpha=0.05\); if it is exactly 0.01, the test rejects at \(\alpha=0.01\).

Key idea: For any fixed p-value, increasing \(\alpha\) can change a decision from “fail to reject” to “reject,” but decreasing \(\alpha\) cannot change a rejection into a failure to reject. The data and p-value stay the same; only the decision rule changes.

Worked Examples: Comparing the Same Evidence with Different Cutoffs

Worked Example: A Result That Changes Between 0.05 and 0.01

A fictional transit office tests whether the true proportion of registered riders who use a new trip-planning feature differs from 0.40. Let \(p\) be the true proportion of registered riders who use the feature. The hypotheses are \(H_0:p=0.40\) and \(H_a:p\ne0.40\). A valid one-proportion \(z\)-test has already been carried out, and its reported p-value is 0.0370. Compare decisions at \(\alpha=0.10\), 0.05, and 0.01.

State: The question is whether the population proportion differs from 0.40. The same two-sided test and the same p-value, 0.0370, will be used for all three comparisons.

Plan: The test’s conditions were checked before its p-value was reported. The 200 riders were randomly selected from a registry of 4,000, so the Random condition is met and the 10% condition is met because \(200\leq0.10(4{,}000)=400\). Under \(H_0\), the expected numbers of feature users and nonusers are \(200(0.40)=80\) and \(200(0.60)=120\). Both are at least 10, so the Large Counts condition is met. Now compare 0.0370 with each stated significance level.

Do: At \(\alpha=0.10\), \(0.0370\leq0.10\), so reject \(H_0\). At \(\alpha=0.05\), \(0.0370\leq0.05\), so reject \(H_0\). At \(\alpha=0.01\), \(0.0370>0.01\), so fail to reject \(H_0\).

Significance levelComparisonDecision
0.100.0370 is less than 0.10Reject \(H_0\)
0.050.0370 is less than 0.05Reject \(H_0\)
0.010.0370 is greater than 0.01Fail to reject \(H_0\)

Conclude: At the 0.10 and 0.05 significance levels, the sample provides convincing evidence that the proportion of registered riders who use the feature differs from 0.40. At the 0.01 level, the sample does not provide convincing evidence for that difference under the test’s decision rule. The p-value is 0.0370 in every sentence; the stated strength of the decision changes because the cutoff changes.

Worked Example: A P-Value Below All Three Levels

A fictional school district examines whether the proportion of families who complete an online annual survey is greater than its historical benchmark of 0.60. Let \(p\) be the true proportion of families in the district who complete the survey. The hypotheses are \(H_0:p=0.60\) and \(H_a:p>0.60\). The completed test reports a p-value of 0.0080. Compare decisions at the three significance levels.

State: This is a right-tailed test of whether the population proportion is greater than 0.60. Its p-value is 0.0080.

Plan: The test’s conditions were verified. Families were randomly selected from a complete district list, so the Random condition is met. The sample included 300 families out of 6,000, and \(300\leq0.10(6{,}000)=600\), so the 10% condition is met. Under \(H_0\), the expected numbers of families who complete and do not complete the survey are \(300(0.60)=180\) and \(300(0.40)=120\). Both are at least 10, satisfying the Large Counts condition. Compare the fixed p-value with each \(\alpha\).

Do: The comparisons are \(0.0080\leq0.10\), \(0.0080\leq0.05\), and \(0.0080\leq0.01\). Therefore, reject \(H_0\) at all three significance levels.

Conclude: At each of the three stated levels, the sample provides convincing evidence that the proportion of families who complete the online survey is greater than 0.60. This example has no change in decision across the specified levels because the p-value is below even the smallest cutoff. It still does not prove that the population proportion exceeds 0.60; the conclusion is about the evidence provided by the test.

Worked Example: Equality at the 0.05 Cutoff

A fictional public library tests whether the proportion of visitors who borrow an audiobook differs from 0.25. Let \(p\) be the true proportion of visitors to the library who borrow an audiobook. The two-sided test reports a p-value of exactly 0.0500. Determine the decisions at \(\alpha=0.10\), 0.05, and 0.01.

State: The hypotheses are \(H_0:p=0.25\) and \(H_a:p\ne0.25\). The p-value to compare with each significance level is 0.0500.

Plan: Assume the library’s random-sampling method and the test conditions have been checked and that the reported p-value is appropriate for this two-sided test. Apply the rule that a p-value less than or equal to \(\alpha\) leads to rejection. Pay particular attention to equality at 0.05.

Do: Since \(0.0500\leq0.10\), reject \(H_0\) at 0.10. Since \(0.0500=0.05\), reject \(H_0\) at 0.05. Since \(0.0500>0.01\), fail to reject \(H_0\) at 0.01.

Conclude: At the 0.10 and 0.05 levels, the test provides convincing evidence that the proportion of library visitors who borrow an audiobook differs from 0.25. At the 0.01 level, it does not provide convincing evidence of a difference. The decision at 0.05 is rejection, not failure to reject, because equality is included in the rejection rule.

What It Means When the Decision Changes

A change in decision across significance levels is not a contradiction. The test is applying different rules to the same measure of evidence. A less stringent cutoff, such as 0.10, allows rejection for a wider range of p-values than a more stringent cutoff, such as 0.01.

In the first worked example, the p-value fell between 0.01 and 0.05. That location explains the pattern: the result met the rejection rule at 0.05 and 0.10, but not at 0.01. It is often clearer to report both the p-value and the decisions than to report only that a result was “significant.” For instance: “The p-value was 0.0370; we reject at the 0.05 level but fail to reject at the 0.01 level.”

The word significant is always relative to a stated \(\alpha\). A result can be statistically significant at 0.05 and not significant at 0.01. Neither label changes the size of the observed difference, its practical importance, or the p-value’s interpretation. As you learned in Statistical Significance Versus Practical Significance, a decision from a test does not by itself determine whether an effect matters in practice.

Also keep the two kinds of probability separate. The p-value is calculated under the assumption that \(H_0\) is true and describes how often results at least as extreme as the observed result would occur under the test model. The significance level is the preselected cutoff for the decision rule and, when the null hypothesis is true, the long-run probability of a Type I error for that rule. Neither number is the probability that \(H_0\) is true.

Common Mistakes and AP Exam Tips

  • Changing the p-value when \(\alpha\) changes: The p-value comes from the test and data. Reuse the same p-value for every comparison; change only the cutoff.
  • Reversing the comparison: Reject when \(p\leq\alpha\), not when \(p\geq\alpha\). A p-value smaller than the cutoff meets the rejection rule.
  • Forgetting equality: If \(p=\alpha\), reject \(H_0\). For example, a p-value of exactly 0.05 leads to rejection at \(\alpha=0.05\).
  • Writing “accept \(H_0\)” when the p-value exceeds \(\alpha\): Write “fail to reject \(H_0\).” As discussed in Why We Never Accept the Null Hypothesis, failure to reject does not establish that the null hypothesis is true.
  • Reporting a decision without the context: State what the evidence says about the population proportion and the alternative. For a right-tailed test, connect rejection to evidence that the proportion is greater than the null value; for a two-sided test, connect it to evidence that the proportion differs.
  • Selecting a favorable \(\alpha\) after seeing the result: A question may ask for comparisons at several levels, but a real study should not choose whichever cutoff gives the desired conclusion. State the level used for the planned decision.
AP Exam Tip: Show each comparison explicitly, including the equality case when relevant. Then state the decision at each level and give a contextual conclusion. If decisions differ, identify the p-value’s position relative to the cutoffs; do not imply that the data or evidence changed.

Key Takeaway

A fixed p-value can lead to different decisions at different significance levels because each \(\alpha\) sets a different cutoff. Compare carefully, use rejection when the p-value is less than or equal to \(\alpha\), and describe every conclusion in context.

Key takeaway: The p-value stays fixed while the decision cutoff changes. For a p-value between 0.01 and 0.05, for example, reject at 0.05 and 0.10 but fail to reject at 0.01. Equality with \(\alpha\) leads to rejection.

Check Your Understanding

For each situation, compare the stated p-value with each significance level and explain what the decision means.

  1. A two-sided test has a p-value of 0.024. What are the decisions at \(\alpha=0.10\), 0.05, and 0.01?
  2. A test has a p-value of 0.0100. What is the decision at \(\alpha=0.01\), and why does equality matter?
  3. A test has a p-value of 0.12. At which, if any, of the three levels 0.10, 0.05, and 0.01 would you reject \(H_0\)?
  4. In a proportion test about whether the share of customers using a store’s mobile checkout is greater than a benchmark, the p-value is 0.043. Write a decision and contextual evidence conclusion at \(\alpha=0.05\), then at \(\alpha=0.01\).
  5. Why is it misleading to report only “the result is significant” when the p-value is 0.037 and the significance level has not been stated?