Tutorials › AP Statistics › Choosing a Significance Level Before Testing

P-values and conclusions for proportions · Tutorial 484 of 1000

Choosing a Significance Level Before Testing

Choose alpha before collecting or examining data by considering the consequences of a false rejection, then use that fixed cutoff to make the test decision.

Intermediate 10 min read

What You'll Learn

  • Explain why choosing alpha after seeing the data undermines a test’s decision rule
  • Interpret alpha as the long-run probability of a Type I error when the null hypothesis is true
  • Compare the consequences of using significance levels of 0.01, 0.05, and 0.10
  • Select a significance level by considering the consequences of a false rejection
  • Apply a preselected alpha to a one-proportion z-test and explain the decision in context
  • Distinguish a significance level from a p-value and from the probability that the null hypothesis is true

Why Choose Alpha Before Seeing the Results?

In Comparing P-Value to the Significance Level, we used \(\alpha\) as the cutoff for deciding whether a p-value is small enough to reject \(H_0\). That comparison only has its intended meaning when the cutoff is chosen before the data are examined. Choosing alpha in advance keeps the rule from shifting to suit a result we already know.

The significance level is also connected to a particular risk: rejecting a null hypothesis that is actually true. In a test, that incorrect decision is called a Type I error. Alpha describes the chance of making this error in repeated use of a testing procedure when the null hypothesis is true. In AP Statistics, this is the key consequence to consider when choosing a significance level.

Definition: The significance level, \(\alpha\), is the probability of a Type I error: rejecting \(H_0\) when \(H_0\) is true. It is a property of the testing rule under the null hypothesis, not the probability that \(H_0\) is true.

For example, a test conducted at \(\alpha=0.05\) uses a rule designed to have a 5% chance of rejecting \(H_0\) when \(H_0\) is true, under the assumptions of the test. In repeated use when the null hypothesis is true, such a rule would falsely reject about 5% of the time in the long run. That does not mean there is a 5% chance that the null hypothesis is true after seeing the data.

The p-value and alpha play different roles. The p-value is calculated from the observed data, assuming \(H_0\) is true. Alpha is the cutoff set in advance. After the p-value is found, compare it with the preselected alpha: if the p-value is less than or equal to \(\alpha\), reject \(H_0\); if it is greater, fail to reject \(H_0\).

What Different Significance Levels Mean

Common choices are 0.01, 0.05, and 0.10. These are not three different ways to calculate a p-value. They are three different thresholds for how much risk of a Type I error the test’s users are willing to accept. A smaller alpha requires stronger evidence against \(H_0\) before the test rejects it.

Significance levelLong-run Type I error rate when \(H_0\) is trueWhat the cutoff implies
\(\alpha=0.01\)About 1 in 100 testsStrict cutoff; stronger evidence is needed to reject \(H_0\)
\(\alpha=0.05\)About 5 in 100 testsCommon middle-ground cutoff
\(\alpha=0.10\)About 10 in 100 testsLess strict cutoff; weaker evidence can lead to rejection

These long-run rates assume the null hypothesis is true and the testing procedure’s conditions are met. They do not predict that exactly 1, 5, or 10 out of every 100 tests will be false rejections. In any particular set of tests, the actual number can vary.

Choosing among these levels is a decision about consequences, not a calculation that the data can settle afterward. If a false rejection could prompt an expensive shutdown, an unnecessary intervention, or a serious public claim, a stricter cutoff such as 0.01 may be appropriate. If a test is being used for an early, low-stakes screen where follow-up investigation is planned, a less strict cutoff such as 0.10 might be considered. A choice of 0.05 is common, but “common” does not mean it is automatically right for every question.

The stricter cutoff has a tradeoff: it is harder for a p-value to meet the rejection rule. That can mean a real departure from \(H_0\) is less likely to lead to rejection. Thus, reducing the chance of a false rejection can make it harder to find evidence of a genuine effect. The choice should reflect the study’s purpose and the consequences of the decisions, not just a preference for one number.

Key principle: Decide how much risk of a Type I error is acceptable before examining the test results. Then keep that alpha fixed when comparing it with the p-value.

How to Choose a Level Before Testing

A useful choice begins with the decision that might follow a rejection. Ask what someone would do if the data led to rejecting \(H_0\), and what harm could result if that decision were based on a false rejection. Also consider what might be lost by setting a cutoff so strict that meaningful evidence often does not meet it.

1
Identify the decision and the population claim.
State what rejecting \(H_0\) would imply in context. Be clear about which population proportion the test concerns.
2
Consider the consequences of a false rejection.
Ask what could happen if \(H_0\) were true but the test rejected it. More serious consequences generally support a stricter cutoff.
3
Select alpha before seeing the results.
Choose a level that suits the context and record it as part of the testing plan. The choice must not depend on the p-value.
4
Conduct the test and apply the fixed rule.
Check conditions, calculate the p-value, and compare it with the alpha already chosen. Explain the decision in context.

This reasoning does not make alpha a measure of how important a result is. As explained in Statistical Significance Versus Practical Significance, a statistically significant result can still be too small to matter practically. Alpha sets the threshold for a test decision; it does not say whether the observed difference is large, useful, or worth acting on.

Worked Examples: Choosing and Using Alpha

Worked Example: A Strict Cutoff for a Sensor Alert

A fictional water-monitoring team is evaluating whether a sensor’s alert rate is higher than its established baseline of 2% during routine tests. A false claim of an increased alert rate could trigger an expensive shutdown and investigation. Before the new data are collected, the team chooses \(\alpha=0.01\). In a random sample of 500 test cycles, 18 produce alerts. Use a one-proportion \(z\)-test.

State: Let \(p\) be the true proportion of routine test cycles in this setting that produce an alert. The hypotheses are \(H_0:p=0.02\) and \(H_a:p>0.02\). The team selected \(\alpha=0.01\) in advance because a false rejection could lead to costly action.

Plan: Use a one-proportion \(z\)-test. The 500 cycles were randomly selected, so the Random condition is met. They were selected without replacement from 20,000 cycles, and \(500\leq0.10(20{,}000)=2{,}000\), so the 10% condition is met. Under \(H_0\), the expected alert count is \(np_0=500(0.02)=10\), and the expected no-alert count is \(n(1-p_0)=500(0.98)=490\). Both are at least 10, so the Large Counts condition is met.

Do: The sample proportion, null standard error, and test statistic are:

$$ \hat{p}=\frac{18}{500}=0.036 $$
$$ SE_0=\sqrt{\frac{0.02(0.98)}{500}} =\sqrt{0.0000392} \approx0.006261 $$
$$ z=\frac{0.036-0.02}{0.006261} \approx2.556 $$

Because \(H_a\) is right-tailed, the p-value is the area to the right of \(z=2.556\), approximately 0.0053, rounded to four decimal places. Compare it with the preselected level:

$$ 0.0053<0.01 $$

Conclude: Reject \(H_0\). At the 0.01 significance level, the sample provides convincing evidence that the true proportion of routine test cycles in this setting that produce an alert is greater than 2%. The conclusion is evidence for the alternative, not proof that the baseline is wrong.

The choice of alpha was made because of the consequences of a false alarm, not because the team knew what p-value the sample would produce. The result meets even this strict cutoff. A different sample could lead to a different p-value and decision, but the planned alpha should remain unchanged.

Worked Example: Why a Borderline P-Value Cannot Choose Alpha

A fictional packaging company plans to test whether the proportion of packages with a label error exceeds its 3% target. Before sampling, the company chooses \(\alpha=0.05\), judging that this is an appropriate balance for its quality review. The test produces a p-value of 0.032. Assume the conditions for the test have been met.

With the preselected value, \(0.032<0.05\), so reject \(H_0\). At the 0.05 significance level, the data provide convincing evidence that the true proportion of packages with a label error exceeds 3%.

Suppose someone then argues that the company should use \(\alpha=0.01\) because that would avoid rejecting \(H_0\). That is not a sound reason to change the level: the proposed change is being made after the p-value is known. If the company had selected 0.01 before sampling because a false rejection would have especially serious consequences, then \(0.032>0.01\) would mean fail to reject \(H_0\). But it should not switch to 0.01 now simply to obtain that decision.

The data and p-value are the same in both comparisons. The different decisions follow from different cutoffs. In a real test, however, the cutoff should be the one set before examining the data.

Worked Example: A Less Strict Cutoff for an Early Pilot

A fictional school district is piloting a new reminder message to encourage families to complete a voluntary online form. The pilot is intended to decide whether to investigate the message further, not to make a high-stakes policy change. Before collecting responses, the district chooses \(\alpha=0.10\), deciding that it can tolerate a higher chance of a false lead at this exploratory stage. The one-proportion test of whether the completion proportion exceeds its 40% benchmark produces a p-value of 0.080. Assume the test conditions are met.

Since \(0.080<0.10\), reject \(H_0\). At the 0.10 significance level, the data provide convincing evidence that the true completion proportion for the population represented by this pilot is greater than 40%. This decision supports further investigation; it does not establish that the reminder will produce a useful improvement in every setting.

If the same p-value were compared with \(\alpha=0.05\), then \(0.080>0.05\), so the decision would be to fail to reject \(H_0\). The 0.10 choice allows weaker evidence to pass the cutoff, which is consistent with a low-stakes exploratory purpose. It would not automatically be suitable for a consequential final decision.

Worked Example: Interpreting Alpha Across Many Tests

Suppose a quality team plans to conduct 1,000 separate tests in settings where each null hypothesis is true. For a simple illustration, assume each test uses a procedure operating at its stated significance level. How many false rejections would the team expect in the long run at each of three common choices?

At \(\alpha=0.01\), the expected count is \(1{,}000(0.01)=10\). At \(\alpha=0.05\), it is \(1{,}000(0.05)=50\). At \(\alpha=0.10\), it is \(1{,}000(0.10)=100\). These are long-run expected counts, not promises about exactly how many false rejections will occur in one set of 1,000 tests.

This comparison illustrates why alpha matters when false alarms have costs. A more permissive threshold can produce more false rejections across repeated testing when the null hypotheses are true. It does not mean that any particular rejected null hypothesis has a probability equal to alpha of being true.

Common Mistakes and AP Exam Tips

  • Choosing alpha after seeing the p-value: This lets the result influence the rule used to judge that same result. A full-credit response identifies the significance level as selected before the data are examined.
  • Calling alpha the probability that \(H_0\) is true: Alpha is the probability of a Type I error under the assumption that \(H_0\) is true. It is not a probability assigned to \(H_0\) after observing the sample.
  • Treating 0.05 as mandatory: It is a common convention, not a universal requirement. Explain why a level fits the consequences and purpose of the study when asked to justify a choice.
  • Assuming a smaller alpha is always better: A smaller alpha reduces the chance of a false rejection when \(H_0\) is true, but it also makes rejection harder. The choice involves the study’s purpose and the consequences of decisions.
  • Changing alpha to get a preferred conclusion: If the p-value is 0.032, do not switch from 0.05 to 0.01 after seeing it just to avoid rejection. Use the value chosen in advance.
  • Confusing statistical significance with practical importance: Alpha helps determine a test decision. The size and real-world importance of a difference require attention to context as well.
AP Exam Tip: When asked why a significance level is appropriate, connect it to the consequences of a Type I error in the specific setting. State that alpha is chosen before examining the data, and do not describe it as the probability that the null hypothesis is true.

Key Takeaway

Choose alpha before examining the results because it sets the test’s tolerance for falsely rejecting a true null hypothesis. A smaller level such as 0.01 is stricter; 0.05 is common; and 0.10 is more permissive. The right choice depends on the context and the consequences of a false rejection. Once selected, keep alpha fixed and compare the p-value with it.

Key takeaway: Alpha is a preselected cutoff and a long-run Type I error rate when \(H_0\) is true. Choose it before testing by considering the consequences of a false rejection; never adjust it afterward to obtain a preferred decision.

Check Your Understanding

Answer each question using the distinction between a preselected significance level and a p-value calculated from data.

  1. In your own words, what does \(\alpha=0.01\) describe when \(H_0\) is true?
  2. A study’s p-value is 0.032. Explain why researchers should not decide afterward to use \(\alpha=0.01\) simply to avoid rejecting \(H_0\).
  3. Why might a team facing serious consequences from a false rejection choose \(\alpha=0.01\) rather than \(\alpha=0.10\)?
  4. A p-value is 0.080 and the study selected \(\alpha=0.10\) before collecting data. What is the decision, and what does it mean in context?
  5. Does \(\alpha=0.05\) mean there is a 5% probability that \(H_0\) is true? Explain what the 5% refers to instead.