Tutorials › AP Statistics › The Significance Level Alpha Explained

P-values and mean-inference conclusions · Tutorial 744 of 1000

The Significance Level Alpha Explained

See how a significance level chosen in advance sets the evidence threshold for a test about a population mean.

Intermediate 10 min read

What You'll Learn

  • Explain why alpha is selected before looking at the results.
  • Describe alpha as a threshold for judging a p-value.
  • Connect alpha to the long-run risk of rejecting a true null hypothesis.
  • Explain why the alternative hypothesis determines which tail or tails use alpha.
  • Compare how different preselected alpha levels affect the evidence required.
  • Recognize why changing alpha after seeing a p-value is not sound practice.

Alpha Is a Threshold Chosen Before the Results

In “Common Misinterpretations of the P-Value,” you learned that a p-value is the probability, assuming the null hypothesis is true, of obtaining a test statistic at least as extreme as the observed statistic in the direction or directions specified by the alternative hypothesis. A p-value alone does not label evidence as convincing or not convincing. For that judgment, a test uses a significance level, written \(\alpha\), as a threshold.

The key planning decision is to choose \(\alpha\) before examining the test results. That way, the standard for evidence is not adjusted to fit the p-value after it is known. For example, a researcher might plan to use \(\alpha=0.05\), or might choose a smaller value if falsely flagging a difference would have especially serious consequences.

Definition: The significance level, \(\alpha\), is the decision threshold selected before analyzing the data. It sets how small a p-value must be to count as statistically significant evidence against the null hypothesis. When the test and its assumptions are appropriate, \(\alpha\) is the probability of a Type I error—the error of rejecting a true null hypothesis—or, in some settings, an upper bound on that probability.

An \(\alpha\) of 0.05 means a 5% significance level. In repeated use of a properly calibrated test when the null hypothesis is true, the procedure would reject the null about 5% of the time in the long run. This is a feature of the testing procedure under the null assumption; it is not a 5% probability that the null hypothesis is true for the particular study.

Alpha is also not the p-value. The p-value comes from the observed data, the test statistic, the degrees of freedom for a t test, and the alternative hypothesis. Alpha is a standard chosen in advance. Changing alpha does not change the data or the calculated p-value; it changes the threshold used to judge that p-value.

Choosing Alpha Depends on the Consequences

There is no single significance level that is automatically right for every investigation. The familiar value \(\alpha=0.05\) is a convention, not a rule of nature. A researcher should consider what could follow from a false alarm—rejecting a true null—and how much evidence is needed before taking action.

A smaller alpha makes the test more demanding: the p-value must be smaller to cross the threshold. This lowers the chance of a Type I error when the null is true, but it can also make it harder to detect a real difference. A larger alpha makes the threshold less demanding, which may help detect a real difference but allows a greater chance of a false alarm under the null. In practical planning, the balance depends on the context and the consequences of each kind of mistake.

Worked Example: Setting a Threshold Before Testing a Water-Quality Mean

A fictional town plans to sample water from a local system and test whether the true mean concentration of a substance differs from a specified reference value. A false alarm could prompt a costly investigation, but failing to identify a meaningful departure could also matter for public decisions. The analysts decide, before collecting or examining the sample, to use \(\alpha=0.01\).

This choice means the planned test will require stronger evidence against the null hypothesis than a test using \(\alpha=0.05\). The choice does not assert that the null is probably true, and it does not guarantee that the test will make no false alarms. Rather, if the null is true and the test’s assumptions are met, the planned procedure has a 1% Type I error rate (or no more than 1%, depending on the procedure).

The analysts should record this threshold in their plan before seeing the sample result. They should not wait for a p-value and then decide that 0.01, 0.05, or another value is most convenient. A planned alpha makes the evidential standard transparent and guards against moving the goalposts after seeing the data.

Alpha and the Test’s Tail or Tails

The alternative hypothesis determines which direction or directions count as evidence against the null. Alpha sets the overall probability threshold for that test. For a one-sided test, the full \(\alpha\) is assigned to the tail specified by the alternative. For a two-sided test, the threshold is split equally between the two tails, so each tail has area \(\alpha/2\).

For a t test, this allocation also determines the critical t value or values. A critical value marks the boundary of the region of test-statistic outcomes that would be considered sufficiently unusual under the null at the chosen alpha level. The t distribution used depends on the degrees of freedom. The alternative must be chosen based on the research question before looking at the data; alpha does not decide whether the test should be one-sided or two-sided.

Worked Example: How the Same Alpha Sets Different T Cutoffs

Suppose a researcher plans a t test with 19 degrees of freedom and chooses \(\alpha=0.05\). Compare the critical values for a right-sided alternative with those for a two-sided alternative. These cutoffs show how the preselected significance level is allocated; they are not calculated from an observed sample result.

For a right-sided test, all 0.05 of the null distribution’s tail area is in the right tail. The critical value is the 95th percentile of the \(t\) distribution with 19 degrees of freedom:

$$ t^*=\operatorname{invT}(0.95,19)\approx1.729. $$

For a two-sided test, the 0.05 total tail area is split into 0.025 in each tail. The critical values are the 2.5th and 97.5th percentiles:

$$ -t^*=-\operatorname{invT}(0.975,19)\approx-2.093, \qquad t^*=\operatorname{invT}(0.975,19)\approx2.093. $$

The two-sided test needs a more extreme statistic in either direction to reach the same overall \(\alpha=0.05\), because its tail area is divided between both directions. These values are rounded to three decimal places and agree with t critical values from a calculator. The choice between the right-sided and two-sided test must come from the question being asked, not from whichever cutoff is easier to cross.

Using Alpha as a Threshold in a Mean Test

The p-value approach and the critical-value approach express the same threshold idea. With the p-value approach, compare the p-value with the preselected \(\alpha\). With the critical-value approach, compare the test statistic with the critical value or values determined by \(\alpha\), the alternative, and the degrees of freedom. In either approach, alpha must be fixed before the results are examined.

Worked Example: A Two-Sided Test About Mean Seedling Height

A fictional greenhouse takes a random sample of 25 seedlings to test whether the true mean height differs from 17.5 centimeters. The sample mean is 18.4 centimeters, and the sample standard deviation is 3 centimeters. A plot shows no strong skewness or outliers. The sample is less than 10% of the seedlings in the target population, and the greenhouse planned a two-sided test at \(\alpha=0.05\) before collecting the sample.

1
State.
Let \(\mu\) be the true mean height, in centimeters, of all seedlings in the target population. The hypotheses are \(H_0:\mu=17.5\) centimeters and \(H_a:\mu\ne17.5\) centimeters.
2
Plan and check conditions.
Use a one-sample t test. The random sample supports inference to the target population. Because the sample is less than 10% of the population, the 10% condition supports treating observations as independent. With \(n=25\), check the data’s shape; the plot shows no strong skewness or outliers, so using a t procedure is reasonable. The planned significance level is \(\alpha=0.05\).
3
Do.
Calculate the standard error and t statistic. There are \(25-1=24\) degrees of freedom. For a two-sided test at \(\alpha=0.05\), the critical values are approximately \(-2.064\) and \(2.064\).
$$ SE_{\bar{x}}=\frac{s}{\sqrt{n}} =\frac{3}{\sqrt{25}} =0.6\text{ centimeter}, \qquad t=\frac{\bar{x}-\mu_0}{SE_{\bar{x}}} =\frac{18.4-17.5}{0.6} =1.50, \qquad df=24. $$

The observed statistic, \(t=1.50\), is between the two critical values. Equivalently, it is not far enough into either tail to cross the threshold set by \(\alpha=0.05\).

4
Conclude in context.
At the preselected 0.05 significance level, fail to reject \(H_0\). These data do not provide convincing evidence that the true mean height of seedlings in the target population differs from 17.5 centimeters.

The conclusion is tied to the planned threshold. It does not prove that the mean is exactly 17.5 centimeters. It also does not mean there is a 5% chance that the null hypothesis is true. Alpha describes the long-run Type I error rate of the test procedure when the null is true, while the p-value describes the evidence in the particular data under the null model.

Why You Must Not Pick Alpha After Seeing the P-Value

Imagine that a researcher planned \(\alpha=0.05\), obtained a p-value of 0.032, and then changed the plan to \(\alpha=0.01\) because the new threshold would avoid rejecting the null. The opposite maneuver is also a problem: changing a planned \(\alpha=0.01\) to 0.05 after seeing a p-value of 0.032 to obtain a rejection. In either case, the threshold is being chosen to suit the result rather than to set a standard in advance.

Worked Example: One P-Value and Two Planned Thresholds

A fictional sports scientist tests whether a training program changes the true mean recovery time for a defined group of athletes. A valid two-sided t test produces a p-value of 0.032, rounded. Compare what the threshold means under two plans that would have to be chosen before the data were examined.

If the study plan specified \(\alpha=0.05\), the p-value is below the planned threshold. If the plan instead specified \(\alpha=0.01\), the same p-value is above that threshold. The p-value itself remains 0.032 in both comparisons: the test result has not changed, only the preselected evidential standard has.

These comparisons illustrate why the study plan matters. A researcher should not choose between 0.05 and 0.01 after seeing 0.032. The appropriate threshold depends on the context and on the planned balance between false alarms and missed differences. State the chosen \(\alpha\) and its rationale before analyzing the results.

Common Mistakes and AP Exam Tips

  • Choosing alpha after seeing the p-value: This makes the evidential standard depend on the result. A full-credit response identifies the significance level as chosen in advance.
  • Calling alpha the probability the null is true: Alpha concerns the long-run probability of a Type I error when the null is true. It is not a probability assigned to the hypothesis.
  • Confusing alpha with the p-value: Alpha is the planned threshold; the p-value is calculated from the observed data under the null model. They play different roles.
  • Putting all alpha in one tail for a two-sided test: A two-sided test splits the total alpha equally between its two tails. A one-sided test places it in the direction named by the alternative.
  • Choosing the alternative to match the observed result: The question and research plan determine the direction of the alternative before the data are examined. Do not switch from a two-sided to a one-sided test because the sample result points one way.
  • Treating 0.05 as mandatory: It is a common convention, not a universal requirement. A different threshold can be appropriate when justified by the consequences in context and chosen in advance.
  • Claiming a smaller alpha always makes a test better: A smaller alpha reduces the chance of a Type I error under the null but makes it harder to identify a real difference. Explain the tradeoff rather than calling one threshold best in every situation.

For clear AP communication, name the chosen significance level, say that it was set before the results were examined, and connect it to the threshold for evidence against \(H_0\). When explaining its meaning, distinguish the long-run Type I error rate from the probability described by a p-value.

Key takeaway: Choose \(\alpha\) before examining the data. It is the threshold for judging evidence against the null hypothesis and represents the long-run Type I error rate of the test when the null is true (or an upper bound on that rate). The context can justify a stricter or less strict threshold, but the p-value should not determine alpha after the fact.

Check Your Understanding

For each question, distinguish the planned threshold from the evidence calculated from the sample.

  1. In your own words, explain what a significance level of 0.01 means for a test procedure when the null hypothesis is true.
  2. Why should a researcher select \(\alpha\) before looking at the test result?
  3. A test has a two-sided alternative and \(\alpha=0.06\). How much null-distribution tail area is allocated to each tail?
  4. Explain one tradeoff involved in using a smaller significance level.
  5. A student says, “The p-value was 0.032, so we should use \(\alpha=0.05\).” What is the problem with this reasoning?