Tutorials › AP Statistics › Statistically Significant Differences and Chance Variation

Experimental design · Tutorial 196 of 1000

Statistically Significant Differences and Chance Variation

Learn to compare an observed treatment difference with differences that chance assignment alone could produce, and explain what statistical significance does—and does not—show.

Beginner 10 min read

What You'll Learn

  • Explain why randomized treatment groups can differ by chance even when the treatment has no effect.
  • Choose a statistic that measures the treatment difference in context.
  • Describe how a randomization simulation models chance assignment under a no-effect assumption.
  • Estimate how often simulated differences are at least as extreme as an observed difference.
  • Interpret statistical significance without confusing it with practical importance or certainty.

Differences Can Happen by Chance

In Choosing Among Completely Randomized, Block, and Matched Pairs Designs, you learned how chance assigns experimental units to treatments in different designs. Random assignment helps create a fair comparison, but it does not guarantee that the treatment groups will be alike in every way. Even if a treatment has no effect, the units assigned to different groups can differ just by chance, and their outcomes can differ as a result.

So when an experiment produces a difference, the key question is not simply whether the groups had different results. It is whether the observed difference is unusually large compared with differences that could plausibly result from the random assignment process alone. A randomization simulation can help answer that question.

Definition: An observed difference is statistically significant when it would be unusual to see a difference that large, or larger, from chance assignment alone under the assumption that the treatment has no effect. Statistical significance is evidence against that no-effect assumption; it is not proof that the treatment caused the difference.

Modeling Chance Assignment

Begin by choosing a statistic that measures the difference relevant to the question. For a comparison of two group means, one possible statistic is the treatment-group mean minus the comparison-group mean. For two proportions, it could be the treatment-group proportion minus the comparison-group proportion. The sign tells which group had the larger statistic, while the size tells how far apart the groups were.

Next, imagine that the treatment had no effect on any experimental unit. Under that assumption, the outcomes would stay with the same units, but treatment labels could be reassigned according to the experiment’s original random assignment plan. Repeating this process creates a randomization distribution: the distribution of the statistic across assignments that could have occurred by chance under the no-effect assumption.

Key idea: A randomization simulation must reproduce the assignment process used in the experiment. Keep the treatment group sizes the same for a completely randomized design; assign treatments within blocks for a block design; and preserve the within-pair assignment structure for a matched pairs design.

Compare the observed statistic with the simulated randomization distribution. If simulated differences often reach or exceed the observed difference, then chance assignment alone is a plausible explanation for the result. If differences as large as the observed one are rare in the simulation, the data provide evidence that the no-effect assumption does not adequately explain the result.

When the question is whether the treatments differ in either direction, compare the absolute sizes of the simulated and observed differences. That counts simulated results that are at least as far from zero as the observed statistic, whether the simulated result favors the treatment or the comparison group. If the question predicted a direction in advance, the comparison can instead focus on results at least as large in that direction.

$$ \text{Estimated proportion at least as extreme} = \frac{\text{number of simulated statistics at least as extreme as observed}} {\text{number of simulated assignments}} $$

This estimated proportion is an approximation to the p-value. A p-value is the probability, assuming the no-effect model and the random assignment process are correct, of obtaining a statistic as extreme as or more extreme than the one observed. A small p-value indicates that the observed result would be unusual under that model. It is not the probability that the no-effect assumption is true, and it does not give the probability that the result happened “just by chance.”

A simulation is an approximation because it uses a limited number of assignments. More repetitions generally give a more stable estimate, but do not change the observed result. Also, the simulation needs to respect the design: for example, randomly shuffling labels across all units would not correctly model a blocked experiment if the original assignments were made separately within blocks.

Worked Example: A Difference That Is Unusual Under Chance Assignment

Worked Example: Comparing Two Seedling Light Schedules

In a fictional experiment, 40 seedlings are assigned at random, 20 to each of two light schedules. The response is growth in centimeters over a fixed period. Schedule A’s seedlings have a mean growth of 12.4 cm, and Schedule B’s have a mean of 10.8 cm. A simulation of 1,000 assignments, keeping 20 seedlings in each group, produces 37 differences with an absolute value at least as large as the observed one.

State: The treatment is light schedule, and the response is seedling growth in centimeters. The no-effect assumption is that neither schedule affects a seedling’s growth. We want to know whether the observed difference is unusual if chance assignment alone is responsible for group differences.

Plan: The experiment used a completely randomized design, so simulate assignments by repeatedly assigning 20 of the 40 seedlings to Schedule A and the remaining 20 to Schedule B. Keep the observed outcomes attached to their seedlings, and calculate the difference in means for each assignment. Because the question asks whether the schedules differ in either direction, compare absolute differences. This simulation matches the assignment plan and group sizes.

Do: The observed mean difference, Schedule A minus Schedule B, is

$$ 12.4 - 10.8 = 1.6 \text{ cm} $$

The simulated proportion at least as extreme is \(37/1000 = 0.037\), or 0.037 to three decimal places. Thus, in this simulation, about 3.7% of the chance assignments produced a mean difference at least 1.6 cm from zero in either direction.

Conclude: A difference as large as the observed 1.6 cm difference was uncommon under the no-effect model in this simulation. The result is statistically significant using the conventional 0.05 benchmark, if that benchmark was selected for the analysis. It provides evidence against the idea that chance assignment alone explains the difference, and suggests that the light schedules may have different effects on seedling growth. It does not prove that the schedules caused the difference or establish how important a 1.6 cm difference is in practice.

The simulation’s conclusion depends on the experiment’s random assignment. Random selection from a population is a different process, as explained in Why Random Selection Matters. Assignment supports a fair treatment comparison; it does not by itself establish that these seedlings represent a broader population.

Worked Example: A Difference That Is Plausible Under Chance

Worked Example: Two Ways to Practice a Music App

A fictional experiment randomly assigns 60 students to practice with either App A or App B, with 30 students in each group. The response is the number of correctly identified notes on a short assessment. The mean is 18.6 notes for App A and 17.9 notes for App B. In 1,000 simulated assignments following the same plan, 224 differences in absolute value are at least as large as the observed difference.

State: The treatments are the two practice apps, and the response is the number of correctly identified notes. The no-effect assumption is that the apps do not change students’ results. We will assess whether the observed difference is unusually large under chance assignment alone.

Plan: Simulate the original completely randomized design by repeatedly assigning 30 of the 60 students to App A and the other 30 to App B. For every assignment, calculate App A’s mean minus App B’s mean. Use absolute values because a difference favoring either app would count as evidence of a difference.

Do: The observed difference is \(18.6 - 17.9 = 0.7\) correctly identified notes. The estimated proportion of simulated differences at least as extreme is \(224/1000 = 0.224\), or 0.224 to three decimal places.

Conclude: Differences at least as large as 0.7 notes occurred in about 22.4% of the simulated assignments under the no-effect model. This is not unusual at the conventional 0.05 benchmark, so the result is not statistically significant at that benchmark. The experiment does not provide convincing evidence that the apps produce different average assessment results. That is not proof that their effects are identical; a real difference might exist but not be clearly revealed by these data.

“Not statistically significant” does not mean that the observed groups had exactly the same results. Here, the means differ by 0.7 notes. The point is that differences of this size were not rare under the chance-only model used in the simulation.

Worked Example: Respecting Matched Pairs

Worked Example: Comparing Two Study Sounds in Matched Pairs

In a fictional matched pairs experiment, 12 students each complete a comparable memory task once with Sound A and once with Sound B. The order is randomly assigned for each student. Define each student’s difference as the score with Sound A minus the score with Sound B. The 12 differences, in points, are \(2, 1, 0, 3, 1, 2, -1, 2, 1, 0, 3, 1\). A simulation of 1,000 legal within-pair assignments produces 42 mean differences with an absolute value at least as large as the observed one.

State: The treatments are the two study sounds, and the response is memory-task score. The no-effect assumption is that neither sound changes any student’s score. The relevant statistic is the mean of the within-student differences, so each student serves as their own comparison.

Plan: Because this is a matched pairs design, simulate the random treatment order within each student’s pair. Equivalently, for each simulated assignment, keep or reverse the sign of each student’s difference according to a chance process, then calculate the mean difference. Do not shuffle all 24 scores into two unrelated groups; that would ignore the pairing in the design.

Do: The observed differences sum to \(2+1+0+3+1+2-1+2+1+0+3+1=15\) points. Their mean is \(15/12=1.25\) points, so the observed mean difference favors Sound A by 1.25 points. Of 1,000 simulated legal assignments, 42 have an absolute mean difference at least 1.25. The estimated proportion is \(42/1000=0.042\), or 0.042 to three decimal places.

Conclude: Under the no-effect model, a mean difference at least this far from zero was uncommon in the simulation. The result is statistically significant at the conventional 0.05 benchmark, if that benchmark was selected. It provides evidence that the sounds may affect memory-task scores differently. The experiment’s matched pairs structure is essential to this comparison, and statistical significance alone does not tell us whether a 1.25-point difference matters for students.

Statistical Significance Is Not Practical Importance

Statistical significance describes how surprising an observed result would be under a specified chance-only model. Practical importance asks whether the size of the difference matters in the situation. These are different questions. A small difference can be statistically significant when data are sufficiently informative, while a potentially important difference might not be statistically significant if the experiment has substantial variation or too few experimental units.

The statistic’s units and context matter. A difference of 1.6 cm in seedling growth, 0.7 correctly identified notes, and 1.25 memory-task points are not interchangeable quantities. Report the observed difference in context even when discussing its statistical significance. A small estimated tail proportion does not tell you the size or practical value of the treatment effect.

Conditions for a randomization comparison:
  • The experiment used random assignment, so chance assignments provide a meaningful model for treatment labels.
  • The simulation follows the actual assignment plan, including group sizes and any blocks or matched pairs.
  • The statistic measures the treatment difference asked about, and the comparison counts simulated outcomes at least as extreme in the relevant direction.

These conditions describe the basis for the chance comparison; they do not guarantee that measurements are perfect or that every feature of a study is free of problems. As discussed in Sources of Variability in Collected Data, data can vary for more than one reason. A randomization simulation specifically examines variation from treatment assignment under its no-effect model.

Common Mistakes and AP Exam Tips

  • Assuming any observed group difference proves a treatment effect. Random assignment can produce unequal groups and unequal outcomes. Compare the observed statistic with what chance assignments could produce.
  • Ignoring the experimental design in a simulation. For matched pairs, preserve pairs; for blocks, reassign within blocks. A simulation that uses a different assignment process does not model the experiment correctly.
  • Interpreting a p-value as the probability the treatment has no effect. State it as a probability of results at least as extreme as the observed result, assuming no treatment effect and the random assignment model.
  • Calling a non-significant result proof of no effect. Say there is not convincing evidence of a difference under the chosen analysis. Do not claim that the treatments are identical.
  • Equating statistical significance with practical importance. Report the size and units of the observed difference, then discuss practical importance separately.
  • Forgetting which direction counts as extreme. If either treatment could produce the larger response, count simulated differences at least as extreme in either direction. If a direction was specified in advance, describe the comparison accordingly.

For a strong AP response, identify the observed statistic and explain how the randomization simulation follows the design. Then compare the observed result with simulated results and state what that comparison says about chance variation under the no-effect assumption. Use “statistically significant” to describe unusualness under that model—not certainty, practical importance, or proof.

Key takeaway: Random assignment can produce differences even when treatments have no effect. An observed difference that is rare in a correctly designed randomization simulation is evidence against a chance-only explanation, but statistical significance is not proof and does not determine practical importance.

Check Your Understanding

Use the ideas in this tutorial to answer each question.

  1. In a randomized experiment, why can two treatment groups have different mean responses even if the treatment has no effect?
  2. A simulation produces 1,000 assignments, and 18 produce a difference at least as extreme as the observed difference. What estimated proportion is at least as extreme?
  3. Why should a randomization simulation for a matched pairs experiment preserve the pairs?
  4. In context, what does it mean if the observed difference is statistically significant under the no-effect model?
  5. Why does a statistically significant result not necessarily mean that the treatment difference is practically important?