Compare an Observed Result with a Simulated Distribution
In Reading a Dotplot of Simulation Results, you learned to count simulated results in a specified tail. Now we use that idea to ask a bigger question: does an observed result seem surprising if a particular chance model is true?
Suppose someone flips a coin 10 times and gets 9 heads. That result is possible even for a fair coin, but is it common or rare? To answer, we can simulate many sets of 10 flips under a fair-coin model and compare the observed count with the simulated distribution of head counts.
Here, the statistic is the number of heads in 10 flips. The model says that each flip has probability 0.5 of heads and that the flips are independent. One complete trial is 10 flips, and one result from a trial is its head count. This follows the simulation planning ideas from Setting Up a Simulation Model and Describing a Simulation in Words.
A simulation could represent each flip with one random digit: digits 0–4 mean heads and 5–9 mean tails. Generate 10 digits for one trial, count the heads, and repeat the full trial many times. As explained in Common Errors in Designing Simulations, the complete 10-flip trial—not an individual flip—must be repeated and counted.
Choose the Relevant Tail
The tail is the part of the simulated distribution containing results at least as extreme as the observed result. Which tail to count depends on the question. If the question is whether the coin produced unusually many heads, count results equal to or greater than the observed number. If it asks about unusually few heads, count results equal to or less than the observed number.
Include results equal to the observation. For 9 heads, a result of 9 is at least as extreme as 9 in the “many heads” direction, and 10 is even more extreme. Counting only results greater than 9 would incorrectly leave out the observed result itself.
This estimate is not the exact probability. As discussed in Estimating a Probability from Simulation Results and How Many Trials Are Enough, simulation results vary from run to run. More trials generally make a relative-frequency estimate less variable, but do not guarantee any one run will be closer to the exact probability.
Worked Example: Is 9 Heads in 10 Flips Surprising?
A student flips a coin 10 times and observes 9 heads. To evaluate this result, a hypothetical simulation repeats 10 flips 1,000 times, using a fair-coin model for every flip. The simulated head counts are summarized below.
| Heads in 10 flips | Number of simulated trials |
|---|---|
| 0 | 1 |
| 1 | 8 |
| 2 | 45 |
| 3 | 117 |
| 4 | 207 |
| 5 | 249 |
| 6 | 202 |
| 7 | 115 |
| 8 | 43 |
| 9 | 11 |
| 10 | 2 |
State: We want to decide whether 9 heads is an unusual result if the coin is fair and the 10 flips are independent. The observed statistic is 9 heads.
Plan: Because the question asks about unusually many heads, count simulated trials with 9 or more heads. The simulation already follows the model of 10 independent fair-coin flips per trial. We also check that the table accounts for all 1,000 trials.
Do: There are 11 simulated trials with 9 heads and 2 with 10 heads, so 13 trials have at least 9 heads. The frequencies in the table sum to 1,000. The estimated tail probability is:
Conclude: In this hypothetical simulation, about 1.3% of sets of 10 fair-coin flips produced 9 or more heads. That is a small proportion, so 9 heads appears unusual under the fair-coin model. It does not prove the coin is unfair: a rare result can still happen by chance, and the conclusion depends on the model and the quality of the simulation.
What “Unusual” Means in This Comparison
“Unusual” is a judgment about how much of the simulated distribution lies at least as far into the relevant tail as the observation. If only a small fraction of simulated trials reach or exceed the observed result, then the result would not commonly occur under the model. If a large fraction do, the result is not especially surprising under that model.
The simulation itself gives a relative frequency, not a universal boundary between unusual and ordinary. For example, the estimate 0.013 in the worked example describes the simulated tail proportion. It does not, by itself, say what action to take or establish that the model is wrong. Later work on significance will formalize how a small tail probability is used as evidence against a chance model.
Keep the comparison tied to the model being tested. Saying “9 heads is unusual” without qualification is incomplete. A careful statement says that 9 or more heads appears unusual if the coin is fair and the flips are independent, based on the simulated tail proportion. The simulation evaluates results under those assumptions; it does not independently verify them.
Worked Example: Is 8 Heads Also Unusual?
Use the same hypothetical distribution of 1,000 simulated sets of 10 fair-coin flips. Suppose a different student observes 8 heads and asks whether this is unusually many heads.
State: The question concerns unusually many heads, so the relevant direction is the upper tail. The observed statistic is 8.
Plan: Count every simulated result of 8, 9, or 10 heads. Include the observed value, 8, and all larger head counts. Divide this total by the 1,000 simulated trials.
Do: The table has 43 trials with 8 heads, 11 with 9 heads, and 2 with 10 heads. Thus, 56 trials produced at least 8 heads:
Conclude: The estimated probability of 8 or more heads under the fair-coin model is 0.056, or 5.6%. In this simulation, 8 heads is less unusual than 9 heads, because more simulated trials reached 8 or more. The result is not as rare in this model as the 9-or-more outcome.
Notice that the tail changes with the observed value. For an observation of 8, counting only 9 and 10 would answer a different question: how often the simulation produced more than 8 heads. The “at least as extreme” rule includes 8 itself.
One Direction or Either Direction?
Sometimes the question predicts a direction in advance, such as whether a coin produces unusually many heads. Then the appropriate comparison is one-sided: count results in that direction. Other questions ask whether the result is surprising in either direction—unusually many or unusually few heads. For those questions, both tails may matter.
The fair-coin model distribution is centered at 5 heads, though the empirical simulated distribution need not be exactly centered there. If the observed result is 9 heads and the question is whether the result differs unusually from the center in either direction, results at least as far from 5 as 9 are 9 and 10 heads, along with 1 and 0 heads. This compares equal distances from the center: 9 is four above 5, while 1 is four below 5.
Worked Example: Count Both Tails for an Either-Direction Question
A question asks whether 9 heads in 10 flips is unusual in either direction under the fair-coin model—not specifically whether there are unusually many heads. Use the same simulated distribution.
State: The observation, 9 heads, is four heads above the center of 5. For this either-direction question, we count results at least four away from 5: 0, 1, 9, or 10 heads.
Plan: Add the simulated frequencies in both tails, including results equally far from the center as the observation. Divide by the total number of simulated trials.
Do: There was 1 trial with 0 heads, 8 with 1 head, 11 with 9 heads, and 2 with 10 heads. The total is 22 trials:
Conclude: In this simulation, 2.2% of trials were at least as far from 5 heads as the observed result in either direction. That is a small simulated proportion, so 9 heads appears unusual for this two-direction question under the fair-coin model. The two-direction estimate differs from the 1.3% upper-tail estimate because it also counts unusually few heads.
Common Mistakes and AP Exam Tips
- Counting the wrong tail. If the question asks about unusually many heads, count outcomes at or above the observation. If it asks about unusually few, count outcomes at or below it. If it asks about either direction, explain which outcomes in both tails qualify.
- Leaving out the observed value. “At least as extreme” includes the observed statistic. For 9 heads, the upper-tail count includes both 9 and 10.
- Dividing by the wrong total. Divide the qualifying simulated trials by all completed trials, not only by the trials in one part of the table. In the first example, the denominator is 1,000.
- Calling the estimate an exact probability. A simulation provides an estimate that can vary from run to run. Say “the estimated proportion in this simulation” or “about 1.3%,” rather than claiming the simulation establishes an exact probability.
- Claiming that a rare result proves the model false. A rare result may provide evidence against the model, but it remains possible under the model. State the conclusion conditionally and in context.
- Choosing a direction after seeing the result without explanation. The direction must match the question. If the question concerns either unusually many or unusually few, account for both tails rather than reporting only the tail that makes the result seem rarer.
A full-credit explanation identifies the model, names the observed statistic, justifies the tail or tails, shows the count and denominator, and interprets the simulated proportion in context. For example: “Assuming a fair coin and independent flips, 13 of 1,000 simulated sets had at least 9 heads, an estimated proportion of 0.013. Thus, 9 heads appears unusual under this model, although it does not prove the coin is unfair.”
Check Your Understanding
Use the simulated distribution in the first worked example unless a question gives different information.
- For an observation of 10 heads, how many simulated trials are at least as extreme in the “unusually many heads” direction? What is the estimated tail proportion?
- For an observation of 7 heads and a question about unusually many heads, which rows of the table should be counted? Find the estimated tail proportion.
- For an observation of 2 heads and a question about unusually few heads, which rows should be counted? Find the estimated tail proportion.
- Why should an explanation say “unusual under the fair-coin model” rather than simply “the coin is unfair”?
- If a question asks whether 9 heads is unusual in either direction, why is counting only 9 and 10 heads not enough?