From a Simulation Description to a Dotplot
In Describing a Simulation in Words, you learned how to specify a chance model, define one trial, and decide what to record. When many trials are simulated, their recorded results can be displayed in a dotplot. Reading that plot lets you estimate how often the simulation produces results at least as extreme as an observed count.
The key is to identify the statistic shown on the horizontal axis, locate the observed result, and decide which direction or directions count as “at least as extreme.” Then count the simulated results in those parts of the plot. As in Estimating a Probability from Simulation Results, divide the number of qualifying trials by the total number of simulated trials.
Usually, each dot represents the statistic recorded from one simulated trial. If several trials produce the same count, their dots are stacked above that count. A dotplot’s shape helps you see which results occur often or rarely, but the estimate comes from counting the relevant dots—not from judging how large an area looks.
Choose the Tail Before Counting
“Extreme” depends on the question. If the question is whether an observed count is unusually high, count simulated results equal to or greater than the observed count. If it is whether the count is unusually low, count results equal to or less than the observed count. The observed result itself is included because the question asks for results at least as extreme.
For a question about unusually different results in either direction, use both tails. When the simulation model says the statistic should be centered at a particular value, compare how far each simulated result is from that center. A two-sided rule includes simulated values at least as far from the center as the observed result.
This is a relative frequency from the completed simulations, so it estimates a probability under the stated model; it is not the exact probability. If the simulation is modeling a null situation, this tail proportion estimates how often chance alone produces a result at least as extreme as the observed one under that model. The decision about which results count as extreme must come from the question, not from which tail happens to have more dots.
Worked Example: Count an Upper Tail
A fair coin is tossed 30 times, and the observed number of heads is 21. To judge how unusual that count would be if the coin were fair, a student simulates 200 sets of 30 fair-coin tosses. The dotplot records the number of heads in each simulated set. Its stacks have these heights:
| Simulated heads | Number of dots |
|---|---|
| 10 | 1 |
| 11 | 2 |
| 12 | 4 |
| 13 | 8 |
| 14 | 15 |
| 15 | 25 |
| 16 | 33 |
| 17 | 38 |
| 18 | 30 |
| 19 | 21 |
| 20 | 13 |
| 21 | 6 |
| 22 | 3 |
| 23 | 1 |
| 24 | 0 |
State: We want to estimate the chance, under the fair-coin model, of getting at least 21 heads in 30 tosses. Because the concern is an unusually high count, the relevant results are 21 or more heads.
Plan: Use the simulated dotplot. Count dots at 21, 22, 23, and any larger values, then divide by all 200 simulated trials. Each dot represents one complete set of 30 simulated tosses.
Do: The upper-tail count is \(6+3+1+0=10\) dots. The frequencies in the entire table sum to 200, so the estimate is:
Conclude: In this simulation, 5% of the simulated sets produced at least 21 heads. Thus, the estimated probability of getting 21 or more heads in 30 tosses under the fair-coin model is 0.05, or 5%. This is an estimate from 200 simulations, not an exact probability.
The six dots at 21 count because the observed result is included. Counting only values greater than 21 would answer a different question: the chance of getting more than 21 heads.
Count the Lower Tail for an Unusually Small Result
For a low observed count, the relevant part of the dotplot is on the low-value side. The same relative-frequency calculation applies, but the direction of the comparison changes. Make that choice before counting, and include the observed value and any values below it.
Worked Example: Estimate a Lower-Tail Probability
A spinner is expected to land on blue 40% of the time. In 20 observed spins, it lands on blue only 3 times. A simulation follows that chance model for 500 trials, each trial consisting of 20 spins. The dotplot displays the following groups of simulated blue counts:
| Blue counts shown in the dotplot | Number of dots |
|---|---|
| 0–3 | 8 |
| 4–5 | 52 |
| 6–7 | 130 |
| 8–9 | 156 |
| 10–11 | 100 |
| 12–13 | 43 |
| 14–15 | 10 |
| 16–20 | 1 |
The observed count of 3 is low, so count simulated results at or below 3. The first group contains exactly those results. Its 8 dots are the qualifying simulations, and all group frequencies sum to \(8+52+130+156+100+43+10+1=500\).
The simulation estimates a 0.016 probability, or 1.6%, of getting 3 or fewer blue results in 20 spins if the spinner lands on blue with probability 0.40 on each spin. This describes the lower tail under the stated model. It does not say that the spinner’s actual probability is exactly 0.40; it reports what the model’s simulations produced.
Because the dotplot groups values 0 through 3 together, there is no need to separate their individual frequencies to answer this question. If the observed value were 5 instead, the first two groups would be relevant, for a total of \(8+52=60\) dots out of 500.
For Two-Sided Questions, Count Both Tails
Sometimes the question is whether a result is unusually far from what the chance model predicts, in either direction. In that case, counting only results larger than the observed value misses possible results that are just as extreme on the low side. First identify the model’s null center for the statistic; then compare distances from that center.
For a difference in counts, the null center is the expected difference under the null model; it is zero when the design makes the expected counts equal, as with equal group sizes. If that center is zero and the observed difference is 6, results at least as far from zero include differences of 6 or more and differences of −6 or less. The plus or minus signs indicate direction; the distance from zero determines extremeness for this two-sided rule.
Worked Example: Count Both Tails of a Difference
In a hypothetical experiment, participants are randomly assigned in equal numbers to one of two programs. The observed statistic is the number meeting a goal in Program A minus the number meeting it in Program B. The observed difference is \(+6\). Under a model in which the programs have no effect, the simulation repeatedly assigns the outcome labels by chance and records the difference in counts. A dotplot summarizes 400 simulated differences:
| Difference | Dots | Difference | Dots |
|---|---|---|---|
| −9 | 1 | 2 | 40 |
| −8 | 2 | 3 | 25 |
| −7 | 5 | 4 | 12 |
| −6 | 12 | 5 | 5 |
| −5 | 20 | 6 | 1 |
| −4 | 32 | 7 | 1 |
| −3 | 42 | 8 or more | 0 |
| −2 | 50 | — | — |
| −1 | 52 | — | — |
| 0 | 48 | — | — |
| 1 | 52 | — | — |
A two-sided question asks for simulated differences at least as far from zero as the observed difference of \(+6\). Therefore, count differences of \(+6\) or greater and differences of \(−6\) or less. The negative side contributes \(1+2+5+12=20\) dots. The positive side contributes \(1+1=2\) dots.
The simulation estimates a 0.055, or 5.5%, chance of a difference at least as far from zero as the observed difference, if the no-effect model is correct. The observed \(+6\) is included in the positive tail. The simulation happened to produce more qualifying dots on the negative side than the positive side; count both sides anyway, because the question is about distance in either direction.
The full set of frequencies sums to 400, so the denominator includes every simulated assignment, not just the values near the tails. As in Statistically Significant Differences and Chance Variation, a small tail proportion can indicate that an observed result would be unusual under the chance model. The estimate itself does not explain why the observed difference occurred.
Common Mistakes and AP Exam Tips
- Counting the wrong tail. A question about “at least” a high count calls for the upper tail; “at most” a low count calls for the lower tail. State the comparison rule before counting.
- Leaving out the observed value. “At least as extreme” includes the observed result. Use greater than or equal to, or less than or equal to, as appropriate.
- Using only one tail for a two-sided question. When either direction is evidence against the model, count qualifying values on both sides of the null center.
- Dividing by the number of dots in the tail. The denominator is the total number of simulated trials. The tail count is the numerator.
- Confusing dots with individual outcomes inside a trial. A dot represents one recorded result from one complete simulated trial. It does not represent one coin toss or one spinner spin in the examples above.
- Calling the estimate exact. A finite simulation gives a relative frequency that estimates a model probability. Different runs can produce different estimates, as discussed in How Many Trials Are Enough.
- Judging the tail by appearance alone. A dotplot makes patterns visible, but a probability estimate requires counting the relevant dots and dividing by the total.
For a complete AP-style explanation, name the observed count or statistic, state which simulated values count as at least as extreme, show the tail count and total number of simulations, and interpret the resulting proportion in context under the model. For a two-sided comparison, also state the center and explain how far from it a simulated result must be.
Check Your Understanding
For each question, use the stated observed result and simulation summary to identify the qualifying dots and calculate the estimated probability.
- A fair die is rolled 12 times. The observed number of sixes is 5. In 300 simulated sets of 12 rolls, 18 sets have 5 sixes and 7 have more than 5. What is the estimated probability of at least 5 sixes?
- A model predicts a count of 10. The observed count is 7, and a dotplot of 250 simulated counts has 4 dots at 7 or below. What lower-tail probability does the simulation estimate?
- A two-sided statistic is centered at zero, and the observed value is −4. A dotplot has 6 simulated values at or below −4 and 9 at or above 4, out of 500 trials. Estimate the probability of a result at least as far from zero.
- Why should the observed value itself be included when counting results “at least as extreme”?
- A dotplot contains 40 dots in the relevant tail and 800 dots in total. Write the estimated probability and explain what it means under the simulation model.