Tutorials › AP Statistics › What Is a Probability Simulation

Estimating probability by simulation · Tutorial 201 of 1000

What Is a Probability Simulation

See how repeated simulated outcomes can estimate a probability, and why results tend to settle near the model’s probability as the number of trials grows.

Beginner 8 min read

What You'll Learn

  • Define a probability simulation and identify its trial, outcome, and event of interest
  • Calculate a simulated probability as a relative frequency
  • Compare fair-coin results after 10, 100, and 1,000 simulated flips
  • Explain why longer simulations tend to give more stable estimates without guaranteeing an exact result
  • Design a chance-based simulation for a multi-step event
  • Recognize why different runs of the same simulation can produce different estimates

How Can Repeated Chance Outcomes Estimate a Probability?

In Mixed Practice: Designing and Critiquing Experiments, you practiced evaluating how researchers assign treatments to experimental units. A probability simulation is different: it uses a chance process to imitate outcomes from a model and help estimate how likely an event is. Instead of conducting a real experiment or waiting for many real-world events, we generate outcomes according to the model and record what happens.

For example, suppose we want to estimate the chance of getting heads when flipping a fair coin. We could flip a real coin many times, or we could use random numbers to imitate the flips. If the simulation is set up to give heads and tails equal chances, the proportion of simulated flips that are heads estimates the probability of heads.

Definition: A probability simulation is a chance-based process that imitates a real or model situation. A trial is one repetition of the process. A simulated probability estimate is the relative frequency of the event of interest: the number of trials in which the event occurs divided by the total number of trials.

The simulation does not make every outcome predictable. It creates outcomes that vary by chance, as real outcomes would. The key is that the chance process must represent the model we want to study. For a fair coin, the simulation must give heads and tails equal chances on each flip.

A simulation estimates a probability because, over many repetitions, the relative frequency of an event tends to settle near its probability in the model. For a fair coin, the model probability of heads is \(0.5\). A short sequence may have a proportion quite different from \(0.5\), but a much longer sequence tends to give a more stable estimate. This tendency does not mean that every longer run is closer than every shorter run, or that a run must eventually produce exactly half heads.

A Fair-Coin Simulation at Three Trial Counts

Imagine using a calculator’s random integer feature to imitate a fair coin. For each simulated flip, generate an integer from 0 to 9. Let 0 through 4 represent tails and 5 through 9 represent heads. There are five digits assigned to each outcome, so each has probability \(0.5\) in this model. Repeat the process once per flip.

The table shows one possible set of cumulative results from a simulated sequence. The outcomes are invented for illustration; they are not results from a real study. The first 10 flips are shown, and the later rows give the cumulative number of heads in that same sequence.

Number of simulated flipsHeads observedRelative frequency of heads
106\(6/10 = 0.600\)
10053\(53/100 = 0.530\)
1,000507\(507/1{,}000 = 0.507\)

In this run, the estimate moves from \(0.600\) after 10 flips to \(0.530\) after 100 and \(0.507\) after 1,000. The last estimate is closer to the model probability \(0.5\) than the first one is. That illustrates how a larger number of repetitions can produce a more stable estimate of the model probability.

However, the relative frequency does not have to move closer to \(0.5\) every time another flip is added. If the next several flips are all heads, for instance, the proportion might move away from \(0.5\) for a while. The general pattern concerns what tends to happen across many trials, not a rule that every individual step improves the estimate.

Worked Example: Estimating the Chance of Heads

A student simulates 10 fair-coin flips by generating random digits and assigning 0 through 4 to tails and 5 through 9 to heads. The simulated sequence is H, T, H, H, T, H, T, H, H, T. Estimate the probability of heads from this run and compare it with the model probability for a fair coin.

Identify the trial and event: One trial is one simulated coin flip. The event of interest is getting heads. There are 10 trials, and heads occurs 6 times in the sequence.

Calculate the estimate: Divide the number of simulated flips that resulted in heads by the total number of simulated flips.

$$ \text{Estimated probability of heads} = \frac{\text{number of heads}}{\text{number of flips}} = \frac{6}{10} = 0.600 $$

The estimated probability is \(0.600\), or 60%. The model probability for heads is \(0.5\), or 50%, so this short simulation gives an estimate \(0.100\) above the model probability. That difference does not show that the coin model is wrong: only 10 flips were simulated, and short runs can vary noticeably by chance.

If the student continues the same simulation and records many more flips, the relative frequency may settle closer to \(0.5\). It could still be above or below \(0.5\), and the amount of improvement is not guaranteed at each added flip.

Simulating a Multi-Flip Event

A simulation can also estimate the probability of an event involving several outcomes together. In that case, one trial of the simulation must include all the steps needed to decide whether the event occurred. For example, to study a three-flip event, one trial consists of three simulated flips—not one flip.

A convenient way to imitate the fair coin is to generate a random integer from 1 to 6 for each flip. Assign 1, 2, or 3 to tails and 4, 5, or 6 to heads. Each outcome is assigned to three of the six possible integers, so each simulated flip has a \(0.5\) chance of being heads. Use a fresh random integer for every flip.

Worked Example: Estimating the Chance of Exactly Two Heads

A student wants to estimate the probability of getting exactly two heads in three flips of a fair coin. The student carries out 800 simulated trials. In each trial, the student generates three random integers from 1 to 6, with 1–3 representing tails and 4–6 representing heads. Exactly two heads occur in 309 of the 800 trials.

Define a trial and the event: One trial consists of three simulated flips. The event of interest is that exactly two of those three flips are heads. The simulation produced 309 trials in which this event occurred, out of 800 total trials.

Calculate the simulated estimate:

$$ \text{Estimated probability} = \frac{\text{number of trials with exactly two heads}}{\text{total number of trials}} = \frac{309}{800} = 0.38625 \approx 0.3863 $$

The simulation estimates the probability as \(0.3863\), or about 38.63%. As a check on how close this estimate is, there are 8 equally likely head-and-tail sequences for three fair flips. Exactly two heads occurs in 3 of them, so the model probability is \(3/8 = 0.375\). The simulated estimate is \(0.01125\) above that value. A simulation estimate need not equal the model probability exactly, even when the chance process is set up correctly.

The estimate is based on 800 repetitions of the whole three-flip process. Counting 2,400 individual flips as though they were 2,400 separate trials for the event “exactly two heads in three flips” would be incorrect: the event can only be checked after completing a group of three flips.

Why Repeated Simulations Can Still Differ

A simulation uses chance, so running it again can produce a different set of outcomes and a different relative frequency. The model probability stays the same, but the estimate from a particular run varies. This is why it is useful to report both the simulated estimate and the number of trials used to obtain it.

The following example shows how separate batches of fair-coin flips can vary even when every batch uses the same chance model. Comparing batches also helps distinguish a single run’s result from the long-run tendency.

Worked Example: Comparing Five Batches of Coin Flips

A student runs five separate simulations, each containing 100 fair-coin flips. The numbers of heads are 46, 51, 54, 48, and 52. Estimate the probability of heads in each batch and for all five batches combined.

Calculate each batch’s relative frequency: Divide each batch’s number of heads by 100. The five estimates are \(46/100 = 0.46\), \(51/100 = 0.51\), \(54/100 = 0.54\), \(48/100 = 0.48\), and \(52/100 = 0.52\). The estimates differ from batch to batch, even though the same fair-coin model was used each time.

Combine the results: There are \(5 \times 100 = 500\) simulated flips. The total number of heads is \(46 + 51 + 54 + 48 + 52 = 251\). Thus, the relative frequency for all 500 flips is:

$$ \frac{251}{500} = 0.502 $$

The combined estimate is \(0.502\), close to the model probability \(0.5\). The five individual estimates range from \(0.46\) to \(0.54\). Pooling more simulated flips gives one estimate based on more trials, but it does not erase the fact that individual batches can differ. Another set of five batches could have different head counts and a different combined estimate.

Set Up a Simulation That Answers the Right Question

A simulation is only useful when its chance process matches the situation being modeled and the event is counted correctly. Before generating outcomes, describe what one trial represents, how chance produces its outcomes, and what counts as a success. A success is simply an outcome that meets the event of interest; it does not have to be a desirable result.

1
Define one trial.
State exactly what is repeated. For a single-coin-flip question, one trial is one flip. For an event involving three flips, one trial is a complete set of three flips.
2
Choose a chance process that matches the model.
Specify how random numbers or another chance device represent the possible outcomes, and make sure the assignments give outcomes the intended chances.
3
Record whether the event occurs.
Apply the event rule to each completed trial. Keep a count of the trials where the event occurs and the total number of trials.
4
Calculate and interpret the estimate.
Divide the event count by the total number of trials. Describe the result as an estimate from the simulation, not as a guarantee about what will happen in one future trial.

These steps make the simulation reproducible in principle and make its result interpretable. If the chance assignments do not match the model, or if the event is counted incorrectly, a large number of simulated trials will not fix the problem.

Common Mistakes and AP Exam Tips

  • Calling a simulated estimate the exact probability. A relative frequency from a finite simulation is an estimate. Say, “The simulation estimates the probability at \(0.507\),” rather than claiming the probability is exactly \(0.507\).
  • Using an unclear trial. State the complete process that is repeated. In the three-flip example, a trial is three flips together, not one flip.
  • Dividing by the wrong total. Use the number of completed trials as the denominator. For 800 three-flip trials, the denominator for the event “exactly two heads in three flips” is 800.
  • Assuming the relative frequency must improve after every trial. It may move toward or away from the model probability as outcomes are added. A larger number of trials tends to give a more stable estimate; it does not guarantee an improvement at each step.
  • Claiming every large run will be close to the model probability. Long runs tend to produce estimates near the model probability, but chance variation remains. A careful conclusion describes a tendency, not a certainty for every run.
  • Forgetting to match the chance process to the model. If a fair coin is being simulated, heads and tails need equal chances. Explain how the random numbers are assigned so that this is clear.

For a full-credit explanation, name the event, give the number of times it occurred and the number of trials, show the relative-frequency calculation, and interpret the result as an estimate in context. When explaining why a longer simulation helps, say that the relative frequency tends to settle near the model probability—not that it must equal the probability or move closer on every step.

Key takeaway: A probability simulation imitates a chance model through repeated trials. The event’s relative frequency estimates its model probability, and more trials tend to make the estimate more stable, while chance variation means different runs can still give different results.

Check Your Understanding

Answer each question using the ideas of trials, events, and relative frequency.

  1. A student simulates 40 fair-coin flips and records 23 heads. What is the estimated probability of heads? Show the calculation and interpret it.
  2. To estimate the probability of getting at least one head in four fair-coin flips, what should one trial of the simulation consist of?
  3. A simulation estimates an event’s probability as \(0.47\) after 100 trials and \(0.51\) after 1,000 trials. Does this show that every longer simulation must be closer to the model probability? Explain.
  4. In 600 simulated trials of a three-flip event, the event occurs 219 times. Calculate the simulated probability estimate.
  5. Explain why two separate simulations of 500 fair-coin flips can produce different relative frequencies of heads, even when both use the same chance process.