Tutorials › AP Statistics › Simulating Repeated Trials Until Success

Estimating probability by simulation · Tutorial 210 of 1000

Simulating Repeated Trials Until Success

Design a simulation that counts boxes until every prize has appeared, with a clear stopping rule for each trial.

Beginner 10 min read

What You'll Learn

  • Define one trial as buying boxes until every prize type has appeared at least once.
  • Assign labels to prizes and use random values to represent each box.
  • Keep repeated prize labels in the count of boxes, even when they add no new prize type.
  • Apply a stopping rule that ends a trial as soon as the collection is complete.
  • Separate consecutive trials by resetting the collection while continuing the random stream.
  • Estimate an event probability using completed trials and their recorded box counts.

When Does a Trial End?

In Sampling With and Without Replacement in Simulations, you learned to apply a rule consistently as random labels are generated. A collection problem needs one more decision: exactly when is one trial complete? If each cereal box contains one randomly selected prize, a trial can represent buying boxes until every prize type has been collected.

Suppose there are three prize types. We can label them 1, 2, and 3. A new box may contain a prize already collected, so a repeated label does not add a new type to the collection. But it still represents another box bought and must be included in the count. The trial ends at the first box after which all three labels have appeared.

Definition: A stopping rule states exactly when a simulated trial ends. In this collection simulation, one trial begins with no prizes collected and stops as soon as every prize type has appeared at least once. The outcome recorded for the trial is the number of boxes bought, including boxes containing repeated prizes.

The stopping rule is based on completing the collection, not on buying a predetermined number of boxes. Some trials may end quickly, while others take longer. A repeat can delay completion, but it does not restart a trial or get left out of its box count.

Set Up the Collection Simulation

Assume that each box independently contains one of the three prizes, and that each prize is equally likely. Use labels 1, 2, and 3 for the prizes. The command randInt(1,3,n) generates \(n\) labels, with each label equally likely. Read the labels in order, tracking which prize types have appeared in the current trial.

This model treats each box as a new chance to receive any of the three prizes. A prize label can therefore appear more than once. In the terminology from the previous tutorial, repeats are kept: they represent real box purchases, even though they do not increase the number of distinct prize types collected. At the start of every new trial, reset the collection to empty.

1
Assign labels.
Use 1, 2, and 3 to represent the three prize types.
2
Define one trial.
Begin with no prizes collected. One trial represents buying boxes until all three prize types have appeared.
3
Generate box outcomes in order.
Each random label represents the prize in the next box. Keep repeated labels and count each one as a box.
4
Apply the stopping rule.
Stop the trial immediately when labels 1, 2, and 3 have all appeared. Record how many labels—and therefore boxes—the trial used.
5
Start another trial.
Reset the collection to empty, then continue generating labels for the next trial.

The model must match the situation being imitated. Here, the labels are equally likely because the example assumes equally likely prizes. If the prizes had unequal chances, the digit assignments would need to represent those probabilities, as in Simulating Events with Unequal Probabilities. The stopping rule would still be the same: stop when each prize type has appeared.

Follow One Trial Carefully

Suppose a random generator produces the labels 2, 1, 2, 3, 1. Start with no prizes. The first box gives prize 2; the second gives prize 1. The third box gives prize 2 again. That repeat does not add a new prize, but it is still the third box. The fourth box gives prize 3, so all three prize types have now appeared. Stop there. The fifth label belongs to the next part of the random stream and is not used in this trial.

Worked Example: Record One Completed Collection

Three prize types are labeled 1, 2, and 3. A simulated trial produces the stream 2, 1, 2, 3, 1. Determine how many boxes are bought before the collection is complete.

After box 1, the collection is \(\{2\}\). After box 2, it is \(\{1,2\}\). Box 3 contains prize 2 again, so the collection remains \(\{1,2\}\), but the box count increases to 3. Box 4 contains prize 3, giving the complete collection \(\{1,2,3\}\). The stopping rule is met at box 4, so the trial outcome is 4 boxes.

Check: The fifth label, 1, is not part of this trial because the collection was already complete after box 4. If a new trial begins, its collection resets to empty, and that next label can be used for the new trial.

This example shows why the stopping point must be stated precisely. Stopping after a fixed number of labels could leave a prize missing or could count extra boxes after the collection was already complete. A clear stopping rule avoids both errors.

Simulate Several Complete Trials

To simulate several trials using one stream, keep reading labels until the current trial is complete. Then reset the collection and continue from the next unused label. Do not start a new trial after a fixed number of generated labels unless that is where the current collection has actually become complete.

Once the trials are complete, you can estimate the probability of an event about the number of boxes. For example, the event might be “the collection takes at least four boxes.” Count the completed trials for which that event occurs, then divide by the total number of completed trials. As in What Is a Probability Simulation, this relative frequency is an estimate from the simulation, not a guarantee of the exact model probability.

Worked Example: Estimate the Chance of Needing at Least Four Boxes

Use this stream to simulate four trials of collecting all three prize types:

$$ 1,\ 2,\ 1,\ 3,\quad 2,\ 2,\ 3,\ 1,\quad 3,\ 1,\ 2,\quad 1,\ 1,\ 2,\ 2,\ 3 $$

Read left to right and start each trial with an empty collection. Trial 1 uses 1, 2, 1, 3. Prize 3 first appears on the fourth box, so the trial outcome is 4. Trial 2 uses 2, 2, 3, 1. The repeated 2 counts as a box, and prize 1 completes the collection on box 4, so this outcome is also 4.

Trial 3 uses 3, 1, 2 and completes after 3 boxes. Trial 4 uses 1, 1, 2, 2, 3 and completes after 5 boxes. The trial boundaries and counts are:

TrialLabels usedBoxes until all prizes appearAt least 4 boxes?
11, 2, 1, 34Yes
22, 2, 3, 14Yes
33, 1, 23No
41, 1, 2, 2, 35Yes

Three of the four completed trials required at least four boxes. The simulated relative frequency is:

$$ \text{Simulated relative frequency} = \frac{\text{completed trials needing at least 4 boxes}} {\text{total completed trials}} = \frac{3}{4} = 0.75 $$
Conclude: In these four simulated collections, 3 of 4 required at least four boxes, giving a simulated relative frequency of 0.75. This is an estimate from this run; a different run could produce a different relative frequency.

What If the Stream Ends Before a Trial Is Complete?

A trial is not complete just because the supplied list of random labels ends. If one or more prize types are still missing, continue generating labels from the random source. Do not record the partial count as a completed trial or include it in the denominator when calculating a simulated relative frequency.

Worked Example: Continue an Incomplete Trial

Use labels 1, 2, and 3 for the prizes. A supplied stream contains 1, 2, 1, 3, 2, 2, 1, 2. Simulate as many complete trials as possible, then continue the stream with additional generated labels 3, 3, 1, 2.

Trial 1 begins with labels 1, 2, 1, 3. All three prize types have appeared after the fourth box, so trial 1 is complete with an outcome of 4 boxes. Trial 2 begins with the remaining original labels 2, 2, 1, 2. After these four boxes, prizes 1 and 2 have appeared, but prize 3 has not. The original stream ends here, so trial 2 is incomplete—not a completed outcome of 4 boxes.

Continue generating labels for trial 2. The next label is 3, which completes the collection after a total of 5 boxes. The following label, also 3, starts trial 3 because the collection resets. Trial 3 then uses 3, 1, 2 and completes after 3 boxes. The completed outcomes are therefore 4, 5, and 3 boxes. The incomplete portion was not counted as a completed trial; its labels were continued until the stopping rule was met.

Check: Trial 2 used four labels from the original stream and one additional label, for five boxes total. The next generated label starts trial 3 only after trial 2 is complete. Each new trial gets its own empty collection.

Common Mistakes and AP Exam Tips

  • Stopping after a preset number of boxes. Unless all prize types have appeared, the collection is not complete. State the stopping rule in terms of collecting every prize, not an arbitrary box count.
  • Leaving repeated prizes out of the count. A repeated label does not add a new prize type, but the box was still bought. Count every generated label used before the stopping point.
  • Continuing after the collection is complete. Stop immediately when the final missing prize appears. Later labels belong to later trials.
  • Carrying prizes into the next trial. Each new trial represents a fresh collection process. Reset the list of collected prize types to empty.
  • Counting an unfinished trial as complete. If the stream ends while a prize is missing, continue generating labels. Do not include that partial trial in a relative-frequency calculation.
  • Changing the prize model midway through a run. The label assignment and chance process must stay consistent with the stated assumptions for every box and every trial.
  • Calling the simulated relative frequency an exact probability. Report the event count and total number of completed trials, then describe the quotient as an estimate from that simulation.

For a full-credit simulation description, name the prize labels, explain how each random value represents one box, define a trial as a complete collection process, and state the stopping rule. Also explain that repeats count as boxes, that each new trial starts with no prizes collected, and what you will do if the stream ends before completion. When reporting an estimate, use completed trials in both the event count and the total.

Key takeaway: In a “buy boxes until all prizes are collected” simulation, one trial begins with an empty collection and stops as soon as every prize type has appeared. Count every box, including boxes with repeated prizes. Reset the collection for each new trial, and only use completed trials when estimating a probability.

Check Your Understanding

Use prize labels 1, 2, and 3. Each generated label represents the prize in one box.

  1. A trial produces 2, 1, 2, 3, 1. How many boxes belong to the trial, and which label is not used before the stopping point?
  2. A trial produces 3, 3, 2, 3. Is the collection complete? What prize type is still missing?
  3. Why does a repeated label still count toward the number of boxes bought?
  4. In a set of 20 completed trials, 13 required at least four boxes. Calculate the simulated relative frequency for that event.
  5. A stream ends after two prize types have appeared in the current trial. Should that partial trial be included in the total number of completed trials? Explain what happens next.