One Hat, Two Different Sampling Rules
In Simulating Events with Equal Probabilities and Simulating Events with Unequal Probabilities, each random value represented an outcome in a trial. When a trial involves drawing more than one name from a hat, the next question is whether a name is returned to the hat before the next draw. That choice changes how repeated labels should be handled.
Suppose a hat contains four names: Ana, Ben, Cora, and Diego. Label them 1, 2, 3, and 4, respectively. A generated label represents the name with that label. For this tutorial, one trial is drawing two names. We will use the event “Ana is selected at least once” to compare the two methods.
Both methods use the same labels. The difference is not a new assignment of labels; it is the rule for what to do when a label appears again during a trial. Decide on that rule before generating outcomes, and apply it consistently.
Set Up the Two Simulations
Use randInt(1,4,n) to generate \(n\) random labels, with each label equally likely. For drawing with replacement, use exactly two generated labels in each trial. For drawing without replacement, accept the first label, then accept the next label that is different from it. If a label repeats within that trial, skip it and continue generating labels until two distinct names have been selected.
A skipped label is still part of the random stream, but it is not an additional name selected in that trial. Start a fresh trial with no names yet selected. In particular, a label used in an earlier trial is not automatically skipped in a later trial: each trial represents a new repetition of the chance process.
Use 1 for Ana, 2 for Ben, 3 for Cora, and 4 for Diego. One trial is two draws.
With replacement, return each name to the hat. Without replacement, do not return a selected name.
Keep every label with replacement. Without replacement, skip a repeat within the current trial and generate another label.
Record two draws with replacement, or two distinct selected names without replacement.
For the event “Ana is selected at least once,” count a completed trial if it includes label 1.
How Repeated Labels Affect a Trial
Consider the generated labels 2, 2, 4. With replacement, the first two labels are the two draws: Ben is selected twice. The third label belongs to a later part of the random stream and is not needed for this trial. Without replacement, the first 2 selects Ben, but the second 2 repeats a name already selected in this trial. Skip that repeated label; then 4 selects Diego, completing the trial with Ben and Diego.
This difference follows the meaning of the two sampling rules. With replacement, selecting Ben does not prevent Ben from being selected again. Without replacement, a selected name is no longer available during that trial. The same generated labels can therefore produce different completed selections under the two rules.
Worked Example: Interpret One Stream Under Both Rules
A random-number generator produces the stream 2, 2, 4, 1. Use it to complete one trial of two draws under each rule. The labels are 1 for Ana, 2 for Ben, 3 for Cora, and 4 for Diego.
With replacement: Use the first two labels, 2 and 2. The selections are Ben and Ben. Both draws count, even though the same name appears twice. The event “Ana is selected at least once” does not occur.
Without replacement: The first 2 selects Ben. Skip the next 2 because Ben has already been selected in this trial. The next label, 4, selects Diego, so the completed trial is Ben and Diego. The event does not occur. The final label, 1, is not used because the trial is already complete.
Complete Several Trials Without Replacement
For a simulation without replacement, determine where each trial ends by counting two accepted, distinct labels—not by taking a fixed number of generated labels. A trial may use more than two generated labels if repeats occur. When a repeat is skipped, continue from the next label in the stream; do not restart the stream or accidentally count the repeated name as a second selection.
Worked Example: Six Trials Without Replacement
Use the following stream to simulate six trials, each consisting of two names drawn without replacement. Estimate the relative frequency of trials in which Ana is selected at least once.
Read from left to right and reset the set of selected names at the start of each trial. In trial 1, labels 2 and 4 select Ben and Diego. In trial 2, label 1 selects Ana; the next 1 repeats Ana, so skip it, and label 3 completes the trial with Cora. Trial 3 uses 4 and 2, selecting Diego and Ben.
Trial 4 begins with 3, selecting Cora. The next 3 repeats Cora, so skip it; label 1 completes the trial with Ana. Trial 5 uses 2 and 4, selecting Ben and Diego. Trial 6 uses 1 and 4, selecting Ana and Diego. The complete tally is:
| Trial | Generated labels used | Selected names | Ana selected? |
|---|---|---|---|
| 1 | 2, 4 | Ben, Diego | No |
| 2 | 1, 1, 3 | Ana, Cora | Yes |
| 3 | 4, 2 | Diego, Ben | No |
| 4 | 3, 3, 1 | Cora, Ana | Yes |
| 5 | 2, 4 | Ben, Diego | No |
| 6 | 1, 4 | Ana, Diego | Yes |
There are 3 trials with Ana selected among 6 completed trials. Therefore:
The stream contains 14 labels because two trials each had a repeated label to skip. All six trials are complete: the accepted labels are 2 and 4 for trial 1; 1 and 3 for trial 2; 4 and 2 for trial 3; 3 and 1 for trial 4; 2 and 4 for trial 5; and 1 and 4 for trial 6. If a supplied stream ended before a trial had two distinct labels, it would be incomplete; the simulation would need to continue with additional random labels.
Complete Several Trials With Replacement
With replacement, every draw is made from all four names. Each trial uses exactly two generated labels, and a repeated label is not skipped. A name appearing twice is a possible outcome of the model, not an error in the simulation.
Worked Example: Six Trials With Replacement
Use the following 12 labels to simulate six trials of two draws with replacement. Estimate the relative frequency of trials in which Ana is selected at least once.
Divide the labels into consecutive pairs because each trial uses exactly two draws. Trial 1 is 2, 4 (Ben and Diego); trial 2 is 1, 1 (Ana and Ana); trial 3 is 4, 2 (Diego and Ben); trial 4 is 3, 1 (Cora and Ana); trial 5 is 2, 2 (Ben and Ben); and trial 6 is 1, 4 (Ana and Diego).
Ana appears at least once in trials 2, 4, and 6. The simulated relative frequency is:
The two examples happened to produce the same relative frequency, but that does not make the sampling rules equivalent. A different random run could give different counts. The rules also allow different outcomes: two identical names can be selected in a with-replacement trial, but not in a without-replacement trial.
Why the Model Probabilities Differ
For these four names and two draws, the model probabilities of selecting Ana at least once differ between the two rules. This helps explain what the simulations are estimating. It does not mean a short simulated relative frequency must match the model probability exactly.
Without replacement, there are 12 equally likely ordered pairs of distinct names: 4 choices for the first draw and then 3 remaining choices for the second. Six of those pairs include Ana: Ana can be first with any of 3 other names second, or second with any of 3 other names first. Thus the model probability is \(6/12=0.50\).
With replacement, there are 16 equally likely ordered pairs because each of the two draws has 4 possible names. Seven pairs include Ana: 4 start with Ana, and 3 more end with Ana but do not start with Ana. The pair Ana, Ana is counted only once. Thus the model probability is \(7/16=0.4375\). These counts are for the specific event “Ana is selected at least once”; another event can have a different probability under each rule.
A simulation’s relative frequency can be above or below its model probability, especially with only a few trials. As in The Law of Large Numbers in Simulations, across many repeated trials the relative frequency tends to settle near the probability for the model being simulated. Keep the replacement rule fixed throughout a simulation; otherwise the trials do not all represent the same chance process.
Common Mistakes and AP Exam Tips
- Skipping repeats in both methods. Skip a repeated label only when simulating without replacement. With replacement, keep the repeat as a selected name.
- Counting a skipped label as a selection. In a without-replacement trial, a repeat does not fill the second selection. Continue generating labels until the trial has two distinct names.
- Stopping after two generated labels without replacement. Two generated labels may be the same, so they may not complete a two-name trial. Define completion by the number of accepted, distinct names.
- Carrying selected names into the next trial. The selected-name list resets at the start of each new trial. A name chosen in one trial can be selected again in a later trial.
- Using a fixed number of random values for a whole without-replacement run. Repeats can require extra values. Continue the stream until every planned trial is complete, and make sure the supplied stream has enough values to do so.
- Counting duplicate names as two different people. With replacement, Ana and Ana are two draws but only one distinct person. Be precise about whether the event concerns a name appearing at least once, a particular draw, or two different people.
- Calling the simulated relative frequency the exact model probability. State the number of event trials and the total number of completed trials, then describe the quotient as an estimate from that run.
For a clear AP-style description, name the trial, identify the labels, state whether names return to the hat, and explain exactly what happens when a repeated label appears. When simulating without replacement, note that repeats are skipped and that random labels are generated until each trial is complete. Report the event count and divide by the number of completed trials.
Check Your Understanding
Use labels 1 through 4 for four names, and suppose one trial consists of drawing two names.
- A without-replacement trial begins with labels 3, 3, 1. Which two names are selected, and how many generated labels were used?
- A with-replacement trial begins with labels 3, 3. What are the two selections, and should either label be skipped?
- In a without-replacement simulation, a stream ends after only one distinct name has been selected for the current trial. Is that trial complete? What should happen next?
- Across 12 completed trials, Ana appears at least once in 5 trials. Calculate and interpret the simulated relative frequency.
- Why can two simulations of the same event produce different relative frequencies, even when they use the same replacement rule?