Tutorials › AP Statistics › Simulating Sample Proportions with Repeated Samples

Sampling distributions for proportions · Tutorial 403 of 1000

Simulating Sample Proportions with Repeated Samples

Practice reading the center, spread, and shape of 100 simulated sample proportions, and learn what a dot plot can—and cannot—tell you.

Intermediate 9 min read

What You'll Learn

  • Identify what each dot represents in a simulated sampling distribution of p-hat.
  • Describe a dot plot’s center using its cluster, median, and most common value.
  • Summarize spread with the range and interquartile range.
  • Recognize unimodal, roughly balanced patterns and notice possible skew or unusual values.
  • Explain how sample size determines the possible steps between sample proportions.
  • Interpret a simulated tail percentage as a result from that run, not a guarantee.

Reading a Dot Plot of Repeated Sample Proportions

In What a Sampling Distribution of p-hat Is, you learned that each random sample produces its own sample proportion, \(\hat{p}\). A simulation makes many samples of the same size under a stated model and records every resulting \(\hat{p}\). A dot plot of those results lets you inspect how the sample proportions cluster and vary.

The key is to describe the simulated statistics, not the individual people or items in the samples. If a dot represents one simulated sample, a pile of dots at a value means that value occurred often in that particular simulation. The plot gives evidence about how repeated samples behaved under the simulation’s model; a different simulation run may look somewhat different.

Definition: A simulated sampling distribution dot plot displays the sample-proportion values calculated from repeated simulated samples. Each dot represents one simulated sample and its \(\hat{p}\).

We will examine 100 simulated samples, each of size \(n=25\), from a population model with \(p=0.40\). Success means an individual has a specified characteristic. The simulation records the number of successes in each sample and divides by 25 to obtain \(\hat{p}\). Thus, every dot is one of 100 sample proportions.

A Dot Plot of 100 Simulated Values

The table below is a compact dot plot. Each black dot represents one simulated sample; dots are placed in a row for readability. For example, the seven dots at \(\hat{p}=0.28\) mean that seven of the 100 simulated samples had that sample proportion.

Successes \(X\)Sample proportion \(\hat{p}=X/25\)Dots in the simulationFrequency
40.16●1
50.20● ●2
60.24● ● ● ●4
70.28● ● ● ● ● ● ●7
80.32● ● ● ● ● ● ● ● ● ●10
90.36● ● ● ● ● ● ● ● ● ● ● ● ●13
100.40● ● ● ● ● ● ● ● ● ● ● ● ● ● ●15
110.44● ● ● ● ● ● ● ● ● ● ● ● ● ●14
120.48● ● ● ● ● ● ● ● ● ● ● ●12
130.52● ● ● ● ● ● ● ● ●9
140.56● ● ● ● ● ●6
150.60● ● ● ●4
160.64● ●2
170.68●1

Before describing the plot, check that it includes all 100 simulated samples:

$$ 1+2+4+7+10+13+15+14+12+9+6+4+2+1=100 $$

That check matters: frequencies should account for every simulated repetition. If they do not add to 100, the display is incomplete or a count has been copied incorrectly.

Describe the Center

The center is the location around which the sample proportions cluster. For a quick visual description, look for the busiest region and the middle of the plotted values. A numerical summary can make the description more precise. The mode is the value that occurs most often; the median is the middle value when the observations are ordered.

In this plot, the most common value is \(\hat{p}=0.40\), with 15 dots. The 50th and 51st ordered values are both 0.40: the cumulative count is 37 through 0.36 and reaches 52 at 0.40. Therefore, the median is 0.40 as well. Both summaries place the center of this simulated distribution near 0.40.

As established in Normal Models for Sample Means and Proportions, the sampling distribution of \(\hat{p}\) is centered at the population proportion \(p\) when the stated sampling conditions apply. Here the simulated model uses \(p=0.40\), and the plot’s center is close to that value. Do not expect every simulation’s median, mode, or average to equal \(p\) exactly.

Reading the center: Describe where most dots cluster, then use a statistic such as the median or mode if a more precise summary is useful. Compare the simulated center with \(p\) cautiously; simulation results vary from run to run.

Describe the Spread

Spread describes how far apart the simulated sample proportions are. Two useful summaries are the range, the largest value minus the smallest, and the interquartile range (IQR), the third quartile minus the first quartile. The range describes the full extent of the observed simulation, while the IQR describes the width of the middle half.

The smallest plotted value is 0.16 and the largest is 0.68, so the range is \(0.68-0.16=0.52\). To find the quartiles from the ordered 100 values, use the middle of the lower 50 values for \(Q_1\) and the middle of the upper 50 for \(Q_3\). Positions 25 and 26 both have value 0.36, so \(Q_1=0.36\). Positions 75 and 76 both have value 0.48, so \(Q_3=0.48\). Therefore:

$$ \text{IQR}=Q_3-Q_1=0.48-0.36=0.12 $$

In this run, the middle half of simulated sample proportions spans 0.12, or 12 percentage points. The overall range is wider because it includes the most extreme values in this particular set of 100 results. Neither the range nor the IQR is a guarantee about what every future sample proportion will be.

Describe the Shape

To describe shape, look for the number of peaks, the balance of the tails, and any gaps or unusually distant values. A distribution is unimodal when it has one main peak. It is roughly symmetric when the values and frequencies on either side of its center are broadly balanced. A longer tail toward larger values suggests right skew; a longer tail toward smaller values suggests left skew.

The dot plot has one main peak around 0.40–0.44. Frequencies generally taper away from that region in both directions. The high-value side extends slightly farther, so one reasonable description is: unimodal and roughly mound-shaped, with a slight extension toward larger values. That description acknowledges the overall pattern without claiming that the two sides match perfectly.

A dot plot can look roughly mound-shaped without proving that a normal model is appropriate. This is a display of 100 simulated values, and random variation can create small bumps or imbalances. Describe the shape you can see; do not label it normal merely because it has a central cluster and tails.

Shape checklist: Identify the number of peaks, compare the tails on either side of the center, and note clear gaps or isolated values. Use cautious language such as “roughly symmetric” when the pattern is not exact.

Worked Example: Summarize the 100-Sample Dot Plot

Worked Example: Summarize the 100-Sample Dot Plot

A student is asked to describe the center, spread, and shape of the simulated sample proportions in the table. The model used \(p=0.40\), and every simulated sample had \(n=25\).

Center: The most frequent value is 0.40, appearing 15 times. The 50th and 51st ordered values are also 0.40, so the median is 0.40. The sample proportions are centered near 0.40, close to the model’s population proportion.

Spread: The observed range is:

$$ 0.68-0.16=0.52 $$

The first and third quartiles are 0.36 and 0.48, respectively, giving:

$$ \text{IQR}=0.48-0.36=0.12 $$

Shape: There is one main peak around 0.40–0.44, and frequencies generally decrease toward both ends. The distribution is roughly mound-shaped, with a slight extension toward larger proportions.

A complete contextual summary is: “For the 100 simulated random samples of 25 individuals under the model \(p=0.40\), the sample proportions are centered near 0.40, with a range from 0.16 to 0.68 and an IQR of 0.12. Their distribution is unimodal and roughly mound-shaped, with a slight extension toward larger values.” This describes the simulation results without claiming that every possible sample must fall in the observed range.

Worked Example: Translate a Dot Plot into Sample Counts

Worked Example: Translate a Dot Plot into Sample Counts

For the same simulation, find the proportion of simulated samples with \(\hat{p}\geq0.52\). Then explain what that proportion means.

Count the dots meeting the rule. The values 0.52, 0.56, 0.60, 0.64, and 0.68 have frequencies 9, 6, 4, 2, and 1. Their total is:

$$ 9+6+4+2+1=22 $$

Divide by the number of simulated samples. There are 100 repetitions:

$$ \frac{22}{100}=0.22=22\% $$

Thus, 22% of the simulated samples in this run had a sample proportion of at least 0.52. This is a simulation proportion for the stated model and event. It is not a claim that exactly 22% of all possible samples must meet the rule, or that the population proportion is 0.52.

Worked Example: Read the Steps Between Possible Values

Worked Example: Read the Steps Between Possible Values

The plot’s sample proportions include 0.36 and 0.40 but no values between them. Explain why, and find the change in \(\hat{p}\) caused by one additional success in a sample of 25.

For a sample of size 25, the sample proportion is the success count divided by 25. Adding one success changes the proportion by:

$$ \frac{1}{25}=0.04 $$

For instance, nine successes give \(9/25=0.36\), while ten successes give \(10/25=0.40\). The gap between those possible values is \(0.40-0.36=0.04\), exactly the change from one additional success. Therefore the dot plot has discrete steps of 0.04; it is not expected to contain every decimal value between 0 and 1.

This spacing comes from the sample size, not from a plotting mistake. With a larger sample size, one additional success would change \(\hat{p}\) by a smaller amount. In this plot, each individual dot still represents a whole simulated sample, even when several samples have the same proportion and appear together.

Common Mistakes and AP Exam Communication

A strong description connects the visual pattern to the simulated statistic and gives specific evidence from the plot. Avoid descriptions that merely say “the graph is normal” or “most values are average.” Explain what is being summarized and how the display supports your statement.

  • Confusing dots with individuals. Each dot represents one simulated sample proportion, not one person or item. Say “22 simulated samples had \(\hat{p}\geq0.52\),” not “22 individuals had the characteristic.”
  • Calling the simulation center the exact population proportion. The plot’s center is a summary of one run. The model specifies \(p=0.40\); the sample proportions vary, and summaries from a run need not equal 0.40 exactly.
  • Reporting only the range as typical spread. The range depends on the two most extreme values in these 100 results. The IQR describes the middle half and can give a complementary view.
  • Calling a mound-shaped plot normal. A visual resemblance alone does not establish a normal model. Describe the observed shape instead, and avoid overstating what 100 simulated outcomes establish.
  • Forgetting the sample size when reading gaps. With \(n=25\), possible \(\hat{p}\) values step by 0.04. Values between those steps cannot occur for this sample size.
  • Treating a simulated percentage as a guarantee. A result such as 22% describes this set of 100 repetitions. Another simulation run can produce a different percentage.
AP Exam Tip: Name the statistic represented by each dot, then describe the center, spread, and shape with values from the plot. Use wording such as “in these 100 simulated samples” when reporting a frequency or percentage.

Key Takeaway

A dot plot of repeated sample proportions summarizes how \(\hat{p}\) varied across simulated samples of the same size. Read its center from the main cluster, its spread from measures such as the range and IQR, and its shape from the peaks and tails. Keep the description tied to the simulation rather than treating its results as a guarantee.

Key takeaway: Each dot is one simulated \(\hat{p}\). Describe where the dots cluster, how far they spread, and the shape they form; then state that the summary describes this simulation under its model.

Check Your Understanding

Use the 100-sample dot plot in this tutorial unless a question gives another setup.

  1. How many simulated samples produced \(\hat{p}=0.44\), and what does each dot at that value represent?
  2. Find the observed range of the simulated sample proportions. What does this range describe?
  3. The 50th and 51st ordered values are both 0.40. What is the median, and how does it compare with the mode?
  4. Describe the plot’s shape using its peak and the behavior of its tails.
  5. Why are there no simulated sample proportions between 0.36 and 0.40 when \(n=25\)?
  6. What does the 22% result for \(\hat{p}\geq0.52\) describe, and why is it not a guarantee about all possible samples?