Tutorials › AP Statistics › Reading a Simulated Sampling Distribution of x-bar

Sampling distributions for means · Tutorial 615 of 1000

Reading a Simulated Sampling Distribution of x-bar

Practice translating a dotplot or histogram of simulated sample means into estimates of center, spread, and probability.

Intermediate 9 min read

What You'll Learn

  • Identify what one dot or histogram count represents in a simulation of sample means.
  • Estimate the center of a simulated sampling distribution using its pattern or a frequency-weighted mean.
  • Describe spread using the range and, when possible, a calculated standard deviation.
  • Estimate probabilities by finding the fraction of simulated sample means in a specified region.
  • Explain how the number of simulation repetitions and histogram bins affect an estimate.
  • Compare simulated sampling distributions for different sample sizes.

What Each Dot Represents

In “Sampling Distribution of a Mean From a Skewed Population,” you saw how repeated samples produce a distribution of sample means. A dotplot or histogram of those simulated means lets you inspect that distribution directly. The graph is not a plot of individual observations: each dot or recorded value represents the mean from one simulated sample of a fixed size.

Suppose a simulation repeatedly samples \(n=12\) observations from a population and calculates \(\bar{x}\) each time. If the simulation is repeated 500 times, the resulting graph contains 500 sample means. Its horizontal axis gives possible values of \(\bar{x}\); its vertical axis shows how often, or how densely, those means occurred.

Definition: A simulated sampling distribution of \(\bar{x}\) is the distribution of the sample means recorded across repeated simulated samples of the same size and sampling method. It approximates the sampling distribution of \(\bar{x}\), which is the probability distribution of the sample means from all possible samples under that method.

The graph can help you estimate three things. Its center is the typical or balancing value of the simulated means. Its spread describes how far those means tend to fall from the center. A probability can be estimated by finding the fraction of simulated means in a region of interest.

As in “Mean of the Sampling Distribution of x-bar” and “Standard Deviation of the Sample Mean,” the theoretical sampling distribution is centered at \(\mu\) and, for independent observations, has standard deviation \(\sigma/\sqrt{n}\). A simulation gives an approximation from a finite number of repetitions, so its observed center and spread will not necessarily equal those theoretical values exactly.

Reading Center and Spread

Start by looking at the overall pattern. For a roughly symmetric dotplot, the center is near the balance point, with roughly similar amounts of the distribution on either side. For a skewed graph, the balance point and the location of the highest stack may differ. A visual estimate should describe the scale and units of the horizontal axis.

If a graph gives the number of simulated means at each exact value, you can calculate its frequency-weighted mean. Multiply each value by its count, add those products, and divide by the total number of simulated means. This gives the exact mean of the values recorded in that simulation. It is an estimate of the theoretical center, not necessarily the theoretical center itself.

$$ \text{Simulated mean of the sample means} =\frac{\sum(\text{sample-mean value})(\text{frequency})}{\text{number of repetitions}}. $$

Spread can first be described visually: report where most means fall, note the approximate range, and look for unusually distant values. If exact values and frequencies are available, you can also calculate the standard deviation of the simulated means. Treating the simulation results as an empirical distribution, divide the weighted sum of squared deviations from their mean by the number of repetitions, then take the square root. That describes the spread in the simulation; it estimates the theoretical standard deviation \(\sigma_{\bar{x}}\).

Be careful when a histogram groups values into intervals. A bar’s height tells how many simulated means fall in a bin, but not their exact positions inside it. You can estimate the center or standard deviation using bin midpoints, but those calculations are approximations. State when your result is approximate, especially if the bins are wide.

Estimating Probabilities from Relative Frequencies

To estimate a probability, match the event to a region on the graph and count the simulated means in that region. Divide that count by the total number of repetitions. For instance, the fraction of simulated means above 30 is an estimate of \(P(\bar{x}>30)\) for the stated population, sample size, and sampling method.

Formula: The estimated probability of an event involving \(\bar{x}\) is the number of simulated sample means satisfying the event divided by the total number of simulated samples.

A dotplot with 200 simulated means can only give relative frequencies in increments of \(1/200=0.005\). With more repetitions, estimates often become more stable, but they are still subject to simulation variability. Also check the event’s inequality carefully: “greater than 26” excludes values equal to 26, while “at least 26” includes them. For histogram bins, decide how to treat values at a boundary using the bin definitions or the information given in the question.

These are estimated probabilities for sample means—not probabilities for individual observations. Keep the statistic and sample size in view. A graph for \(n=12\) does not directly estimate the probability for \(n=30\); those sample means come from a different sampling distribution.

Worked Examples: Interpreting Simulated Means

Worked Example: Estimate the Center and Spread from a Dotplot

A simulation records 100 sample means for samples of 10 bottles. Each \(\bar{x}\) is the mean fill volume in milliliters. The dotplot’s exact values and frequencies are shown below. Estimate the center and describe the spread; then calculate the standard deviation of these 100 recorded means.

Simulated \(\bar{x}\) (mL)Frequency
184
2012
2222
2428
2620
2810
304

State. The graph summarizes 100 simulated sample means, each based on \(n=10\) bottles. We want the center and spread of those simulated means, not the spread of individual bottle fill volumes.

Plan. Use the frequency-weighted mean for the simulated center. Then calculate the empirical standard deviation from the weighted squared deviations. The frequencies sum to \(4+12+22+28+20+10+4=100\).

Do. The weighted mean is

$$ \bar{x}_{\text{sim}} =\frac{18(4)+20(12)+22(22)+24(28)+26(20)+28(10)+30(4)}{100} =\frac{2388}{100}=23.88\text{ mL}. $$

This is close to 24 mL. The values range from 18 to 30 mL, and most of the simulated means lie between 20 and 28 mL. To calculate the standard deviation, first find the weighted mean of the squared values:

$$ \frac{18^2(4)+20^2(12)+22^2(22)+24^2(28)+26^2(20)+28^2(10)+30^2(4)}{100} =\frac{57832}{100}=578.32. $$

Subtract the square of the simulated mean to get the variance, then take the square root:

$$ s_{\text{sim}}^2=578.32-(23.88)^2=8.0656\text{ mL}^2, \qquad s_{\text{sim}}=\sqrt{8.0656}\approx2.840\text{ mL}. $$

As a check on the variance, the weighted sum of squared deviations from 23.88 is \(806.56\); dividing by 100 gives \(8.0656\), the same result.

Conclude. The simulated means are centered at about 23.88 mL, with a standard deviation of about 2.840 mL and a range of 18 to 30 mL. These describe this set of 100 simulation results. The mean and standard deviation approximate the center and standard deviation of the theoretical sampling distribution, but the simulation may not match them exactly.

Worked Example: Estimate a Probability from a Dotplot

A simulation of 200 sample means for \(n=8\) records the values in the table. Use the relative frequencies to estimate \(P(\bar{x}\geq26)\) and \(P(20\leq\bar{x}<26)\). The sample means are in minutes.

Simulated \(\bar{x}\) (minutes)Frequency
158
1822
2040
2255
2442
2621
289
303

State. Each of the 200 entries is a mean from a simulated sample of eight observations. We will estimate each probability by the proportion of recorded means in the event.

Plan. For \(\bar{x}\geq26\), include 26 and all larger listed values. For \(20\leq\bar{x}<26\), include 20, 22, and 24, but exclude 26. Divide each event count by 200.

Do. For the first event, the count is \(21+9+3=33\), so

$$ P(\bar{x}\geq26)\approx\frac{33}{200}=0.165. $$

For the second event, the count is \(40+55+42=137\), so

$$ P(20\leq\bar{x}<26)\approx\frac{137}{200}=0.685. $$

Conclude. The simulation estimates that about 16.5% of sample means are at least 26 minutes, and about 68.5% are at least 20 but less than 26 minutes. These estimates apply to the simulated sampling setup with \(n=8\); they are not exact probabilities and do not describe individual observations.

Worked Example: Compare Spread for Two Sample Sizes

Two simulations each produce 100 sample means of a water-use measurement, in liters. The first uses \(n=10\); the second uses \(n=40\). The table gives the values and frequencies. Compare their centers and spreads, and use the simulations to estimate the chance that the sample mean is below 16 liters.

\(\bar{x}\) (L)Frequency for \(n=10\)Frequency for \(n=40\)
12 or 155 at 125 at 15
14 or 1610 at 1410 at 16
16 or 1720 at 1620 at 17
183030
20 or 1920 at 2020 at 19
22 or 2010 at 2210 at 20
24 or 215 at 245 at 21

State. Each column represents 100 simulated sample means from its own fixed sample size. We compare the empirical centers and standard deviations, and then count the means below 16 L.

Plan. Both frequency patterns are symmetric around 18 L, so each simulated mean is 18 L. Calculate spread from squared distances from 18, weighted by frequency, and divide by 100 before taking the square root.

Do. For \(n=10\), the weighted squared-distance average is

$$ \frac{10(6^2)+20(4^2)+40(2^2)+30(0^2)}{100} =\frac{840}{100}=8.4, \qquad s_{\text{sim}}=\sqrt{8.4}\approx2.898\text{ L}. $$

For \(n=40\), the corresponding calculation is

$$ \frac{10(3^2)+20(2^2)+40(1^2)+30(0^2)}{100} =\frac{210}{100}=2.1, \qquad s_{\text{sim}}=\sqrt{2.1}\approx1.449\text{ L}. $$

As a check, the \(n=10\) distribution has squared distances 36, 16, and 4 occurring 10, 20, and 40 times: \(360+320+160=840\). For \(n=40\), the corresponding total is \(90+80+40=210\). The chance below 16 L is \(5+10=15\) of 100 for \(n=10\), or \(0.15\); for \(n=40\), it is 5 of 100, or \(0.05\).

Conclude. Both simulated distributions are centered at 18 L, but the \(n=40\) means have less spread: their simulated standard deviation is about 1.449 L, compared with 2.898 L for \(n=10\). In these simulations, fewer \(n=40\) means fall below 16 L. This comparison is consistent with the earlier result that increasing sample size reduces the variability of \(\bar{x}\); the simulated proportions are estimates and could differ in another run.

Common Mistakes and AP Exam Tip

  • Reading a dot as an individual observation: Each dot is one simulated sample mean. Name \(\bar{x}\) and its sample size when describing the graph.
  • Calling the simulated center the population mean without qualification: The simulated mean is an estimate based on the recorded repetitions. Say “the simulated means are centered near…” unless the theoretical population mean is separately given.
  • Using a count instead of a relative frequency for probability: A count of 33 does not mean a probability of 33. Divide by the total of 200 to get 0.165.
  • Mixing up strict and inclusive boundaries: “At least 26” includes 26; “greater than 26” does not. Identify which plotted values belong before adding frequencies.
  • Confusing spread of sample means with spread of individual data: The graph’s horizontal spread is the variability of \(\bar{x}\), not the variability of \(X\). Include units for the statistic shown.
  • Reporting false precision from a histogram: If values are binned, their exact positions are unknown. Use midpoint-based results only as approximations and avoid reporting more precision than the graph supports.

For full-credit communication, state what one graph value represents, identify the number of repetitions, show the relevant frequencies and denominator for a probability estimate, and interpret the result in context. When comparing sample sizes, distinguish a pattern in the simulations from a guaranteed outcome.

Key takeaway: Read a simulated sampling distribution as a collection of sample means. Estimate its center and spread from the horizontal pattern, and estimate a probability by dividing the number of simulated means in the event by the total number of repetitions.

Check Your Understanding

Use the ideas in this tutorial to answer the questions.

  1. A simulation contains 400 dots, each representing a sample mean for \(n=15\). What does one dot represent, and what does the total of 400 represent?
  2. In a simulation of 250 sample means, 47 are greater than or equal to 32 minutes. Estimate the probability of \(\bar{x}\geq32\) minutes.
  3. A histogram’s bars group sample means into intervals. Why is a center calculated from the bin midpoints only an approximation?
  4. Two simulated sampling distributions have centers near 50 seconds, but one is visibly narrower. What does that tell you about the variability of the simulated sample means?
  5. Explain why a simulated probability from 100 repetitions may vary from one simulation run to another.