What Does “95% Confident” Mean Over Repeated Samples?
In “Interpreting a Confidence Interval for a Mean,” you learned how to describe the endpoints of a one-sample t interval in context. The confidence level explains something different: how the interval-making method performs when it is used repeatedly. A 95% confidence level is a long-run capture rate, not a statement that 95% of individual observations fall inside an interval.
Imagine repeatedly taking random samples of the same size from the same population and calculating a 95% one-sample t interval from each sample. The population mean \(\mu\) is fixed. The sample changes from repetition to repetition, so the sample mean, sample standard deviation, and interval endpoints can change. Each resulting interval either contains \(\mu\) or misses it.
The word “about” matters. In a finite simulation, the proportion of intervals that capture \(\mu\) will vary. Even if a procedure has a 95% long-run capture rate, one set of 20 intervals might include only 18 captures, or all 20. The confidence level describes the procedure’s long-run performance, not a guarantee about the exact results of a particular batch of repetitions.
This is the repeated-sampling meaning behind saying, “We are 95% confident.” Once one interval has been calculated, \(\mu\) is still a fixed population value, and that particular interval either contains it or does not. The 95% refers to the success rate of the method across repeated samples—not a 95% probability that the fixed \(\mu\) is in this one calculated interval.
A Simulation Picture: Fixed Mean, Moving Intervals
A simulation makes the long-run idea visible. Suppose a computer repeatedly creates random samples from a population with a known mean of 50 minutes. For each sample, it calculates a 95% one-sample t interval for the population mean. Picture the intervals as horizontal line segments and the true mean, 50 minutes, as a vertical reference line.
An interval that crosses the reference line captures the population mean. An interval entirely to the left or right misses it. The reference line does not move; the interval segments do. Here is a small, fictional set of 12 intervals from such a simulation:
| Simulation run | 95% interval (minutes) | Does it contain \(\mu=50\)? |
|---|---|---|
| 1 | (48, 55) | Yes |
| 2 | (45, 52) | Yes |
| 3 | (51, 57) | No |
| 4 | (42, 49) | No |
| 5 | (46, 53) | Yes |
| 6 | (49, 54) | Yes |
| 7 | (44, 51) | Yes |
| 8 | (52, 58) | No |
| 9 | (47, 50) | Yes |
| 10 | (43, 48) | No |
| 11 | (49.5, 55) | Yes |
| 12 | (46, 52) | Yes |
In this small set, 8 of the 12 intervals contain 50. The observed capture proportion is
or about 66.67%. That result does not contradict the 95% confidence level. Twelve repetitions are far too few to expect the observed proportion to be exactly, or necessarily very close to, 95%. If many more repetitions are run under the same conditions, the overall capture proportion should tend to settle near the method’s long-run capture rate.
The picture also shows why “95% of the intervals contain \(\mu\)” does not mean “95% of the measurements are between the endpoints.” Each line segment is an interval estimating the population mean. It is not a display of individual measurements.
How the Repeated-Sampling Simulation Works
A simulation needs a population model or a set of population values, a sample size, and a procedure to repeat. For example, a computer could generate many independent samples of size \(n=16\) from a Normal population with mean 50 minutes. For each sample, it would calculate \(\bar{x}\) and \(s\), then use the one-sample t interval formula with \(df=15\) and the 95% critical value. The simulation would record whether each interval contains 50.
Keep the population, sample size, sampling design, and confidence level the same for every repetition.
Use the specified random sampling process. Different samples generally produce different sample means and standard deviations.
Calculate a 95% interval from each sample and mark whether it contains the fixed population mean.
Divide the number of intervals that contain \(\mu\) by the total number of intervals. With many repetitions, this proportion should be near 0.95 if the procedure’s conditions are met.
The simulation must represent a situation where the interval procedure is appropriate. As discussed in “Building a One-Sample t Interval by Hand,” check that the data come from a random sample or suitable randomized process, that observations are independent (including the 10% condition when sampling without replacement), and that the population shape or sample size supports a t procedure. A t interval has its exact stated coverage when the observations are independent and Normally distributed; when sampling without replacement from a finite population, the 10% condition supports treating observations as independent, so coverage is generally treated as approximate.
The simulation’s capture proportion is an estimate of the procedure’s long-run performance. It is not a new confidence interval for \(\mu\), and it is not a probability assigned to \(\mu\) after seeing one sample.
Worked Examples
Worked Example: Reading a Small Simulation
A fictional simulation uses repeated random samples and a 95% t interval to estimate the mean time, in minutes, for a certain task. The true population mean used by the simulation is 50 minutes. In 12 runs, 8 intervals contain 50 and 4 do not.
State. The fixed parameter is \(\mu=50\) minutes. Each run produces one interval, and the question is what the simulation says about capture.
Plan. Count the intervals that contain the fixed mean and divide by the total number of intervals. This gives the observed capture proportion for these 12 repetitions.
Do. There are 8 captures among 12 intervals, so
The simulation’s observed capture rate is about 66.67%.
Conclude. This small set of runs captured the true mean in 8 of 12 intervals. It does not show that the procedure has only a 66.67% confidence level. With just 12 repetitions, the observed rate can differ substantially from the 95% long-run capture rate.
In the interval picture, the four intervals that miss 50 lie entirely on one side of the vertical reference line. The other eight cross it or have 50 as an endpoint, so they count as captures.
Worked Example: Interpreting a Larger Simulation
A fictional computer simulation repeats a suitable sampling process 1,000 times. It calculates a 95% t interval for the population mean from each sample. In this run of the simulation, 948 intervals contain the fixed true mean and 52 miss it.
State. The simulation estimates the long-run capture performance of this 95% interval procedure for the stated population and sampling setup.
Plan. Calculate the proportion of simulated intervals that contain the true mean, then compare that observed proportion with the procedure’s nominal 95% confidence level.
Do. The capture proportion is
Thus, 94.8% of the intervals in this particular simulation captured the mean.
Conclude. The observed simulation result, 94.8%, is close to the 95% long-run capture rate. It need not equal 95.0% exactly: a simulation has its own random variation. The result illustrates the confidence level as a property of the repeated interval procedure, not as a guarantee that exactly 950 of these 1,000 intervals must capture the mean.
Worked Example: Explaining One Calculated Interval
A random sample of 20 trail sections is used to calculate a 95% confidence interval of \((3.2,\ 4.6)\) kilometers for the true mean length of all trail sections in a defined park system. Assume the conditions for a one-sample t interval are met.
State. The parameter is the true population mean trail-section length, \(\mu\), measured in kilometers. The calculated interval is 3.2 to 4.6 kilometers.
Plan. Give the standard contextual interpretation of this particular 95% interval, then explain what the confidence level would mean if the sampling and interval procedure were repeated many times.
Do. The interval interpretation is: “We are 95% confident that the true mean length of all trail sections in this park system is between 3.2 and 4.6 kilometers.” In a repeated-sampling picture, the population mean stays fixed, while each new random sample produces a potentially different interval.
Conclude. The 95% level means that the interval procedure would capture the true mean in about 95% of repeated samples in the long run, under the stated conditions. It does not mean that there is a 95% probability that this already-calculated interval contains the fixed true mean.
Common Mistakes and AP Exam Tips
- Saying the mean has a 95% chance of being in this particular interval. The population mean is fixed after the sample is taken; the interval either captures it or misses it. Full-credit wording connects 95% to the long-run success rate of the method.
- Claiming every batch of intervals must have exactly 95% captures. A finite simulation varies by chance. Say “about 95% in the long run,” not “exactly 95 out of every 100.”
- Confusing intervals with individual data values. The capture rate counts intervals that contain \(\mu\). It does not count observations that fall between an interval’s endpoints.
- Letting the population mean appear to move. In the simulation picture, \(\mu\) is the fixed reference value. The sample and its interval change from repetition to repetition.
- Forgetting the conditions. Long-run coverage describes a procedure applied under appropriate sampling and interval conditions. Do not claim that any interval-making method automatically captures the mean at its stated level.
- Treating one simulation as proof of exact coverage. A simulated proportion is an observed result that illustrates long-run performance. Its closeness to 95% depends partly on the number of repetitions and random variation.
For an AP Statistics response, use a sentence such as: “If we repeatedly took random samples of the same size from this population and calculated a 95% t interval each time, about 95% of those intervals would contain the true population mean.” This identifies the repetition, the procedure, the long-run rate, and the parameter being captured.
Check Your Understanding
Use the repeated-sampling meaning of confidence level to answer each question.
- In a simulation, 47 of 50 intervals contain the fixed population mean. Calculate the observed capture proportion and explain why it need not equal 95% exactly.
- In a repeated-sampling picture, which changes from run to run: the population mean, the sample, or the interval endpoints?
- Explain why “There is a 95% probability that the true mean is in this interval” is not the correct long-run interpretation of one calculated interval.
- A student says a 95% confidence interval means 95% of individual measurements lie between the endpoints. Identify the error.
- Write a sentence explaining what a 90% confidence level means for repeated samples and intervals.