When a Condition Fails, What Changes?
In Common Errors in Verifying Conditions, you practiced matching each condition check to the proposed procedure. This tutorial looks at what can happen when a one-proportion \(z\)-interval’s Large Counts condition fails: we use repeated simulated samples to examine how often intervals capture the true population proportion.
A 95% confidence level describes a method’s long-run capture rate when the method’s conditions are appropriate. In repeated use, about 95% of intervals should contain the fixed population proportion. But a formula producing an interval does not guarantee that its stated confidence level is achieved. When the Normal approximation is poor, the actual capture rate can differ substantially from 95%.
A simulation can make this long-run idea visible. The simulator sets a population proportion \(p\), repeatedly generates samples of size \(n\), constructs the same interval each time, and checks whether each interval contains the known \(p\). In real inference, \(p\) is unknown; it is known in the simulation only so we can evaluate the method.
For the one-proportion \(z\)-interval, the Large Counts check uses the observed successes and failures, \(x=n\hat{p}\) and \(n-x\), as explained in Large Counts Condition for Confidence Intervals. In a simulation, the model’s expected counts, \(np\) and \(n(1-p)\), help explain why the sample proportions may be far from Normal. Those expected counts do not replace the interval’s observed-count condition check.
How to Read a Capture-Rate Simulation
A simulation compares two percentages: the method’s stated confidence level and the observed fraction of simulated intervals that contain the true \(p\). The simulation does not change the confidence level. It investigates how the method performs under the specified population and sample size.
Choose the known population proportion \(p\) and sample size \(n\). For each simulated sample, generate \(n\) independent success-or-failure outcomes with success probability \(p\).
For every sample, calculate \(\hat{p}=x/n\) and use the same one-proportion \(z\)-interval method, such as the 95% interval \(\hat{p}\pm1.96\sqrt{\hat{p}(1-\hat{p})/n}\).
Count an interval as a capture if its endpoints contain the fixed \(p\). Otherwise, record a miss. Do this for every repetition.
Divide the number of captures by the total number of repetitions, then compare that estimate with the nominal 95% confidence level.
The simulation’s result will vary somewhat from run to run. A simulation with more repetitions usually gives a more stable estimate of the method’s capture rate, but no finite run has to equal the exact long-run rate.
Worked Examples
Worked Example: A Rare Success and a Low Capture Rate
A simulation models a population in which 5% of items have a particular defect, so \(p=0.05\). It repeatedly draws samples of \(n=20\) items and constructs a 95% one-proportion \(z\)-interval for each sample. Explain the condition concern and interpret a hypothetical simulation run in which 6,390 of 10,000 intervals contain \(p\).
State: The question is whether the usual 95% interval method captures the true defect proportion \(p=0.05\) about 95% of the time under this model. We will compare a simulated capture rate with 95% and examine the Large Counts issue.
Plan: In every repetition, the interval uses the observed sample proportion. For a sample with \(x\) defects, the observed counts are \(x\) successes and \(20-x\) failures. The interval’s Large Counts condition requires both counts to be at least 10. The model’s expected counts also help explain the shape of the sampling distribution, but they are not the formal interval check.
Do: Under the simulation model, the expected numbers of defects and nondefects are:
The expected number of defects is only 1, so the sampling distribution of \(\hat{p}\) is not well approximated by a symmetric Normal distribution. In any particular repetition, the interval’s observed-count condition also fails whenever \(x<10\) or \(20-x<10\).
The hypothetical run’s estimated capture rate is:
That run captures the true \(p\) in about 63.90% of its intervals, not about 95%. The gap is substantial. The simulation illustrates how a failed Large Counts condition can be associated with a distorted capture rate for this interval method.
As a check on the illustration, the exact capture probability for this particular model can be calculated from the binomial distribution. If \(X\) is the number of defects in one sample, the ordinary 95% interval contains \(0.05\) when \(X=1,2,3,\) or \(4\). For \(X=0\), the interval is \([0,0]\); for \(X=5\), its lower endpoint is already above \(0.05\). Thus:
The exact rate, about 63.89%, is close to the hypothetical simulation estimate of 63.90%. A real simulation run might give a somewhat different percentage because of random variation.
Conclude: The simulated capture rate is far below the nominal 95% level in this setting. The Large Counts condition warns that the Normal-based interval should not be treated as a reliable 95% procedure here. This example does not mean every interval will miss; it shows that the method’s long-run performance can be poor when its conditions fail.
Worked Example: Comparing With a Setting That Meets Large Counts
Now suppose the model has \(p=0.50\), with samples of \(n=20\), and the same 95% interval is used. A hypothetical run of 10,000 repetitions produces 9,587 intervals that contain \(p\). Compare this result with the rare-success example.
For the simulation model, the expected counts are:
Both expected counts equal 10. For any one interval, however, check the observed counts \(x\) and \(20-x\). The 95% Wald interval contains \(0.50\) for sample counts \(x=6\) through \(x=14\); for example, at \(x=6\), \(\hat{p}=0.30\), and the upper endpoint is about \(0.501\). At \(x=5\), the upper endpoint is about \(0.440\), so that interval misses \(0.50\).
The simulated capture rate is:
This is close to 95%, unlike the roughly 64% capture rate in the first example. The difference is consistent with the conditions: the second model is much more balanced, and its expected counts meet the Large Counts threshold. It is not a guarantee that every 95% interval method will capture exactly 95% in every setting.
For this model, the exact probability of capture is \(P(6\leq X\leq14)\), where \(X\) has a binomial distribution with \(n=20\) and \(p=0.50\). By symmetry, the excluded tails are \(X\leq5\) and \(X\geq15\):
The exact capture rate, about 95.86%, agrees closely with the hypothetical simulation result. In both examples, the simulation estimates a method’s long-run performance; it does not assign a probability to the truth of any one observed interval.
Worked Example: Why One Interval Does Not Tell the Whole Story
Return to the rare-defect model with \(p=0.05\) and \(n=20\). Consider two possible simulated samples: one has \(x=0\) defects, and another has \(x=1\). Calculate each 95% Wald interval and decide whether it captures \(p\).
For \(x=0\), \(\hat{p}=0/20=0\). The interval is:
This interval misses the true \(p=0.05\). It also has zero width because the estimated standard error is zero. The observed counts are 0 successes and 20 failures, so the Large Counts condition fails.
For \(x=1\), \(\hat{p}=1/20=0.05\). The standard error and interval are:
This interval contains \(0.05\), so it captures the true proportion. Its lower endpoint is negative, which is not a possible population proportion; that is another warning sign for this Normal-based method in a small-count setting. The observed counts are 1 success and 19 failures, so Large Counts fails for this interval as well.
The first interval misses and the second captures, but neither single result tells us the long-run capture rate. That rate comes from examining all the intervals across many repetitions. The simulation in the first example showed that, despite some captures, the overall rate can be far below the advertised 95%.
What a Failed Condition Does—and Does Not—Say
A failed Large Counts condition means that the Normal approximation required for the usual one-proportion \(z\)-interval is not adequately supported. It does not mean that the sample proportion is automatically wrong, that the interval must miss \(p\), or that every interval from the procedure will have the same defect. Instead, the method’s stated long-run capture rate is no longer assured by the conditions.
The impact depends on the setting. In the rare-success example, the distribution of sample proportions is concentrated near zero with a long right tail, unlike the symmetric Normal model used to motivate the interval. A symmetric interval based on \(\hat{p}\) can therefore miss more often than its nominal confidence level suggests. In other situations, the actual capture rate may be distorted in a different direction. Do not assume that every condition failure produces exactly the same error.
A simulation is a useful way to investigate performance, but it does not repair the condition failure. If the observed counts in your actual sample do not satisfy Large Counts, do not present the usual interval as though its stated confidence level were justified. Explain the failed condition and follow the method appropriate to the problem and course expectations.
Common Mistakes and What Full Credit Requires
- Claiming every interval misses. A failed condition warns that the method’s performance may be unreliable; it does not determine the result of each individual interval. State what the condition failure means for the method, not for every repetition.
- Calling one simulation run the exact capture rate. The proportion of captures in a finite run is an estimate. Report the number of captures and repetitions, calculate their ratio, and describe it as a simulated estimate.
- Using expected counts as the interval’s formal check. The one-proportion interval check uses the observed counts \(x\) and \(n-x\). In a simulation, \(np\) and \(n(1-p)\) describe the model’s expected counts; explain their role without substituting them for the observed-count check.
- Confusing confidence level with a single interval’s probability. The 95% confidence level refers to the long-run capture rate of the method under appropriate conditions. It does not say there is a 95% probability that one particular interval contains the fixed \(p\).
- Concluding that a close simulated rate proves the conditions are met. A finite simulation can produce a rate near 95% by chance. Check the procedure’s conditions from the sample and study design; do not use a favorable simulation result as a substitute.
Key Takeaway
Repeated sampling reveals a consequence of violating Large Counts: a one-proportion \(z\)-interval’s actual capture rate can differ substantially from its stated confidence level. In the rare-success simulation, the estimated rate was about 64%, while a balanced comparison was close to 96%. The condition check is a warning about the method, not a prediction about any one interval.
Check Your Understanding
Answer each question using the distinction between an individual interval and a method’s long-run capture rate.
- A simulation produces 7,420 capturing intervals out of 10,000. Calculate the simulated capture rate and compare it with a nominal 95% confidence level.
- For an interval based on \(n=30\) and \(x=4\), identify the observed success and failure counts. Does the Large Counts condition hold?
- A student says, “Large Counts failed, so this particular interval definitely misses the true proportion.” Explain the error.
- In a simulation model with \(p=0.10\) and \(n=50\), calculate the expected success and failure counts. Explain why those values do not replace the observed-count check for each interval.
- What does a confidence level describe, and how does a capture-rate simulation estimate it?