Tutorials › AP Statistics › Hypothesis Testing With Simulated Sampling Distributions

One-proportion hypothesis tests · Tutorial 474 of 1000

Hypothesis Testing With Simulated Sampling Distributions

Model repeated samples as if the null proportion were true, then estimate a p-value by finding the fraction of simulated results at least as extreme as the observed result.

Intermediate 10 min read

What You'll Learn

  • Set up simulated samples using the null proportion as the success probability.
  • Decide which simulated counts are at least as extreme as the observed count.
  • Estimate a p-value with the proportion of simulations meeting that extremeness rule.
  • Adapt the counting rule for left-tailed, right-tailed, and two-sided tests.
  • Explain why a simulated p-value is an estimate and how more repetitions affect it.

Simulate What the Null Hypothesis Predicts

In Sketching the P-Value Region for a Test, you identified the part of a null model that counts as at least as extreme as the observed result. A simulation offers another way to estimate the area of that region. Instead of calculating an area under a Normal curve, repeatedly generate samples as if the null hypothesis were true and count how often their results are at least as extreme as the observed one.

For a one-proportion test, the null hypothesis supplies the success probability \(p_0\). In every simulated sample, use the original sample size \(n\), and make each observation a success with probability \(p_0\) and a failure with probability \(1-p_0\). Record the number of successes, or equivalently the sample proportion, for each repetition. This collection of simulated results is a simulated sampling distribution under the null hypothesis.

Definition: A simulated p-value is the fraction of simulated samples generated under \(H_0\) whose results are at least as extreme as the observed result, according to the alternative hypothesis.

If \(R\) samples are simulated and \(E\) of their results are counted as extreme, the estimated p-value is \(E/R\). This is an estimate, not an exact probability: a different run of the simulation may produce a slightly different fraction. The simulation estimates the probability described by the p-value; it does not change the meaning of that probability.

$$ \text{simulated p-value}=\frac{\text{number of simulated results at least as extreme as observed}}{R} $$

Set Up the Simulation and Define “Extreme”

The simulation must represent the null hypothesis, not an estimate chosen from the observed data. For example, if \(H_0:p=0.40\), each simulated observation has a 0.40 chance of success. Keep \(n\) fixed across repetitions because the observed sample size is fixed. Within each repetition, simulate \(n\) independent success-or-failure outcomes, then calculate the simulated number of successes \(x_{\mathrm{sim}}\) or proportion \(\hat{p}_{\mathrm{sim}}=x_{\mathrm{sim}}/n\).

The alternative hypothesis determines which simulated outcomes count. Because \(n\) is fixed, you can compare success counts or sample proportions; they give the same ordering. Include a result tied with the observed result: “at least as extreme” includes equality.

AlternativeCount a simulated result when...
\(H_a:p<p_0\)\(x_{\mathrm{sim}}\leq x_{\mathrm{obs}}\), or \(\hat{p}_{\mathrm{sim}}\leq\hat{p}_{\mathrm{obs}}\)
\(H_a:p>p_0\)\(x_{\mathrm{sim}}\geq x_{\mathrm{obs}}\), or \(\hat{p}_{\mathrm{sim}}\geq\hat{p}_{\mathrm{obs}}\)
\(H_a:p\ne p_0\)The simulated result is at least as far from \(p_0\) as the observed result: \(\lvert\hat{p}_{\mathrm{sim}}-p_0\rvert\geq\lvert\hat{p}_{\mathrm{obs}}-p_0\rvert\)

For a two-sided test, compare distances from the null value in either direction. In counts, the rule is \(\lvert x_{\mathrm{sim}}-np_0\rvert\geq\lvert x_{\mathrm{obs}}-np_0\rvert\). Since success counts are whole numbers, the two tails need not contain matching numbers of possible outcomes. Do not assume the simulated two-sided p-value is always twice a simulated one-tail fraction.

This method directly simulates the discrete outcomes predicted by \(H_0\). It does not require a Normal curve or the Large Counts condition used for a one-proportion \(z\)-test. The Random condition and, when appropriate, the 10% condition still matter for whether the sample supports inference about the target population. Also, the simulated outcomes should follow the null model: independent binary outcomes with success probability \(p_0\).

Worked Examples

Worked Example: Estimate a Right-Tailed P-Value

A fictional community garden randomly selects 100 members from a membership list of 3,000. In the sample, 48 members say they compost food scraps. Test whether the true proportion of members who compost is greater than 0.40, using \(\alpha=0.05\). A simulation of 5,000 samples, each of size 100 generated with success probability 0.40, produces 326 samples with at least 48 successes.

State: Let \(p\) be the true proportion of members on this list who compost food scraps. The hypotheses are \(H_0:p=0.40\) and \(H_a:p>0.40\).

Plan and check conditions: Use a simulation under the null hypothesis to estimate the p-value. The membership list was randomly sampled, so the Random condition is met. The sample was drawn without replacement from 3,000 members, and \(0.10(3{,}000)=300\); since \(100\leq300\), the 10% condition is met. For the simulation, model each of the 100 outcomes as an independent success with probability 0.40 under the null. The null expected counts are \(100(0.40)=40\) successes and \(100(0.60)=60\) failures. The simulation itself does not need the Large Counts condition for a Normal approximation.

Do: The observed sample proportion is:

$$ \hat{p}_{\mathrm{obs}}=\frac{48}{100}=0.48 $$

Because the alternative is right-tailed, count simulated samples with 48 or more successes. The reported simulation has \(E=326\) such samples out of \(R=5{,}000\):

$$ \text{simulated p-value} =\frac{326}{5{,}000} =0.0652 $$

Conclude: Assuming the true proportion is 0.40, the simulation estimates that about 0.0652 of samples of 100 members would have 48 or more composters. Since \(0.0652>0.05\), fail to reject \(H_0\). The sample does not provide convincing evidence that more than 0.40 of members on this list compost food scraps.

Worked Example: Estimate a Left-Tailed P-Value

A fictional language program randomly selects 80 current participants from a list of 2,000. Of those selected, 39 say they use a language-learning app at least four days a week. The question is whether the true proportion of participants who use the app that often is less than 0.60. In a simulation of 4,000 samples of 80 outcomes generated with success probability 0.60, 116 samples have 39 or fewer successes.

Let \(p\) be the true proportion of current participants on the list who use the app at least four days a week. The hypotheses are \(H_0:p=0.60\) and \(H_a:p<0.60\). The sample is random, and \(80\leq0.10(2{,}000)=200\), so the 10% condition is met. Simulate 80 independent outcomes per repetition, with success probability 0.60 under \(H_0\). The null expected counts are \(80(0.60)=48\) successes and \(80(0.40)=32\) failures.

The observed sample proportion is:

$$ \hat{p}_{\mathrm{obs}}=\frac{39}{80}=0.4875 $$

For the left-tailed alternative, simulated counts at or below the observed count are extreme. Thus, \(E=116\) and \(R=4{,}000\):

$$ \text{simulated p-value} =\frac{116}{4{,}000} =0.0290 $$

If the significance level is 0.05, then \(0.0290<0.05\), so reject \(H_0\). The simulated results provide convincing evidence that the true proportion of current participants on this list who use the app at least four days a week is less than 0.60.

Worked Example: Count Both Tails in a Two-Sided Test

A fictional wildlife center randomly selects 50 registered volunteers from a list of 1,000. Fifteen say they have seen a particular bird species during volunteer shifts. Test whether the true proportion of volunteers who have seen the species differs from 0.20. In a simulation of 10,000 samples of 50 outcomes generated with success probability 0.20, 795 samples have results at least as far from the null expectation as the observed result.

Let \(p\) be the true proportion of registered volunteers who have seen the species during a shift. The hypotheses are \(H_0:p=0.20\) and \(H_a:p\ne0.20\). The volunteers were randomly selected, and \(50\leq0.10(1{,}000)=100\), so the Random and 10% conditions are met. For the simulation, generate 50 independent outcomes per repetition with success probability 0.20 under the null.

The observed proportion is \(15/50=0.30\). Under \(H_0\), the expected number of successes is \(np_0=50(0.20)=10\). The observed count is 5 above that expectation. Therefore, simulated counts at least as far from 10 as 15 are \(x_{\mathrm{sim}}\leq5\) or \(x_{\mathrm{sim}}\geq15\). The simulation reports 795 such results:

$$ \text{simulated p-value} =\frac{795}{10{,}000} =0.0795 $$

Since \(0.0795>0.05\), fail to reject \(H_0\) at the 0.05 significance level. The simulated results do not provide convincing evidence that the true proportion of registered volunteers who have seen the species differs from 0.20. The simulation counted departures in either direction because the alternative is two-sided.

Simulation Variation and Common Mistakes

A simulated p-value can change a little from run to run because the simulated samples are random. Increasing the number of repetitions generally makes the estimate less variable, though it does not make it exact. A useful approximate measure of the simulation’s own variability is the standard error of the estimated fraction:

$$ SE_{\mathrm{sim}}\approx \sqrt{\frac{\hat{q}(1-\hat{q})}{R}}, \qquad \hat{q}=\frac{E}{R} $$

For the garden simulation, \(\hat{q}=0.0652\) and \(R=5{,}000\), so \(SE_{\mathrm{sim}}\approx\sqrt{0.0652(0.9348)/5{,}000}\approx0.0035\), rounded. This describes random variation in the simulation estimate; it is not a new test statistic and does not replace interpreting the p-value in context.

  • Simulating with \(\hat{p}\) instead of \(p_0\). The p-value asks what results are plausible if the null hypothesis is true. Use the null proportion as the simulated success probability.
  • Changing the sample size. Each repetition should have the original \(n\). Otherwise, the simulated results do not represent the sampling distribution for the observed study.
  • Letting the observed direction choose the tail. Use \(H_a\) to decide what counts as extreme. A left-tailed test counts results at or below the observed result, even if that result happens to be above \(p_0\).
  • Counting only results more extreme, but not tied. Include simulated results equal to the observed count or equally distant from the null. The rule is “at least as extreme.”
  • Calling the estimate exact. State that the simulation estimates the p-value. Report the fraction and the number of repetitions so the estimate can be checked.
  • Using only one tail for a two-sided question. For \(H_a:p\ne p_0\), count simulated proportions at least as far from \(p_0\) as the observed proportion, on either side.
AP Exam Tip: Describe the null model, keep the observed sample size fixed, identify the rule for counting extreme simulated outcomes from \(H_a\), and calculate \(E/R\). Interpret that fraction as an estimated probability assuming \(H_0\) is true, then compare it with \(\alpha\) to make the test conclusion in context.

Key Takeaway

A simulation turns the null hypothesis into a repeatable model: generate many samples under \(p_0\), apply the alternative’s extremeness rule to every result, and use the fraction counted to estimate the p-value. The observed data determine the cutoff for “as extreme”; the null hypothesis determines how the simulated samples are generated.

Key takeaway: Simulate with \(p_0\) and the original \(n\); count outcomes at least as extreme as observed in the direction or directions specified by \(H_a\); divide by the number of repetitions.

Check Your Understanding

Answer each question using the logic of a simulated sampling distribution under the null hypothesis.

  1. For \(H_0:p=0.35\) with \(n=120\), what success probability and sample size should each simulated repetition use?
  2. A left-tailed test observes 42 successes. Which simulated success counts should be included in the extreme-result total?
  3. A two-sided test has \(np_0=30\) and observes 36 successes. State the rule for which simulated success counts are at least as extreme.
  4. In 2,000 repetitions, 74 simulated outcomes meet the extremeness rule. Calculate the simulated p-value.
  5. Why can two correct runs of the same simulation report slightly different p-value estimates?