What a Randomization Distribution Shows
In a two-proportion test, the observed difference \(\hat{p}_1-\hat{p}_2\) is compared with differences that could occur if the null hypothesis were true. Earlier tutorials used a pooled standard error and a \(z\)-distribution to make that comparison. A randomization simulation offers another approach: repeatedly rearrange outcomes in a way that represents the null model, calculate the difference for each rearrangement, and examine the resulting distribution.
The simulated distribution is a model of the differences expected under the null, not a distribution of possible values of the unknown parameters. Its center should be near zero when the null hypothesis says \(p_1=p_2\). The simulation helps answer a specific question: if the null model were true, how often would a difference at least as extreme as the observed one occur?
Building a Distribution Under Equal Proportions
For two independent samples with binary outcomes, a pooled permutation simulation represents the null hypothesis \(H_0:p_1=p_2\). Under this null, both populations have the same probability of success. Once we condition on the total number of observed successes in the two samples, the success labels can be treated as exchangeable between the groups. In other words, under the common-proportion model, no observed success label is more likely to be assigned to one sample than the other.
To carry out the simulation, combine the observed success and failure labels, shuffle them, and deal them into two groups with the original sample sizes \(n_1\) and \(n_2\). For each shuffle, calculate the simulated difference \(D^*=\hat{p}_1^*-\hat{p}_2^*\), keeping the original group order. Repeat this many times. The pooled proportion is \((x_1+x_2)/(n_1+n_2)\), but each shuffle uses the fixed total number of observed successes rather than generating a new total.
Calculate \(D_{\text{obs}}=\hat{p}_1-\hat{p}_2\) in the group order used for the hypotheses.
Pool the observed success and failure labels, then randomly assign them to groups of sizes \(n_1\) and \(n_2\).
For every shuffle, calculate \(D^*=\hat{p}_1^*-\hat{p}_2^*\). The collection of these values forms the simulated randomization distribution.
Count the simulated differences at least as extreme as \(D_{\text{obs}}\), in the direction or directions specified by \(H_a\).
The phrase “at least as extreme” includes ties. For \(H_a:p_1>p_2\), count simulated differences \(D^*\ge D_{\text{obs}}\). For \(H_a:p_1<p_2\), count \(D^*\le D_{\text{obs}}\). For \(H_a:p_1\ne p_2\), count values with \(|D^*|\ge|D_{\text{obs}}|\), which includes both tails. The direction comes from the alternative hypothesis, not from a direction chosen after seeing the sample results.
This is a simulation estimate, so its value can vary somewhat from run to run. A larger number of shuffles usually makes the estimate more stable, but it does not change the observed difference or the hypotheses. If none of \(B\) simulations is as extreme as the observed result, do not conclude that the p-value is zero: those simulations only show that the estimated tail proportion is below \(1/B\). The AP-style estimate is the number of extreme simulated results divided by the number of simulations.
Random Assignment and the Null Being Tested
The phrase “randomization test” also arises with randomized experiments, but the justification depends on which null hypothesis is being tested. In an experiment, one valid randomization procedure holds each person’s observed outcome fixed and reallocates the treatment labels according to the experiment’s assignment process. This procedure tests the sharp null hypothesis that treatment has no effect on any individual unit: each person would have had the same outcome under either assignment.
That sharp null is stronger than merely saying the treatment and control population proportions are equal. Equality of population proportions alone does not tell us that every individual’s outcome would remain unchanged if their assignment changed. If the goal is to test only equality of proportions using a pooled permutation simulation, the procedure instead relies on an exchangeability or common-distribution assumption: under the null, outcomes in the two groups have the same distribution. For a binary outcome, this means the same success probability. Keep the tested null and the simulation method aligned.
Worked Examples
Worked Example: A Two-Sided Test of Recycling Participation
Question: In a fictional study, independent random samples of 50 households are selected from each of two neighborhoods. In Neighborhood 1, 31 households report participating in a recycling program; in Neighborhood 2, 21 do. A simulation shuffles the 52 success labels and 48 failure labels into groups of 50, repeating this process 1,000 times. Estimate the p-value for testing whether the population proportions differ.
State: Let \(p_1\) and \(p_2\) be the true proportions of households in Neighborhood 1 and Neighborhood 2, respectively, that participate in the recycling program. The hypotheses are \(H_0:p_1=p_2\) and \(H_a:p_1\ne p_2\).
Plan: Each neighborhood contributes an independent random sample, and each household is classified as a success or failure for the same outcome. Under the null, both populations have a common success proportion, so the pooled outcomes can be shuffled between fixed groups of 50. The simulation estimates the two-sided p-value by counting simulated differences at least as far from zero as the observed difference.
Do: The sample proportions are \(\hat{p}_1=31/50=0.62\) and \(\hat{p}_2=21/50=0.42\). Thus, the observed difference is \(D_{\text{obs}}=0.62-0.42=0.20\). The pooled proportion is \((31+21)/(50+50)=52/100=0.52\). In the 1,000 shuffles, suppose 85 produce \(|D^*|\ge0.20\). The estimated p-value is:
For example, a summary of this illustrative simulation might place 42 results at or below \(-0.20\), 135 between \(-0.20\) and \(-0.10\), 645 from \(-0.10\) through \(0.10\), 135 between \(0.10\) and \(0.20\), and 43 at or above \(0.20\). These bins total \(42+135+645+135+43=1{,}000\), and the two tails contain \(42+43=85\) results. The estimated p-value is therefore the proportion in both tails, not just the proportion in the direction of the observed difference.
Conclude: Under the common-proportion null model, about 8.5% of the simulated reallocations produced a difference at least as far from zero as the observed difference of 0.20. The simulation estimates the two-sided p-value as 0.085. This describes how unusual the sample difference is under the null model; it is not the probability that the null hypothesis is true.
Worked Example: A One-Sided Test of Shade-Tree Survival
Question: In a fictional survey, independent random samples of 60 residents are asked whether a newly planted street tree near their home survived its first summer. Thirty-nine residents in Area A and 27 in Area B report that it did. Researchers want to know whether the survival proportion is higher in Area A. A pooled permutation simulation produces 2,000 differences. Estimate the p-value if 54 simulated differences are at least as large as the observed difference.
State: Let \(p_1\) be the true proportion of residents in Area A whose newly planted street tree survived its first summer, and \(p_2\) the corresponding proportion in Area B. Test \(H_0:p_1=p_2\) against \(H_a:p_1>p_2\).
Plan: The samples are independent random samples, each response is success or failure for the same characteristic, and the alternative specifies that Area A has the higher proportion. The pooled simulation keeps the 66 observed successes and 54 failures fixed, shuffling them into two groups of 60. For the upper-tail alternative, count simulated differences \(D^*\ge D_{\text{obs}}\).
Do: The sample proportions are \(\hat{p}_1=39/60=0.65\) and \(\hat{p}_2=27/60=0.45\). The observed difference is \(D_{\text{obs}}=0.65-0.45=0.20\). Of the 2,000 simulated differences, 54 are at least \(0.20\). The estimated p-value is:
The simulation’s left tail is not included in this count. Although differences in the opposite direction could be far from zero, they are not as extreme in the direction required by \(H_a:p_1>p_2\).
Conclude: If the common-proportion null model were true, about 2.7% of the simulated reallocations produced a difference of at least 0.20 in favor of Area A. The simulation estimates the upper-tail p-value as 0.027. The direction of the tail follows the research question, rather than being selected simply because the observed difference is positive.
Worked Example: A Randomized Experiment and the Sharp Null
Question: In a fictional experiment, 40 seedlings are randomly assigned, 20 to a new soil treatment and 20 to standard soil. After a set period, 14 treatment seedlings and 6 standard-soil seedlings meet a defined growth target. A simulation reallocates the 20 treatment labels among the 40 seedlings 2,000 times, keeping each seedling’s observed outcome fixed. Suppose 113 reallocations produce a treatment-minus-standard difference at least as large as the observed one. Interpret the estimated p-value and identify the null being tested.
State: Let \(p_T\) and \(p_C\) be the proportions of seedlings that would meet the growth target under treatment and standard soil, respectively. The observed difference is in the direction \(p_T-p_C\). The stated label-reallocation procedure tests the sharp null that the soil treatment has no effect on the growth outcome of any individual seedling.
Plan: The seedlings were randomly assigned, so the treatment labels can be reallocated according to the experiment’s assignment process. Under the sharp null, every seedling’s observed outcome would be unchanged under either assignment; therefore, the outcomes can be held fixed while the labels are shuffled. This is not justified merely by the weaker claim \(p_T=p_C\).
Do: The observed proportions are \(\hat{p}_T=14/20=0.70\) and \(\hat{p}_C=6/20=0.30\). The observed treatment-minus-standard difference is \(0.70-0.30=0.40\). The simulation produces 113 differences at least \(0.40\), so the estimated upper-tail p-value is:
Conclude: Under the sharp null of no individual treatment effect, about 5.65% of the simulated assignments produced a treatment-minus-standard difference at least as large as 0.40. This is an estimated p-value for the sharp null, not automatically a test of every possible null that says only the two population proportions are equal.
Common Mistakes and AP Exam Tips
- Counting the wrong tail: For a one-sided alternative, count simulated differences in its specified direction. For a two-sided alternative, count both tails using absolute differences.
- Forgetting ties: “At least as extreme” includes simulated statistics equal to the observed statistic, not just values strictly beyond it.
- Changing group order: Calculate every simulated difference as Group 1 minus Group 2. Reversing the order changes the sign and can send you to the wrong tail.
- Shuffling without preserving group sizes: Each simulated allocation must retain the original sample sizes. Otherwise, the statistic is not being generated under the same design.
- Calling a simulated proportion the exact p-value: \(K/B\) estimates the p-value from the simulated runs. State the number of extreme results and the total number of simulations.
- Claiming that a zero count means a zero p-value: If no simulation is as extreme, the result only indicates that the tail event was not seen in those \(B\) trials. More simulations may reveal such outcomes.
- Confusing the sharp null with equal proportions: Reallocating fixed outcomes in a randomized experiment tests no individual treatment effect. If testing equality of proportions with a pooled permutation method, state the common-distribution or exchangeability assumption.
Key Takeaway
A randomization distribution shows the differences in sample proportions that can arise under a stated null model. Its tail area estimates the p-value: count simulated differences at least as extreme as the observed difference, using the tail or tails specified by the alternative. The validity of the simulation depends on matching the reallocation method to the null hypothesis.
Check Your Understanding
Answer each question using the group order and null model stated.
- For \(H_a:p_1>p_2\), the observed difference is 0.12. Which simulated differences count toward the p-value estimate?
- A two-sided simulation generates 1,500 differences, and 96 have absolute value at least as large as the observed difference. Calculate and interpret the estimated p-value.
- Why should a pooled permutation simulation keep the two sample sizes and the total number of observed successes fixed?
- In a randomized experiment, what null hypothesis is tested by reallocating treatment labels while keeping every unit’s observed outcome fixed?
- A simulation produces zero results at least as extreme as the observed difference in 500 runs. What can you conclude about the estimated p-value, and what should you avoid claiming?