Testing a Treatment Effect in an Experiment
A two-proportion test can help assess whether an observed difference between treatment and control groups is convincing evidence of a treatment effect. In One-Sided Two-Proportion Test Example, the question was whether the treatment success proportion was higher than the control success proportion. Here, the emphasis is on the experimental design: random assignment can support a cause-and-effect conclusion when the test provides convincing evidence of a difference.
A test result does not, by itself, establish that a difference was caused by the treatment. The study design matters. Random assignment helps create comparable treatment and control groups, so a difference in outcomes can reasonably be attributed to the assigned condition. But random assignment is not the same as random sampling: it does not, by itself, justify generalizing the result to people outside the experiment.
Choose the alternative to match the research question, not the observed sample results. If the question predicts a higher success proportion under treatment, use \(H_a:p_T>p_C\). If it asks whether the treatment has any effect, use \(H_a:p_T\ne p_C\). These hypotheses and the pooled test calculation are covered in Hypotheses for Comparing Two Population Proportions, Pooled Proportion and Why It Is Used, and Test Statistic for Two Proportions.
Conditions and What Random Assignment Supports
Before calculating a test statistic, check the conditions for a two-proportion \(z\)-test. In an experiment, establish that units were randomly assigned to treatment and control and that the groups consist of independent units. If the experiment uses the same units in both conditions or pairs units together, a two-proportion test for independent groups is not appropriate. For sampling without replacement, also check the 10% condition when it applies. Finally, check the Large Counts condition using expected successes and failures under the null hypothesis, as explained in Large Counts Using Expected Successes and Failures in Each Group.
Random assignment supports a causal interpretation because chance determines which experimental units receive each condition. If a test rejects equal proportions, the result can provide convincing evidence that the treatment caused a difference in the outcome for the experimental units. Whether the result can also be generalized depends on how those units were selected. As discussed in Scope of Inference Based on How Data Were Collected, random sampling can support generalization to the population sampled; random assignment supports a causal conclusion. One design feature does not substitute for the other.
A Four-Step Test with an Experimental Conclusion
A complete response follows the State, Plan, Do, Conclude structure. In the conclusion, report the decision about \(H_0\), describe the evidence about the true proportions in context, and let the design guide the causal claim. If the p-value is not small, do not claim that the treatment had no effect. If it is small, do not automatically generalize beyond the units represented by the study.
Define \(p_T\) and \(p_C\), then write the null and alternative hypotheses that match the question.
Identify the random assignment, check group independence and the 10% condition when relevant, and verify all four pooled expected counts.
Calculate the pooled proportion, pooled standard error, \(z\)-statistic, and p-value for the specified alternative.
Compare the p-value with \(\alpha\), make the correct decision, and state what the evidence supports about the treatment effect and the scope of the conclusion.
Worked Examples
Worked Example: A Randomized Study of Tutoring Reminders
Question: In a fictional experiment, 200 students are randomly selected from a college roster of 5,000 students and randomly assigned, 100 per group. One group receives a new tutoring-session reminder; the other receives the usual message. During the study, 65 students in the new-reminder group attend a scheduled session, compared with 45 in the usual-message group. Does the experiment provide evidence that the new reminder increases attendance? Use \(\alpha=0.05\).
State: Let \(p_T\) be the true proportion of students who would attend a scheduled session under the new reminder, and \(p_C\) the true proportion who would attend under the usual message. Test \(H_0:p_T=p_C\) against \(H_a:p_T>p_C\). The one-sided alternative matches the question’s prediction that the reminder increases attendance.
Plan: The students were randomly assigned to the two message conditions, and each student was in only one group, supporting random assignment and independent groups. The random sample of 200 is less than 10% of the roster of 5,000, so the 10% condition is met. Under \(H_0\), the pooled proportion is \((65+45)/(100+100)=110/200=0.55\). The expected successes are \(100(0.55)=55\) in each group, and the expected failures are \(100(0.45)=45\) in each group. All four expected counts are at least 10, so the Large Counts condition is met.
Do: The sample proportions are \(\hat{p}_T=65/100=0.65\) and \(\hat{p}_C=45/100=0.45\). The observed difference is \(0.65-0.45=0.20\). The pooled standard error and test statistic are:
For \(H_a:p_T>p_C\), the p-value is the standard normal area to the right of \(2.843\): \(P(Z\ge2.843)\approx0.0022\), rounded to four decimal places. A TI-84 2-PropZTest with \(x_T=65,n_T=100,x_C=45,n_C=100\), and the \(p_T>p_C\) alternative gives a test statistic and p-value that round to these values.
Conclude: If the true attendance proportions under the two message conditions are equal, the probability of getting a sample difference at least as large in the treatment-favoring direction as the observed difference is about \(0.0022\). Since \(0.0022<0.05\), we reject \(H_0\). The experiment provides convincing evidence that the new reminder caused a higher attendance proportion among students represented by the study. Because students were randomly selected from the college roster as well as randomly assigned, generalizing to that roster’s student population is also supported by the design.
Worked Example: A Randomized Wellness Program Without Convincing Evidence
Question: In a fictional wellness experiment, 240 people who volunteered for a program are randomly assigned in equal numbers to a daily walking prompt or to general wellness information. By the end of the study, 82 of the 120 people assigned to the prompt meet a specified weekly activity goal, compared with 74 of the 120 assigned to general information. Is there evidence that the prompt increases the success proportion? Use \(\alpha=0.05\).
Let \(p_T\) and \(p_C\) be the true proportions of these volunteers who would meet the activity goal under the prompt and general-information conditions, respectively. Test \(H_0:p_T=p_C\) against \(H_a:p_T>p_C\). The groups are independent because each volunteer was assigned to just one condition. Random assignment was used, and the volunteers were not sampled without replacement from a defined target population, so a 10% condition for such sampling is not applicable. The pooled proportion is \((82+74)/(120+120)=156/240=0.65\). The expected successes are \(120(0.65)=78\) per group, and the expected failures are \(120(0.35)=42\) per group. All four expected counts exceed 10.
The sample proportions are \(\hat{p}_T=82/120\approx0.6833\) and \(\hat{p}_C=74/120\approx0.6167\), giving a difference of approximately \(0.0667\). The pooled standard error is:
Using the unrounded sample proportions, the test statistic is:
The p-value is the area to the right of \(1.082\): \(P(Z\ge1.082)\approx0.1397\), rounded to four decimal places. Since \(0.1397>0.05\), we fail to reject \(H_0\). These data do not provide convincing evidence that the walking prompt increases the true activity-goal success proportion for volunteers like those in the experiment. Random assignment would support a causal conclusion if the evidence were convincing, but it does not turn this non-significant result into evidence of a treatment effect. The volunteers were not randomly sampled, so the result also should not be generalized automatically to all people.
Worked Example: Evidence of a Difference in a Randomized Campus Experiment
Question: In a fictional campus experiment, 300 dining-hall volunteers are randomly assigned in equal numbers to receive either a personalized reusable-cup reminder or the usual dining-hall information. During the following week, 90 of the 150 students assigned to the reminder bring a reusable cup, compared with 65 of the 150 students assigned to usual information. Is there evidence that the reminder changes the proportion who bring a cup? Use \(\alpha=0.05\).
Let \(p_T\) be the true proportion of the experiment’s volunteer students who would bring a reusable cup under the personalized reminder, and \(p_C\) the corresponding proportion under usual information. Because the question asks whether the reminder changes the proportion in either direction, test \(H_0:p_T=p_C\) against \(H_a:p_T\ne p_C\).
The students were randomly assigned to separate groups, supporting random assignment and independence. They were volunteers rather than a random sample of all campus students, so a 10% condition for sampling without replacement from a target population does not apply. The pooled proportion is \((90+65)/(150+150)=155/300\approx0.5167\). The expected successes in each group are \(150(155/300)=77.5\), and the expected failures in each group are \(150(145/300)=72.5\). All four expected counts are at least 10.
The sample proportions are \(\hat{p}_T=90/150=0.60\) and \(\hat{p}_C=65/150\approx0.4333\). The pooled standard error and test statistic are:
For the two-sided alternative, the p-value includes standard normal outcomes at least as far from zero as \(2.888\) in either direction: \(2P(Z\ge2.888)\approx0.0039\), rounded to four decimal places. Since \(0.0039<0.05\), we reject \(H_0\). The experiment provides convincing evidence that the personalized reminder caused a change in the reusable-cup proportion among the volunteer students studied; because the observed treatment proportion is higher, the evidence points toward an increase. The volunteer sample does not, by itself, justify generalizing this causal conclusion to all campus students.
Common Mistakes and AP Exam Tips
- Confusing random assignment with random sampling: Random assignment supports a causal conclusion; random sampling supports generalization to the population sampled. A study can have one, both, or neither.
- Claiming causation just because the p-value is small: A small p-value gives evidence against the null model. The random assignment is what makes a cause-and-effect conclusion reasonable. Without it, describe an association rather than a causal effect.
- Overgeneralizing a randomized experiment: If participants volunteered or were recruited by convenience, do not claim the result applies to everyone. Name the experimental units in the conclusion.
- Changing the alternative after seeing the data: Choose \(>\), \(<\), or \(\ne\) from the research question before inspecting which sample proportion is larger.
- Using separate sample proportions in the test standard error: For a test of equal population proportions, use the pooled proportion and pooled standard error. The calculation differs from the standard error used for a two-proportion confidence interval.
- Saying “there is no effect” after failing to reject: A large p-value is not proof of equality. Say that the experiment did not provide convincing evidence of the stated effect.
- Ignoring the units and context: Describe the proportions as proportions of experimental units with the defined outcome under each assigned condition, and preserve treatment-minus-control order.
Key Takeaway
A two-proportion \(z\)-test evaluates whether the observed treatment-control difference is unusual under a null model of equal proportions. Random assignment makes a causal interpretation of convincing evidence reasonable, but the participant-selection method determines whether that conclusion can extend beyond the experimental units.
Check Your Understanding
Assume the two-proportion \(z\)-test conditions are met unless a question asks you to identify a condition.
- A randomized experiment compares a new study-planning prompt with usual instructions. Define \(p_T\) and \(p_C\), and write hypotheses for testing whether the prompt increases the success proportion.
- In a randomized experiment, participants volunteer and are then randomly assigned. What kind of conclusion can random assignment support, and what does volunteering fail to establish?
- A study randomly selects participants from a roster and randomly assigns them to two conditions. Which design features support generalization, and which support a causal conclusion?
- For a two-proportion test with \(H_a:p_T\ne p_C\), the test statistic is \(z=2.10\). Which standard normal areas make up the p-value?
- A randomized experiment gives \(p=0.08\) for a test at \(\alpha=0.05\). State the decision and a careful conclusion without claiming the treatment has no effect.