From a Treatment Comparison to a Causal Interval
A two-sample t interval can estimate the difference between two means when the groups come from a randomized experiment. The calculation is the same unpooled interval introduced in “The Form of a Two-Sample t Interval.” What changes is how the study design shapes the conclusion: random assignment can support interpreting a difference in outcomes as an effect of the treatments.
First choose and state the subtraction order. In this tutorial, group 1 receives treatment A and group 2 receives treatment B, so the parameter is \(\mu_A-\mu_B\). A positive value means the mean response under treatment A is higher than the mean response under treatment B. Keep that order consistent in the sample difference, interval, and interpretation.
The two treatment groups must be independent, unpaired groups. If each subject receives both treatments, the data are paired and the paired t procedure applies instead. As discussed in “Conditions for a Two-Sample t Interval,” check the design, independence, and distribution or sample size in each group before calculating the interval.
Calculate the Interval, Then Interpret the Design
For independent groups, the unpooled two-sample t interval uses the difference in sample means as its center. It estimates the variability separately in each group and uses Welch degrees of freedom. This procedure does not combine the groups’ standard deviations into a pooled estimate.
Here, \(t^*\) is chosen for the confidence level using the Welch degrees of freedom, and the square-root expression is the standard error. The interval’s units are the same as the response variable’s units. The arithmetic estimates a difference in means; the experiment’s random assignment is what supports describing that difference as a treatment effect.
Random assignment and random sampling answer different questions. Random assignment makes the treatment groups comparable on average, supporting cause-and-effect conclusions. Random sampling helps make a sample representative of a larger population, supporting generalization. An experiment can have random assignment without random sampling, so it can support a causal conclusion about its participants without justifying a claim about everyone in a broader population.
In an experiment using volunteers, there may be no random sample from a defined population, so a sampling-based 10% check may not apply. Still, consider whether subjects’ responses can reasonably be treated as independent—for example, whether one subject’s treatment or outcome could influence another’s. If that is doubtful, say so rather than assuming the condition is met.
Worked Examples
Worked Example: A Randomized Study-Skills Experiment
In an invented experiment, 32 volunteers are randomly assigned to use either a focus timer (treatment A) or a standard timer (treatment B) while completing a sorting task. The response is the number of items sorted correctly in 10 minutes. Each group has 16 subjects. The focus-timer group has \(\bar{x}_A=42.6\) and \(s_A=6.0\); the standard-timer group has \(\bar{x}_B=37.2\) and \(s_B=6.0\). The outcome distributions are approximately Normal with no apparent outliers. Construct and interpret a 95% confidence interval for \(\mu_A-\mu_B\).
State: The parameter is \(\mu_A-\mu_B\), the mean number of items sorted correctly under the focus timer minus the mean under the standard timer for the experimental units represented by this study. We want a 95% confidence interval, with focus timer minus standard timer as the subtraction order.
Plan: Use an unpooled two-sample t interval. The groups consist of different subjects, so the observations are unpaired. The subjects were randomly assigned, and the described distributions support the t procedure for both groups. These volunteers were not described as a random sample from a wider population, so random assignment supports a causal conclusion for the experiment but not automatic generalization to all students or workers.
Do: The observed difference in sample means is \(42.6-37.2=5.4\) correctly sorted items. The standard error is:
Because the two sample sizes and standard deviations are equal, the Welch degrees of freedom are \(30\). For a 95% interval, \(t^*\approx2.042\). The margin of error is \(2.042(2.121)\approx4.332\), so the interval is:
Conclude: We are 95% confident that, for the experimental units represented by this study, using the focus timer rather than the standard timer increases the mean number of correctly sorted items by between about 1.07 and 9.73 items in 10 minutes. Because treatment was randomly assigned, the difference can be interpreted causally for these participants. The volunteers were not a random sample, so this experiment alone does not establish the same effect for all people who might use a timer.
Worked Example: An Interval That Includes No Difference
In a second invented experiment, 50 volunteers are randomly assigned to take either a short guided breathing break (treatment A) or a quiet seated break (treatment B) before a memory task. The response is the number of points earned. Each group has 25 participants. The guided-break group has a mean of 72 points and a standard deviation of 10 points; the quiet-break group has a mean of 69 points and a standard deviation of 10 points. Both distributions are roughly symmetric with no apparent outliers. Find and interpret a 95% confidence interval for \(\mu_A-\mu_B\).
State: The parameter is the true mean memory-task score under the guided break minus the true mean score under the quiet break for the experimental units represented by the study, measured in points.
Plan: Use an unpooled two-sample t interval. Different participants are in the two groups, and assignment was random. The sample sizes are below 30, but the described distributions are roughly symmetric without apparent outliers, supporting the procedure. As in the first example, the volunteers are not established as a random sample from a wider population.
Do: The sample mean difference is \(72-69=3\) points. The standard error and degrees of freedom are:
For 95% confidence and 48 degrees of freedom, \(t^*\approx2.011\). The margin of error is \(2.011(2.828)\approx5.687\), giving:
Conclude: We are 95% confident that the mean score under the guided break is between about 2.69 points lower and 8.69 points higher than the mean under the quiet break for the experimental units represented by this study. Since zero is in the interval, the data do not pin down the direction of the mean difference at this confidence level. This does not prove the treatments have the same effect. Random assignment still allows a causal interpretation of the estimated difference, but the interval’s uncertainty leaves several mean differences plausible.
Worked Example: Unequal Group Sizes in a Randomized Experiment
An invented experiment tests a short practice activity before a science quiz. Researchers randomly assign 54 volunteers to the activity or to a usual-review group. The response is quiz points. In the activity group, \(n_A=30\), \(\bar{x}_A=18.4\), and \(s_A=4.5\). In the usual-review group, \(n_B=24\), \(\bar{x}_B=15.1\), and \(s_B=3.6\). Plots show no extreme outliers, and both groups’ distributions are reasonably symmetric. Construct a 95% interval for \(\mu_A-\mu_B\) and state what the experiment supports.
State: The parameter is the mean quiz score under the short practice activity minus the mean quiz score under usual review for the experimental units represented by the study, in points.
Plan: Use an unpooled two-sample t interval. The groups contain different subjects, treatment was randomly assigned, and both distributions are reasonably symmetric without extreme outliers. The sample sizes are 30 and 24; neither group is being treated as a random sample from a larger population. Thus, random assignment supports a causal conclusion for the participants, while generalization beyond them needs additional justification.
Do: The sample mean difference is \(18.4-15.1=3.3\) points. Calculate the standard error from the separate variance contributions:
Welch’s degrees of freedom are approximately \(52.0\), so for a 95% interval \(t^*\approx2.007\). The margin of error is \(2.007(1.102)\approx2.212\). Therefore:
Conclude: We are 95% confident that the short practice activity increases the mean quiz score by about 1.09 to 5.51 points compared with usual review for the experimental units represented by this study. The interval is entirely above zero, and random assignment supports attributing the mean difference to the assigned activity for these participants. The volunteer design does not by itself justify generalizing the estimated effect to all students.
What the Causal Interpretation Does—and Does Not—Say
Random assignment helps prevent pre-existing differences between groups from systematically favoring one treatment. It does not guarantee that the groups will have exactly the same characteristics; chance imbalances can occur. The interval reflects uncertainty in estimating the mean difference, while the randomized design supports attributing that difference to the treatments rather than to a deliberately chosen group difference.
A causal conclusion is about a difference in mean response. It does not say that every individual benefits, that every participant’s response changes by an amount inside the interval, or that the treatment works for every setting. The interval also does not measure how much individual responses vary around their group means.
For an experiment that also uses a random sample from a clearly defined population, the design can support both causal inference and generalization to that population, subject to the study’s conditions. If subjects are volunteers or otherwise not randomly sampled, describe the causal result in relation to the participants studied and avoid claiming broad representativeness.
Common Mistakes and AP Exam Tips
- Confusing assignment with sampling: Random assignment supports cause and effect; random sampling supports generalization. State which process the experiment actually used.
- Reversing the subtraction order: If the interval is for treatment A minus treatment B, describe every endpoint as A’s mean response minus B’s mean response. Reversing the order reverses the signs and switches the endpoints.
- Calling a positive interval “proof”: An interval entirely above zero gives evidence that the mean response under A is higher than under B. In a randomized experiment, the difference can be interpreted as a treatment effect for the experimental units; it is not proof that every person responds positively.
- Interpreting an interval containing zero as proof of no effect: Such an interval includes zero as a plausible mean difference, but it also includes nonzero differences. Say the interval does not establish a clear direction at that confidence level.
- Leaving out units and context: State what is being measured and compare treatment A with treatment B. “The interval is 1 to 10” is not a complete interpretation.
- Claiming universal generalization: A randomized experiment with volunteers can support causal conclusions about its participants, but random assignment alone does not make those participants representative of everyone else.
A strong AP response identifies the parameter and subtraction order, names the two-sample t interval, addresses the design and conditions, shows the standard error and interval, and interprets both endpoints in context. Then explain what random assignment permits you to say causally—and whether the recruitment process supports generalization.
Check Your Understanding
Use the treatment-minus-control order in your answers, and distinguish conclusions about causation from conclusions about generalization.
- A randomized experiment estimates a 95% interval of \((2.4,\ 7.1)\) points for treatment A minus treatment B. Interpret the interval and explain what random assignment contributes.
- An experiment randomly assigns volunteers to two treatments, but does not randomly sample them. What kind of conclusion is supported, and what kind is not automatically supported?
- A 95% interval for a mean treatment difference is \((-1.8,\ 5.2)\) minutes. Does this prove there is no treatment effect? Explain what the interval does say.
- In a two-treatment experiment, the same participants receive both treatments. Should the analysis use an independent two-sample t interval? Explain which design feature matters.
- Why does a confidence interval for a difference in mean response not describe the treatment effect for every individual participant?