Why the Way Groups Are Formed Matters
A two-sample t test can provide evidence that two population means differ. But the test statistic and p-value do not, by themselves, tell us whether a treatment caused the difference. That depends on how the people or other observational units entered the groups.
In a randomized experiment, researchers assign units to treatments using chance. In an observational study, researchers record which group units belong to without assigning that group or treatment. The same numerical difference between group means can therefore support different conclusions, depending on the design.
Earlier, in “Paired t Conclusions About Causation and Generalization,” we distinguished random assignment from random sampling. Here we apply that distinction to an independent two-sample comparison. Random assignment addresses whether a difference can be attributed to a treatment; random sampling addresses whether results can be generalized to a population. One does not automatically provide the other.
What Random Assignment Adds
Suppose an experiment compares Treatment A with Treatment B and measures the same quantitative response for each participant. Let \(\mu_A\) and \(\mu_B\) denote the true mean responses for the groups assigned to A and B in the setting of the experiment. A two-sample t test can assess whether the observed difference in sample means is convincing evidence of a difference in these treatment-group means.
Because assignment is random, characteristics that could affect the response—such as experience, health, or motivation—are not deliberately concentrated in one treatment group. Chance assignment tends to distribute both measured and unmeasured characteristics across groups, though it does not guarantee that every characteristic will be exactly balanced in a particular experiment. This makes a treatment explanation more credible than it would be if participants chose their own treatment.
If the result is statistically significant, and the experiment was carried out appropriately, the difference is evidence that the treatment assignment affected the mean response in the study setting. The causal claim is about the treatments as implemented and the units studied. A small p-value is not proof, and random assignment does not repair problems such as substantial noncompliance, differential loss of participants, or treatment groups influencing one another.
Random assignment alone does not establish that the participants represent a broader population. If volunteers are randomly assigned to treatments, the experiment can support a causal conclusion for those participants, but generalizing to all eligible people requires additional justification. Conversely, a random sample can represent a population well, but if people in the sample chose their own treatments, the comparison is still observational.
| How the study works | What the design can support |
|---|---|
| Random assignment to treatments | A cause-and-effect conclusion about treatment differences, if the experiment is well conducted |
| Random sampling from a population | Generalization to that population, if the sampling process and other conditions are appropriate |
| Both random sampling and random assignment | Potentially both generalization to the sampled population and a causal conclusion |
| Neither process | Usually a cautious description of the studied groups; broad causal or population claims are not justified by design alone |
Why Observational Comparisons Do Not Establish Causation
In an observational study, the groups may differ in ways besides the factor being compared. A confounding variable is a variable associated with group membership and with the response, making it difficult to separate the effect of the group factor from the influence of that variable.
For example, people who choose to use a study app may also spend more time studying. If app users have higher scores, the difference could reflect the app, extra study time, or both. A two-sample t test can measure how surprising a difference like this would be if the population means were equal. It cannot determine which explanation produced the difference.
Statistical significance does not remove confounding. It describes evidence against a null hypothesis under the assumptions of the test; it does not turn self-selection into random assignment. For observational data, a careful conclusion describes a difference or association between the groups, not an effect caused by group membership.
Worked Examples
Worked Example: A Significant Difference in a Randomized Experiment
A hypothetical school recruits 40 volunteers to compare two short reading-practice programs. Researchers randomly assign 20 volunteers to Program A and 20 to Program B. After a fixed period, each student takes the same reading-comprehension assessment. Program A’s group has a mean score of 24 points with a standard deviation of 4 points; Program B’s group has a mean score of 19 points with a standard deviation of 4 points. Higher scores indicate better performance. Suppose the study’s prespecified question is whether Program A produces a higher mean score.
State: Let \(\mu_A\) be the true mean assessment score for participants assigned to Program A in this experiment, and let \(\mu_B\) be the corresponding mean for participants assigned to Program B. The hypotheses are \(H_0:\mu_A-\mu_B=0\) and \(H_a:\mu_A-\mu_B>0\).
Plan and conditions: A two-sample t test is appropriate for comparing the two independent group means. The students were randomly assigned, which supports using the experiment to investigate a causal treatment difference. Each student contributes one score and is in only one group, so the groups are unpaired; assume students’ results do not affect one another. Suppose dotplots show both groups to be roughly symmetric with no strong outliers. These features support the t procedure. The volunteers were not randomly sampled, so the design does not by itself justify generalizing to all students.
Do: The observed difference is \(\bar{x}_A-\bar{x}_B=24-19=5\) points. The standard error is
The test statistic is \(t=5/1.2649\approx3.953\). The Welch degrees of freedom are \(df=38\), since the sample sizes and standard deviations are equal. For the right-tailed alternative, a calculator gives \(p\approx0.0002\), rounded.
Conclude: Since \(p\approx0.0002\) is below a significance level of \(0.05\), there is convincing evidence that the mean assessment score is higher for participants assigned to Program A than for those assigned to Program B. Because treatment was randomly assigned, the difference supports a causal conclusion about the programs for participants in this experiment. The volunteer recruitment limits how confidently we can generalize that effect to a wider student population.
Worked Example: A Significant Difference in an Observational Study
A hypothetical school wants to compare end-of-term mathematics scores for students who chose to enroll in an optional tutoring program and students who did not. From the school records, researchers independently select random samples of 30 students from each group. Tutoring students have a mean score of 78 points with a standard deviation of 10 points; students who did not enroll have a mean of 72 points with a standard deviation of 10 points. Assume the group plots are roughly symmetric without strong outliers, students’ scores can be treated as independent, and each sample is less than 10% of its corresponding group in the records.
Let \(\mu_1\) be the mean score among students who chose tutoring and \(\mu_2\) the mean among students who did not. For a two-sided comparison, use \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). The random samples and stated conditions support using a two-sample t test to compare the means of these two groups.
The difference in sample means is \(78-72=6\) points. The standard error is
Thus \(t=6/2.5820\approx2.324\). With equal sample sizes and standard deviations, the Welch degrees of freedom are \(df=58\). The two-sided p-value is approximately \(0.0237\), rounded. At the \(0.05\) significance level, there is convincing evidence of a difference in mean scores between students who chose tutoring and those who did not.
This is not evidence that tutoring caused the difference. Students chose whether to enroll, so the groups might differ in motivation, prior achievement, available study time, or other factors related to scores. The random samples help represent the two recorded groups, but they do not make tutoring assignment random. The defensible conclusion is an association in mean scores, not a causal effect of tutoring.
Worked Example: Random Assignment but No Significant Difference
In a hypothetical experiment, 32 volunteers are randomly assigned to one of two reminder systems for a month, with 16 participants in each group. The response is the number of minutes spent on a planned activity per day. The new reminder group has a mean of 42 minutes and a standard deviation of 5 minutes; the standard reminder group has a mean of 40 minutes and a standard deviation of 5 minutes. The research question is whether the mean times differ in either direction.
Let \(\mu_1\) and \(\mu_2\) be the true mean daily activity times for participants assigned to the new and standard reminders, respectively. Use \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). Assume each participant contributes one independent response and that both group plots are reasonably symmetric with no strong outliers. Random assignment supports a causal comparison; volunteer recruitment does not establish broad generalizability.
The observed difference is \(42-40=2\) minutes. The standard error is
The test statistic is \(t=2/1.7678\approx1.131\), with Welch \(df=30\). The two-sided p-value is approximately \(0.267\), rounded. Since this exceeds \(0.05\), fail to reject \(H_0\). There is not convincing evidence of a difference in mean daily activity time between participants assigned to the two reminders. Random assignment means a significant result could have supported a causal claim, but this nonsignificant result does not establish that the reminders have identical effects or that there is no effect.
Common Mistakes and AP Exam Tips
When interpreting a two-sample test, state both what the statistical result says and what the design permits you to conclude. Keep the evidence for a mean difference separate from the reasoning about causation and generalization.
- Writing “significant, so it caused the difference” without checking the design: First establish whether treatments were assigned randomly. If people chose their groups, describe an association or difference, not a cause-and-effect result.
- Confusing random assignment with random sampling: Assignment supports causal interpretation; sampling supports generalization. A study may have one, both, or neither.
- Assuming random assignment balances every characteristic exactly: Chance assignment tends to distribute characteristics across groups, but an imbalance can occur. Say it helps control confounding on average; do not claim it guarantees identical groups.
- Assuming a large or small p-value answers every question: A p-value does not diagnose confounding, establish practical importance, prove a causal claim, or show that the sample represents a population. Interpret it together with the design and conditions.
- Claiming that a nonsignificant result proves no effect: The correct test language is “fail to reject” and “there is not convincing evidence of a difference.” The experiment may simply provide limited evidence about the size of any effect.
- Making a broad population claim from volunteers: Random assignment of volunteers can support a causal conclusion for the experiment, but it does not turn volunteers into a random sample of everyone who might use the treatment.
Check Your Understanding
For each situation, separate what the test may show from what the study design allows you to conclude.
- Researchers randomly assign volunteers to two exercise plans. The test finds a significant difference in mean weekly exercise time. What causal conclusion is supported, and what limitation remains if the volunteers were not randomly sampled?
- People choose whether to use a budgeting app. App users have a significantly higher mean savings amount. Why does the significant result not establish that the app caused the difference?
- A study randomly samples students from a school but compares students who chose two different courses. Which kind of inference does random sampling support, and why is a causal claim still unwarranted?
- In a randomized experiment, a two-sided test has \(p=0.31\). What is an appropriate conclusion, and what should you avoid claiming?
- Can random assignment guarantee that the two groups have identical prior experience? Explain how random assignment helps while avoiding an overstatement.