Significance Is Not the Same as Causation
A statistically significant result can be compelling evidence that two population means differ. But significance alone does not tell us why they differ. In particular, when people or other study units choose their own groups rather than being randomly assigned, a difference may reflect pre-existing differences between the groups, not an effect of the condition being studied.
As in earlier tutorials on two-sample t tests, a small p-value can lead us to reject a null hypothesis of equal population means. The interpretation must also account for how the data were collected. A test describes evidence about a difference; the study design determines whether that difference can be attributed to a treatment.
An observational study records what happens to units in groups that already exist or that the units choose for themselves. Researchers do not assign the explanatory condition. A randomized experiment assigns units to treatment groups by chance. As introduced in “Randomized Experiment Versus Observational Comparison,” random assignment helps make the groups comparable at the start, on average, including with respect to factors researchers have not measured.
A confounding variable is a factor related to both the group or condition being compared and the response. It can make it difficult to separate the relationship between the condition and the response from the effect of that other factor. Random assignment helps guard against confounding; a statistically significant p-value does not remove it.
What a Significant Observational Comparison Supports
Suppose an observational study compares a quantitative response in two naturally occurring groups. A two-sample t test may assess whether the data provide convincing evidence of a difference between the groups’ population means, when the sampling design and t-procedure conditions support that inference. If the p-value is at or below the chosen \(\alpha\), we reject the null hypothesis of equal means.
The conclusion should describe a difference or association between the group means. It should not say that membership in one group caused the difference. For example, a finding that students who attend tutoring have higher mean scores does not, by itself, show that tutoring raised their scores. Students who attend may differ in prior achievement, motivation, available time, or other factors.
A small p-value does not measure the size or practical importance of a difference, either. It describes how unusual a test statistic at least as extreme as the observed statistic would be if the null hypothesis were true. As in “Interpreting a P-Value From a t Test in Context,” it is not the probability that the null hypothesis is true, nor the probability that a causal explanation is correct.
Worked Example: A Significant Difference in Tutoring Outcomes
Worked Example: A Significant Difference in Tutoring Outcomes
A fictional school randomly selects 50 students who attended optional mathematics tutoring and, separately, 50 students who did not attend. The students chose whether to attend; the school did not assign tutoring. At the end of the term, the tutoring group’s mean exam score is 82 points with a standard deviation of 10 points. The non-tutoring group’s mean is 77 points with a standard deviation of 10 points. Assume each group contains more than 500 students, and plots show no severe skewness or extreme outliers. Test for a difference in population mean scores at \(\alpha=0.05\).
Let \(\mu_1\) be the true mean exam score, in points, for students at this school who attended optional mathematics tutoring during the term. Let \(\mu_2\) be the true mean exam score, in points, for students at this school who did not attend. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
Use an unpooled two-sample t test. The students were randomly sampled separately from the two groups, so the samples are random. Each sample is less than 10% of its group, supporting the 10% condition. The groups contain different students, and the plots show no severe skewness or extreme outliers; with 50 students per group, a t procedure is reasonable. However, tutoring attendance was not randomly assigned, so the design does not support a causal conclusion.
The estimated difference is \(82-77=5\) points. The standard error and test statistic are
The data provide convincing evidence of a difference in mean exam scores between students who attended optional mathematics tutoring and those who did not at this school. The observed difference favors the tutoring group, but because students chose whether to attend, this study does not establish that tutoring caused the higher scores.
Prior achievement could be a confounding variable: it may be related to students’ decisions to seek tutoring and to their exam scores. Motivation or available study time could also matter. The significant result is evidence of a group difference, but the test cannot determine which explanation produced it.
Worked Example: A Non-Significant Result Is Not Proof of No Effect
Worked Example: A Non-Significant Result Is Not Proof of No Effect
In a fictional observational survey, researchers randomly sample 36 adults who report using a phone during the hour before bed and 36 adults who do not. The adults chose their own habits. The first group reports a mean of 5.8 hours of sleep per night, with a standard deviation of 1.2 hours; the second reports 6.1 hours, also with a standard deviation of 1.2 hours. Assume each group is large enough for the 10% condition, the samples are independent, and plots show no severe skewness or extreme outliers. Test for a difference in population mean sleep time at \(\alpha=0.05\).
State: Let \(\mu_1\) be the true mean sleep time, in hours per night, for adults in the target setting who use a phone during the hour before bed. Let \(\mu_2\) be the true mean sleep time for adults there who do not. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).
Plan and check conditions: Use an unpooled two-sample t test. The adults were randomly sampled from the two naturally occurring groups; the samples are independent, and each is less than 10% of its group. The sample sizes and plots support using a t procedure. Since phone habits were not assigned, this is an observational comparison and cannot establish causation.
Do: The observed difference is \(5.8-6.1=-0.3\) hours. The standard error and test statistic are
The Welch degrees of freedom are \(df=70\), and the two-sided p-value is approximately \(0.293\), rounded. Since \(0.293>0.05\), fail to reject \(H_0\).
Conclude: The data do not provide convincing evidence of a difference in mean sleep time between adults who use a phone during the hour before bed and those who do not in the target setting. This result does not prove that the population means are equal, and it does not show that phone use has no effect. The study is observational, so even a significant result would not establish a causal effect.
For instance, work schedules could be related both to late phone use and to sleep duration. A non-significant test does not remove possible confounding, just as a significant test does not.
Worked Example: Why Random Assignment Changes the Conclusion
Worked Example: Why Random Assignment Changes the Conclusion
A fictional greenhouse randomly assigns 60 similar seedlings to one of two lighting conditions, with 30 seedlings in each group. After a fixed growing period, seedlings under a new lamp have a mean height of 18.4 centimeters and a standard deviation of 3 centimeters. Seedlings under the usual lamp have a mean height of 16.0 centimeters and a standard deviation of 3 centimeters. Assume the measurements in each group show no severe skewness or extreme outliers. Test at \(\alpha=0.05\) whether the lighting conditions caused a difference in mean height for these experimental seedlings.
State: Let \(\mu_1\) be the mean height, in centimeters, after the growing period for the experimental seedlings assigned to the new lamp. Let \(\mu_2\) be the mean height for the experimental seedlings assigned to the usual lamp. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
Plan and check conditions: Use an unpooled two-sample t test. The seedlings were randomly assigned, meeting the random-assignment condition. Each seedling received one lighting condition, so the groups are independent; the plots show no severe skewness or extreme outliers. These conditions support the t procedure. Unlike the tutoring and phone examples, this randomized experiment can support a cause-and-effect conclusion about the lighting treatments for the experimental seedlings.
Do: The observed difference is \(18.4-16.0=2.4\) centimeters. The standard error and test statistic are
The Welch degrees of freedom are \(df=58\). The two-sided p-value is approximately \(0.003\), rounded. Since \(0.003<0.05\), reject \(H_0\).
Conclude: The data provide convincing evidence that the lighting conditions caused a difference in mean height for the experimental seedlings. Because lighting condition was randomly assigned, the experiment provides evidence that the new lighting condition caused a difference in mean height for the seedlings in this experiment. This causal conclusion comes from the design, not simply from the small p-value.
Common Mistakes and AP Exam Tips
- Writing “caused” after any significant result: A small p-value is not a substitute for random assignment. In an observational study, say the data provide evidence of a difference or association between groups.
- Confusing random sampling with random assignment: Random sampling concerns how units are selected; random assignment concerns how units are placed into groups. A study can use random samples and still be observational.
- Assuming a test accounts for confounding: A two-sample t test measures evidence about a difference in means under its assumptions. It does not control for differences such as prior achievement, health, or motivation that were not part of the analysis.
- Claiming no difference after failing to reject: Failure to reject means the data do not provide convincing evidence for the alternative at the chosen \(\alpha\). It does not prove equal means or establish that a treatment has no effect.
- Calling a result important because it is significant: Statistical significance is a decision based on a p-value and \(\alpha\). It does not, by itself, describe whether the difference is large or practically meaningful.
- Leaving the design out of the conclusion: A full-credit conclusion names the comparison and response in context, states the reject-or-fail-to-reject decision, and limits causal language to a design with random assignment.
A useful final check is to ask, “Who chose or assigned the groups?” If participants chose their groups or researchers merely recorded existing groups, describe an association. If researchers randomly assigned treatments and the experiment was conducted appropriately, a significant difference can support a causal conclusion about the treatments.
Check Your Understanding
For each situation, distinguish what the test result says from what the study design allows you to conclude.
- A random sample finds a significant difference in mean weekly exercise time between people who voluntarily use a fitness app and those who do not. Can the researchers conclude that using the app caused the difference? Explain.
- In an observational study, adults who choose a nutrition course have a significantly different mean blood pressure from adults who do not take the course. Name one possible confounding variable and explain how it might be related to both group membership and blood pressure.
- Explain the difference between random sampling and random assignment in one or two sentences.
- A two-sample t test in an observational study has \(p=0.21\) at \(\alpha=0.05\). What decision should be made, and why would it be incorrect to say the population means are proved equal?
- Researchers randomly assign plants to two fertilizers and find a significant difference in mean growth. Which feature of the study supports a causal conclusion: the small p-value or the random assignment? Explain.