Tutorials › AP Statistics › Randomized Experiment Versus Observational Comparison

Two-sample t hypothesis tests · Tutorial 731 of 1000

Randomized Experiment Versus Observational Comparison

Distinguish evidence of a treatment effect from an association between groups by asking how people entered those groups.

Intermediate 10 min read

What You'll Learn

  • Explain how random assignment supports a cause-and-effect conclusion about a treatment.
  • Distinguish random assignment from random sampling and identify what each allows.
  • Explain why a significant observational comparison does not establish causation.
  • Apply the distinction when interpreting two-sample t test results.
  • Identify how confounding, participant selection, and study limitations affect conclusions.

Why the Way Groups Are Formed Matters

A two-sample t test can provide evidence that two population means differ. But the test statistic and p-value do not, by themselves, tell us whether a treatment caused the difference. That depends on how the people or other observational units entered the groups.

In a randomized experiment, researchers assign units to treatments using chance. In an observational study, researchers record which group units belong to without assigning that group or treatment. The same numerical difference between group means can therefore support different conclusions, depending on the design.

Earlier, in “Paired t Conclusions About Causation and Generalization,” we distinguished random assignment from random sampling. Here we apply that distinction to an independent two-sample comparison. Random assignment addresses whether a difference can be attributed to a treatment; random sampling addresses whether results can be generalized to a population. One does not automatically provide the other.

Definition: Random assignment uses chance to place study units into treatment groups. When a well-designed randomized experiment finds a statistically significant difference in mean response, that result can support a cause-and-effect conclusion about the treatments. Random sampling uses chance to select units from a population; it supports generalizing results to that population. Neither process should be mistaken for the other.

What Random Assignment Adds

Suppose an experiment compares Treatment A with Treatment B and measures the same quantitative response for each participant. Let \(\mu_A\) and \(\mu_B\) denote the true mean responses for the groups assigned to A and B in the setting of the experiment. A two-sample t test can assess whether the observed difference in sample means is convincing evidence of a difference in these treatment-group means.

Because assignment is random, characteristics that could affect the response—such as experience, health, or motivation—are not deliberately concentrated in one treatment group. Chance assignment tends to distribute both measured and unmeasured characteristics across groups, though it does not guarantee that every characteristic will be exactly balanced in a particular experiment. This makes a treatment explanation more credible than it would be if participants chose their own treatment.

If the result is statistically significant, and the experiment was carried out appropriately, the difference is evidence that the treatment assignment affected the mean response in the study setting. The causal claim is about the treatments as implemented and the units studied. A small p-value is not proof, and random assignment does not repair problems such as substantial noncompliance, differential loss of participants, or treatment groups influencing one another.

Random assignment alone does not establish that the participants represent a broader population. If volunteers are randomly assigned to treatments, the experiment can support a causal conclusion for those participants, but generalizing to all eligible people requires additional justification. Conversely, a random sample can represent a population well, but if people in the sample chose their own treatments, the comparison is still observational.

How the study worksWhat the design can support
Random assignment to treatmentsA cause-and-effect conclusion about treatment differences, if the experiment is well conducted
Random sampling from a populationGeneralization to that population, if the sampling process and other conditions are appropriate
Both random sampling and random assignmentPotentially both generalization to the sampled population and a causal conclusion
Neither processUsually a cautious description of the studied groups; broad causal or population claims are not justified by design alone

Why Observational Comparisons Do Not Establish Causation

In an observational study, the groups may differ in ways besides the factor being compared. A confounding variable is a variable associated with group membership and with the response, making it difficult to separate the effect of the group factor from the influence of that variable.

For example, people who choose to use a study app may also spend more time studying. If app users have higher scores, the difference could reflect the app, extra study time, or both. A two-sample t test can measure how surprising a difference like this would be if the population means were equal. It cannot determine which explanation produced the difference.

Statistical significance does not remove confounding. It describes evidence against a null hypothesis under the assumptions of the test; it does not turn self-selection into random assignment. For observational data, a careful conclusion describes a difference or association between the groups, not an effect caused by group membership.

Key distinction: A significant two-sample t test addresses whether the data provide convincing evidence of a difference in mean responses. The study design determines whether that difference can reasonably be attributed to a treatment. Random assignment can support that attribution; an observational comparison alone cannot.

Worked Examples

Worked Example: A Significant Difference in a Randomized Experiment

A hypothetical school recruits 40 volunteers to compare two short reading-practice programs. Researchers randomly assign 20 volunteers to Program A and 20 to Program B. After a fixed period, each student takes the same reading-comprehension assessment. Program A’s group has a mean score of 24 points with a standard deviation of 4 points; Program B’s group has a mean score of 19 points with a standard deviation of 4 points. Higher scores indicate better performance. Suppose the study’s prespecified question is whether Program A produces a higher mean score.

State: Let \(\mu_A\) be the true mean assessment score for participants assigned to Program A in this experiment, and let \(\mu_B\) be the corresponding mean for participants assigned to Program B. The hypotheses are \(H_0:\mu_A-\mu_B=0\) and \(H_a:\mu_A-\mu_B>0\).

Plan and conditions: A two-sample t test is appropriate for comparing the two independent group means. The students were randomly assigned, which supports using the experiment to investigate a causal treatment difference. Each student contributes one score and is in only one group, so the groups are unpaired; assume students’ results do not affect one another. Suppose dotplots show both groups to be roughly symmetric with no strong outliers. These features support the t procedure. The volunteers were not randomly sampled, so the design does not by itself justify generalizing to all students.

Do: The observed difference is \(\bar{x}_A-\bar{x}_B=24-19=5\) points. The standard error is

$$ SE_{\bar{x}_A-\bar{x}_B} =\sqrt{\frac{4^2}{20}+\frac{4^2}{20}} =\sqrt{1.6} \approx 1.2649 $$

The test statistic is \(t=5/1.2649\approx3.953\). The Welch degrees of freedom are \(df=38\), since the sample sizes and standard deviations are equal. For the right-tailed alternative, a calculator gives \(p\approx0.0002\), rounded.

Conclude: Since \(p\approx0.0002\) is below a significance level of \(0.05\), there is convincing evidence that the mean assessment score is higher for participants assigned to Program A than for those assigned to Program B. Because treatment was randomly assigned, the difference supports a causal conclusion about the programs for participants in this experiment. The volunteer recruitment limits how confidently we can generalize that effect to a wider student population.

Worked Example: A Significant Difference in an Observational Study

A hypothetical school wants to compare end-of-term mathematics scores for students who chose to enroll in an optional tutoring program and students who did not. From the school records, researchers independently select random samples of 30 students from each group. Tutoring students have a mean score of 78 points with a standard deviation of 10 points; students who did not enroll have a mean of 72 points with a standard deviation of 10 points. Assume the group plots are roughly symmetric without strong outliers, students’ scores can be treated as independent, and each sample is less than 10% of its corresponding group in the records.

Let \(\mu_1\) be the mean score among students who chose tutoring and \(\mu_2\) the mean among students who did not. For a two-sided comparison, use \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). The random samples and stated conditions support using a two-sample t test to compare the means of these two groups.

The difference in sample means is \(78-72=6\) points. The standard error is

$$ SE_{\bar{x}_1-\bar{x}_2} =\sqrt{\frac{10^2}{30}+\frac{10^2}{30}} =\sqrt{\frac{200}{30}} \approx 2.5820 $$

Thus \(t=6/2.5820\approx2.324\). With equal sample sizes and standard deviations, the Welch degrees of freedom are \(df=58\). The two-sided p-value is approximately \(0.0237\), rounded. At the \(0.05\) significance level, there is convincing evidence of a difference in mean scores between students who chose tutoring and those who did not.

This is not evidence that tutoring caused the difference. Students chose whether to enroll, so the groups might differ in motivation, prior achievement, available study time, or other factors related to scores. The random samples help represent the two recorded groups, but they do not make tutoring assignment random. The defensible conclusion is an association in mean scores, not a causal effect of tutoring.

Worked Example: Random Assignment but No Significant Difference

In a hypothetical experiment, 32 volunteers are randomly assigned to one of two reminder systems for a month, with 16 participants in each group. The response is the number of minutes spent on a planned activity per day. The new reminder group has a mean of 42 minutes and a standard deviation of 5 minutes; the standard reminder group has a mean of 40 minutes and a standard deviation of 5 minutes. The research question is whether the mean times differ in either direction.

Let \(\mu_1\) and \(\mu_2\) be the true mean daily activity times for participants assigned to the new and standard reminders, respectively. Use \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). Assume each participant contributes one independent response and that both group plots are reasonably symmetric with no strong outliers. Random assignment supports a causal comparison; volunteer recruitment does not establish broad generalizability.

The observed difference is \(42-40=2\) minutes. The standard error is

$$ SE_{\bar{x}_1-\bar{x}_2} =\sqrt{\frac{5^2}{16}+\frac{5^2}{16}} =\sqrt{3.125} \approx 1.7678 $$

The test statistic is \(t=2/1.7678\approx1.131\), with Welch \(df=30\). The two-sided p-value is approximately \(0.267\), rounded. Since this exceeds \(0.05\), fail to reject \(H_0\). There is not convincing evidence of a difference in mean daily activity time between participants assigned to the two reminders. Random assignment means a significant result could have supported a causal claim, but this nonsignificant result does not establish that the reminders have identical effects or that there is no effect.

Common Mistakes and AP Exam Tips

When interpreting a two-sample test, state both what the statistical result says and what the design permits you to conclude. Keep the evidence for a mean difference separate from the reasoning about causation and generalization.

  • Writing “significant, so it caused the difference” without checking the design: First establish whether treatments were assigned randomly. If people chose their groups, describe an association or difference, not a cause-and-effect result.
  • Confusing random assignment with random sampling: Assignment supports causal interpretation; sampling supports generalization. A study may have one, both, or neither.
  • Assuming random assignment balances every characteristic exactly: Chance assignment tends to distribute characteristics across groups, but an imbalance can occur. Say it helps control confounding on average; do not claim it guarantees identical groups.
  • Assuming a large or small p-value answers every question: A p-value does not diagnose confounding, establish practical importance, prove a causal claim, or show that the sample represents a population. Interpret it together with the design and conditions.
  • Claiming that a nonsignificant result proves no effect: The correct test language is “fail to reject” and “there is not convincing evidence of a difference.” The experiment may simply provide limited evidence about the size of any effect.
  • Making a broad population claim from volunteers: Random assignment of volunteers can support a causal conclusion for the experiment, but it does not turn volunteers into a random sample of everyone who might use the treatment.
Key takeaway: A two-sample t test evaluates evidence about a difference in mean responses. Random assignment is the design feature that can make a significant treatment-group difference support causation. An observational comparison, even with a small p-value, supports an association rather than a cause-and-effect conclusion.

Check Your Understanding

For each situation, separate what the test may show from what the study design allows you to conclude.

  1. Researchers randomly assign volunteers to two exercise plans. The test finds a significant difference in mean weekly exercise time. What causal conclusion is supported, and what limitation remains if the volunteers were not randomly sampled?
  2. People choose whether to use a budgeting app. App users have a significantly higher mean savings amount. Why does the significant result not establish that the app caused the difference?
  3. A study randomly samples students from a school but compares students who chose two different courses. Which kind of inference does random sampling support, and why is a causal claim still unwarranted?
  4. In a randomized experiment, a two-sided test has \(p=0.31\). What is an appropriate conclusion, and what should you avoid claiming?
  5. Can random assignment guarantee that the two groups have identical prior experience? Explain how random assignment helps while avoiding an overstatement.