Tutorials › AP Statistics › Paired t Conclusions About Causation and Generalization

Paired data and paired t procedures · Tutorial 699 of 1000

Paired t Conclusions About Causation and Generalization

Use the study design—not the paired t result alone—to decide whether a conclusion can be generalized, interpreted causally, both, or neither.

Intermediate 9 min read

What You'll Learn

  • Distinguish random sampling from random assignment in a paired study
  • Decide whether a paired t result supports a causal conclusion
  • Identify the population, if any, to which results can be generalized
  • Evaluate studies with only random sampling, only random assignment, both, or neither
  • Write conclusions that match the study design and the statistical evidence

What a Paired t Result Can—and Cannot—Tell You

A paired t procedure analyzes the differences within pairs. As “Common Mistakes With Paired t Procedures” emphasizes, the calculation uses one consistently defined difference per pair. But even a correctly calculated t statistic and p-value do not, by themselves, tell you whether a result applies to a wider population or whether one condition caused a change. Those conclusions depend on how the data were collected.

Two design features matter especially: random sampling and random assignment. Random sampling is about who enters the study. Random assignment is about which treatment or condition participants receive, or the order in which they receive conditions. They serve different purposes and are not interchangeable.

Key idea: Random sampling supports generalizing from the sample to the population from which it was selected. Random assignment supports a cause-and-effect conclusion about the treatments being compared. A paired t result does not create either kind of support if the design did not provide it.

Two Design Features, Two Different Conclusions

A random sample is selected using a chance process from a defined population. If the sampling process is carried out well, the sample can represent that population, supporting generalization from the sample to that population. The population must be named carefully: a sample drawn from volunteers at one school does not automatically represent all students.

Random assignment uses chance to allocate treatments or treatment order. In a paired experiment, participants may each receive both conditions, with the order randomly assigned. If the study is otherwise well conducted, random assignment helps make treatment conditions comparable and supports attributing a difference to the treatment rather than to pre-existing differences between groups. In a crossover study, the treatment order, possible period effects, and carryover must also be considered.

A design can have either feature, both, or neither. Randomly selecting participants does not mean they were randomly assigned to treatments. Randomly assigning treatments does not mean the participants were randomly selected from a broad population.

Design featureWhat it supportsWhat it does not establish by itself
Random sampleGeneralizing to the population from which the sample was selectedThat a treatment caused the observed difference
Random assignmentA cause-and-effect conclusion about the treatments for the experimental unitsGeneralizing to a broad population that was not randomly sampled
BothPotentially, a causal conclusion that generalizes to the sampled-from populationThat the study was free of every source of bias or limitation
NeitherDescribing the observed sample and its paired differencesBroad generalization or a secure causal conclusion

This table describes what the design supports, not what a small p-value proves. As in earlier tutorials on significance tests, a small p-value can provide convincing evidence against a null hypothesis under the test’s assumptions. It cannot repair biased sampling, create random assignment, or show that results apply to people who were not represented in the study.

Worked Example: Random Sample and Randomized Treatment Order

Worked Example: Two Study Conditions for Commuters

A transit researcher randomly selects 30 people from a register of 300 regular bus commuters. Each commuter completes a short attention task once after listening to a particular audio track and once in silence. The order is randomly assigned, and the sessions are held on separate days. Let \(d=\text{errors after audio}-\text{errors in silence}\). A paired t test reports a two-sided p-value of \(0.012\) for testing \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\), at \(\alpha=0.05\).

Solution. The p-value is below \(0.05\), so the test rejects \(H_0\). The data provide convincing evidence of a nonzero mean difference in task errors between the audio and silence conditions. Because \(d\) is audio minus silence, a negative mean difference would mean fewer errors after audio, while a positive mean difference would mean more.

The study uses random sampling from the register of regular bus commuters. If the sample was obtained as described and the paired t conditions are reasonable, the evidence can be generalized to the population represented by that register—not automatically to all commuters or all people.

The study also randomly assigns the order of the two conditions. That supports a cause-and-effect interpretation of the difference between listening to the audio track and working in silence, provided the crossover design is conducted appropriately. For example, the sessions should be arranged so that the first condition does not unduly affect performance in the second. Here, the design supports both generalization to the sampled-from commuter population and a causal conclusion about the conditions, within the study’s limits.

Worked Example: Random Assignment Without a Random Sample

Worked Example: Two Interfaces Tested by Volunteers

Twenty-four students volunteer from one high school to compare two versions of a vocabulary app. Every student uses both versions, with the order randomly assigned. The outcome is the number of correctly answered questions in a timed practice session. Define \(d=\text{score with Version B}-\text{score with Version A}\). A paired t test reports a one-sided p-value of \(0.028\) for \(H_a:\mu_d>0\), using \(\alpha=0.05\).

Solution. Since \(0.028<0.05\), the test rejects \(H_0:\mu_d=0\). The data provide convincing evidence that Version B produces a higher mean score than Version A for the population to which this experiment’s inference applies. The positive direction matches the definition of \(d\).

The volunteers were not randomly sampled from all high-school students. Therefore, the result does not automatically generalize to all high-school students, students in other schools, or all app users. The most direct description is about the students who volunteered, with any broader application requiring caution.

The order of the two versions was randomly assigned. That supports attributing a difference in scores to the app version rather than to a fixed order in which everyone used the apps, assuming the sessions and other study conditions were suitable. Thus, random assignment supports a causal conclusion about the comparison for these experimental participants; it does not make the volunteer group representative of a larger population.

Worked Example: Random Sampling Without Random Assignment

Worked Example: Garden Measurements Before and After a City Advisory

A city analyst randomly selects 40 gardens from a register of 400 community gardens. For each garden, the analyst records weekly water use before and after a city conservation advisory is introduced. Let \(d=\text{after}-\text{before}\), in liters per week. A paired t test reports a one-sided p-value of \(0.020\) for \(H_a:\mu_d<0\), at \(\alpha=0.05\).

Solution. Since \(0.020<0.05\), the test rejects \(H_0:\mu_d=0\). The data provide convincing evidence that the mean after-minus-before water-use difference is below zero for the population of gardens represented by the register. In context, that is evidence of a decrease in mean weekly water use after the advisory.

Because the gardens were randomly sampled from the register, generalization to the population represented by that register is supported, assuming the sampling and paired t conditions are reasonable. The conclusion should not automatically extend to gardens outside that population.

The analyst did not randomly assign gardens to receive the advisory or to remain without it. The before-and-after pairing controls for some garden-to-garden variation, but it does not rule out other explanations for the change. Weather, seasonal patterns, or another water-conservation effort could have affected use during the same period. The study supports evidence of a mean change, but it does not establish that the advisory caused the decrease.

Matching the Conclusion to the Design

When reading a paired study, separate the statistical conclusion from the design conclusion. First determine what the paired t test says about the mean difference. Then ask two distinct questions: Was there random sampling from a defined population? Was there random assignment of treatments or order?

1
Identify the sampling process.
Was there a random sample from a specific population? If so, name that population as the one to which generalization may be supported. If not, limit generalization and describe the sample or the population actually represented.
2
Identify the assignment process.
Were treatments or treatment order randomly assigned? If so, a causal interpretation may be supported, provided the experiment was appropriately conducted. If not, describe an observed difference or association rather than claiming that a treatment caused it.
3
Combine the two judgments.
Random sampling and random assignment can support both generalization and causation. Having only one supports only the corresponding conclusion. Having neither calls for a conclusion limited to the observed data, with appropriate caution.

In paired studies, the pair structure is important for analyzing variability, but it does not substitute for either design feature. Measuring the same units twice can reduce the effect of differences between units. It cannot make a convenience sample representative, and it cannot turn an observational before-and-after comparison into a randomized experiment.

Common Mistakes and AP Exam Tips

  • Claiming that random assignment makes a sample representative: Random assignment concerns treatment allocation, not how participants were selected. State which population, if any, was randomly sampled.
  • Claiming that random sampling proves causation: A random sample supports generalization, but it does not rule out confounding in a treatment comparison. Look for random assignment before making a causal claim.
  • Calling every before-and-after change a treatment effect: Pairing links each unit’s measurements, but time-related changes or other events could explain the difference if the treatment was not randomly assigned.
  • Generalizing beyond the sampling frame: Name the population from which the sample was selected. Do not quietly expand a sample of volunteers from one school into a claim about all students.
  • Letting a small p-value do the work of the design: Statistical significance addresses evidence about a parameter under the test assumptions. It does not establish representativeness or prove a causal mechanism.
  • Writing a conclusion that is too broad: A full-credit conclusion connects the test result to the mean difference, uses the study’s population limits, and makes a causal claim only when random assignment supports it.

A careful AP response might say: “The paired t test provides convincing evidence that the mean difference is not zero for the population represented by the random sample. Because treatment order was randomly assigned, the study also supports a causal interpretation of the treatment comparison.” If the sample was not random, remove the unsupported population generalization. If treatment was not randomly assigned, describe evidence of a change or difference without saying the treatment caused it.

Key takeaway: Random sampling answers “To whom can the result generalize?” Random assignment answers “Can the difference be attributed to the treatment?” Evaluate these separately, then keep the paired t conclusion within the limits of both the statistical evidence and the study design.

Check Your Understanding

For each situation, decide what random sampling and random assignment support, and state an appropriately limited conclusion.

  1. A random sample of 50 clinic patients completes two paired treatment conditions in randomly assigned order. Which design features support generalization and causation?
  2. Twenty volunteers from one sports team are randomly assigned to treatment order in a paired recovery study. Can the result automatically be generalized to all athletes? Explain.
  3. A random sample of lakes is measured before and after a regional policy begins, but no lakes are randomly assigned to the policy. What can random sampling support, and what causal claim is not established?
  4. A study uses neither random sampling nor random assignment, but its paired t test has a small p-value. What can the p-value support, and what does it not fix?
  5. Why does measuring the same individuals twice not, by itself, establish that the second condition caused a change?