Let the Study Design Set the Limits
In “Writing a Complete Inference Response,” you practiced connecting a statistical conclusion to the question, procedure, and context. One more check is essential: does the study design support the kind of claim the conclusion makes? A result may describe an association, support a cause-and-effect conclusion, or allow generalization to a population. Those are different claims, and a study does not automatically support all three.
Two features of a study are especially important. Random sampling concerns how individuals enter the study; a probability sample can support generalizing to the population from which the sample was drawn. Random assignment concerns how participants are placed into treatment groups; it can support a cause-and-effect conclusion about the treatments for the participants studied. Neither feature does the work of the other.
An association means that values or outcomes differ in a pattern across groups or variables. In an observational study, researchers measure variables without assigning treatments. The results can show an association, but other variables may help explain it. For example, an association between exercise and sleep might also reflect age, health, or work schedules. A possible alternative explanation is not proof that confounding occurred; it is a reason to avoid claiming that one measured variable caused the other.
In an experiment, researchers impose treatments. If participants are randomly assigned to treatments, differences in outcomes can be attributed to the treatments more plausibly because the assignment process helps balance other factors across groups. Random assignment supports a causal interpretation; it does not make every observed difference automatically convincing evidence of a treatment effect. The size and uncertainty of the observed results still matter.
Generalization asks a separate question: who does the sample represent? A random sample from a defined population can support conclusions about that population, subject to the study’s other limitations. A sample of volunteers, one classroom, or customers who chose to respond does not automatically represent a wider population. Even a carefully randomized experiment with volunteers does not, just by being an experiment, establish that its results apply to everyone.
A Quick Design-to-Conclusion Map
Before writing a conclusion, identify how people were selected and how groups were formed. The following map summarizes what those features can support. It describes the design’s potential, not a guarantee that the evidence is strong enough for a particular claim.
| How participants were selected | How groups were formed | Conclusion the design can support |
|---|---|---|
| Random sample from a defined population | Variables observed; no treatment assigned | An association in the population may be supported; a cause-and-effect claim is not. |
| Volunteers or another nonrandom group | Random assignment to treatments | A cause-and-effect conclusion may be supported for the participants; broad generalization is not automatic. |
| Random sample from a defined population | Random assignment to treatments | Causal conclusions may be supported for the population represented by the sample, if the evidence warrants them. |
| Volunteers or another nonrandom group | Variables observed; no treatment assigned | Describe the observed association among those studied; do not claim causation or broad generalization. |
These distinctions also help you spot overclaiming. “The variables were associated in the sample” is narrower than “one variable caused the other.” “The participants improved” is narrower than “the treatment works for all teenagers.” Good statistical wording makes the scope visible.
Worked Example: Sleep and Quiz Scores
Worked Example: What Can a School Survey Say?
Scenario. A fictional school randomly selects 120 students from its enrollment of 2,400. Students report whether they usually sleep at least eight hours on school nights, and the school records a recent quiz score. In the sample, 50 students report at least eight hours of sleep, with a mean score of 84 points. The other 70 students have a mean score of 78 points.
What the results show. The difference between the two sample means is \(84-78=6\) points. In this sample, students reporting at least eight hours of sleep had a mean quiz score 6 points higher than the mean for students reporting less sleep. That is a description of an observed association between reported sleep category and quiz score.
Weak conclusion. “Getting eight hours of sleep causes students to score 6 points higher on quizzes, and this is true for teenagers everywhere.” This sentence makes two unsupported leaps. Sleep was measured, not assigned, so the study cannot establish that sleep caused the score difference. The random sample came from one school, so it does not represent all teenagers.
Stronger conclusion. “Among the 120 students randomly sampled from this school, those who reported at least eight hours of sleep had a sample mean quiz score 6 points higher than those who reported less sleep. The survey shows an association, not that sleep caused the difference. Because the students were randomly sampled from the school, the results may support generalizing the association to this school’s enrollment, but not to teenagers generally.”
Why this wording fits. Random selection supports generalizing to the school population, while the observational design supports describing an association rather than asserting causation. The arithmetic reports a sample difference; without an inference analysis, it does not establish that the population difference is statistically convincing. A careful answer does not turn a sample result into stronger evidence than was calculated.
Worked Example: A Randomized App Trial
Worked Example: What Can a Volunteer Experiment Say?
Scenario. Sixty students volunteer for a fictional study of a study-planning app. Researchers randomly assign 30 students to use the app for four weeks and 30 to follow their usual study routine. At the end, the app group has a mean quiz score of 82 points and the usual-routine group has a mean of 77 points. The observed difference is \(82-77=5\) points.
What the design supports. The students were volunteers, not a random sample of all students. However, they were randomly assigned to the two study conditions. That assignment supports a causal interpretation of a difference between the groups for these participants, provided the study was carried out as planned and the outcome comparison is credible. The observed 5-point difference is a sample result; by itself, it does not tell us whether there is convincing evidence of a treatment effect beyond chance variation.
Weak conclusion. “The app raises quiz scores by 5 points for students everywhere.” This overstates both the numerical result and the scope. Five points is the observed difference between sample means, not necessarily the treatment’s exact effect. And the volunteers do not automatically represent all students.
Stronger conclusion. “In this experiment, the volunteers assigned to use the app had a sample mean quiz score 5 points higher than the volunteers assigned to their usual routine. Because the researchers randomly assigned treatment, the experiment can support a cause-and-effect interpretation for these participants if the analysis provides convincing evidence of a difference. The volunteer sample does not justify generalizing the result to all students.”
Why this wording fits. It gives random assignment its proper role without treating it as a guarantee of a real or important effect. It also keeps the population claim limited to the people studied. To report convincing evidence, the writer would need appropriate statistical evidence, not just the two means.
Worked Example: A Random Sample and Random Assignment
Worked Example: A Community Reminder Program
Scenario. A fictional community center has 1,000 adult members. Researchers randomly select 100 members to evaluate a reminder program. They randomly assign 50 to receive the reminders and 50 to receive the center’s usual communications. By the end of the month, 36 people in the reminder group and 24 in the usual-communications group attend a scheduled class.
What the results show. The sample attendance proportions are \(36/50=0.72\), or 72%, for the reminder group and \(24/50=0.48\), or 48%, for the usual-communications group. The observed difference is \(0.72-0.48=0.24\), or 24 percentage points. These calculations describe the sample; they do not, by themselves, establish the strength of evidence for a population effect.
Weak conclusion. “Reminders definitely make all adults attend more classes.” “Definitely” is too certain, and “all adults” goes far beyond the defined population. The study concerns members of this community center, not adults generally.
Stronger conclusion. “In the sample, the proportion attending was 24 percentage points higher among members assigned to receive reminders than among those assigned usual communications. Random assignment supports a cause-and-effect conclusion about the reminder program, and random sampling supports generalizing to the center’s adult members. A claim that the program increases attendance in the member population would still require statistical evidence that the observed difference is convincing.”
Why this wording fits. This design includes both features: random selection from the center’s membership and random assignment to the two communication conditions. That combination can support both generalization to the defined membership and a causal conclusion about the program. It does not extend the conclusion to adults outside that population, and the design alone does not settle whether the numerical difference is more than chance variation.
Revise Claims by Checking Their Scope
A useful way to review a conclusion is to underline the words that describe cause, certainty, and population. Words such as “caused,” “proved,” “all,” and “always” often signal a claim that needs close checking. Then ask what the study actually did. Was the variable measured or assigned? Were people selected randomly from a population? Does the conclusion name that same population?
Decide whether the sentence describes an association, asserts causation, generalizes to a population, or combines these claims.
Look for random sampling and random assignment separately. Do not use one as a substitute for the other.
Use “was associated with” for an observed relationship when causation is not supported. Name the participants or population the design can represent.
Report what the sample showed, and use “convincing evidence” only when the statistical analysis supports it. Avoid treating an observed difference as proof.
A cautious conclusion is not a vague conclusion. It can state the direction and size of a sample result, name the groups and outcome, and explain what the design allows. It is cautious because it does not claim more than the data and study design support.
Common Mistakes and AP Exam Tips
- Confusing random sampling with random assignment. Random sampling supports generalization to the population sampled; random assignment supports cause-and-effect reasoning. Full-credit wording identifies which feature the study actually used.
- Treating an association as proof of cause. If researchers only measured two variables, say they were associated. Do not write that one caused the other, because other variables may help explain the pattern.
- Generalizing from volunteers to everyone. Random assignment among volunteers does not make them a random sample. Limit the conclusion to the participants or explain why broader generalization is not justified.
- Calling a sample difference an exact treatment effect. A difference between sample means or proportions is an observed result. Do not claim the population effect equals that number unless the analysis and wording justify an estimate, and do not call it convincing evidence without appropriate statistical evidence.
- Using “proved” or “definitely.” Statistical conclusions are supported by evidence, not proved with absolute certainty. State the result and its limits instead of using absolute language.
- Leaving the population vague. “The results apply to people” does not identify who the study represents. Name the population actually sampled, such as the members of a specified center or the students enrolled at a particular school.
For a strong AP response, connect each part of the conclusion to a design feature or result: state the observed pattern, identify whether treatment was randomly assigned, identify whether participants were randomly sampled, and limit the claim accordingly. If statistical evidence is not provided, describe the sample result without inventing a level of certainty.
Check Your Understanding
For each situation, decide what conclusion is justified and identify wording that would overclaim.
- A random sample of employees at one company reports weekly exercise hours and self-rated energy. Can the study establish that exercise causes higher energy? To whom might an association be generalized?
- Researchers recruit volunteers for a nutrition study and randomly assign them to two meal plans. What does random assignment support, and what does the volunteer recruitment not establish?
- A random sample from a defined town is randomly assigned to two versions of a public-information message. Which features support generalization and which support a causal conclusion?
- A randomized study reports a 4-point difference between sample means but gives no inference results. What can a careful conclusion say, and what should it avoid claiming?
- Revise this claim: “The program proved that every student will improve because participants assigned to the program had a higher sample average.”