Tutorials › AP Statistics › Random Assignment Versus Random Sampling Conditions

Conditions for mean inference · Tutorial 654 of 1000

Random Assignment Versus Random Sampling Conditions

Distinguish the roles of random sampling and random assignment so your conclusions about population means and treatment effects match how the data were collected.

Intermediate 10 min read

What You'll Learn

  • Explain how random sampling supports generalizing from a sample to its population.
  • Explain how random assignment supports a cause-and-effect conclusion about treatments.
  • Distinguish random selection of participants from random allocation to groups.
  • Identify what can be concluded when a study uses both, one, or neither.
  • Match the scope of a conclusion to the study’s sampling and assignment design.
  • Separate these design questions from the other conditions for mean inference.

Two Kinds of Randomness, Two Different Conclusions

In “Checking the Random Condition,” we identified random sampling and random assignment as different ways a chance process can enter a study. Here we focus on why that difference matters. Random sampling and random assignment are not interchangeable: one supports generalizing results to a population, while the other supports concluding that a treatment caused a difference.

A study may randomly select people from a population, randomly assign participants to treatments, do both, or do neither. To decide what its results justify, ask two separate questions: How were the study units chosen? And how were they placed into groups? The first concerns sampling; the second concerns assignment.

Definition: Random sampling uses a chance process to select units from a defined population for a study. It supports generalizing from the sample to that population. Random assignment uses a chance process to allocate study units to treatment groups. It supports a cause-and-effect conclusion about the treatments for the study units.

A randomized design does not automatically settle every question about a mean procedure. As explained in “Why Conditions Matter in Mean Inference,” randomness is one part of a larger conditions check. Independence and the Normal/Large Sample condition must also be considered when they apply. Here, the central task is to connect each type of randomization to the conclusion it supports.

Random Sampling Supports Generalization

Suppose researchers want to estimate or compare mean outcomes for a particular population. If they select a random sample from that population, the chance-based selection helps make the sample representative of the population, in the sense needed for inference. This supports using the sample results to draw a conclusion about the population from which the sample was selected.

The target population matters. If a random sample is drawn from all eligible residents of a town, the design supports generalizing to those residents—not automatically to residents of other towns, all adults, or people who were not eligible. State the population that the sampling process actually represents.

Random sampling by itself does not establish that one condition caused a higher or lower mean. In an observational study, people may already belong to the groups being compared. Differences in their outcomes could be related to other differences between the groups, not just the group label. Random selection helps with the reach of a conclusion, not with identifying cause and effect.

Random Assignment Supports Causation

In an experiment, researchers impose treatments and use random assignment to decide which study units receive each one. Because assignment is determined by chance rather than by participants’ characteristics or researchers’ choices, the groups tend to be comparable before treatment. Random assignment helps balance other variables across treatment groups over the long run, including variables that researchers may not have measured.

If the groups’ mean responses differ after the treatments, random assignment supports attributing the difference to the treatments rather than to a systematic pre-existing difference between the groups. This supports a cause-and-effect conclusion for the study units in the experiment, provided the study was carried out appropriately.

Random assignment does not guarantee that the groups are identical in every characteristic in one particular experiment. Chance differences can occur. Nor does random assignment make a sample representative of a larger population. If participants volunteered or were recruited by convenience, the randomized experiment may support a causal claim about those participants without supporting broad generalization to people who did not take part.

Key distinction: Random sampling helps answer, “To which population can we generalize?” Random assignment helps answer, “Can we attribute a difference to the treatment?” A study needs the relevant design feature to support each conclusion.

A Two-Question Conclusion Audit

A useful technique is a two-question conclusion audit. Before writing a conclusion about means, separately identify the sampling process and the group-formation process. Then match each process to the claim it supports. This prevents a common error: using the word “random” without saying what was randomized.

1
Identify the study units and target population.
Name what one observation represents and the population, if any, from which the units were selected.
2
Audit selection.
Ask whether a chance process selected units from that population. If so, state which population the sample can represent.
3
Audit group formation.
Ask whether a chance process assigned units to treatments. If so, the design supports a cause-and-effect conclusion about the treatments for the study units.
4
Limit the conclusion to what the design supports.
Do not claim generalization without suitable random sampling or causation without random assignment. Check the other conditions for the mean procedure separately.
Random sampling?Random assignment?What the design supports
YesNoGeneralizing to the sampled population may be supported; a causal conclusion is not established by sampling alone.
NoYesA causal conclusion for the study units may be supported; generalizing to a broader population is not established by assignment alone.
YesYesBoth generalization to the sampled population and a causal conclusion about the treatments may be supported.
NoNoNeither broad generalization nor a causal conclusion is supported by these design features alone.

These are distinct design questions, not competing choices. A study can use random sampling first and random assignment afterward. It can also use one without the other. When describing a mean comparison, be precise: “randomly selected” describes how units entered the study; “randomly assigned” describes how they entered treatment groups.

Worked Examples

Worked Example: Random Sample, Observational Comparison

A fictional county health team randomly selects 80 households from a list of county households. It asks each household to report the average number of hours spent outdoors per week and whether the household has a garden. The team compares the mean outdoor hours for households with gardens and those without gardens. The households were not assigned to have or not have gardens.

State. Let \(\mu_G\) be the mean weekly outdoor hours for county households with gardens, and let \(\mu_N\) be the mean for county households without gardens. The question is whether the design supports generalizing the comparison and whether it supports a causal conclusion.

Plan. Use the two-question conclusion audit. Check how households were selected, then check whether garden status was randomly assigned. These checks address the scope of the conclusion; the conditions for a two-sample t procedure, including independence and shape, must also be assessed separately.

Do. The households were randomly selected from a county household list, so the sampling design supports generalizing the comparison to the households represented by that list. However, the team did not assign households to have gardens. Garden status was observed, and other factors—such as yard size or interest in outdoor activities—could be related to both having a garden and spending time outdoors.

Conclude. If the other conditions for mean inference are met, the team may use the sample to draw a conclusion about the difference in mean weekly outdoor hours between county households with and without gardens. The study does not establish that having a garden causes households to spend more time outdoors because garden status was not randomly assigned.

Worked Example: Random Assignment of Volunteers

A fictional sleep center recruits 60 volunteers who want to try a new evening relaxation routine. Using a random process, it assigns 30 volunteers to the routine and 30 to their usual evening habits. After four weeks, the center compares the groups’ mean nightly sleep duration. The volunteers were not randomly selected from a larger population.

State. Let \(\mu_R\) be the mean nightly sleep duration for the volunteers assigned to the relaxation routine, and let \(\mu_U\) be the mean for the volunteers assigned to usual habits. We will identify which conclusions the design supports.

Plan. Check selection and assignment separately. Random selection would support generalization to a population; random assignment would support a causal conclusion about the treatments for the study units. Also, the two-sample t conditions must be checked before using that procedure for the difference in means.

Do. The participants volunteered, so they were not randomly selected from all people who might use the routine. The sample may differ from that wider population. However, the center randomly assigned the 60 volunteers to the two groups. That chance process supports treating the groups as comparable for a causal comparison of the routine and usual habits among the study volunteers.

Conclude. If the mean difference is supported by an appropriate analysis and the other conditions are met, the center may conclude that the assigned routine caused a difference in mean nightly sleep duration for these volunteers. Random assignment alone does not justify generalizing that effect to all people, because the volunteers were not randomly sampled from that population.

Worked Example: Random Sampling and Random Assignment

A fictional regional research team randomly selects 120 adults from a defined list of residents who meet the study’s eligibility rules. It then randomly assigns the selected adults to use one of two hydration reminders for a month. At the end of the month, the team compares the groups’ mean daily water intake.

State. Let \(\mu_A\) and \(\mu_B\) be the mean daily water intake, in the study’s target population, under reminders A and B. The design includes both a random sample and random assignment, so we assess the support each provides.

Plan. The sample-selection process determines whether results can be generalized to the population represented by the eligibility list. The treatment-assignment process determines whether a difference in the groups’ means can be attributed to the reminders. Check the other conditions for a two-sample t procedure separately.

Do. The adults were randomly selected from the defined eligible population, supporting generalization to that population, subject to the study’s participation and implementation details. The selected adults were also randomly assigned to the two reminders. That assignment supports a causal comparison of reminder A with reminder B for the study units.

Conclude. If the remaining conditions for mean inference are satisfied, the team may use the results to make a causal conclusion about the difference in mean daily water intake between the reminders, and may generalize that conclusion to the population represented by the sampling frame. The claim should not extend automatically to people outside the eligibility rules or the population represented by the list.

Worked Example: Neither Kind of Randomization

A fictional school newsletter asks readers to volunteer for a survey about daily screen time. The survey compares mean screen time for students who already use a study-planning app and students who do not. The newsletter neither randomly selects students nor assigns app use.

State. Let \(\mu_A\) and \(\mu_N\) represent the mean daily screen time for students who use the app and those who do not. We need to decide whether the design supports generalization or causation.

Plan. Check whether the students were randomly selected and whether app use was randomly assigned. Then restrict the conclusion to what those design features support.

Do. Students chose whether to respond, so the respondents are a volunteer group rather than a random sample of all students. The app groups also formed through students’ existing choices, not random assignment. Students who use the app may differ in other ways that relate to screen time.

Conclude. The comparison describes the responding students, but the design alone does not support generalizing to all students or concluding that app use caused a difference in mean screen time. A calculator or statistically significant result cannot replace random sampling or random assignment.

Common Mistakes and AP Exam Tips

  • Using “random” without naming what was randomized. State whether units were randomly selected, randomly assigned, or both. The two phrases support different claims.
  • Claiming random assignment makes a sample representative. Assignment creates treatment groups; it does not make volunteers representative of a larger population. Explain that random sampling, not assignment, supports generalization.
  • Claiming random sampling proves causation. Random selection does not control which group people belong to. Without random assignment, an observed difference in an observational comparison may be associated with other group differences.
  • Making a broader claim than the sampling frame permits. Name the population actually represented by the sampling process. Do not generalize from eligible residents to everyone if the study excluded some groups.
  • Treating a causal conclusion as universal. Random assignment supports cause-and-effect conclusions for the study units. Generalizing that effect beyond them requires a suitable sampling design and a clearly defined population.
  • Assuming either kind of randomness satisfies every inference condition. Randomization addresses a particular design concern. For a mean procedure, also check independence and the Normal/Large Sample condition as appropriate.

A strong AP response ties the design to the claim: “The participants were randomly assigned to the two treatments, so the design supports a cause-and-effect conclusion for these participants. Because they volunteered rather than being randomly selected from the target population, the results cannot automatically be generalized to that population.” For a randomly selected observational sample, say that the sample supports generalizing to its population, but the lack of random assignment does not establish causation.

Key takeaway: Keep the two questions separate: random sampling supports generalization to a population, and random assignment supports causation for the study units. State exactly what was randomized, identify the population and study units, and make no claim beyond what the design supports.

Check Your Understanding

For each situation, identify whether random sampling, random assignment, both, or neither is described. Then state which conclusion the design supports.

  1. A random sample of apartment residents reports whether they own a bicycle and their weekly cycling time. Does this design establish that bicycle ownership causes more cycling?
  2. Volunteers are randomly assigned to two stretching routines. What does random assignment support, and what does it not establish about generalization?
  3. A random sample of eligible library members is randomly assigned to receive one of two reminder messages. Which two kinds of conclusions may the design support?
  4. Students who choose to join a coding club are compared with students who do not, using a survey distributed to volunteers. What limits apply to generalization and causation?
  5. In your own words, distinguish “randomly selected” from “randomly assigned.”