Tutorials › AP Statistics › Checking Conditions for Data from an Experiment

Conditions for one-proportion inference · Tutorial 453 of 1000

Checking Conditions for Data from an Experiment

Distinguish the role of random assignment from random sampling, and match an experiment’s conclusions to the people who took part.

Intermediate 9 min read

What You'll Learn

  • Explain when random assignment can support inference about experimental units without random sampling
  • Distinguish causal conclusions from generalizations to a wider population
  • Identify the population to which an experiment’s results can apply
  • Check randomization and Large Counts conditions for a one-proportion inference
  • Recognize how volunteer enrollment, nonresponse, and missing outcomes limit conclusions

Random Assignment and Random Sampling Answer Different Questions

In Conditions for a Survey With Nonresponse, you saw that the way people enter a study can affect whom its results represent. Experiments add another kind of random process: researchers may randomly assign participants to treatments. Random sampling and random assignment are both important, but they do different jobs.

Random sampling uses chance to select people from a population. It supports generalizing results from the sample to that population. Random assignment uses chance to allocate study participants to treatment groups. It supports a fair comparison of treatments and, when the experiment is well conducted, causal conclusions about the participants in the experiment.

Key distinction: Random sampling supports generalization to a population. Random assignment supports a cause-and-effect conclusion for the experimental units. Random assignment does not, by itself, make volunteers representative of a wider population.

For a one-proportion inference, imagine that a study randomly assigns some participants to receive a particular treatment and records whether each has a defined outcome. Because assignment was random, the treatment group is a chance-selected group from the study participants. Its sample proportion can support inference about the proportion of those participants who would have that outcome under the treatment, assuming the study is otherwise conducted appropriately and outcomes are obtained as planned.

This use of random assignment does not mean that the treatment group is a random sample of everyone who could ever receive the treatment. The inference is tied to the experimental units who took part. If the researcher wants a causal conclusion about treatment effectiveness, the study needs an appropriate comparison, such as a control group. A one-proportion result for one group alone does not show that the treatment caused the observed outcome.

What the Randomization Supports—and What It Does Not

When a study has no random sample, ask what was randomized. If participants were randomly assigned to treatment conditions, that assignment supports a randomization-based inference about the experimental units. For a one-proportion procedure focused on one treatment group, the group’s outcomes can be used to estimate the outcome proportion for the study participants under that treatment.

The target is not automatically the general public, all patients with a condition, or everyone eligible for the treatment. It is limited to the experimental units represented by the study and the treatment condition being considered. If participants volunteered, for example, they may differ from people who did not volunteer. Random assignment among volunteers does not remove that difference.

When both random sampling and random assignment are used, each process supports a different part of the conclusion. Random sampling can support generalization to the population from which the sample was drawn. Random assignment can support a causal comparison among the experimental units. In a carefully designed experiment, both kinds of conclusion may be reasonable, but each depends on its own evidence.

Conditions to assess for a one-proportion inference from an experiment:
  • Randomization: Participants were randomly assigned to the treatment group whose outcome proportion is being analyzed. This supports inference about the experimental units, not automatically a broader population.
  • Independence: Consider how the assignment and study were conducted. If a random sample was also taken without replacement from a finite population, check the 10% condition as described in Independence and the 10 Percent Condition. The 10% condition is not a substitute for checking how assignment occurred.
  • Large Counts: For a one-proportion \(z\)-interval, check that the observed success and failure counts are each at least 10, as in Large Counts Condition for Confidence Intervals. For a one-proportion \(z\)-test, use the null proportion to check expected counts, as in Checking the Success-Failure Condition for Tests.
  • Outcome collection: Check whether outcomes were recorded for the assigned participants. Missing outcomes or substantial loss to follow-up can weaken the basis for inference.

The Large Counts condition addresses whether a Normal-based procedure is appropriate; it does not establish that participants represent a population. Likewise, random assignment does not guarantee that every other condition is met. Keep the condition checks separate and explain what each one supports.

Worked Examples

Worked Example: Volunteers Randomly Assigned to a Treatment

A clinic recruits 80 volunteers to try a new reminder program. The volunteers are randomly assigned: 40 receive the program and 40 receive standard reminders. At the end of the study, 31 of the 40 people assigned to the new program attend a scheduled follow-up appointment. Assess whether the conditions support a one-proportion \(z\)-interval for the proportion of the study participants who would attend under the new program, and state the scope of the conclusion.

The treatment group was created by random assignment, so the group of 40 is a chance-selected group from the 80 participants. This supports inference about the study participants under the new program. It does not make the 80 volunteers a random sample of all clinic patients.

For the interval’s Large Counts condition, there are 31 successes (attended) and \(40-31=9\) failures (did not attend). The success count meets the threshold, but the failure count does not:

$$ n\hat{p}=31\geq10 \qquad\text{and}\qquad n(1-\hat{p})=40-31=9<10 $$

The Large Counts condition fails, so a one-proportion \(z\)-interval is not justified by the usual Normal approximation. The random assignment supports the relevant randomization argument, but it cannot fix the small failure count. Also, the observed proportion \(31/40=0.775\) describes the people assigned to the program; it is not automatically an estimate for all clinic patients.

A careful condition statement is: “The 80 volunteers were randomly assigned to the two reminder conditions, so the new-program group was randomly formed from the study participants. This supports inference about the participants’ outcomes under the program, not generalization to all clinic patients. However, the group has 31 successes and 9 failures, so the Large Counts condition for a one-proportion \(z\)-interval is not met. I would not use that interval procedure here.”

Worked Example: Random Assignment Does Not Make a Volunteer Sample Representative

A group of 120 students volunteers for a study of a new study-planning app. Researchers randomly assign 60 volunteers to use the app and 60 to use their usual planning method. At the end, 42 students assigned to the app submit all required weekly plans. Assess the conditions for a one-proportion \(z\)-interval for the app group and describe what the result could represent.

The app group was formed through random assignment from the 120 volunteers. That supports inference about these 120 study participants under the app condition. The volunteers were not described as a random sample of students, so the results cannot automatically be generalized to all students at the school or to students elsewhere.

The app group has 42 successes and \(60-42=18\) failures. Both observed counts are at least 10:

$$ n\hat{p}=42\geq10 \qquad\text{and}\qquad n(1-\hat{p})=60-42=18\geq10 $$

Thus, the Large Counts condition for a one-proportion \(z\)-interval is met. The random assignment supports inference about the study participants, and the observed counts support the Normal approximation. If a one-proportion interval is calculated, its target must be described as the proportion of these study participants who would submit all required plans under the app condition—not the proportion of all students who would do so.

Because a usual-method group was also created, researchers could use a comparison of the groups to investigate whether the app caused a difference for the volunteers. That is a different inference question from estimating the app group’s one outcome proportion. Random assignment can support a causal comparison, but it does not erase the limits on generalizing from volunteers.

Worked Example: A Random Sample and Random Assignment Support Different Parts of the Conclusion

A research team randomly selects 400 households from a town’s list of 20,000 households. The selected households agree to take part in a trial and are randomly assigned in equal numbers to a new water-filter instruction or the usual instruction. In the new-instruction group, 168 of 200 households correctly install the filter. Assess the conditions for a one-proportion \(z\)-interval and explain what kinds of conclusions the design supports.

1
State.
Let \(p\) be the proportion of the 400 participating households that would correctly install the filter if assigned the new instruction. This parameter concerns the study participants under that instruction.
2
Plan.
For a one-proportion \(z\)-interval, assess the random processes, any relevant independence condition, the observed success and failure counts, and whether outcomes were recorded for the assigned households. Keep generalization and causal conclusions distinct.
3
Do.
The 400 households were randomly selected from the town’s list. The sample is no more than 10% of the 20,000 listed households because \(400\leq0.10(20{,}000)=2{,}000\). This supports the independence condition for sampling without replacement. The households were then randomly assigned to instructions, supporting a causal comparison for the participating households. In the new-instruction group, there are 168 successes and \(200-168=32\) failures, and both counts are at least 10. The Large Counts condition is met. The description states that outcomes were recorded for all 200 households in that group.
4
Conclude.
The conditions described support a one-proportion \(z\)-interval for the proportion of participating households that would correctly install the filter under the new instruction. Random selection supports generalizing to the town households represented by the list, while random assignment supports a causal comparison of instructions among participating households. The sample-selection process does not establish that households that declined participation are represented in the experiment.

The observed proportion in the new-instruction group is \(168/200=0.84\), or 84%. That is the sample result, not the interval’s full conclusion. The interval would estimate the relevant proportion while accounting for sampling variability. The design’s random sample and random assignment should still be described separately: the first concerns whom the results can represent, and the second concerns whether differences between instruction groups can be attributed to the instructions.

When Assignment or Follow-Up Is Incomplete

A study description should tell you how assignment was carried out, not merely say that it was an experiment. If researchers let participants choose their treatment, treatment groups may differ for reasons other than the treatment. Calling the study an experiment does not make the assignment random. Look for a clear statement that chance determined who received which condition.

Also notice the difference between assignment and what participants actually do. If some participants do not follow their assigned treatment, the original random assignment still matters, but a simple analysis grouped by the treatment people chose to follow may no longer preserve that random assignment. Do not silently treat self-selected treatment use as random.

Missing outcomes create another concern. If participants assigned to a treatment are more or less likely to provide outcome data depending on their results, the observed group may not represent everyone assigned to that treatment. As in Conditions for a Survey With Nonresponse, report who supplied data and assess how missing responses could affect the conclusion. Random assignment at the start of a study does not guarantee that later respondents or completers remain a representative group.

Finally, random assignment is not the same as random sampling even when both are present. A random sample of willing participants may support generalization to its sampling population, while random assignment supports a causal comparison for the participants who entered the experiment. State the target population and the experimental units explicitly so that the reader can see the boundary of each conclusion.

Common Mistakes and AP Exam Tips

  • Claiming that random assignment makes volunteers representative. It randomly forms treatment groups from the volunteers; it does not randomly select those volunteers from a wider population.
  • Using “random” without saying what was randomized. Specify whether people were randomly sampled, randomly assigned, or both. These methods support different conclusions.
  • Claiming a causal effect from one treatment group’s proportion alone. A single group’s outcome proportion describes outcomes under that condition. A causal conclusion requires an appropriate comparison between treatment conditions.
  • Generalizing beyond the study units without evidence. Name the participants or the population represented by a random sample. Do not say “all patients” or “all students” just because an experiment used random assignment.
  • Treating the Large Counts condition as proof of representativeness. Large success and failure counts support the Normal approximation. They do not establish that the participants represent a target population.
  • Checking the 10% condition for the wrong reason. It applies when a random sample is drawn without replacement from a finite population. Random assignment alone is not a random sample from that population.
  • Ignoring missing outcomes or noncompliance. Random assignment supports the design, but missing data or treatment choices can complicate the analysis. State the limitation supported by the study description.
AP Exam Tip: Write one sentence about each random process. For example: “The participants were volunteers, so they were not shown to be a random sample of all patients. They were randomly assigned to treatment groups, which supports a causal comparison among the participating volunteers.” Then check the other conditions for the specific inference procedure.

Key Takeaway

Random assignment can provide the randomization needed for inference about experimental units even when researchers did not take a random sample. It supports causal comparisons among those units, while random sampling supports generalization to a population. Neither process replaces the other, and neither removes the need to check the remaining conditions for a one-proportion procedure.

Key takeaway: Ask two questions: “Who was randomly selected?” and “Who was randomly assigned?” Use random sampling to justify population generalization and random assignment to justify causal conclusions. If there was no random sample, keep the conclusion tied to the experiment’s participants.

Check Your Understanding

For each situation, identify what random assignment supports and the appropriate scope of the conclusion.

  1. Forty-five volunteers are randomly assigned to a new exercise plan, and 36 meet a stated activity goal. Calculate the success and failure counts. Does the Large Counts condition for a one-proportion \(z\)-interval hold?
  2. A study randomly assigns volunteer patients to a new treatment or a standard treatment. Does random assignment alone justify generalizing results to all patients with the condition? Explain.
  3. A random sample of 300 households is selected from 10,000 households and then randomly assigned to two product instructions. State what random sampling supports and what random assignment supports.
  4. Researchers randomly assign participants to two treatments, but only some participants provide outcome data. Why should the condition assessment discuss missing outcomes?
  5. Write a brief condition assessment for a study that randomly assigns 100 volunteers to a treatment group and records 88 successes and 12 failures. Include the scope of inference and the Large Counts check.