A Quick Audit for Experimental-Design Errors
In Conclusions an Experiment Can Support, you learned that random assignment and random selection serve different purposes. When reading a proposed experiment, it is easy to notice that something seems “random” without identifying what chance is actually doing. It is also easy to see two groups and assume that the design is a fair experiment. A careful audit checks the design’s parts separately.
Start by asking what receives a treatment, how treatments are assigned, what comparison will be made, and whether any groups are formed before assignment. These questions help identify three frequent errors: no random assignment, no useful comparison group, and confusion between blocking and random selection.
This audit builds on Experimental Units, Factors, and Treatments and Explanatory and Response Variables in Experiments. First identify the units and treatment levels; then trace what happens to each unit. Do not decide that a study is randomized merely because its participants were chosen at random, or because the description uses the word “group.”
Error 1: A Treatment Is Given Without Random Assignment
Random assignment uses chance to decide which experimental units receive each treatment. As explained in Why Random Assignment Is the Key to Causation, this helps make treatment groups comparable, on average, with respect to potential lurking variables. If the researcher or participants choose who receives which treatment, the groups may differ in other ways before the treatment begins.
A study can include two treatment groups and still lack random assignment. For example, people who select a new program may differ from people who continue their usual routine. Comparing the groups can describe a difference, but the design does not rule out other explanations for that difference. As in Confounding in Experiments, a background variable that differs systematically between groups may be entangled with the treatment.
A useful diagnosis states what the procedure actually did, not just that the groups are “unfair.” Say that treatment assignment was not random, identify how groups were formed, and explain that differences between groups could be due to pre-existing differences as well as the treatment. Do not conclude that the treatment had no effect; the design simply does not isolate its causal effect.
Worked Example: A Voluntary Stretching Plan
A recreation center wants to study whether a 15-minute stretching plan affects flexibility. It invites 48 adult members to participate. Members who want to try the plan join the stretching group; those who do not want to change their routine form the comparison group. After six weeks, staff measure each member’s flexibility using the same test.
Audit the units and treatments: The experimental units are the 48 participating members. The treatments are the stretching plan and the usual routine. Flexibility after six weeks is the response.
Identify the error: There are two groups to compare, but members choose their own groups. Treatment assignment is not random. People who volunteer for stretching may already exercise more often, have different flexibility, or be more motivated to improve.
Explain the consequence: If the stretching group has higher flexibility at the end, the difference could reflect the plan, pre-existing differences, or both. A consistent test helps make measurement comparable, but it does not make the treatment groups comparable. The study can describe an association between group membership and flexibility; it does not establish that stretching caused a difference.
Suggest a correction: After recruiting the participants, assign them to the stretching plan or usual routine by a chance process. That change would support a causal comparison for these participants if the plan is followed and outcomes are measured consistently. It would not, by itself, make the participants representative of all adults.
Error 2: There Is No Useful Comparison
A treatment effect is assessed by comparing outcomes under different treatment conditions. A control group, as discussed in Control Groups and Comparison Groups, provides a baseline for that comparison. The control condition might be a placebo, the usual practice, or another appropriate treatment. The best choice depends on the question.
If every unit receives the same treatment and the researcher only measures outcomes afterward, there is no group receiving a different condition at the same time. A high or low outcome alone does not show what would have happened without the treatment. A before-and-after measurement on the same people provides a comparison over time, but changes could also reflect other events or changes during that period. It is not the same as randomly assigning units to treatment and comparison conditions.
“Comparison group” does not mean that the comparison must always be an untreated group. An experiment comparing two active treatments has a treatment contrast, even if neither group is called a control group. Also, a matched-pairs design can provide a comparison within each pair or within each unit, as introduced in Matched Pairs Design. The audit asks whether the design actually compares outcomes under relevant conditions—not whether it uses a particular label.
Worked Example: A New Study Playlist for Everyone
A school learning team wants to know whether a playlist improves concentration during independent study. It asks 32 students to use the playlist for two weeks and records each student’s concentration score at the end. The team plans to report whether the scores are high.
Audit the units and treatment: The experimental units are the 32 students. The treatment is using the playlist during independent study, and the response is the concentration score. However, all students receive the same treatment.
Identify the error: The design has no comparison condition. A high concentration score after two weeks does not reveal whether the playlist improved concentration. The students might have had high scores without it, or another change during the two weeks might explain the result.
Suggest a correction: Assign some students by chance to use the playlist and the rest to study without it, keeping the study period and score measurement the same for both groups. The two groups would then provide a treatment comparison. If researchers want to compare the playlist with another kind of music, that second condition could serve as the comparison; a no-playlist condition is not the only possible choice.
State the limit of the original plan: The original plan can describe the students’ scores after using the playlist. It cannot determine whether the playlist caused an improvement, because it does not compare those outcomes with outcomes under another condition.
Be precise when a description includes a before-and-after measure. Measuring the same units twice creates a within-unit comparison, but by itself it does not answer whether the treatment caused the change. Time, practice with the measurement, or another event may also affect the response. A design with a concurrent comparison condition can help address these alternative explanations.
Error 3: Confusing Blocking With Random Selection
Blocking and random selection both involve groups, but they answer different questions. In a randomized block design, researchers form blocks of similar experimental units and then randomly assign treatments separately within each block. Blocking organizes treatment assignment. In stratified random sampling, researchers divide a population into strata and randomly select individuals from every stratum. Stratification organizes sample selection.
The same characteristic—such as grade level, starting skill, or age group—could be used to form blocks in an experiment or strata in a sampling plan. The purpose and the next step determine which method is being used. If the researcher is assigning treatments to experimental units within the groups, those groups are blocks. If the researcher is selecting sample members from the population groups, those groups are strata.
Blocking does not automatically mean that the experimental units were randomly selected from a population. A researcher may recruit volunteers, divide them into blocks based on an important characteristic, and randomly assign treatments within each block. That is a blocked randomized experiment, but it is not a random sample. Conversely, randomly selecting a sample does not create blocks or randomly assign treatments.
Worked Example: Blocking Students by Starting Skill
A tutoring team recruits 40 volunteers to compare two vocabulary practice methods. Before practice begins, it classifies each volunteer as having lower or higher starting vocabulary skill. Within each skill group, a chance process assigns half the students to flashcards and half to a spaced-retrieval app. The team compares vocabulary scores after four weeks.
Audit the groups: Starting skill is used to form groups of similar experimental units before treatment assignment. Within each group, the students are randomly assigned to both practice methods. These are blocks, and the design is a randomized block design.
Explain what the design does: Blocking accounts for starting skill when comparing the methods. Each method is used by students in both starting-skill blocks, so the comparison is not simply between lower-skill students using one method and higher-skill students using the other. Random assignment within each block supports a cause-and-effect comparison of the methods for these volunteers.
Identify what it does not do: The volunteers were recruited; they were not randomly selected from all students. The blocks do not turn them into a random sample. The results do not automatically generalize to all students, even though treatment assignment was randomized.
State the distinction: The study uses random assignment within blocks, not random selection within strata. If the team had instead randomly selected students from each starting-skill group to form a sample, that would be stratified random sampling. It would not, on its own, assign those students to practice methods.
How to Repair a Design Description
When you identify an error, explain both the problem and its consequence. Then describe a repair that addresses that specific problem. Adding the word “random” is not enough; name what chance selects or assigns. The following distinctions can help keep a repair focused.
| Problem in the description | What to check | Possible repair |
|---|---|---|
| Treatment groups are formed by choice or convenience. | Was each experimental unit assigned to a treatment using chance? | Use a chance process to assign units to treatment conditions. |
| All units receive the same condition and are measured only afterward. | What other condition provides a meaningful comparison? | Add an appropriate comparison condition and assign units to conditions by chance. |
| Groups are formed based on a characteristic, but their purpose is unclear. | Are researchers selecting people for a sample or assigning treatments within groups? | Describe the sampling method and treatment-assignment method separately. |
| Blocks are formed, but assignment within each block is not described. | Does each block include units assigned to each treatment by chance? | Randomly assign units to the treatment conditions separately within each block. |
A repair should also preserve the study’s question. For example, if the aim is to compare two practice methods, replacing one with a different question or response does not fix the assignment flaw. A strong description says what the units are, how the relevant treatment conditions differ, how chance is used, and what outcome will be compared.
Common Mistakes and AP Exam Tips
- Calling a study randomized because participants were randomly selected. Random selection chooses who enters the sample; random assignment decides which treatment each experimental unit receives. State which process occurred and what it supports.
- Calling a study an experiment just because there are two groups. Check whether the researcher imposed treatments and assigned units to them. If people chose their own existing conditions, the comparison is observational rather than a randomized experiment.
- Saying a treatment “worked” when there is no comparison. A post-treatment outcome alone does not show what would have happened under another condition. A full-credit response identifies the missing comparison and explains why the observed outcome cannot isolate a treatment effect.
- Assuming the comparison must be “no treatment.” A comparison may be usual practice, a placebo, or another active treatment. Name the actual conditions being compared.
- Calling blocks a random sample. Blocks are formed to organize treatment assignment. A random sample requires a chance-based selection procedure from a defined population or sampling frame.
- Claiming that blocking removes all differences. Blocking accounts for the characteristic used to form the blocks. Random assignment within blocks helps with other potential lurking variables, but the design does not guarantee identical groups or eliminate every limitation.
- Using vague language such as “the groups were random.” Say whether individuals were randomly selected, treatments were randomly assigned, or units were randomly assigned within blocks. These descriptions have different meanings.
For an AP response, name the design feature that is missing or misidentified, describe the actual procedure, and connect that feature to the conclusion. For instance: “Students chose their own practice method, so the study did not randomly assign treatments. Differences in motivation could help explain the score difference, so the study does not establish that the method caused it.” This is more informative than simply writing “there may be bias.”
Check Your Understanding
For each situation, identify the design issue and explain what conclusion is limited.
- A randomly selected sample of residents is divided between two wellness programs, but each resident chooses a program. Did random selection or random assignment occur? What can the design not establish?
- Every participant receives a new reminder app, and researchers record app satisfaction only after one month. What comparison is missing, and what can the results describe?
- Researchers group experimental units by age and randomly assign both treatments within each age group. Are the groups blocks or strata? Explain.
- A researcher randomly selects students from each grade, then gives every selected student the same study plan. Which chance process occurred, and is there a treatment comparison?
- Rewrite this claim more carefully: “The treatment caused improvement because the study randomly selected participants.” State what random selection supports and what information about assignment is still needed.