When Uneven Conditions Complicate an Experiment
In Controlling Variables Across Treatment Groups, you learned to keep relevant background conditions as consistent as practical. This matters because a treatment comparison is hard to interpret when groups also experience different conditions. If those conditions differ systematically with treatment, the design may confound the treatment effect with the effect of the uneven condition.
As introduced in Confounding Variables Explained, a confounding variable is related to both the explanatory variable and the response, making their separate relationships difficult to distinguish. In an experiment, a treatment can be confounded with a background condition when the condition changes along with the treatment. If the response differs between groups, the experiment alone cannot tell whether the treatment, the condition, or both contributed to the difference.
The key issue is not simply that groups or experimental units differ. Individuals may vary in experience, ability, or other characteristics even in a well-designed experiment. The concern here is that a condition is built into the procedure in a way that tracks treatment. For example, if every unit receiving Treatment A is tested in a quiet room and every unit receiving Treatment B is tested in a noisy room, treatment and room are inseparable in that comparison.
A useful diagnostic question is: Could the treatment be compared while the condition stayed the same? If the answer is no because every treatment is tied to its own condition, a difference in the response has more than one possible explanation. A table can make this pattern visible.
| Group | Treatment | Testing condition |
|---|---|---|
| A | New study routine | Quiet room |
| B | Usual study routine | Noisy room |
In this table, study routine and room condition vary together. The design does not compare the new routine with the usual routine under the same room condition. Therefore, even if the group responses differ, the study cannot isolate which of these two differences accounts for the gap.
A Diagnostic for Confounding
When reading an experiment description, identify the treatment factor and its levels, then list important conditions that might influence the response. For each condition, check whether all treatment groups experience it comparably. A condition does not need to be identical in every detail, but it should not systematically favor or disadvantage one treatment group.
State which treatment levels the experiment is comparing and what response is measured.
Look for differences in location, timing, equipment, instructions, staff, duration, or measurement procedure.
Ask whether the same condition occurs with multiple treatments, or whether each treatment is always paired with a different condition.
If treatment and condition move together, state that any response difference could be due to the treatment, the condition, or both.
Keep the condition consistent across treatment groups when practical, while preserving the planned treatment difference.
This diagnostic focuses on the design, not just on the final response values. Even if the observed responses happen to be similar, a confounded design remains flawed: it has not cleanly isolated the treatment comparison. And if responses differ, the difference does not reveal which of the linked factors caused it.
Worked Example: A Study Routine Tested in Different Rooms
Worked Example: Routine and Room Are Confounded
A fictional school compares a new ten-minute study routine with students’ usual review routine. The response is a score on a 10-point vocabulary quiz. Every student using the new routine studies in a quiet library room. Every student using the usual routine studies in a busy common area. Suppose the new-routine group’s scores are 8, 9, 9, and 10, while the usual-routine group’s scores are 5, 6, 6, and 7.
State: The treatment factor is study routine, with the new routine and usual review as treatments. The response is quiz score. Room noise is a background condition that could affect students’ ability to study.
Plan: The comparison should give both treatment groups comparable study conditions. But in the described plan, the new routine is always paired with the quiet room, and the usual routine is always paired with the busy room. Treatment and room condition therefore change together.
Do: The new-routine group’s mean score is \((8+9+9+10)/4=36/4=9\). The usual-routine group’s mean is \((5+6+6+7)/4=24/4=6\). The observed difference in means is \(9-6=3\) points. This is a description of these fictional results; it does not identify the cause of the difference.
Conclude about the design: The new routine may have contributed to the higher scores, but the quiet room may also have helped. Because each routine was used in only one room condition, the experiment cannot separate the effect of study routine from the effect of room noise. A fairer plan would test both routines in the same room under the same schedule and other relevant conditions.
The numerical difference does not repair the design. Calculating a difference in means describes what happened in these groups; it cannot untangle factors that the experiment always paired together. As discussed in Association Versus Causation in Study Conclusions, a cause-and-effect conclusion depends on how the experiment was designed, not simply on whether the observed results differ.
Worked Example: A Plant Treatment With Uneven Light
Worked Example: Fertilizer and Sunlight Are Linked
A fictional gardening club tests whether a new fertilizer affects plant growth. It applies the new fertilizer to six plants placed beside a sunny window. Six plants receiving no fertilizer are placed on a shaded shelf. After four weeks, the response is each plant’s increase in height, in centimeters.
State: The treatment factor is fertilizer use, with new fertilizer and no fertilizer as the two treatments. The response is increase in plant height. Sunlight is a background condition that could also affect growth.
Plan: The club should make sunlight exposure comparable across the two treatment groups. Instead, all fertilized plants receive more direct sunlight and all untreated plants receive less. Thus, treatment and light exposure are tied together in the design.
Do: Suppose the fertilized plants have a mean increase of 8 centimeters and the untreated plants have a mean increase of 4 centimeters. That 4-centimeter difference describes the observed groups. It cannot establish that fertilizer alone produced the difference, because the groups also had different sunlight conditions.
Conclude about the design: Fertilizer and sunlight are confounded in this experiment. To study the fertilizer comparison more fairly, the club could keep location and light conditions consistent while applying the different fertilizer treatments. It should also use multiple plants in each treatment group, as replication requires.
Having multiple plants in each group does not remove this confounding. Replication helps reveal variation among units receiving each treatment, but every fertilized plant is still in the sunny location and every untreated plant is still in the shade. Replication and consistent conditions address different design concerns.
Worked Example: A New App Tested at a Different Time
Worked Example: App Type and Time of Day Are Confounded
A fictional technology club compares two reminder apps and records whether participants complete a short daily task. Participants using App A receive reminders in the morning, while participants using App B receive reminders late in the evening. The club wants to know whether the reminder design affects completion.
State: The treatment factor is reminder app, with App A and App B as treatments. The response is task completion. Time of day is a background condition because participants’ availability and routines may differ at different times.
Plan: If each app is used at only one time of day, app type is always paired with timing. A completion-rate difference could reflect the app, the time of the reminder, or both. The design does not compare the two apps at a common time.
Do: Suppose 18 of 20 participants using App A complete the task, while 12 of 20 using App B do so. The observed completion proportions are \(18/20=0.90\) and \(12/20=0.60\), a difference of 0.30. These values describe the groups, but the comparison cannot assign that difference to app design alone.
Conclude about the design: App type and time of day are confounded because each app is used at a different time. To isolate the app comparison more clearly, the club could send both apps’ reminders at the same time under otherwise consistent procedures. It should avoid claiming that App A caused higher completion from this design alone.
Confounding Is About the Design, Not Just the Results
An uneven condition creates a design problem when it is linked to the treatment groups. The outcome might be a numerical measurement, such as quiz score or plant growth, or a categorical response, such as task completed or not completed. In either case, the core issue is the same: the procedure does not provide a comparison that separates treatment from condition.
Random assignment, introduced in Why Random Assignment Is the Key to Causation, helps make treatment groups comparable on average with respect to other characteristics. But it does not make a deliberately uneven protocol harmless. For example, randomly assigning people to two groups would not resolve the problem if the researchers then always tested one treatment in the morning and the other in the evening. The condition still tracks treatment by design.
A confounded experiment does not prove that the treatment had no effect. It means the experiment cannot determine the treatment’s effect separately from the linked condition. Be precise: the design makes a treatment-only explanation uncertain; it does not show that the treatment is ineffective or that the condition definitely caused the response difference.
Common Mistakes and AP Exam Tips
- Naming a condition without explaining the link. Saying “room noise is a problem” is incomplete. Explain that one treatment group always used the quiet room and the other always used the noisy room, so treatment and noise are confounded.
- Claiming the treatment caused the observed difference. A response difference does not resolve confounding. A full-credit explanation says the difference could be due to the treatment, the uneven condition, or both.
- Assuming random assignment fixes every design flaw. Random assignment helps balance characteristics across groups, but it does not erase conditions deliberately paired with treatment. Describe the actual procedure, not just the presence of random assignment.
- Calling all group differences confounding. Units can differ in many ways without a condition being systematically assigned alongside treatment. Identify the particular condition and show how it tracks the treatments.
- Suggesting a fix that removes the treatment contrast. The goal is not to give both groups the same treatment. Keep the treatment levels different while making the background condition comparable across groups.
- Overstating the conclusion. Do not say the study proves the condition caused the outcome. Say the design cannot separate the possible effects of the treatment and condition.
A concise AP response can follow this pattern: “The groups differ in both [treatment] and [condition]. Because [condition] is always paired with [treatment level], a response difference could be due to either factor or both; their effects are confounded.” Then suggest a change, such as using the same room, schedule, or measurement procedure for every treatment group.
Check Your Understanding
For each situation, identify whether a background condition is confounded with treatment and explain what the design can or cannot establish.
- A fictional study compares two tutoring methods. One method is always taught in a quiet classroom and the other is always taught in a hallway. Name the treatment and the potentially confounded condition.
- A garden experiment compares two seed varieties. Both varieties are grown on the same balcony using the same watering schedule. Is the location automatically confounded with seed variety? Explain.
- A fictional experiment compares two notification sounds. One is tested using a new phone model and the other using an older model. Explain why an outcome difference cannot automatically be attributed to sound.
- Why does using many experimental units per treatment not, by itself, remove a condition that is always paired with one treatment?
- Write one sentence explaining how a study that compares a morning reminder with an evening reminder could be redesigned to isolate reminder type more clearly.