Why an Experiment Needs a Comparison
In Completely Randomized Design and Carrying Out Random Assignment With a Number Table or Calculator, you learned how chance can assign experimental units to treatment groups. The next question is what those groups should receive. A group given the treatment of interest is difficult to interpret on its own: an outcome might have changed even if the treatment had not been given.
A comparison group provides a reference point, or baseline, for judging the treatment group’s outcome. Researchers compare the groups’ responses to ask whether the treatment group had a different outcome under the conditions of the experiment. A comparison does not mean that the groups must be identical in every way. As explained in Why Random Assignment Is the Key to Causation, random assignment helps make groups comparable on average, but chance can still produce differences between the particular groups.
The words “control group” do not always mean “a group that gets nothing.” The right control depends on the research question. If researchers want to know whether a new treatment works better than doing nothing beyond ordinary conditions, a no-treatment or usual-conditions control may be suitable. If the practical question is whether a new treatment improves on an existing treatment, the existing treatment can be the control. That is an active comparison.
A useful control also helps the researchers make the comparison fair. Other parts of the study should be kept as similar as practical: for example, the length of the study, the measurement process, and the conditions under which participants receive their assigned treatment. The treatment being compared should be the important planned difference between the groups. Random assignment supports that goal, but it does not replace careful planning.
What a Baseline Helps Reveal
Suppose every seedling in an experiment grows taller during the study. That growth alone does not show that a fertilizer caused the growth. Seedlings may grow as time passes even without added fertilizer. A group grown under the same conditions without the experimental fertilizer gives researchers a baseline for judging whether the fertilized group’s growth differs.
More generally, a control group helps reveal what happens under a comparison condition while the treatment group receives the treatment of interest. This matters because outcomes can change for reasons other than the treatment, such as ordinary development, seasonal changes, or events that affect both groups. If both groups are exposed to the same general conditions, a difference in their responses is more informative than the treatment group’s response by itself.
A baseline comparison is not a promise that the control group will have no change. A control group could improve, worsen, or stay about the same. Its purpose is to show the outcome under the chosen comparison condition. The researchers then compare that outcome with the outcome under the treatment of interest.
This distinction is easy to miss. If researchers measure participants before and after a program, the first measurement is a baseline measurement. But if everyone receives the program, there is no control group to show how outcomes might have changed over the same period without that program. A comparison group addresses that question by following another group under a different assigned condition.
Choosing a Control Condition
The comparison condition should match the question the researchers want to answer. Ask: “Compared with what?” The answer identifies what the control group needs to receive. A no-treatment control can be appropriate when researchers want to compare the treatment with ordinary conditions. An active comparison is more appropriate when the question is whether the new option is better than an existing option.
Be specific about what the researchers want to learn, such as whether a new program performs better than the usual program.
Decide what comparison condition answers that question: no added treatment, usual conditions, or another treatment.
Use the same response measurement and, where practical, similar schedules and study conditions for all groups.
Describe which conditions were compared and what difference in outcomes would address the research question.
The control condition should not be selected merely because it is easy to describe. It needs to be relevant to the question. For instance, a new tutoring format compared with no tutoring answers a different question from the same format compared with the school’s existing tutoring program. Both designs may be useful, but they do not test the same claim.
Worked Example: A No-Added-Treatment Control
Worked Example: Testing a Soil Amendment
A fictional horticulture team wants to know whether a soil amendment increases the growth of young tomato plants over 21 days. The team has 40 similar seedlings and plans to randomly assign 20 to receive soil with the amendment and 20 to receive the same soil without it. Both groups will get the same amount of water and light. The response is each seedling’s increase in height, in centimeters, over the 21 days.
State: The treatment of interest is soil with the amendment. The control condition is the same soil without the amendment. The question is whether seedlings assigned to the amendment condition have a different average increase in height from seedlings assigned to the control condition.
Plan: Randomly assign the 40 seedlings to the two conditions. Keep the watering schedule, light, containers, study length, and height-measurement method the same. The planned treatment difference is whether the soil contains the amendment.
Do: At the end of 21 days, measure each seedling’s height increase. Compare the responses from the amendment group with those from the no-amendment control group. No outcome values are needed to describe the design: the control’s role is to show how much seedlings grew under the same general conditions without the amendment.
Conclude about the design: This control provides a baseline for ordinary growth under the study conditions. If the groups’ outcomes differ, the comparison is more informative about the amendment than observing growth in the amendment group alone. Random assignment helps support a cause-and-effect comparison for these seedlings, but it does not guarantee that the groups will be identical or make them a random sample of all tomato plants.
Worked Example: An Active Comparison
Worked Example: Comparing Two Reading Programs
A fictional school team wants to know whether a new small-group reading program leads to better reading assessment results than the school’s existing program. The team has 60 students whose families have agreed to participate. It randomly assigns 30 students to the new program and 30 to the existing program. Both programs run for six weeks, and the same assessment is used for both groups at the end.
State: The new program is the treatment of interest. The existing program is the active comparison and serves as the control condition. The question is whether the students assigned to the new program have different assessment results from students assigned to the program the school already uses.
Plan: Randomly assign students to the two programs. Use the same study period and assessment process for both groups. The existing program is a relevant baseline because the team wants to know whether the new option improves on the school’s current approach, not merely whether students can improve while receiving instruction.
Do: After six weeks, collect the assessment results and compare the groups. If the new-program group has better results, that difference addresses how the new program performed relative to the existing program in this experiment. It does not directly establish how either program compares with receiving no reading instruction.
Conclude about the design: A control group can receive a treatment. Here, the existing program provides an active comparison. Because students were randomly assigned, the team can use the group comparison to study the effect of assignment to the new rather than the existing program for these participants, provided the programs and measurements are carried out as planned.
Worked Example: A Measurement Baseline Is Not a Control Group
Worked Example: Evaluating a School Wellness Workshop
A fictional school plans a wellness workshop and wants to study its effect on a short self-reported stress score. It randomly assigns 50 volunteers to attend the workshop and 50 to continue their usual advisory activities. Both groups complete the same stress survey just before the study and again four weeks later.
State: The workshop group receives the treatment of interest. The usual-advisory group is the control group because it provides a comparison condition. The first survey is a pre-treatment baseline measurement; it is not itself the control group.
Plan: Randomly assign the volunteers to the workshop or usual-advisory condition. Have both groups complete the survey at the same two times using the same instructions. The researchers can consider the initial measurements when describing where participants began, and compare the groups’ later outcomes to study the workshop relative to usual advisory activities.
Do: Record the initial and four-week scores for each participant. Then examine how the outcomes for the workshop group compare with those for the control group. The first survey provides information about participants before the assigned conditions begin; the control group provides information about outcomes during the study under the usual-advisory condition.
Conclude about the design: The two ideas serve different purposes. The pre-treatment survey measures an initial condition, while the control group supplies a comparison over the same four weeks. If the workshop group changes, the control group helps the researchers judge whether a similar change occurred under usual advisory activities. The design does not make an initial survey a substitute for a comparison group.
Common Mistakes and AP Exam Tips
- Assuming “control” means “nothing.” A control group may receive usual conditions or an active alternative treatment. Name what it actually receives.
- Reporting only the treatment group’s outcome. Improvement in the treatment group alone does not show what would have happened under another condition. Explain why the comparison group provides a useful baseline.
- Calling a pre-treatment measurement a control group. A measurement is not a group. Say “baseline measurement” for the initial score and “control group” for the group assigned to the comparison condition.
- Choosing a comparison that does not match the question. Comparing a new program with no program does not answer whether it is better than the current program. State the specific comparison the research question requires.
- Claiming random assignment makes groups exactly equal. Random assignment helps create comparable groups on average, but the groups in one experiment may differ by chance. Do not say it guarantees identical starting characteristics.
- Using a vague conclusion. Full-credit communication names the treatment and control conditions, the response being compared, and the population or participants to which the conclusion applies. For example: “The experiment compares the assessment results of students assigned to the new reading program with those assigned to the school’s existing program.”
A strong AP response identifies the control condition in context and explains why it is the relevant baseline. If the control receives another treatment, call it an active comparison and state what question that comparison answers. If there is also a pre-treatment measurement, name it separately.
Check Your Understanding
Answer each question by identifying the comparison condition and explaining its role.
- A plant experiment compares plants given a new nutrient mix with plants grown under the same conditions without the mix. Which group is the control, and what baseline does it provide?
- A study compares a new phone-based language course with the course students currently use. Why can the current course be an active control?
- Researchers measure participants’ fitness before a program and again later, but everyone receives the program. Is the first measurement a control group? Explain.
- A team wants to know whether a new after-school tutoring plan is better than the school’s current tutoring plan. Which is the more relevant control condition: no tutoring or the current plan? Explain.
- Why does random assignment not guarantee that the treatment and control groups have exactly the same characteristics?