Why Replication Matters
In Blinding and Double-Blind Experiments, you learned how keeping specified people unaware of treatment assignments can limit the influence of expectations. Another important feature of a well-designed experiment is replication: using multiple experimental units for each treatment. If a treatment is given to only one unit, an unusual result for that unit can be mistaken for a treatment effect.
In Experimental Units, Factors, and Treatments, you learned to identify the experimental units as the individuals or objects that receive the assigned treatment. That definition tells you what to count when you assess replication. If a treatment is assigned to 20 plants, there are 20 experimental units in that treatment group. If it is assigned to one plant and that plant is measured 20 times, there is still only one experimental unit.
Units receiving the same treatment will often have different responses, even when the experiment is carried out carefully. People may begin with different characteristics, and objects may differ in ways the researchers did not measure. Replication lets researchers see whether a pattern appears across several units or depends heavily on one unusual unit. It does not make all units identical, but it helps keep an individual unit’s distinctive outcome from standing in for an entire treatment group.
Count the Units That Receive the Treatment
To check replication, start with the assignment: What receives a treatment? That is the experimental unit, and it determines the count. Then find how many distinct units receive each treatment. Do not count every measurement, test, or observation as though it came from a different unit.
For example, suppose researchers apply a fertilizer to one garden bed and measure the height of 15 plants growing in that bed. If the fertilizer was assigned to the whole bed, the bed is the experimental unit. The 15 plants are observations within that bed, but they do not provide 15 independent replications of the fertilizer treatment. Conditions shared by the bed—such as its soil and location—could influence all 15 plants.
By contrast, if researchers assign treatments separately to 15 individual plants, then each plant is an experimental unit. That design has 15 units receiving their assigned treatments, provided the assignment and treatment are indeed applied at the plant level. The study’s description of how treatment is assigned matters; the researcher cannot decide afterward to count smaller observations as units just because there are many of them.
Repeated measurements can still be useful. Measuring a unit several times may help describe its response more carefully, reveal changes over time, or reduce the impact of a single imprecise reading. But those measurements share the same unit and many of its characteristics. They do not provide the same information as assigning the treatment to more units.
Why More Units Are Different from More Measurements
Imagine testing a setting on a single phone by measuring its battery use repeatedly. All the readings come from that phone. Its battery condition, age, and hardware are shared across the measurements. If its battery behaves unusually, many repeated readings might consistently reflect that one phone’s behavior. Testing more phones gives information about variation across phones as well.
This distinction matters when a study’s goal is to compare treatments for a broader collection of units. With several units per treatment, researchers can examine whether responses under that treatment are fairly consistent or vary considerably from unit to unit. If only one unit receives a treatment, its outcome cannot show how outcomes under that treatment might vary across units.
Replication also makes a treatment comparison less dependent on a single unusual unit. Suppose one experimental unit has an unusually high response for reasons unrelated to the treatment. With only one unit in its group, that response is the group’s entire result. With multiple units, the unusual outcome is one part of the group’s evidence, rather than the only outcome.
Replication does not guarantee that treatment groups will be perfectly balanced. As discussed in Why Random Assignment Is the Key to Causation, random assignment helps make groups comparable on average. Replication and random assignment have different roles: random assignment uses chance to allocate units, while replication provides multiple units under each treatment. A strong design uses both when possible.
Planning for Replication
When planning an experiment, researchers should identify the experimental units and treatments before deciding how much replication the design needs. A clear plan states how many units will receive each treatment and how the treatment will be assigned. The units should be distinct at the level where treatment is applied.
There is no single number of units that is always enough. Practical limits, the expected amount of unit-to-unit variation, and the question being studied all matter. The central principle at this stage is not to treat many readings from a few units as though they were many independently treated units. Researchers should use multiple units for each treatment when feasible.
State exactly what receives each treatment condition.
Use the definition from Experimental Units, Factors, and Treatments: the units are the individuals or objects to which treatments are assigned.
Count distinct units assigned to each treatment, not repeated readings taken from those units.
Each treatment should be applied to multiple experimental units so that the comparison is not based on just one unit per treatment.
Worked Example: Testing a Water-Filter Design
Worked Example: Many Samples from One Filter or Many Filters?
A fictional student engineering team wants to compare two water-filter designs. In its first plan, the team builds one filter of each design and runs 20 water samples through each filter, recording the clarity after every run.
State: The factor is filter design, with two treatments: design A and design B. The response is the clarity of the filtered water. The filter is the experimental unit if the design is assigned to and built as a particular filter. The 20 runs through a filter are repeated observations from that filter.
Plan: The original plan has only one experimental unit receiving each treatment. The team should build multiple filters of each design and assign each filter to its design condition. For example, it could build 12 filters of each design and test each filter using the same procedure. It may take several readings from each filter, but it must keep the distinction between filters and readings clear.
Do: In the original plan, each design is represented by one filter. If design A’s filter performs unusually well or design B’s filter has a flaw, that difference could be mistaken for a general difference between designs. In the revised plan, 12 distinct filters receive each design, so the team can compare outcomes across multiple filters in each treatment group. The repeated water samples can describe the performance of each filter, but they do not increase the number of filters.
Conclude about the design: The revised plan has replication because each filter design is used on multiple experimental units. The original plan does not gain replication merely by passing many samples through one filter of each design. The team should describe both counts: how many filters were tested per design and how many readings were taken per filter.
Worked Example: Comparing Two Study-Room Conditions
Worked Example: Students or Classrooms as the Units?
A fictional school wants to compare two study-room conditions: quiet music and no music. One proposal assigns one classroom to each condition, then has 18 students in each room complete a short task. A second proposal assigns individual students to conditions, with 18 students in each group.
State: The factor is study-room condition, and the response is each student’s task score. In the first proposal, the whole classroom receives the assigned condition, so each classroom is the experimental unit. In the second proposal, the condition is assigned to individual students, so each student is an experimental unit.
Plan: The first proposal has only one classroom per treatment. The 18 student scores in each classroom are useful observations, but all students in a room share that classroom’s setting. To obtain replication in a classroom-level design, the school would need multiple classrooms assigned to each condition. Alternatively, it could assign conditions to individual students if students can receive their assigned conditions separately without undermining the plan.
Do: The first proposal cannot separate a difference between the two study conditions from a difference between the two particular classrooms. For example, one room might be brighter or quieter for reasons beyond the assigned music condition. The second proposal has multiple individual experimental units per treatment, assuming the assignment and treatment are genuinely at the student level. The assignment level determines which count is appropriate.
Conclude about the design: The school should not claim 18 replications per condition for the first proposal just because it records 18 scores in each room. It has one classroom assigned to each condition. A full design description identifies the unit that receives the treatment, then reports the number of those units in each group.
Worked Example: Testing a Phone Battery Setting
Worked Example: Repeated Tests on One Phone
A fictional technology club tests whether a phone’s low-power setting changes battery use during a fixed task. One student uses the same phone for 10 runs with the setting on and 10 runs with it off. The club proposes to call this 10 replications per setting.
State: The factor is the phone setting, with two treatments: low-power mode on and low-power mode off. The response is battery use during the fixed task. In this plan, the same phone is used for both conditions, and repeated runs are made on that phone.
Plan: If the question is about how the setting performs across phones, the club needs multiple phones assigned to each treatment condition, rather than relying on repeated runs from one phone. It should use a consistent task and measurement procedure for the phones. Repeated runs may still be made on each phone to check how stable its readings are, but the club should count phones—not runs—as the replicated units.
Do: The 10 runs in each setting may show how battery use varied across repeated trials on this particular phone. They do not show how outcomes vary from phone to phone. If this one phone has an unusual battery or hardware condition, repeated trials will continue to reflect that phone’s characteristics. Multiple phones per setting would provide replication across phones.
Conclude about the design: The club should not report 10 experimental units per treatment. It has one phone observed repeatedly under each setting. The repeated runs may be informative for this phone, but they do not replace assigning treatments to multiple phones when the study is meant to compare phone performance more generally.
Common Mistakes and AP Exam Tips
- Counting measurements instead of units. Twenty readings from one object are not 20 experimental units. State what received the treatment and count those units.
- Assuming a large number of observations guarantees replication. A dataset can contain many measurements but only a few units assigned to treatment. Explain the assignment level, not just the total number of recorded values.
- Calling each person in a treated group a unit automatically. If a treatment was assigned to a classroom, clinic, or other group as a whole, that group is the experimental unit. Individual measurements within it do not create additional treatment assignments.
- Claiming replication removes all other problems. Replication helps show variation across units, but it does not replace random assignment, appropriate control of other variables, or careful measurement.
- Saying repeated measurements are useless. They can provide useful information about one unit or help assess measurement consistency. The specific point is that they do not count as additional experimental units.
- Leaving the count unclear. A full-credit explanation names the experimental unit and reports how many units receive each treatment. If there are repeated readings, report those separately.
For a strong AP response, identify what receives the assigned treatment, distinguish those units from repeated observations, and explain why multiple units per treatment help keep a treatment comparison from resting on a single unusual unit.
Check Your Understanding
For each situation, identify the experimental unit and decide whether the design has replication. Explain your reasoning.
- A researcher assigns one seed tray to each of two lighting conditions and measures 25 seedlings in each tray. What is the experimental unit, and how many units receive each treatment?
- A team assigns 14 separate seedlings to one light condition and 14 separate seedlings to another. What feature of this plan provides replication?
- A technician tests one printer 12 times with setting A and 12 times with setting B. Are there 12 experimental units per setting? Explain.
- A school assigns each of four classrooms to a learning activity, with two classrooms receiving each activity. What should count as the experimental units?
- In one sentence, explain why repeated measurements on a single unit do not replace multiple units per treatment.