Tutorials › AP Statistics › Randomized Block Design

Experimental design · Tutorial 192 of 1000

Randomized Block Design

See how grouping similar experimental units before random assignment can account for an important source of variation in an experiment.

Beginner 9 min read

What You'll Learn

  • Define a randomized block design and explain why researchers use it.
  • Choose a block variable that is known before treatment and related to the response.
  • Describe random assignment of treatments separately within each block.
  • Use block sizes and treatment counts to plan a balanced assignment.
  • Explain how blocking differs from controlling a condition across all groups.
  • Identify what blocking can and cannot establish about a treatment.

Use a Relevant Difference to Improve an Experiment

In Confounding in Experiments, you learned that a treatment comparison is difficult to interpret when a background condition changes systematically with treatment. One way to plan a more informative comparison is to identify an important source of variation before assigning treatments, group similar experimental units together, and randomize treatments separately within each group.

This approach is called a randomized block design. Blocking does not mean giving every unit the same treatment. Instead, it organizes the random assignment so that each treatment is represented among units with similar values of a relevant characteristic. This can make treatment comparisons less affected by that characteristic.

Definition: In a randomized block design, researchers divide experimental units into groups, called blocks, based on a variable that is related to the response. They then randomly assign treatments to units separately within each block.

A block variable might be a measurement taken before the experiment, such as starting fitness, or a characteristic such as age group or prior experience. The researcher chooses it because units with similar values are expected to have more similar responses than units with quite different values. Blocks can be based on categories that already exist or on ranges of a measured value.

The block variable is not the treatment factor. The treatment factor is what the experiment deliberately changes; the block variable helps organize the assignment. For example, a study could compare two training plans while using athletes’ starting fitness levels to form blocks. Fitness is not what the study assigns—the training plan is.

Plan the Blocks, Then Randomize Within Them

A useful block variable is known before treatment begins and is meaningfully related to the response. Researchers should choose blocks that are practical to form and large enough to assign the planned treatments within them. A block is not useful simply because it creates more groups; the grouping should help make units within a block more alike in a way that matters for the response.

Once blocks are defined, the random assignment happens separately inside each block. If there are two treatments and a block contains 12 units, the researcher might randomly assign 6 units to each treatment. The same process is then carried out in every other block. Equal numbers are often convenient, though the essential feature is that chance determines which units in a block receive which treatments.

1
Name the treatment factor and response.
State what the experiment will assign and what outcome will be measured.
2
Choose a relevant block variable.
Use information available before treatment that is expected to be related to the response.
3
Form the blocks.
Group units with similar values of the block variable. Specify the categories or ranges used.
4
Randomize within each block.
Use a chance process in every block to assign its units to the treatment groups.
5
Apply treatments consistently and measure the response.
Keep other procedures comparable across treatment groups and use the same response measurement process.

Blocking and random assignment have related but different jobs. Blocking puts units with similar values of a relevant characteristic together. Random assignment within each block helps prevent the researcher from choosing which units receive which treatment. As in Why Random Assignment Is the Key to Causation, chance assignment helps make treatment groups comparable on average; blocking also ensures that each treatment is represented within the chosen categories or ranges.

Blocking is part of experimental design, not a method for selecting a sample from a population. It does not, by itself, make the experimental units representative of a broader population. As discussed in Scope of Inference: Four Combinations, random selection and random assignment address different questions.

Worked Example: Comparing Two Rehabilitation Routines

Worked Example: Block on Starting Mobility

A fictional rehabilitation clinic wants to compare two exercise routines for improving patients’ knee mobility over four weeks. The response is the change in knee-flexion range, measured in degrees. The clinic expects starting mobility to be related to the amount of change, so it records each patient’s range before treatment. It has 24 patients: 12 with lower starting mobility and 12 with higher starting mobility.

State: The treatment factor is exercise routine, with Routine A and Routine B as the two treatments. The response is change in knee-flexion range after four weeks. Starting mobility is the proposed block variable because it is measured before treatment and may be related to the response.

Plan: Form two blocks of 12 patients each: lower starting mobility and higher starting mobility. Randomly assign six patients in each block to Routine A and the other six to Routine B. This puts both routines into both starting-mobility categories.

Do: In the lower-mobility block, label the patients 01 through 12. Suppose a chance process selects labels 02, 04, 05, 07, 09, and 12 for Routine A; the other six receive Routine B. In the higher-mobility block, label its patients 01 through 12 separately and use a new chance process to select six for Routine A; the remaining six receive Routine B. The assignment is conducted independently within the two blocks.

Check the design: Each treatment has six patients in each block, so neither routine is assigned only to patients with lower or only to patients with higher starting mobility. Both routines can be compared among patients who began in the same mobility category. The clinic should use the same four-week schedule and the same method for measuring knee-flexion change for both routines.

Conclude about the plan: This is a randomized block design because patients are grouped by starting mobility and then randomly assigned to a routine within each group. If the routines later produce different responses, the design supports a treatment comparison that accounts for the starting-mobility categories. The plan alone does not tell us which routine works better; that requires observing the outcomes.

The important feature is not the particular labels selected for Routine A. A different valid random assignment could select different patients. What makes the design blocked is that the clinic performs a random assignment within each starting-mobility group, rather than assigning all 24 patients together without using the mobility information.

Worked Example: Training Plans for Runners

Worked Example: Block on Previous Race Performance

A fictional running club wants to compare two eight-week training plans. Its response is each runner’s time, in seconds, on a specified course after the program. The club expects previous race performance to be related to the final time. It has 18 runners and groups them into three blocks of six using their previous course times: faster, middle, and slower.

State: The treatments are Training Plan A and Training Plan B. The response is post-program course time in seconds. Previous race performance is the block variable because it is known before the program and is expected to be related to the later time.

Plan and do: Within each block of six runners, randomly assign three to Plan A and three to Plan B. For instance, the club could label the six runners in a block, shuffle six identical slips with three marked A and three marked B, and give one slip to each runner. It repeats this chance process separately for the faster, middle, and slower blocks. The result is nine runners on each plan overall, with three from each performance block on each plan.

Check the design: Both plans are represented in all three performance blocks, and assignment within each block is determined by chance. The club should make the training schedules, course, timing procedure, and other relevant instructions comparable across plans, apart from the assigned training plan.

Conclude about the plan: This design accounts for the runners’ previous performance when organizing assignment. It does not guarantee that the two treatment groups will have identical abilities or that the final times will differ. It creates a fairer basis for comparing the plans than assigning one plan to all faster runners and the other to all slower runners.

A block variable need not be a treatment or something the researcher can control. The club cannot change runners’ earlier performance, but it can use that information to make the treatment assignments within each performance group. The treatment contrast remains the training plans; the blocks help structure that contrast.

Worked Example: Testing Two Watering Schedules

Worked Example: Block on Initial Soil Moisture

A fictional community garden tests two watering schedules on 20 planting beds. The response is the number of kilograms of produce harvested from each bed over a month. Because initial soil moisture may be related to plant growth, the garden measures moisture before the experiment and forms two blocks: 10 beds with lower initial moisture and 10 with higher initial moisture.

State: The treatment factor is watering schedule, with Schedule A and Schedule B. The response is harvest mass in kilograms per bed. Initial soil moisture is the block variable, measured before the treatments begin.

Plan and do: In each block of 10 beds, randomly assign five beds to Schedule A and five to Schedule B. For example, number the beds in each block from 1 to 10, use a chance process to select five labels for Schedule A, and assign the other five to Schedule B. Repeat with a separate random selection in the other block.

Check the design: Both watering schedules are used among beds with lower initial moisture and among beds with higher initial moisture. The garden should keep other relevant procedures, such as the type of seed and harvest measurement, consistent across the schedules.

Conclude about the plan: This is a randomized block design because beds are grouped by a variable expected to relate to harvest and treatments are randomly assigned within each group. Comparing the schedules within both initial-moisture categories can help separate the schedule comparison from differences associated with initial moisture. The design does not establish the harvest results in advance.

What Blocking Helps With—and What It Does Not

Blocking is most useful when the chosen variable is related to the response and the treatment groups can be compared within each block. If starting ability is related to a test score, for example, forming blocks by ability can keep the treatment comparison from depending only on which group happened to receive the more experienced participants. The same general idea applies to initial plant conditions, prior performance, or a baseline health measurement.

Blocking does not remove all differences among experimental units. Two people in the same starting-mobility block can still differ in many ways. Nor does blocking guarantee that treatments will have equal outcomes. It is a design choice that accounts for a particular known source of variation while preserving random assignment.

The block variable should be measured or known before the treatment is applied. Grouping units according to a response measured after treatment would not serve the same planning purpose: the treatment may itself have influenced that response. Also, a researcher should not claim that blocking fixes an unfair procedure if other conditions still track treatment. As in Controlling Variables Across Treatment Groups, relevant procedures should remain comparable where practical.

Key takeaway: A randomized block design forms groups of similar experimental units using a variable related to the response, then randomly assigns treatments separately within each block. Blocking organizes the comparison; random assignment provides the chance-based treatment allocation.

Common Mistakes and AP Exam Tips

  • Calling the block variable a treatment. Starting mobility or previous performance is used to form groups; the exercise routine or training plan is what is assigned. Name these roles separately.
  • Randomly assigning everyone together. If the design is blocked, describe a chance process within each block, not just a single random assignment across all units.
  • Blocking on a variable unrelated to the response. State why the proposed variable is expected to matter for the outcome. A block variable should have a design purpose, not simply create additional categories.
  • Forgetting that each treatment needs to appear in each block. A block design compares treatments within blocks. If one treatment appears only in one block, the intended within-block comparison is missing.
  • Claiming that blocking guarantees a treatment effect or eliminates all variation. Blocking improves the design of the comparison; it does not predetermine the results or make units identical.
  • Confusing random assignment with random selection. Random assignment allocates experimental units to treatments. It does not, by itself, make those units a random sample from a larger population.

For a full-credit description, name the treatments and response, identify a block variable related to the response, state how the blocks are formed, and explain that treatments are randomly assigned within every block. Then say what this organization helps account for without claiming that it guarantees a particular outcome.

Check Your Understanding

For each situation, decide whether the proposed plan is a randomized block design and explain the role of the block variable and random assignment.

  1. A study compares two reading programs. Researchers group students by a reading assessment taken before the study, then randomly assign students within each group to a program. Identify the treatments, response, and block variable.
  2. A team compares two nutrition plans but assigns every experienced athlete to Plan A and every new athlete to Plan B. What feature of a randomized block design is missing?
  3. A garden compares two seed treatments. It groups plots by initial soil acidity and randomly assigns both treatments within each acidity group. Why might initial acidity be a useful block variable?
  4. Does forming blocks by previous performance make the experimental units a random sample of all people who might use the treatment? Explain.
  5. In one sentence, describe how a researcher should assign treatments within each block of 10 units when comparing two treatments equally.