Tutorials › AP Statistics › Mixed Practice: Designing and Critiquing Experiments

Experimental design · Tutorial 200 of 1000

Mixed Practice: Designing and Critiquing Experiments

Work through exam-style experimental design problems and learn to identify flaws, propose practical fixes, and explain what a well-designed experiment can conclude.

Beginner 9 min read

What You'll Learn

  • Identify experimental units, treatments, and responses in unfamiliar scenarios
  • Explain how random assignment, control, and replication strengthen a comparison
  • Recognize confounding and revise a plan to separate treatment effects from other conditions
  • Choose between a completely randomized design and a randomized block design
  • Distinguish what random assignment supports from what random selection supports
  • Write concise, specific explanations suited to experimental design questions

Practice the Decisions Behind an Experimental Design

In Writing a Complete Experimental Design Description, you practiced laying out a study so its treatment comparison can be understood and carried out. This tutorial puts those skills to work in unfamiliar situations. You will decide what a scenario gets right, identify what could undermine its comparison, and write a more effective plan.

Design questions often ask you to do more than name a method. A prompt might ask you to explain why a researcher’s conclusion is not justified, recommend a fix, or describe a complete experiment. A strong answer connects each design choice to the study’s question: who or what receives the treatments, what differs between groups, what is measured, and what conclusions the design could support.

Key idea: Critique a design by tracing how units are assigned, how treatment conditions differ, and how the response is measured. Explain why each identified feature matters, then propose a specific change that improves the comparison.

A Quick Audit for Design Problems

Before recommending a fix, separate the parts of a scenario. As introduced in Experimental Units, Factors, and Treatments, the experimental units are the individuals or objects assigned to treatments; the factor is what researchers deliberately vary; and the treatments are the specific conditions. Then check whether the plan can compare those treatments fairly.

1
Locate the assignment.
Identify what receives a treatment and whether chance, rather than preference or convenience, determines its assignment.
2
Compare the conditions.
Ask whether groups differ in a planned way other than the treatment. Look for differences in setting, schedule, instructions, or who administers the treatment.
3
Check the response and replication.
Identify exactly what will be measured, how it will be recorded, and whether multiple experimental units receive each treatment.
4
Match the conclusion to the design.
Random assignment can support a cause-and-effect conclusion for the experimental units. Generalizing to a wider population depends on how those units were selected.

This audit helps distinguish a design flaw from a limitation on the conclusion. For example, using volunteers may limit generalization but does not automatically prevent a fair treatment comparison if volunteers are randomly assigned. By contrast, letting participants choose their treatments can make the groups systematically different before the treatment begins.

A useful critique is specific and actionable. “There may be bias” does not explain the problem. Instead, name the feature that could affect the outcome and describe how to address it. If one group works in a quiet room and another in a noisy room, for instance, use the same room conditions for both groups or balance the rooms across treatments.

Worked Example: Fixing a Confounded Study-Space Comparison

A school wants to compare studying in silence with studying while instrumental music plays. Thirty-two volunteers agree to participate. The researcher assigns the first 16 volunteers to silence in the library and the last 16 to music in a classroom. Each student studies the same material for 25 minutes, then completes the same 18-question quiz. The researcher plans to compare the number of correct answers, measured in items correct out of 18.

State: The experimental units are the 32 participating students. The factor is study sound condition, with silence and instrumental music as its two levels. The response is each student’s quiz score, measured in items correct out of 18. The intended question is whether the study sound condition affects quiz performance for these participants.

Plan: The current plan does not use random assignment: the first 16 students receive silence and the last 16 receive music. Order of arrival could be associated with other differences, such as class schedule or student habits. In addition, sound condition is confounded with room: every student in silence uses the library, while every student listening to music uses a classroom. A difference in scores could be due to sound condition, room, or both.

Revise the plan as a completely randomized design, as described in Completely Randomized Design. Label the students 01 through 32 and use a chance process to assign 16 distinct labels to silence; assign the remaining 16 students to instrumental music. Have both groups study in the same room under the same lighting and seating conditions. Give everyone the same material, study time, quiz, instructions, and scoring procedure. Replication is present because 16 different students receive each sound condition.

Do: Carry out the study and record one quiz score for each student in items correct out of 18. Compare the score distributions, or compare the groups’ mean quiz scores, to describe how performance differed in this experiment. No results are supplied here, so a numerical difference cannot be calculated.

Conclude: With random assignment and the room condition controlled, the study could provide evidence about a cause-and-effect difference between silence and instrumental music for the 32 participating volunteers. Because they volunteered rather than being randomly selected from a defined population of students, the results do not automatically generalize to all students.

Worked Problems: Choose a Design and Defend It

A design does not have to be complicated to be effective. The right choice depends on the experimental units and on whether a known feature of those units is likely to be related to the response. In Choosing Among Completely Randomized, Block, and Matched Pairs Designs, you learned how those designs differ. The next problem asks you to apply that choice and explain it in context.

Worked Example: Blocking on Typing Experience

A community center plans to compare two keyboard-practice programs for 24 adult learners. Before the experiment, each learner completes a short typing assessment. The center classifies learners as either beginner or experienced, with 12 learners in each category. It will compare a game-based program with a lesson-based program. Each learner will practice for 20 minutes on four days and then take the same typing test. The response is typing speed in words per minute.

State: The experimental units are the 24 adult learners. The factor is practice program, with game-based and lesson-based programs as its two levels. The response is each learner’s typing speed on the final test, measured in words per minute. Prior typing experience is a reasonable blocking variable because it may be related to typing speed.

Plan: Use a randomized block design. Within the 12 beginner learners, use chance to assign 6 to the game-based program and 6 to the lesson-based program. Separately, within the 12 experienced learners, use chance to assign 6 to each program. In total, 12 learners receive each program, and both programs are represented equally within each experience category.

Give both groups the same practice duration, number of practice days, typing content, equipment, test instructions, and final test. The programs themselves should be clearly specified so the planned difference is the method of practice rather than different amounts of practice. The design includes replication because multiple learners receive each program. Blocking does not guarantee that learners within a category are identical; it organizes assignment so that comparisons can be made within experience categories.

Do: Record each learner’s final typing speed in words per minute. Compare the programs’ results while accounting for the two experience blocks—for example, examine the difference between program outcomes among beginners and the difference among experienced learners. No typing-speed results are provided, so the direction or size of a program difference cannot be determined.

Conclude: If carried out as planned, random assignment within experience categories can support a cause-and-effect comparison of the two programs for these participants. The design alone does not establish which program performs better; that requires the measured results. Since the learners are participants from one community center rather than a random sample of all adult learners, broad generalization is not automatically justified.

Notice how the explanation names why experience is used for blocking instead of merely saying that the study “uses blocks.” A block should be based on a characteristic related to the response, and treatments are randomly assigned within each block. This keeps the design’s logic visible to the reader.

Worked Example: Finding the Problem with a Plant Experiment

A student wants to compare two fertilizers on the growth of 20 seedlings. The student puts all 10 seedlings receiving Fertilizer A on the sunny side of a greenhouse bench and all 10 receiving Fertilizer B on the shaded side. Each seedling receives the same amount of water. After three weeks, the student will measure height in centimeters.

The experimental units are the 20 individual seedlings, the factor is fertilizer type, and the response is seedling height after three weeks, measured in centimeters. Both fertilizers are applied to multiple seedlings, so there is replication. The main flaw is that fertilizer type is confounded with bench location. If seedlings on the sunny side grow differently because of light exposure, the student cannot separate a fertilizer effect from a location effect.

A stronger plan would label the seedlings 01 through 20 and use chance to assign 10 to Fertilizer A and 10 to Fertilizer B. Distribute seedlings from both treatment groups across the sunny and shaded parts of the bench, rather than putting each fertilizer in just one location. Apply the same amount of water and fertilizer according to a consistent schedule, and measure every seedling in the same way after three weeks. If location is expected to affect growth, the student could instead form location blocks and randomly assign both fertilizers within each block.

The revised plan can support a cause-and-effect comparison for these 20 seedlings because fertilizer is assigned by chance and location is no longer tied to just one fertilizer. It still would not justify a claim about all plants or all growing conditions without an appropriate basis for generalizing. Merely measuring many seedlings does not fix the original confounding: the treatment assignments and locations also need to be arranged so the comparison is meaningful.

Common Mistakes and AP Exam Tips

  • Calling a study randomized because it has two groups. Two groups alone do not show random assignment. State how chance assigns the experimental units to treatments.
  • Suggesting “use random assignment” without correcting other problems. If treatment is tied to a different room, time, or location, address that feature too. Random assignment is valuable, but a plan should also keep relevant conditions comparable or balance them across treatments.
  • Confusing blocking with random sampling. Blocks organize treatment assignment among experimental units; strata organize selection from a population. Name which process the scenario uses, as in Blocking Versus Stratifying.
  • Claiming random assignment makes volunteers representative. Random assignment supports a causal comparison for the units in the experiment. Random selection is the method relevant to generalizing to a population.
  • Writing a vague response variable. “Growth” or “performance” is not sufficiently precise. State what will be measured and give the measurement units, such as height in centimeters or typing speed in words per minute.
  • Stating a flaw without explaining its consequence. A strong critique completes the reasoning: name the design feature, explain how it could affect the response or comparison, and recommend a concrete fix.
  • Claiming one treatment is better before results exist. A design can show how a question will be investigated. Which treatment performs better depends on the outcomes collected and compared.

For full credit, make each claim do a job. If you identify confounding, name both the treatment and the background condition that vary together. If you recommend random assignment, specify what is assigned and how chance is used. If you discuss a conclusion, connect its scope to the design’s random assignment and the way participants entered the study.

Key takeaway: A strong experimental design creates a clear treatment comparison: identify the units, treatments, and measured response; use chance to assign treatments; provide replication; and control or balance relevant conditions. Then explain conclusions only as broadly as the design allows.

Check Your Understanding

For each scenario, identify the key design issue or propose a plan, and explain your reasoning in context.

  1. A coach lets athletes choose whether to use a new warm-up or their usual warm-up, then compares sprint times. Identify a possible problem and describe one way to improve the treatment comparison.
  2. A researcher compares two reminder apps. One group receives reminders in the morning and the other in the evening. Explain why this plan may not isolate the effect of the app, and suggest a fix.
  3. A garden club compares two watering schedules using 18 plants. Name the experimental units, factor levels, and a suitable response with measurement units. Describe a chance assignment with replication.
  4. A study randomly assigns volunteers at one school to two review methods. State what random assignment supports and what it does not establish about students at other schools.
  5. A researcher divides participants into blocks based on experience, then assigns everyone in one block to Treatment A and everyone in the other block to Treatment B. Explain why this does not provide a randomized comparison of treatments within blocks.