When an Experiment Compares Two Treatments
A randomized experiment can compare two ways of treating the same kind of experimental unit. For example, researchers might randomly assign seedlings to two fertilizers, laboratory samples to two cleaning methods, or volunteers to two exercise programs. If the response is quantitative and the groups consist of different, unpaired units, the comparison of their mean responses calls for a two-sample t procedure.
The design matters as much as the arithmetic. As you learned in “Randomized Experiment Versus Observational Comparison,” random assignment uses chance to place units into treatment groups. It helps make the groups comparable at the start, so a difference in their responses can provide evidence of a treatment effect. Random assignment is not the same as random sampling: it supports cause-and-effect conclusions about the experimental units, but does not by itself make the units representative of a wider population.
Let \(\mu_A\) be the true mean response for experimental units under treatment A, and let \(\mu_B\) be the true mean response for experimental units under treatment B. The parameter of interest is the difference \(\mu_A-\mu_B\). The order matters: stating the treatment named first first makes the sign of the difference and the alternative hypothesis easier to interpret.
Use a two-sided alternative when the question asks whether the treatments produce different mean responses. Use a one-sided alternative only when the research question specifies a direction, such as whether treatment A produces a greater mean response than treatment B. Choose the direction from the question, not from the sample results. As in earlier tutorials on hypotheses for two means, the hypotheses describe population means, not the observed sample means.
The two-sample t statistic compares the observed difference in sample means with its estimated standard error. The AP Statistics default is the unpooled procedure: it does not assume that the two populations have equal variances.
Here, \(\bar{x}_A\) and \(\bar{x}_B\) are the sample means, \(s_A\) and \(s_B\) are the sample standard deviations, and \(n_A\) and \(n_B\) are the group sample sizes. A calculator uses an approximate number of degrees of freedom for this unpooled test. As covered in “Interpreting the Test Statistic for Two Means,” \(t\) measures the observed difference in estimated standard-error units.
Check the Design and Conditions
First decide whether the groups are genuinely independent. In the experiment, different units should be assigned to the two treatments, with no one-to-one pairing built into the design. If the same units receive both treatments, or units are deliberately matched, use the paired t procedure on the differences instead. “Paired t Versus Two-Sample t” and “Spotting Independent Samples in Word Problems” explain how to recognize that distinction.
Then assess whether t inference is reasonable. Random assignment supports the design condition for an experiment. The two groups should consist of distinct experimental units, and observations should be independent within and between groups. If units were also sampled without replacement from a finite population, consider the 10% condition for that sampling. Random assignment alone does not guarantee independent responses if one unit’s outcome can affect another’s.
Finally, consider the response distributions in each group. With small samples, examine the groups separately for strong skewness or outliers. With larger samples, two-sample t procedures are generally more robust to skewness, though severe skewness and extreme outliers remain concerns. A large total sample does not make a very small treatment group large; evaluate both groups.
A randomized experiment can make a significant difference in mean responses evidence that the treatments cause different responses for the experimental units, assuming the experiment was carried out appropriately. It does not automatically justify generalizing to people, objects, or settings beyond those represented by the experiment. That depends on how the units were selected, as explained in “Generalizing Conclusions to a Population.”
Worked Examples: Choosing and Using the Procedure
Worked Example: Comparing Two Fertilizers
A researcher randomly assigns 24 similar seedlings to fertilizer A or fertilizer B, with 12 seedlings in each group. After six weeks, the seedlings assigned to A have a mean growth of 18.6 centimeters and a standard deviation of 3.0 centimeters. Those assigned to B have a mean growth of 16.1 centimeters and a standard deviation of 3.0 centimeters. Separate plots show no strong skewness or outliers. Is there evidence that fertilizer A produces greater mean growth?
State: Let \(\mu_A\) be the true mean growth, in centimeters, for seedlings under fertilizer A, and let \(\mu_B\) be the true mean growth under fertilizer B. The research question is whether A produces greater growth, so \(H_0:\mu_A-\mu_B=0\) and \(H_a:\mu_A-\mu_B>0\).
Plan: Use an unpooled two-sample t test. Growth is quantitative, and different seedlings were randomly assigned to the two treatment groups, so this is an experiment with independent, unpaired groups. The plots show no strong skewness or outliers in either small group. These conditions support the test.
Do: The observed difference in sample means is \(18.6-16.1=2.5\) centimeters. The estimated standard error is \(\sqrt{3.0^2/12+3.0^2/12}=\sqrt{1.5}\), or about \(1.225\) centimeters. Thus:
The unpooled test has approximately 22 degrees of freedom. For the greater-than alternative, the p-value is approximately \(0.0267\), rounded. Assuming there is no difference in true mean growth under the fertilizers, this is the probability of obtaining a t statistic of \(2.041\) or greater in the direction of fertilizer A.
Conclude: At \(\alpha=0.05\), \(0.0267<0.05\), so reject \(H_0\). The experiment provides convincing evidence that fertilizer A causes greater mean six-week growth than fertilizer B for seedlings like those used in the experiment. Because the seedlings were not described as a random sample from a broader population, the result does not automatically generalize to all seedlings.
Worked Example: Comparing Two Air-Filter Designs
In a laboratory experiment, 32 air filters are randomly assigned to a new design or a standard design, with 16 filters per group. After the same test, the new-design filters have a mean of 42 units of particles remaining and a standard deviation of 4 units. The standard-design filters have a mean of 44.5 units and a standard deviation of 4 units. The response is approximately symmetric in both groups, with no outliers. Does the new design reduce the mean amount of particles remaining?
State: Let \(\mu_N\) be the true mean amount of particles remaining, in the stated units, for filters using the new design. Let \(\mu_S\) be the corresponding mean for the standard design. “Reduce” specifies \(H_a:\mu_N-\mu_S<0\), with \(H_0:\mu_N-\mu_S=0\).
Plan: Use an unpooled two-sample t test because the response is quantitative and different filters were randomly assigned to the two designs. The groups are independent and unpaired. Both groups have 16 observations, and their distributions are described as approximately symmetric with no outliers, so the conditions support t inference.
Do: The observed difference is \(42-44.5=-2.5\) units. The estimated standard error is \(\sqrt{4^2/16+4^2/16}=\sqrt{2}\), or about \(1.414\) units. The test statistic is:
The unpooled degrees of freedom are 30. For the less-than alternative, the left-tail p-value is approximately \(0.04364\), rounded. Assuming equal true mean amounts under the two designs, this is the probability of obtaining a t statistic of \(-1.768\) or less.
Conclude: At \(\alpha=0.05\), \(0.04364<0.05\), so reject \(H_0\). The experiment provides convincing evidence that the new filter design causes a lower mean amount of particles to remain than the standard design under the laboratory test conditions.
Worked Example: Comparing Two App Layouts
A product team randomly assigns 22 volunteers to try one of two app layouts, with each volunteer using only one layout. The time to complete a task is recorded in minutes. The first layout has a sample mean of 7.4 minutes and a standard deviation of 1.5 minutes for 10 volunteers. The second has a mean of 8.1 minutes and a standard deviation of 2.0 minutes for 12 volunteers. The plots show no strong skewness or outliers. Is there evidence of a difference in mean completion time?
State: Let \(\mu_1\) and \(\mu_2\) be the true mean completion times, in minutes, for volunteers using layout 1 and layout 2, respectively. The question asks whether the means differ, so use \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
Plan: Use an unpooled two-sample t test. The response is quantitative, and volunteers were randomly assigned to one layout each, so the groups are independent and unpaired. Both groups are small, so the stated shape checks matter; no strong skewness or outliers are present. The conditions support t inference.
Do: The observed difference is \(7.4-8.1=-0.7\) minutes. The estimated standard error is \(\sqrt{1.5^2/10+2.0^2/12}=\sqrt{0.5583}\), or about \(0.7472\) minutes. Therefore:
The unpooled degrees of freedom are approximately 19.8. The two-sided p-value is approximately \(0.3601\), rounded. Assuming the two layouts have equal true mean completion times, this is the probability of obtaining a t statistic at least as far from zero as \(-0.937\) in either direction.
Conclude: At \(\alpha=0.05\), \(0.3601>0.05\), so fail to reject \(H_0\). The experiment does not provide convincing evidence that the two layouts cause different mean task-completion times. This result does not prove that their true mean times are equal.
Common Mistakes and AP Exam Tips
- Choosing by the number of columns alone: Two columns do not guarantee a two-sample t procedure. If each row links the same unit under both treatments, analyze paired differences instead.
- Using a one-sample test against zero: The target is the difference between two population means, not one group mean compared with a fixed benchmark. State both means and the order of subtraction.
- Writing hypotheses about sample means: \(\bar{x}_A\) and \(\bar{x}_B\) are statistics. Hypotheses must describe the population means \(\mu_A\) and \(\mu_B\).
- Claiming random assignment makes a sample representative: Random assignment supports a cause-and-effect conclusion; it does not establish generalizability to a broad population. Be precise about the experimental units and setting.
- Checking only the combined distribution: For small samples, assess shape and outliers in each treatment group, not just in a pooled set of responses.
- Overstating a nonsignificant result: Failing to reject does not show that the means are equal. As in “Why a t Test Never Proves the Null Mean,” say the data do not provide convincing evidence for the stated alternative.
For full credit, connect the design to the procedure in the Plan step, show the unpooled standard-error calculation and test statistic in the Do step, and finish with the decision and an in-context evidence statement. “The treatments are different” is not a sufficient conclusion: identify which mean response is higher or lower when the evidence supports that direction, and say what the experiment allows you to infer.
Check Your Understanding
For each situation, decide whether a two-sample t procedure is appropriate and identify what the design allows you to conclude.
- Researchers randomly assign different plants to two watering schedules and measure height after a month. What feature makes this an independent-groups comparison?
- Participants are randomly assigned to one of two study routines, and the response is the number of minutes needed to complete a practice set. Name the parameter difference the test would address.
- A team randomly assigns filters to two materials and asks whether the mean amount of residue differs. Which alternative hypothesis matches the question?
- In a two-treatment experiment, each participant tries both treatments and the response is measured after each. Should the analysis use two independent samples or paired differences? Explain from the design.
- A statistically significant two-sample t test follows random assignment, but the experimental units were volunteers rather than a random sample. What can the result support, and what does it not automatically support?