Tutorials › AP Statistics › Conditions for a Two-Sample t Test

Two-sample t hypothesis tests · Tutorial 730 of 1000

Conditions for a Two-Sample t Test

Learn how design, sample size, and separate group graphs help you decide whether a two-sample t test is appropriate.

Intermediate 9 min read

What You'll Learn

  • Identify what the study design can and cannot tell you about independence.
  • Check the 10% condition separately for samples drawn without replacement.
  • Decide how closely to inspect each group’s distribution based on its sample size.
  • Use plots to look for skewness, clusters, and outliers in each group.
  • Explain when a two-sample t test is reasonable and when the conditions raise concern.

Why Check Conditions Before Comparing Two Means?

A two-sample t test compares the means of two independent groups using the difference \(\bar{x}_1-\bar{x}_2\). Before calculating a test statistic or p-value, check whether the study design and the distributions of the two groups support the procedure. A calculator can produce a result even when the conditions are not met; the number alone does not make an inference trustworthy.

As in “Conditions for a Two-Sample t Interval,” the design, independence, 10% condition, and distribution shape all matter. Here the focus is on how to assess independence and Normality for each group. Independence is mainly judged from how the data were collected. Normality is assessed by considering sample size and examining each group’s graph, especially when the samples are small.

Definition: The Nearly Normal condition for a two-sample t test means that each group’s data are reasonably compatible with using a t procedure. For small samples, look for distributions that are roughly symmetric and have no strong outliers. Larger samples can usually tolerate more departure from Normality, but an extreme outlier or very unusual shape can still be concerning.

The condition is about the two group distributions separately. Do not combine the groups into one graph to assess Normality, and do not graph differences as you would for paired data. In “Identifying Two Independent Samples” and “Checking Normality of the Differences,” earlier in this course, the distinction between independent groups and paired observations was established. For a two-sample test, assess the observed values in group 1 and group 2 separately.

A Practical Check: Design First, Then Each Group’s Shape

Start with the observational units and the data-collection plan. Ask whether the groups are genuinely independent and whether observations within each group can reasonably be treated as independent. Then check the 10% condition if observations were sampled without replacement. Only after these design checks should you use sample sizes and graphs to assess whether the t procedure is suitable.

1
Describe the design.
Identify how the units entered the study and how they were placed into groups. Determine whether the groups are separate and unpaired, or whether the same units or matched units contribute to both groups.
2
Check independence and the 10% condition.
Explain why observations within each group and across groups can reasonably be treated as independent. If sampling without replacement, check that each sample is less than 10% of its own population.
3
Record both sample sizes.
Judge group 1 and group 2 separately. A large sample in one group does not automatically compensate for a very small sample in the other.
4
Inspect each group’s graph.
For small samples, look carefully for symmetry, skewness, gaps, clusters, and outliers. With larger samples, moderate non-Normality is generally less concerning, but severe skewness or extreme observations still deserve attention.
Conditions:
  • Random design: Data should come from appropriate random samples or a randomized experiment. Identify which kind of random process was used.
  • Independent groups and observations: The groups should not be paired, and observations within each group should be reasonably independent. Use information about the design; a graph cannot establish independence.
  • 10% condition: For each sample drawn without replacement, the sample size should be less than 10% of that sample’s population.
  • Nearly Normal condition: Consider the sample size and inspect the distribution in each group separately. With small samples, look for rough symmetry and no strong outliers. Larger samples are more resistant to moderate departures from Normality.

In AP Statistics, a sample size of about 30 or more in each group is a useful guideline for when the t procedure is generally more resistant to non-Normality. It is not a rule that makes every graph acceptable: substantial skewness or a very extreme outlier can still be problematic. Nor is it a requirement that every sample have at least 30 observations. With smaller samples, the group plots need to provide stronger support for the procedure.

A histogram shows overall shape, while a dotplot or boxplot can make individual observations and possible outliers easier to notice. For small samples, a dotplot is often especially useful because it shows every value. Use a display that preserves the identity of each group, and describe what it shows rather than claiming that a graph proves the population is Normal.

Worked Examples

Worked Example: Small Samples with Reasonable Shapes

A hypothetical community garden program compares the weekly watering time, in minutes, for plots using two different irrigation timers. Separate random samples of plots are selected from a large set of eligible plots: \(n_1=14\) plots use Timer A and \(n_2=16\) use Timer B. Each plot uses only one timer. Dotplots show both groups to be roughly symmetric, with no gaps or isolated values far from the rest. Assess whether a two-sample t test is reasonable.

Random design: The plots were selected randomly from the eligible plots, so the sampling design supports inference to that target group of plots, assuming the rest of the conditions are reasonable.

Independence: Each plot contributes one watering-time observation and uses only one timer, so the groups are unpaired. The garden staff report that plots are managed separately, making it reasonable to treat observations within each group and across groups as independent. Since the samples were drawn without replacement, check the 10% condition for each population: 14 plots are less than 10% of the eligible Timer A plots, and 16 plots are less than 10% of the eligible Timer B plots.

Nearly Normal: Both sample sizes are small, so the plots need careful inspection. The dotplots show no strong skewness, unusual gaps, or outliers in either group. The evidence from each group’s shape supports using a t procedure.

Conclusion: The random sampling, reasonable independence, separate 10% checks, and acceptable shapes in both groups support using a two-sample t test to compare the mean watering times. The graph of one group does not stand in for the other; both groups were checked.

Worked Example: Larger Samples with Some Skewness

In a hypothetical study of two independent delivery zones, random samples of \(n_1=42\) and \(n_2=36\) completed deliveries are selected from large zone-specific records. The quantitative variable is delivery time in minutes. Each delivery belongs to just one zone. Histograms show mild right-skewness in both groups but no isolated extreme values. Assess the conditions for a two-sample t test.

Random design: Each zone’s deliveries were randomly sampled from its own records. This supports inference about the mean delivery time in each zone’s recorded delivery population.

Independence: Different deliveries form the two groups, with no matching of a delivery in one zone to a delivery in the other. The records indicate that no driver completed multiple sampled deliveries during the same shift, so there is no evident repeated-measure link among the selected observations. The group sizes are less than 10% of the corresponding large zone records, satisfying the 10% condition for each sample.

Nearly Normal: Inspect each histogram separately. The distributions have mild right-skewness, but both sample sizes are at least about 30 and neither graph shows an extreme outlier. These relatively large samples make the t procedure more resistant to the moderate skewness observed.

Conclusion: The design and distribution checks support a two-sample t test. The sample sizes help here because they are large in both groups; the conclusion does not rely on treating the two groups as a single combined distribution. If the plots had shown an extreme outlier or severe skewness, the large-sample guideline alone would not settle the concern.

Worked Example: A Strong Outlier in a Small Group

A hypothetical study compares the number of minutes spent on a daily outdoor activity by randomly sampled members of two separate recreation clubs. The first group has \(n_1=10\); the second has \(n_2=12\). Each person belongs to one club and contributes one observation. Both samples are less than 10% of their respective club populations. The first group’s dotplot is fairly balanced, but the second group’s dotplot has most values between 20 and 35 minutes and one value at 110 minutes. Assess the conditions.

Random design: The two samples were randomly selected from their respective club populations, which supports generalizing to those populations if the remaining conditions hold.

Independence: People are not matched across clubs, and each contributes one value. Suppose the study also confirms that participants’ activity times are not linked through a shared household or another obvious sampling cluster. The groups can reasonably be treated as independent. The 10% condition is met separately for both club samples.

Nearly Normal: Both samples are small, so inspect each graph closely. The first group’s shape is not concerning, but the second group has a strong high outlier relative to the rest of its observations. With only 12 observations, that one value may have a substantial effect on the sample mean and standard deviation. The Nearly Normal condition is questionable for the second group.

Conclusion: The sampling and independence checks do not remove the concern about the second group’s shape. The conditions do not adequately support proceeding with the usual two-sample t test without further investigation. Check that the value is not a recording error and describe the limitation; do not simply delete a valid observation or combine the groups to hide it.

Worked Example: Random Assignment Does Not Fix Dependence

A hypothetical experiment tests two study schedules. Researchers randomly assign 24 students to Schedule A and 24 to Schedule B, then record each student’s quiz score. However, the students work in four study rooms, and all students in a given room follow the same schedule and share study materials. Can the ordinary two-sample t test conditions be assumed?

Random design: The schedules were randomly assigned, which is an appropriate randomization process for an experiment.

Independence: Each student is assigned to only one schedule, so the groups are not paired. But the shared rooms and materials may make scores within a room related. Random assignment by itself does not guarantee that all individual observations are independent. The study’s room structure needs to be considered; if the room is the unit that effectively received a schedule, counting all 48 students as independent observations could overstate the information in the data.

Nearly Normal: Even if the two score histograms look roughly symmetric, that would not resolve the dependence concern. A graph describes the observed values; it cannot show whether students’ results are linked by their shared room.

Conclusion: The ordinary two-sample t test conditions cannot be justified from the information given. The experiment has random assignment, but the possible room-level dependence remains. This example shows why condition checking starts with the data-collection and assignment design, not just the sample sizes or graphs.

Common Mistakes and AP Exam Tips

A strong conditions statement gives evidence for each condition in the particular study. Avoid a bare statement such as “the data are Normal” or “the samples are independent.” State how the design and the group plots support—or fail to support—the procedure.

  • Checking only one group’s graph: Examine group 1 and group 2 separately. A reasonable shape in one group does not compensate for a troubling shape in the other, especially when both samples are small.
  • Using a pooled graph: A combined display can conceal differences in shape between the groups. Keep the group identities visible and discuss each distribution on its own.
  • Claiming independence from random assignment alone: Random assignment supports an experimental comparison, but repeated measurements, matching, clustering, or shared conditions can still create dependence. Explain what each unit contributes and how units are linked.
  • Treating 30 as an automatic pass: About 30 observations per group is a helpful guideline, not a guarantee. Mention severe skewness or extreme outliers if the graphs show them.
  • Calling every unusual value an outlier: Describe what the plot shows. If a value is unusually distant from the rest, investigate it and consider its effect; do not remove a valid observation just because it is inconvenient.
  • Forgetting the two 10% checks: When sampling without replacement, compare each sample size with its own population size. Do not check only the combined sample against one population.
  • Confusing paired and independent data: If the same units provide both measurements, or there is deliberate matching, the data are paired. A two-sample t test is for two independent groups, not for the distribution of paired differences.
Key takeaway: Independence is assessed from the study design; it cannot be established by looking at a graph. For Normality, use the sample size and inspect each group separately. Small samples need roughly symmetric distributions without strong outliers, while larger samples can tolerate moderate departures from Normality—but not every extreme shape.

Check Your Understanding

For each situation, think about the design and the two group distributions separately.

  1. Two independent samples have sizes 9 and 34. Why should the smaller group’s graph receive especially careful attention?
  2. A researcher randomly assigns people to two treatments, but each person measures their response three times. Does random assignment alone make all the recorded observations independent? Explain.
  3. When samples are drawn without replacement from two different populations, how should the 10% condition be checked?
  4. Both groups have sample sizes of 40, and one histogram shows mild skewness without extreme values. What does the sample size suggest, and what should still be reported about the graph?
  5. Why is a graph of all observations combined not enough to assess the Nearly Normal condition for a two-sample t test?