Tutorials › AP Statistics › Conditions for a Two-Sample t Interval

Two-sample t confidence intervals · Tutorial 715 of 1000

Conditions for a Two-Sample t Interval

Practice checking each group’s design, population size, and distribution before using a two-sample t interval.

Intermediate 9 min read

What You'll Learn

  • Identify the random-sampling or random-assignment evidence needed for a two-sample t interval.
  • Check the 10% condition separately for samples drawn without replacement from each population.
  • Decide whether each group’s distribution is sufficiently well behaved or its sample is large enough.
  • Use plots and design information to distinguish a condition that is met from one that cannot be verified.
  • Explain how a failed condition affects whether the interval is trustworthy.

Check the Design Before Calculating

In “How Variability Affects the Width of a Two-Sample Interval,” you saw how sample sizes and standard deviations affect an interval’s precision. Before calculating an interval, first ask whether the data support using the procedure at all. For a two-sample t interval, assess the study’s randomness, the 10% condition when sampling without replacement, and the distribution or sample size in each group.

The interval estimates \(\mu_1-\mu_2\), the difference between the two population means, with group 1 minus group 2. As established in “Identifying Two Independent Samples,” the groups must also consist of independent, unpaired observations. A design that links the observations—for example, measuring the same people twice—calls for paired methods instead.

Conditions: Before using an unpooled two-sample t interval, check that the data come from appropriate random samples or a randomized experiment; that observations can reasonably be treated as independent, including the 10% condition for each sample drawn without replacement; and that each group’s population distribution is approximately Normal or its sample is large enough for the t procedure to be reasonably robust.

Randomness and the 10% Condition

For a population-mean interval based on samples, look for a random sample from each population. Random selection helps make the sample representative of its population, supporting generalization to that population. A study description that merely says “participants volunteered” or “people were selected” does not establish random sampling.

In an experiment, subjects may instead be randomly assigned to treatment groups. Random assignment supports a fair comparison of the treatments. It is not the same as random sampling: random assignment alone does not make the subjects representative of a larger population. Keep the purpose of the random process clear when describing what a result can establish.

The 10% condition helps justify treating observations within a sample as independent when sampling without replacement from a finite population. Check it separately for each group: the sample size should be no more than 10% of the population from which that group was drawn. If the population size is \(N_i\), the check is \(n_i\leq0.10N_i\), or equivalently \(10n_i\leq N_i\).

$$ \text{Group 1: } n_1\leq0.10N_1 \qquad\text{and}\qquad \text{Group 2: } n_2\leq0.10N_2 $$

Do not check the combined sample size against just one population. Each group may represent a different population, with a different population size. If a problem does not give enough information to verify the condition, say what population size would be required or state that the condition cannot be confirmed from the information provided. Do not simply assume it is met.

Normality or Large Samples: Check Both Groups

The t procedure works well when the sampling distribution of the difference in sample means is reasonably well behaved. If the sample sizes are small, assess the shape of the data in each group. Dotplots, histograms, or boxplots can reveal strong skewness or outliers. With small samples, evidence that both groups’ distributions are approximately Normal and have no strong outliers supports using the interval.

A commonly used AP Statistics guideline is that samples of at least 30 observations in each group are large enough for the t procedure to be reasonably robust to some departures from Normality. This is a rule of thumb, not a guarantee that every large sample is suitable. Strong skewness or especially extreme outliers can still be concerning, so consider the plots and the context as well as the sample sizes.

Assess the groups separately. A large sample in group 1 does not compensate for a small, strongly skewed sample in group 2. Nor should you combine the observations from both groups into one graph and judge only the combined shape: the procedure estimates a mean for each population, so the shape within each group matters.

Group-by-group check: For each group, record its sample size, whether the sampling or assignment process was random, the relevant population size if sampling without replacement, and the evidence about its distribution. The Normality-or-large-sample condition is supported if both groups have approximately Normal distributions or both samples are large enough, with no severe distributional problem that undermines the method.

A Practical Condition-Checking Routine

Use a consistent order so a condition does not get overlooked. First identify the two populations and confirm that the observations are independent rather than paired. Then describe the random process and check the 10% condition where it applies. Finally, assess each group’s sample size and distribution. If a condition is not established, explain the limitation instead of presenting an unsupported interval as reliable.

1
Identify the design.
Confirm that the groups are independent and unpaired. State whether there are random samples, random assignment, or neither.
2
Check sampling independence.
For each sample drawn without replacement, compare its size with 10% of its own population. Note if the population size is unknown.
3
Check group 1 and group 2.
For each group, use its sample size and a plot or description of its distribution to assess Normality or whether a large sample makes the t procedure reasonable.
4
Give a clear judgment.
Say whether the conditions support a two-sample t interval, and name any condition that is questionable or cannot be verified.

Worked Examples

Worked Example: Small Samples With Reasonable Distributions

An invented study compares the mean number of minutes two types of rechargeable lantern operate on one charge. Researchers take independent random samples of 18 lanterns of type A from a population of 240 and 20 lanterns of type B from a population of 260. Plots of each group show roughly symmetric distributions without apparent outliers. Assess the conditions for a two-sample t interval for \(\mu_A-\mu_B\), the difference in true mean operating time, in minutes.

State: The parameter is \(\mu_A-\mu_B\), the true mean operating time of all type A lanterns minus that of all type B lanterns, in minutes.

Plan: Use the condition checks for an unpooled two-sample t interval. The samples must be random, observations independent, and the Normality-or-large-sample condition supported in both groups.

Do: Both groups are described as independent random samples, so the randomness condition is met and the groups are not paired. Check the 10% condition separately:

$$ \frac{18}{240}=0.075=7.5\%, \qquad \frac{20}{260}\approx0.0769=7.69\% $$

Each sample is less than 10% of its population, so the 10% condition is met for both groups. Neither sample is large, since 18 and 20 are below 30. The plots therefore matter: both are described as roughly symmetric with no apparent outliers, which supports the Normality condition for these small samples.

Conclude: The stated evidence supports using a two-sample t interval for the difference in mean operating times. In particular, randomness, the 10% condition, and the distribution check for each group have been addressed.

Worked Example: One Small Sample Has a Strong Outlier

An invented school study compares the mean minutes per day students at two schools spend reading for pleasure. Researchers take independent random samples of 14 students from a school with 400 students and 16 students from a school with 500 students. The first group’s plot is roughly symmetric with no apparent outliers. The second group’s plot shows strong right skew and one exceptionally high value. Assess whether a two-sample t interval is supported.

State: The target is \(\mu_1-\mu_2\), the true mean daily reading time for all students at school 1 minus the true mean daily reading time for all students at school 2, in minutes.

Plan: Check the random samples, the 10% condition for each school, and the Normality-or-large-sample condition separately for each group.

Do: Both samples are described as random, and the groups are independent. For the 10% condition:

$$ \frac{14}{400}=0.035=3.5\%, \qquad \frac{16}{500}=0.032=3.2\% $$

Both percentages are below 10%. However, the sample sizes, 14 and 16, are small. The first group’s roughly symmetric plot without outliers supports the condition for that group. The second group has strong right skew and an exceptionally high value, so its distribution does not provide the needed support. Having one acceptable group does not offset the problem in the other.

Conclude: The conditions do not adequately support a two-sample t interval based on these data. The concern is the second group’s small sample and strongly skewed distribution with an outlier. The interval should not be presented as trustworthy without addressing that limitation; the random sampling and 10% checks alone are not enough.

Worked Example: Large Samples With Some Skewness

An invented environmental survey compares mean daily household water use in two regions. Independent random samples include 42 households from a region with 900 households and 36 households from a region with 700 households. Plots show some right skew in both groups, but no extreme outliers. Assess the conditions for an interval estimating \(\mu_1-\mu_2\), in liters per household per day.

State: The parameter is \(\mu_1-\mu_2\), the true mean daily water use in region 1 minus that in region 2, in liters per household per day.

Plan: Check the random design, the 10% condition for each region, and whether both groups have large samples with no severe distributional features.

Do: The two samples are independent random samples. Their sampling fractions are:

$$ \frac{42}{900}\approx0.0467=4.67\%, \qquad \frac{36}{700}\approx0.0514=5.14\% $$

Both fractions are below 10%, so the 10% condition is met separately for each region. Both sample sizes are at least 30: \(n_1=42\) and \(n_2=36\). The plots show some skewness, but no extreme outliers. The large sample sizes support using the t procedure despite this moderate departure from Normality.

Conclude: The stated evidence supports a two-sample t interval for the difference in mean daily water use. The conclusion depends on both samples being large and the plots not showing extreme outliers; “large sample” should not be used to ignore severe problems in the data.

Common Mistakes and AP Exam Tips

  • Checking only one group: The Normality-or-large-sample condition applies to both groups. State evidence for each rather than treating the larger sample as a substitute for checking the other.
  • Pooling the sample sizes: Two samples of 18 do not make one sample of 36 for the distribution check. Each sample estimates a separate population mean.
  • Checking the 10% condition against the wrong population: Compare \(n_1\) with population 1 and \(n_2\) with population 2. If a population size is not given, explain what would be needed to verify the condition.
  • Treating “random” as a vague label: Say what was random—selection of the samples or assignment of subjects to groups. These processes support different conclusions.
  • Assuming \(n\geq30\) fixes every distribution problem: Thirty is a useful guideline, not permission to ignore extreme outliers or severe skewness. Mention the plots or data features.
  • Calling a condition met without evidence: Give the actual sample sizes and population sizes for the 10% check, or describe the relevant plot features. A complete answer makes the reasoning visible.

For full credit, identify the parameter and procedure, then address the design, sampling independence, and each group’s distribution in context. If information is missing, state that the condition cannot be confirmed rather than inventing evidence.

Key takeaway: A two-sample t interval needs more than two means and standard deviations. Verify the random design, check the 10% condition for each sample drawn without replacement, and assess Normality or adequate sample size in both groups.

Check Your Understanding

For each question, explain which condition is being checked and what evidence would support your answer.

  1. Independent random samples of 12 and 15 people are taken from populations of 180 and 240 people. Calculate each sampling fraction and decide whether the 10% condition is met.
  2. One group has \(n_1=45\) and the other has \(n_2=18\). Why is it not enough to say the combined sample size is 63?
  3. A study uses volunteer participants who are randomly assigned to two groups. What does random assignment support, and why is it not the same as random sampling?
  4. Both groups have \(n\geq30\), but one plot shows an extreme outlier. What should you say about the Normality-or-large-sample condition?
  5. If the problem gives no population sizes for samples drawn without replacement, how can you discuss the 10% condition accurately?